EconBase
← Back to paper

Partial Identification of Policy-Relevant Treatment Effects with Instrumental Variables via Optimal Transport

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

142,003 characters · 24 sections · 77 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Partial Identification of Policy-Relevant Treatment Effects with Instrumental Variables via Optimal Transport

abstractPolicy-Relevant Treatment Effects (PRTEs) are generally not point-identified under standard Instrumental Variable (IV) assumptions when the instrument generates limited support in treatment propensity. We show that PRTE partial identification in the generalized Roy model can instead be formulated as a Constrained Conditional Optimal Transport (CCOT) problem over the joint conditional law of the potential outcome and the latent resistance. The resulting multidimensional CCOT problem reduces analytically to separable one-dimensional OT problems with product costs, yielding sharp closed-form bounds and avoiding direct solution of the original high-dimensional CCOT problem. We also develop estimation and inference procedures for these bounds: for discrete instruments, we use a Double Machine Learning (DML) approach based on Neyman-orthogonal scores that accommodates high-dimensional covariates while achieving the parametric $\sqrt{n}$ rate and asymptotic normality; for continuous instruments, we explicitly characterize the corresponding nonparametric convergence rates. The framework accommodates covariates, discrete and continuous instruments, and extensions to general treatment settings. In simulations and a bed-net subsidy application, the resulting bounds are substantially tighter than the moment-relaxation method.



Introduction

Instrumental variables (IVs) are a standard tool for identifying causal effects when treatment is endogenous in observational studies. Under the classical IV assumptions, angrist1995identification show that the Local Average Treatment Effect (LATE), defined as the average effect for the subpopulation of compliers, is point-identified. However, many policy questions are not about compliers themselves; they ask how outcomes would change under a counterfactual policy that shifts treatment take-up in the broader population. Such questions are naturally captured by the Policy-Relevant Treatment Effects (PRTEs), a central target in the generalized Roy and Marginal Treatment Effect (MTE) framework heckman1999local,heckman2005structural,heckman2006understanding, which encompass a wide range of causal estimands including the Average Treatment Effect (ATE). Empirical policy evaluations within this framework include carneiro2010evaluating,carneiro2011estimating. Unlike LATE, PRTEs are generally not point-identified when the instrument induces only limited support in the treatment propensity.

This limited-support problem is central in the generalized Roy model heckman2005structural, where treatment selection follows a threshold-crossing mechanism formally equivalent to the classical monotonicity condition angrist1995identification,vytlacil2002independence. In this framework, PRTEs can be written as weighted averages of the MTE, so identification depends on how much of the latent resistance distribution is reached by the propensity score induced by the instrument. When the propensity score lacks full support over the unit interval, PRTEs and related quantities such as the ATE remain only partially identified, and a long line of work has studied such identified sets under various assumptions manski1990nonparametric,manski1997monotone,manski_pepper2000monotone,heckman2001instrumental,magne2018using.

Despite its importance, deriving tight, closed-form bounds for general PRTEs remains a significant challenge. Existing strategies fall into three broad approaches: moment-relaxation over MTR functions magne2018using,mogstad2018identification, second-order stochastic dominance constraints on the MTE function solved by rearrangement marx2024sharp, and infinite-dimensional linear programming over the joint law of outcome and latent resistance solved by sieve approximation han_yang2024computational. Each faces limitations---discarded distributional information, restricted scope to weights on the complier interval under a binary instrument, lack of closed-form solutions or incompatibility with continuous instruments---that we discuss in detail in (ref).

In this paper, we develop a unified Optimal Transport (OT) framework for partial identification in the generalized Roy model, and use it to derive sharp, closed-form bounds for PRTEs. Specifically, we recast PRTE partial identification as a Constrained Conditional Optimal Transport (CCOT) problem, a structured instance within the broader OT framework that is equivalent to the infinite-dimensional LP of han_yang2024computational. Rather than summarizing the observed distributions through moment conditions or imposing distributional constraints on conditional mean (MTE) functions, we place the distributional constraints directly on the joint conditional law of the potential outcome and latent resistance. Then, we exploit the OT structure to prove that this heavily constrained, multidimensional CCOT problem analytically reduces to separable standard one-dimensional OT problems. This reduction completely bypasses computationally intensive linear programming and sieve approximation, yielding explicit, closed-form solutions for the sharp bounds. Because the framework constrains the full joint latent law rather than a conditional mean, it applies in principle to a broader class of partial identification problems in the generalized Roy model; here we focus on PRTEs.

Our main contributions can be summarized as follows.

enumerate• A Unified OT Framework for the Generalized Roy Model: We formulate partial identification in the generalized Roy model as a CCOT problem. The framework imposes distributional constraints directly at the joint law level and we prove that constraints decomposes analytically into separable marginal constraints. • Analytic Decomposition and Closed-Form PRTE Bounds: Leveraging the decomposition, we prove the CCOT problem decomposes into separable one-dimensional OT problems with product costs, yielding explicit, closed-form sharp bounds and bypassing the sieve approximation required by existing computational approaches. • Semi-parametric Estimation and Inference: We develop estimation results for the derived bounds. For discrete instruments, we leverage Double Machine Learning (DML) to construct a debiased estimator, achieving $\sqrt{n}$-consistency and asymptotic normality. For continuous instruments, we explicitly characterize the nonparametric convergence rates. • Generalization to General Treatments: We extend our identification strategy beyond the binary treatment setting to continuous and multi-valued treatments, demonstrating the flexibility of the OT framework. For the multi-valued treatment case, this extension requires an additional algebraic condition on the ranges of the propensity scores induced by the instrument ((ref)). • Empirical Validation: We validate our theoretical results through both synthetic simulations and a real-world empirical application, demonstrating how our closed-form bounds tighten the identified set relative to the moment-relaxation methods.

\paragraph{Organization of this paper.} The remainder of the paper is organized as follows. (ref) introduces the setup and optimal transport background; (ref) presents the main identification theory and closed-form bounds; (ref) extends the framework to general treatments; (ref) develops estimation and inference; and (ref) reports simulation and empirical results.

Related Work

\paragraph{Local IV and MTE.} The Marginal Treatment Effect (MTE), introduced by heckman1999local and systematically developed by heckman2005structural, is the foundational building block of our identification analysis. The MTE is defined as the treatment effect for individuals at a specific quantile of unobserved resistance to treatment, and serves as a unifying object from which all standard IV estimands--including LATE, ATE, ATT, and PRTE--can be recovered as weighted averages heckman2005structural, heckman2006understanding. The formal equivalence between the threshold-crossing selection model and the monotonicity assumption of angrist1995identification is established by vytlacil2002independence, who shows that the two frameworks are equivalent under a continuity condition on the unobservable. Point identification of the MTE is achieved via the Local Instrumental Variables (LIV) estimand heckman1999local, which requires full propensity score support over $[0,1]$---a condition that fails under limited instrument variation, the regime our paper addresses. Empirical MTE-based policy evaluations include carneiro2010evaluating,carneiro2011estimating.

\paragraph{IV with General Treatment.} angrist1995two extends the LATE framework to treatments with multiple discrete ordered levels. heckman2007econometric extends the MTE framework to the ordered choice model. kirkeboen2016field provides methodology for identifying treatment effects in settings with multiple unordered discrete choices--specifically fields of study--by accounting for how individuals self-select into alternatives based on their unobserved comparative advantages. lee2018identifying develop identification results for multivalued treatments. In the continuous treatment setting, florens2008identification use the control function method to identify the ATE and ATT, and imbens2009identification use a similar approach to identify structural equations in models with continuous endogenous variables. While these results establish point identification for certain causal quantities, our framework instead targets partial identification of PRTEs, and we extend our results to the continuous and multi-valued treatment settings in (ref).

\paragraph{Partial Identification.} manski1989anatomy,manski1990nonparametric,manski1997monotone,manski_pepper2000monotone,manskiInferenceRegressionsInterval2002 pioneered partial identification, providing closed-form bounds for the IV problem under a range of assumptions. In the econometric literature, partial identification is often cast as estimating the identified set defined by moment inequalities chernozhukov_hong_tamer2007,rosen2008confidence,romano2008inference,canay2010inference,bugni2010bootstrap; see tamer2010partial for a comprehensive survey. As a special case, our bound recovers the classical Manski bound for ATE. The computational and analytical approaches of han_yang2024computational and marx2024sharp to PRTE partial identification in the MTE framework are discussed in the PRTE Partial Identification paragraph below.

balke_pearl1997bounds develop the linear programming approach for partial identification. Recently, this optimization-based approach has been revisited and generalized using modern techniques balazadeh_meresht_syrgkanis_krishnan2022,guo_yin_wang_jordan2022,duarte2024automated,voronin2025linear,levis2025covariate. In particular, levis2025covariate use covariates to tighten balke_pearl1997bounds's classical bound and develop estimation theory in the IV setting. Several works establish connections between partial identification and robust optimization guo_yin_wang_jordan2022,balazadeh_meresht_syrgkanis_krishnan2022,gao_ge_qian2024,penn_gunderson_bravohermsdorff_silva_watson2025,fan_pass_shi2025,tan2024consistency. The partial identification problem can be formulated as an optimal transport problem with marginal constraints derived from the observational distribution ji_lei_spector2024,lin2025estimation,lin2025tightening. For a recent primer on optimal transport methods for causal inference, see gunsilius2025primer. In a related but distinct setting, fan_guerre_zhu2017partial study partial identification of functionals of the joint distribution of potential outcomes in settings where the conditional marginal distributions are identified---including the latent threshold-crossing model of heckman2005structural under full propensity-score support. oberreynolds2023estimating characterizes sharp bounds on functionals of the joint complier outcome distribution via optimal transport under the classical binary-instrument LATE framework of angrist1995identification, focusing on the identified complier subpopulation rather than the partially identified PRTEs we target here.

\paragraph{PRTE Partial Identification.} Three strategies have emerged for PRTE partial identification under limited support. magne2018using optimize over MTR functions subject to moment constraints mogstad2018identification, which discards distributional information beyond the first moment. han_yang2024computational preserve full distributional content by formulating an infinite-dimensional LP over the joint law of latent states, solved by sieve approximation for discrete instruments. marx2024sharp generalize magne2018using by imposing full distributional content on the MTE function via a second-order stochastic dominance (SSD) constraint, obtaining sharp closed-form bounds by rearrangement bounds; his explicit theorems cover weights on the complier interval under a binary instrument, with extensions beyond compliers requiring rank similarity.

Our work differs from all three in scope: we impose distributional constraints directly on the joint conditional law of the potential outcome and latent resistance rather than on a conditional mean, cast the resulting program as a CCOT problem equivalent to the LP of han_yang2024computational, and show it reduces analytically to separable one-dimensional OT problems with product costs. This delivers (i) sharp closed-form bounds for the full PRTE on all of $[0,1]$ under multi-valued instruments via an interval decomposition across $K\!+\!1$ compliance types, including boundary intervals where only one potential outcome marginal is identified, (ii) extensions to continuous and mixed instruments with sub-problems indexed by the propensity score continuum, and (iii) discrete-instrument inference together with continuous-instrument rate results, which are not available in either marx2024sharp or han_yang2024computational. On the complier interval under a binary instrument, the CCOT solution and Marx's SSD rearrangement coincide. Beyond PRTE, our sharpness result characterizes the identified set at the level of the full joint latent law and sheds light on the joint-identification question raised in marx2024sharp through its reduction to independent one-dimensional marginal constraints.



Preliminaries

IV Model and Policy-Relevant Treatment Effects

For most of this paper, we work with the generalized Roy model heckman2005structural. In the model, the observable variables are treatment $W\in \mathcal{W} = \{0,1\}$, instrument $ Z \in \mathcal{Z}$, covariates $X\in\mathcal{X}$ and the outcome $Y \in \mathcal{Y}$, where $\mathcal{X} \subset\mathbb{R}^{d_X}, \mathcal{Z}$ and $ \mathcal{Y} \subset \mathbb{R}$ are the domains of the corresponding variables. We consider both discrete and continuous instrument settings in this paper. Throughout the paper, we assume the domains of all random variables are compact. The unobservables are the potential outcomes $Y(0), Y(1)$, and the variable $U$ in the selection equation ((ref)).

assumption[Structural Assumptions] We will make the following structural assumptions for our IV model. Throughout, we assume $Y(w) \in \mathcal{Y}$, $Z \in \mathcal{Z}$, and $X \in \mathcal{X}$ almost surely. \begin{enumerate} • (Consistency) $Y = WY(1) + (1-W)Y(0)$. • (Conditional Instrumental Exogeneity) $Z \perp\!\!\!\perp (U,Y(0),Y(1)) \mid X$, where $\perp\!\!\!\perp$ denotes statistical independence. • (Threshold Crossing) The treatment is selected by \begin{align} W = \indicator (U\leqslant p(Z,X)), \end{align} where $p(Z,X) = \mathbb{P}(W=1 \mid Z, X)$. Moreover, $U \mid X \sim \mathrm{Unif}(0,1)$. • (Common Instrument Support) The conditional support of $Z$ given $X$ does not depend on $X$: $\mathrm{Supp}(Z \mid X=x) = \mathcal{Z}$ for all $x \in \mathcal{X}$. \end{enumerate}

It is without loss of generality to assume that $U\sim\text{Unif}(0,1)$ if $U$ is continuously distributed, because we can always normalize its marginal distribution.\footnote{By the probability integral transform, if $U \mid X$ has a continuous CDF $F_{U\mid X}(\cdot \mid x)$, then $\tilde{U} := F_{U\mid X}(U \mid X) \sim \mathrm{Unif}(0,1)$ conditional on $X$. Since $\tilde{U}$ is a strictly increasing function of $U$, the threshold-crossing model is preserved under this reparametrization, and we may work with $\tilde{U}$ in place of $U$ without loss of generality.} The common instrument support condition is essential for the identification of the propensity score $p(z,x)$: without $z \in \mathrm{Supp}(Z \mid X=x)$ for all $x$, the conditional probability $\mathbb{P}(W=1 \mid Z=z, X=x)$ is not well-defined and hence $p(z,x)$ is not identified.

We also require the following boundedness condition on the potential outcomes, which is needed to obtain informative bounds.

assumption[Boundedness of the Potential Outcome] The outcome domain satisfies $\mathcal{Y} \subseteq [y_{\min}, y_{\max}]$ for known constants $y_{\min} < y_{\max}$.

The fundamental building block for evaluating causal parameters in this framework is the MTE, originally introduced by heckman1999local,heckman2005structural. The MTE is defined as the average treatment effect for individuals with covariates $X=x$ and latent variable $U=u$: $$ \text{MTE}(x,u) = \mathbb{E}[Y(1) - Y(0) \mid X=x, U=u]. $$

In policy analysis, we are often interested in the effect of an intervention that shifts the distribution of the instrument $Z$, or more generally, alters the propensity score from a baseline status to a new regime. Let the baseline policy be characterized by the current propensity score $p(Z,X)$ and the corresponding treatment status $W$. Consider an alternative policy that induces a new propensity score $q(Z,X)$ and a new counterfactual treatment status $W^q = \indicator(U \leqslant q(Z,X))$, where we assume the latent variable $U$ is invariant to the policy change, as is standard in the MTE framework.

The PRTE evaluates the average per-person impact of shifting from the baseline policy to the alternative policy. Let $Y^q$ denote the outcome realized under the alternative policy, such that $Y^q = W^qY(1) + (1-W^q)Y(0)$. The PRTE is formally defined as: $$ \text{PRTE} = \frac{\mathbb{E}[Y^q - Y]}{\mathbb{E}[W^q - W]}. $$

Many common treatment effect parameters can be expressed as a weighted average of the MTE. Specifically, the PRTE can be rewritten as: $$ \text{PRTE} = \mathbb{E}_{X,U}[\text{MTE}(X, U) \omega_{\text{PRTE}}(X,U)], $$ where the policy-specific weight function is given by: $$ \omega_{\text{PRTE}}(x,u) = \frac{\mathbb{P}(q(Z,X) \geqslant u \mid X=x) - \mathbb{P}(p(Z,X) \geqslant u \mid X=x)}{\mathbb{E}[q(Z,X) - p(Z,X)]}, $$ provided $\mathbb{E}[q(Z,X) - p(Z,X)] \neq 0$. Following heckman2005structural,magne2018using, we consider the following more general target parameter.

align[align omitted — 119 chars of source]

where $\omega(x,u)$ is an identifiable function. We can obtain different PRTEs by taking

align[align omitted — 147 chars of source]

for given propensity functions $q_0$ and $q_1$, where $q_0 = p$ and $q_1 = q$ correspond to the baseline and alternative propensity score functions, respectively, in the PRTE case. The coefficient hidden inside the proportionality symbol is identifiable. (ref) shows a variety of common causal parameters under different choices of $\omega$.

table[table omitted — 1,156 chars of source]

While our bounds recover the more classical bound for the first 4 quantities in (ref), we will focus in particular on PRTEs. As a leading example, consider a uniform policy shift of the form $q(Z,X) = \min(p(Z,X) + \alpha, 1)$ for some $\alpha > 0$. This corresponds to a policy that uniformly raises the probability of treatment uptake by $\alpha$. In our empirical application ((ref)), the instrument is the offered price of an insecticide-treated bed net dupas2014subsidies, and $q = p + \alpha$ models a uniform price subsidy that lowers the cost of purchase and thereby increases take-up probability by $\alpha$.

heckman2001instrumental shows that if the support of $p(Z,x)$ spans over the entire interval $(0,1)$ for all $x$, then MTE is identifiable. However, because the common support of $p(Z,X)$ is often a strict subset of $(0,1)$, the MTE is generally not identifiable. For instance, when both $X$ and $Z$ are discrete, the supports of $p(Z,x)$ are discrete and this common support assumption is violated. Consequently, the integral defining the PRTE is not point-identified without imposing strong parametric extrapolations.

A Primer on Optimal Transport

To build the framework for our partial identification results, we briefly introduce the relevant concepts from OT theory. OT provides a mathematical framework for coupling two probability distributions at minimum cost, and has emerged as a powerful toolkit for bounding joint distributions when only their marginals are known galichon2016optimal.

Let $\mu$ and $\nu$ be two probability measures on state spaces $\mathcal{A}$ and $\mathcal{B}$, respectively. In the context of partial identification, we often observe the marginal distributions of two variables but not their joint distribution. We define $\Pi(\mu, \nu)$ as the set of all couplings between $\mu$ and $\nu$. Formally, a coupling $\pi \in \Pi(\mu, \nu)$ is a joint probability measure on the product space $\mathcal{A} \times \mathcal{B}$ such that its marginals are exactly $\mu$ and $\nu$.

Given a cost function $c: \mathcal{A} \times \mathcal{B} \to \mathbb{R} \cup \{+\infty\}$ that represents the cost of associating state $a$ with state $b$, the standard Kantorovich optimal transport problem seeks the coupling that minimizes the expected cost:

align[align omitted — 138 chars of source]

Generally, solving ((ref)) in multidimensional spaces requires computationally intensive linear programming. However, a celebrated result in OT literature demonstrates that when the spaces are one-dimensional and the cost function satisfies specific structural conditions--such as supermodularity--the optimal coupling admits a sharp, closed-form solution.

We formalize this one-dimensional closed-form solution in the following theorem, which will be instrumental in deriving the analytic bounds for the PRTE.

theorem[1D Optimal Transport with Product Cost, e.g., villani2021topics, galichon2016optimal] Let $\mu$ and $\nu$ be probability measures on $\mathbb{R}$ with quantile functions $Q_\mu$ and $Q_\nu$. Suppose the cost function is a simple product $c(a,b) = a \cdot b$. Then, the minimum and maximum of the expected cost over all valid couplings are completely characterized by the countermonotonic and comonotonic couplings, respectively: \begin{align*} \min_{\pi \in \Pi(\mu, \nu)} \int_{\mathbb{R}^2} a \cdot b \, \mathrm{d}\pi(a, b) &= \int_0^1 Q_\mu(t) Q_\nu(1-t) \, \mathrm{d}t, \\ \max_{\pi \in \Pi(\mu, \nu)} \int_{\mathbb{R}^2} a \cdot b \, \mathrm{d}\pi(a, b) &= \int_0^1 Q_\mu(t) Q_\nu(t) \, \mathrm{d}t. \end{align*} The optimal coupling is attained by the countermonotonic coupling $(Q_\mu(\xi), Q_\nu(1-\xi))$ for the minimum, and the comonotonic coupling $(Q_\mu(\xi), Q_\nu(\xi))$ for the maximum, where $\xi \sim \mathrm{Unif}(0,1)$.

(ref) provides a closed-form characterization of the optimal value of 1D OT problems. We will leverage this characterization to derive tight bounds for PRTEs later.

Notation

We use $\mathbb{P}$ and $\mathbb{E}$ for probability and expectation, and $\perp\!\!\!\perp$ for conditional independence. $\indicator(\cdot)$ denotes the indicator function. We write $\mathrm{d}$ for the differential or Lebesgue measure. For a probability measure $\mu$ on $\mathbb{R}$, $Q_\mu(t)$ denotes its quantile function. $\Pi(\mu, \nu)$ denotes the set of all couplings of two probability measures $\mu$ and $\nu$. Calligraphic letters $\mathcal{X}, \mathcal{Y}, \mathcal{Z}, \mathcal{W}$ denote domains of the corresponding variables. We write $O_P(\cdot)$ and $o_P(\cdot)$ for stochastic order notation: $X_n = O_P(a_n)$ means $X_n/a_n$ is bounded in probability, and $X_n = o_P(a_n)$ means $X_n/a_n \xrightarrow{P} 0$.



Partial Identification via Optimal Transport

CCOT Formulation of Partial Identification

In this section, we demonstrate how to derive closed-form, tight bounds for the PRTE. We begin by formally mapping the partial identification problem into an optimal transport framework.

The core intuition is to search over all possible joint distributions of the latent and observable variables $(Y(0), Y(1), U, X, Z, W)$ that are compatible with both the observed data and the structural conditions in (ref). Because the potential outcomes are conditionally independent of the instrument given $X$, the identifying power of our model is entirely captured by the joint distributions of $(Y(w), U, X)$ for $w \in \{0, 1\}$.

Let $\pi_w$ denote the joint probability measure of $(Y(w), U, X)$ and $ \pi_w(\cdot\mid x,u) $ be the conditional distribution of $Y(w)$ given $x,u$. Recall that our target parameter is a linear functional of these measures: $$\theta_\omega(\pi_0, \pi_1) = \mathbb{E}_{(X,U,Y(1))\sim\pi_1} [Y(1) \omega(X,U)] - \mathbb{E}_{(X,U,Y(0))\sim\pi_0} [Y(0) \omega(X,U)]. $$ To achieve sharp bounds, we must constrain $\pi_w$ using all available observational information. Let $\mathbb{P}_{obs}$ denote the true observed joint distribution of $(Y, W, Z, X)$. Under the threshold-crossing model in (ref), while $U$ is generally unobserved, its conditional distribution is explicitly tied to the treatment status and the propensity score. Specifically, conditional on $X$ and $Z$, the event $W=1$ perfectly corresponds to $U \leqslant p(Z,X)$, and $W=0$ corresponds to $U > p(Z,X)$.

Consequently, the joint distributions $\pi_0$ and $\pi_1$ fully parameterize the observational distribution of $(X,Z,W,Y)$. To see this, we can link the unknown structural distributions directly to the observed conditional distribution of the data by exploiting the consistency equation $Y = W Y(1) + (1-W) Y(0)$ and the selection mechanism $W = \indicator(U \leqslant p(Z,X))$. For any measurable set $A \subseteq \mathcal{Y}$ and $W=1$, we have:

align*[align* omitted — 123 chars of source]

By the conditional independence assumption $Z \perp\!\!\!\perp (Y(1), U) \mid X$, we can omit the $Z$ in the second conditioning:

align*[align* omitted — 130 chars of source]

Since instruments sharing the same propensity score $p(z,x)$ yield the same right-hand side value, we can replace the conditioning on $Z$ with conditioning on $p(Z)$. Next, we disintegrate the joint measure $\pi_1$ into the conditional distribution of the potential outcome given the covariates and the latent variable. Because $U \mid X \sim \text{Unif}(0,1)$, integrating over the region $U \leqslant p(z,x)$ yields:

align*[align* omitted — 113 chars of source]

Following the exact same sequence of steps for $W=0$, where the observed outcome is $Y(0)$ and the selection mechanism dictates $U > p(z,x)$, we obtain the symmetric result:

align*[align* omitted — 113 chars of source]

Furthermore, by construction, the marginal distribution of $(U, X)$ under $\pi_w$ must equal the known population distribution $\mathbb{P}_{U,X}$.

This allows us to formulate the partial identification of the PRTE as the following infinite dimensional linear programming problem:

equation[equation omitted — 736 chars of source]

We call the optimization in (ref) the Constrained Conditional Optimal Transport (CCOT) problem. The feasible set is constrained by the observational compatibility conditions in (ref); the problem is conditional in that, for fixed covariates $x$ and latent resistance $u$, it seeks a conditional measure $\pi_w(\cdot \mid u, x)$ on $\mathcal{Y}$; and it is a transport problem because the objective is linear in $\pi_w(\cdot \mid u, x)$ with effective cost $c(y, u, x) = y \cdot \omega(x, u)$, so that mass is transported from the latent space $U$ to the outcome space $Y(w)$ conditional on $X$.

To formally guarantee that the CCOT formulation in ((ref)) yields sharp bounds for the PRTE, we must show that any pair of measures $(\pi_0, \pi_1)$ satisfying the CCOT constraints corresponds to a valid data-generating process that satisfies our structural assumptions and perfectly reproduces the observed data. Let $\Gamma(\mathbb{P}_{obs})$ be the set of all measure pairs $(\pi_0, \pi_1)$ satisfying the constraints in ((ref)). The following proposition characterizes the joint distribution of $(Y(0), Y(1), U, X)$, from which observational equivalence of the marginal pair $(\pi_0, \pi_1)$ follows immediately.

proposition[Sharp Identified Set of the Joint Latent Law] Suppose the observed distribution $\mathbb{P}_{obs}$ is generated by a true structural model satisfying (ref). Let $\tilde{\pi}$ be any probability measure on $\mathcal{Y} \times \mathcal{Y} \times [0,1] \times \mathcal{X}$ whose $(Y(w), U, X)$-marginals $\tilde{\pi}_w$ satisfy $(\tilde{\pi}_0, \tilde{\pi}_1) \in \Gamma(\mathbb{P}_{obs})$. Then there exists a joint distribution of the latent and observable variables $(Y(0), Y(1), U, X, Z, W)$ such that: \begin{enumerate} • the joint marginal distribution of $(Y(0), Y(1), U, X)$ is exactly $\tilde{\pi}$; • (ref) holds for $(Y(0), Y(1), U, X, Z, W)$; • the induced distribution of $(Z, X, W, Y)$ exactly matches $\mathbb{P}_{obs}$. \end{enumerate} Conversely, every joint marginal $\tilde{\pi}$ arising from a distribution of $(Y(0), Y(1), U, X, Z, W)$ satisfying (ii)--(iii) has $(Y(w), U, X)$-marginals in $\Gamma(\mathbb{P}_{obs})$ for $w \in \{0, 1\}$.

The content of (ref) is twofold. First, the CCOT constraints act only on the two marginals $\tilde{\pi}_0$ and $\tilde{\pi}_1$: any coupling of these marginals---including every joint law of $(Y(0), Y(1))$ given $(U, X)$---is observationally equivalent. Second, the converse makes the characterization sharp: no joint law outside this set can be generated by a structural model matching $\mathbb{P}_{obs}$. In particular, applying (ref) to the independent coupling $\tilde{\pi} := \tilde{\pi}_0 \otimes_{(U, X)} \tilde{\pi}_1$ shows that every marginal pair $(\tilde{\pi}_0, \tilde{\pi}_1) \in \Gamma(\mathbb{P}_{obs})$ is observationally equivalent to the true data-generating process. Consequently, the identified set for the PRTE target is $\{\theta_\omega(\pi_0, \pi_1) : (\pi_0, \pi_1) \in \Gamma(\mathbb{P}_{obs})\}$, and the maximum and minimum of $\theta_\omega(\pi_0, \pi_1)$ over $\Gamma(\mathbb{P}_{obs})$ are the sharp upper and lower bounds on the PRTE.

remarkThe fact that the moment relaxation approach of magne2018using can yield arbitrarily loose bounds compared to our CCOT formulation is well-recognized; see also han_yang2024computational. A detailed comparison example is given in (ref) in the appendix.
remark[Beyond PRTE: CCOT for general joint functionals] The CCOT perspective is not specific to the PRTE. At a conceptual level, one could consider a bounded measurable cost $f:\mathcal{Y}^2 \to \mathbb{R}$ and the joint target $\theta_f := \mathbb{E}[f(Y(1),Y(0))]$ by replacing the PRTE objective in ((ref)) with $\mathbb{E}_{\pi}[f(Y(1),Y(0))]$ and enlarging the decision variable from the pair $(\pi_0,\pi_1)$ to a joint conditional kernel $\pi({\rm d} y_1,{\rm d} y_0 \mid u,x)$ whose marginals satisfy the same observational constraints. A sharp characterization should again follow from the lifting logic behind (ref). What is special about the PRTE is the separable product cost $y \cdot \omega(x,u)$: it allows the CCOT problem to collapse into one-dimensional OT problems for each potential outcome separately. For a general joint functional, that separability is lost, and the resulting interval-specific problems are instead OT problems over the coupling between the identified marginals of $Y(1)$ and $Y(0)$, with no comparable closed-form reduction in general. Since the present paper is about the PRTE case, we do not pursue this extension here.

Next, we show how to obtain the closed-form solution of the CCOT problem ((ref)). Notice that in ((ref)), the objective is the difference of two linear functionals of $\pi_0$ and $\pi_1$, and the constraints on $\pi_0$ and $\pi_1$ are separate. Therefore, we only need to consider one side of the problem, as the other side follows identically. In what follows, we only consider the $W=1$ side minimization problem, i.e.,

equation[equation omitted — 529 chars of source]

The solution for the maximization problem can be derived similarly. Writing $\theta_\omega = \theta_{\omega,1} - \theta_{\omega,0}$, let $[\underline{\theta}_{\omega,w}, \overline{\theta}_{\omega,w}]$ denote the sharp identified interval for the $w$-th component. Because the feasible sets for $\pi_0$ and $\pi_1$ are Cartesian products, the sharp identified interval for the full target is \[ [\underline{\theta}_\omega, \overline{\theta}_\omega] = [\underline{\theta}_{\omega,1} - \overline{\theta}_{\omega,0},\, \overline{\theta}_{\omega,1} - \underline{\theta}_{\omega,0}]. \] Accordingly, we state the main derivations for $\theta_{\omega,1}$; the formulas for $\theta_{\omega,0}$ are obtained by the same arguments with the untreated analogues.

Closed-Form Bounds without Covariates

To better illustrate the idea, we first derive the results without covariates and generalize them in the next subsection. Therefore, in this subsection, $\omega$ and $p$ do not depend on $x$.

Discrete Instrument Setting

We first consider the setting where there are no covariates and the instrument $Z$ takes values in a finite discrete set $\mathcal{Z}$. Let the unique values of the propensity score be ordered as: \[ \mathrm{Range}(p(z)) := \{p(z) : z \in \mathcal{Z}\} = \{p_1, \dots, p_K\}, \] with the conventions $p_0 := 0$ and $p_{K+1} := 1$, such that $0 = p_0 \leqslant p_1 < p_2 < \dots < p_K \leqslant p_{K+1} = 1$. We allow different instruments to be mapped to the same propensity value. This discretizes the unit interval of the latent variable $U$ into $K+1$ disjoint sub-intervals, $I_i = (p_i, p_{i+1}]$ for $i=0, \dots, K$.

Because the objective function ((ref)) is an integral over $U$, we can additively decompose it across these sub-intervals:

align*[align* omitted — 132 chars of source]

The fundamental advantage of the discrete setting is that the global observational constraints difference out, perfectly isolating the conditional distribution of $Y(1)$ within each interval $I_i$. Recall that from ((ref)),

align[align omitted — 166 chars of source]

The key insight of this work is that ((ref)) can be decomposed into separate marginal constraints, which enables us to convert ((ref)) into a series of seperate standard one-dimensional OT problems. For $i = 1, \dots, K-1$, subtracting the $i$-th constraint from the $(i+1)$-th in ((ref)) isolates the integral over $(p_i, p_{i+1}]$; for $i=0$, the first constraint directly governs $(0, p_1]$. Since cumulative sums recover the original constraints, this invertible transformation shows that the system ((ref)) is equivalent to

align*[align* omitted — 149 chars of source]

where $\mu_{1,i}$ is given by

align[align omitted — 376 chars of source]

Note that for the final interval $I_K = (p_K, 1]$, the potential outcome $Y(1)$ is never observed because no individual in the population has a propensity score high enough to induce treatment when $U > p_K$. By ((ref)), since $\pi_1$ is nonnegative, $\mathbb{P}_{obs}(\mathrm{d}y, W=1 \mid p(Z)=p)$ is non-decreasing in $p$. Note that $\mu_{1,i}(\mathcal{Y}) = 1$ since $\mathbb{P}_{obs}(W=1 \mid p(Z)=p) = p$. Therefore, the $\mu_{1,i}$ are valid distributions. mourifie2017testing provide a simple method to test the monotonicity of $\mathbb{P}_{obs}(\mathrm{d}y, W=1 \mid p(Z)=p)$.

For $0<i<K$, notice that

align*[align* omitted — 360 chars of source]

for any measurable set $A$, where we use (ref).(3) and consistency in the second equality and (ref).(2) in the last equality. Therefore, $\mu_{1,i}$ can be interpreted as the distribution of the potential outcome under treatment for the $i$-th complier group. The $i$-th complier group consists of units that refuse treatment when $p(Z) \leqslant p_i$ and comply only when $p(Z) > p_i$. In particular, when the instrument is binary, this recovers the identification results of imbens1997estimating

Substituting the objective decomposition and the interval-wise constraints derived above, the original problem ((ref)) can be equivalently written as

equation[equation omitted — 206 chars of source]

subject to the $K$ independent marginal constraints

equation[equation omitted — 267 chars of source]

with the support constraint $\text{supp}(\pi_1(\cdot \mid u)) \subseteq \mathcal{Y}$ for all $u \in [0,1]$, and no further distributional constraint on $\pi_1(\cdot \mid u)$ for $u \in I_K = (p_K, 1]$. Crucially, each constraint in ((ref)) involves $\pi_1$ only on the single interval $I_i$, and the objective ((ref)) is additively separable across these intervals. Therefore, the global minimization problem decomposes into $K+1$ independent subproblems, one for each interval.

Specifically, for each interval $i = 0, \dots, K-1$, the data perfectly fixes the marginal distribution of $Y(1)$ as $\mu_{1, i}$ and the marginal distribution of $U$ as uniform on $(p_i, p_{i+1}]$. Let $\nu_i$ denote the uniform probability measure $\text{Unif}(p_i, p_{i+1})$. The set of valid joint distributions for $(Y(1), U)$ conditional on $U \in I_i$ is exactly the set of all couplings $\Pi(\mu_{1, i}, \nu_i)$. Therefore, for each of these $K$ intervals, we solve a standard 1D optimal transport problem:

equation[equation omitted — 164 chars of source]

Conversely, for the final interval $I_K = (p_K, 1]$, the propensity score never reaches a value high enough to induce treatment, meaning the data provides absolutely no constraints on the distribution of $Y(1)$ for this subpopulation. Without distributional constraints, the optimal transport problem degenerates into a trivial pointwise minimization. By (ref), for each $u \in I_K$, the optimal solution assigns all probability mass to $y_{\min}$ if $\omega(u) \geqslant 0$, or to $y_{\max}$ if $\omega(u) < 0$:

equation[equation omitted — 252 chars of source]

Summing the optimal values of the $K$ constrained subproblems in ((ref)) scaled by their respective interval lengths $(p_{i+1} - p_i)$ and adding the trivial unconstrained bound from ((ref)), we obtain the global minimum. Applying (ref) to each subproblem in ((ref)) yields the countermonotonic coupling, which leads directly to the following closed-form sharp bounds.

theorem[Closed-Form Sharp Bounds] Suppose (ref) and (ref) hold, $\mathcal{X} = \emptyset$, and $\mathcal{Z}$ is finite. For $i=0, \dots, K-1$, let $Q_{Y, i}(t)$ be the quantile function of the identified conditional distribution $\mu_{1, i}$ defined in ((ref)). Let $Q_{\omega, i}(t)$ be the quantile function of the random variable $\omega(U)$ where $U \sim \text{Unif}(p_i, p_{i+1})$. The sharp lower bound for the $W=1$ component of the PRTE is given by: \begin{equation} \theta_{\omega, 1} = \sum_{i=0}^{K-1} (p_{i+1} - p_i) \int_0^1 Q_{Y, i}(t) Q_{\omega, i}(1-t) \, \mathrm{d}t \ + \ \int_{p_K}^1 \Big( y_{\min} \max\{0, \omega(u)\} + y_{\max} \min\{0, \omega(u)\} \Big) \mathrm{d}u. \end{equation} Similarly, the sharp upper bound is given by replacing $Q_{\omega,k}(1-t)$ with $Q_{\omega,k}(t)$ in the OT term and swapping $y_{\min}$ and $y_{\max}$ in the trivial tail term.

We defer the formal proof of this theorem to (ref), since the present theorem is a special case ($\mathcal{X}=\emptyset$, every LIV interval degenerate). (ref) provides a closed-form, analytically tractable expression for the sharp bounds, bypassing any need for linear programming and requiring only estimation of the conditional quantile function from the observed data.

As a special case, if we set $ \omega(u) = 1 $, (ref) recovers the classical bound by manski1990nonparametric for $\mathbb{E}[Y(1)]$, which is known to be tight heckman2001instrumental:

align*[align* omitted — 369 chars of source]

(ref) illustrates the three identification regions that compose the bound in (ref); (ref) in the appendix works through a concrete binary-instrument example with a constant propensity score, deriving the explicit three-part decomposition.

figure[figure omitted — 6,714 chars of source]

By (ref), the length of the bound is

align*[align* omitted — 255 chars of source]

which depends on the range of the propensity function and distribution $\mu_{1,i}$. If $p_K$ is close to $1$ and the support of $\mu_{1,i}$ is small, the length of the bound is also small. In particular, if each $Q_{Y,i}$ is constant and $p_K = 1,p_0 = 0$, $\theta_{\omega,1}$ is identifiable from the data.

remark[Decomposition into 1D marginal constraints and marx2024sharp] marx2024sharp raises the open question of characterizing the distribution of treatment effects for the complier group. Although (ref) states sharpness at the joint level of $(Y(0), Y(1), U, X)$, the CCOT constraints defining $\Gamma(\mathbb{P}_{obs})$ reduce to independent one-dimensional marginal constraints on the conditional kernels $\pi_w({\rm d} y \mid u, x)$. Consequently, once these marginals are pinned down, the remaining freedom lies entirely in the coupling between $Y(1)$ and $Y(0)$ given $(U, X)$. In particular, in the binary-instrument, no-covariate setting with propensity scores $p_1 < p_2$, the complier interval is $I_c=(p_1,p_2]$. By ((ref)), the treated-outcome distribution for compliers is identified as $\mu_{1,1}$. Symmetrically, the untreated-outcome distribution for the same complier group is identified as \[ \mu_{0,1}({\rm d} y) := \frac{\mathbb{P}_{obs}({\rm d} y, W=0 \mid p(Z)=p_1)-\mathbb{P}_{obs}({\rm d} y, W=0 \mid p(Z)=p_2)}{p_2-p_1}. \] Hence the identified set for the complier treatment-effect distribution is exactly \[ \mathcal I_{\Delta,c} = \left\{ (y_1-y_0)_{\#}\gamma : \gamma \in \Pi(\mu_{1,1},\mu_{0,1}) \right\}. \] Thus, Marx's question can be restated as follows: once the identified complier marginals are fixed, characterize all possible pushforwards of their couplings under the difference map $(y_1,y_0)\mapsto y_1-y_0$. We do not have a simpler closed-form characterization of $\mathcal I_{\Delta,c}$, but this makes precise that the remaining difficulty is entirely a coupling problem rather than a missing marginal-identification problem.

General Instrument Setting

We now unify consider a general instrument setting under a single regularity condition on the instrument domain. Rather than assuming $\mathcal{Z}$ is finite, we only require it to be compact with finitely many connected components. This framework nests the finite discrete case (each component a singleton), the connected continuous case of the LIV literature heckman1999local, and any intermixture of the two.

assumption[Compactness and Piecewise Continuity] The domain of the instrument $\mathcal{Z}$ is a compact set with finitely many connected components. Furthermore, for almost every $x \in \mathcal{X}$, the conditional propensity score function $z \mapsto p(z,x) = \mathbb{P}(W=1 \mid Z=z, X=x)$ is continuous on each connected component.

Under (ref), the image of each connected component of $\mathcal{Z}$ under $p(\cdot)$ is a closed interval (possibly a singleton if $p$ is constant on the component or if the component is a single point). Merging overlaps and ordering from left to right, the range of the propensity score admits a maximal-interval decomposition

equation[equation omitted — 299 chars of source]

where we explicitly allow degenerate intervals $\underline{p}_k = \overline{p}_k$ corresponding to isolated propensity values, as in the finite discrete case. Adopting the boundary conventions $\overline{p}_0 := 0$ and $\underline{p}_{K+1} := 1$, this decomposition partitions the unit interval of the latent variable $U$ into three types of regions: the LIV intervals $L_k := [\underline{p}_k,\overline{p}_k]$ for $k=1,\dots,K$, the OT gaps $G_k := (\overline{p}_k, \underline{p}_{k+1})$ for $k=0,\dots,K-1$, and the trivial tail $(\overline{p}_K,1]$.

The CCOT problem for the treated ($W=1$) component becomes

equation[equation omitted — 368 chars of source]

Following the same strategy as in (ref), we rewrite ((ref)) into an obviously separable form across the three region types. The objective splits additively as

align*[align* omitted — 363 chars of source]

On each gap $G_k$, the same telescoping argument that produced ((ref)) --- differencing the observational constraint in ((ref)) at $p=\overline{p}_k$ and $p=\underline{p}_{k+1}$ --- fixes the marginal of $Y(1)$ on $G_k$ to be

equation[equation omitted — 275 chars of source]

with the convention $\mathbb{P}_{obs}(\mathrm{d}y, W=1 \mid p(Z)=0) \equiv 0$ so that $k=0$ reduces to $\mu_{1,0}(\mathrm{d}y) = \mathbb{P}_{obs}(\mathrm{d}y, W=1 \mid p(Z)=\underline{p}_1)/\underline{p}_1$. On each non-degenerate LIV interval $L_k$, the constraint in ((ref)) holds for every $p$ in the continuum $[\underline{p}_k,\overline{p}_k]$; differentiating in $p$ pins $\pi_1(\mathrm{d}y\mid u)$ down pointwise via the LIV identity

equation[equation omitted — 186 chars of source]

whose absolute continuity in $p$ follows from the integral representation in ((ref)). The trivial tail $(\overline{p}_K,1]$ carries no constraint. Because cumulative sums of the gap marginals plus integration of the LIV densities recover the original constraint at every $p\in p(\mathcal{Z})$, this transformation is invertible and the two constraint systems are equivalent. Each rewritten constraint involves $\pi_1$ only on its own region and the objective is additively separable, so the global problem decomposes into three types of sub-problems, one for each region type.

enumerate• The OT Gaps $G_k = (\overline{p}_k, \underline{p}_{k+1})$ for $k=0,\dots,K-1$: the marginal of $Y(1)$ on $G_k$ is fixed to $\mu_{1,k}$ of ((ref)) and the $U$-marginal is $\text{Unif}(G_k)$, so the contribution of $G_k$ to the objective is a standard 1D optimal transport problem between $\mu_{1,k}$ and $\text{Unif}(G_k)$, which is solved by (ref). • The LIV Intervals $L_k = [\underline{p}_k, \overline{p}_k]$ for $k=1,\dots,K$: for non-degenerate $L_k$, the LIV identity in ((ref)) point-identifies the conditional mean \[\mathbb{E}[Y(1) \mid U=u] \;=\; \frac{{\rm d}}{{\rm d} u}\mathbb{E}_{obs}[Y W \mid p(Z)=u], \qquad u \in L_k,\] so the contribution of $L_k$ to the objective is uniquely determined and no optimization is required. Degenerate $L_k$ ($\underline{p}_k = \overline{p}_k$) contribute a Lebesgue-null integral and drop out. • The Trivial Tail $u \in (\overline{p}_K, 1]$: for individuals with extreme resistance to treatment, $Y(1)$ is never observed and the data imposes no constraint. Utilizing (ref), the minimization problem degenerates to assigning all mass to $y_{\min}$ where $\omega(u) \geqslant 0$ and to $y_{\max}$ where $\omega(u) < 0$.

Combining the solutions across these regions yields the closed-form sharp bounds for the general instrument setting, which are illustrated in (ref).

theorem[General Closed-Form Sharp Bounds] Suppose (ref), (ref), and (ref) hold, and $\mathcal{X} = \emptyset$. Let $p(\mathcal{Z}) = \bigcup_{k=1}^{K}[\underline{p}_k, \overline{p}_k]$ be the maximal-interval decomposition in ((ref)). For each $k = 0,\dots,K-1$, let $Q_{Y,k}(t)$ be the quantile function of $\mu_{1,k}$ defined in ((ref)), and $Q_{\omega,k}(t)$ the quantile function of $\omega(U)$ for $U \sim \text{Unif}(G_k)$. The sharp lower bound for the $W=1$ component of the target parameter is given by \begin{equation} \begin{split} \theta_{\omega,1}\ =\ & \underbrace{\sum_{k=0}^{K-1}(p_{k+1}-\overline{p}_k)\int_0^1 Q_{Y,k}(t)\,Q_{\omega,k}(1-t)\,\mathrm{d}t}_{OT Bound on the Gaps} + \underbrace{\sum_{k=1}^{K}\int_{p_k}^{\overline{p}_k}\left(\frac{{\rm d}}{{\rm d} u}\mathbb{E}_{obs}[YW \mid p(Z)=u]\right)\omega(u)\,\mathrm{d}u}_{Point-Identified LIV on the Intervals} \\ & \quad\quad + \underbrace{\int_{\overline{p}_K}^{1}\Big(y_{\min}\max\{0,\omega(u)\} + y_{\max}\min\{0,\omega(u)\}\Big)\mathrm{d}u}_{Trivial Lower Bound on the Tail}. \end{split} \end{equation} Similarly, the sharp upper bound is given by replaceing $Q_{\omega,k}(1-t)$ with $Q_{\omega,k}(t)$ in the OT term and swapping $y_{\min}$ and $y_{\max}$ in the trivial tail term.

The proof combines the subtraction-of-constraints decomposition from (ref) (applied region-by-region to identify $\mu_{1,k}$ on each gap) with the LIV-derivative argument on each non-degenerate interval $L_k$, followed by the 1D OT coupling of (ref). We defer the formal proof to (ref), which restates the result in the unified form in the covariate subsection below. (ref) subsumes both classical settings as limit cases. When every LIV interval is degenerate ($\underline{p}_k = \overline{p}_k$ for all $k$), the LIV sum vanishes identically, each $\mu_{1,k}$ coincides with the discrete complier distribution in ((ref)), and ((ref)) reduces exactly to (ref). Conversely, when $K=1$ with $\underline{p}_1 < \overline{p}_1$, the gap sum contains a single OT term on $(0,\underline{p}_1)$, the LIV sum contributes a single integral on $[\underline{p}_1,\overline{p}_1]$, and we recover the connected-continuous case; in particular, we match the identification results of heckman1999local when additionally $\underline{p}_1 = 0$ and $\overline{p}_1 = 1$, so the propensity score has full support with no gaps. The length of the bound depends on the gap widths, the identified gap distributions $\mu_{1,k}$, and the tail beyond $\overline{p}_K$:

align*[align* omitted — 280 chars of source]
figure[figure omitted — 5,772 chars of source]

Closed-Form Bounds with Covariates

Generalizing the previous results to accommodate covariates $X$ is mathematically straightforward. Because our structural assumption (ref) imposes conditional independence $Z \perp\!\!\!\perp (Y(w), U) \mid X$, the global CCOT problem naturally disintegrates into a family of conditional one-dimensional optimal transport problems, indexed by $x \in \mathcal{X}$. We can therefore solve the optimization conditional on $X=x$ and aggregate the resulting conditional bounds over the marginal distribution of $X$ via the law of iterated expectations.

Under (ref), for $\mathbb{P}_{obs}$-a.e.\ $x \in \mathcal{X}$ the conditional range of the propensity score admits a maximal-interval decomposition

equation[equation omitted — 282 chars of source]

with degenerate intervals $\underline{p}_k(x) = \overline{p}_k(x)$ corresponding to isolated conditional propensity values and boundary conventions $\overline{p}_0(x) := 0$ and $\underline{p}_{K_x+1}(x) := 1$. For each $x$, this partitions $[0,1]$ into the conditional LIV intervals $L_k(x) := [\underline{p}_k(x), \overline{p}_k(x)]$, the OT gaps $G_k(x) := (\overline{p}_k(x), \underline{p}_{k+1}(x))$, and the trivial tail $(\overline{p}_{K_x}(x), 1]$. The conditional gap-identified measure of $Y(1)$ on $G_k(x)$ is

equation[equation omitted — 289 chars of source]

for $k=0,\dots,K_x-1$, with the convention $\mathbb{P}_{obs}(\mathrm{d}y, W=1 \mid p(Z,X)=0, X=x) \equiv 0$.

theorem[Conditional General Bound] Suppose (ref), (ref), and (ref) hold. For each $k = 0,\dots,K_x-1$, let $Q_{Y, k \mid x}(t)$ be the quantile function of $\mu_{1,k\mid x}$ defined in ((ref)), and $Q_{\omega, k \mid x}(t)$ the quantile function of $\omega(x, U)$ for $U \sim \text{Unif}(G_k(x))$. The sharp lower bound for the $W=1$ component of the target parameter is \begin{equation} \begin{split} \theta_{\omega,1} = \mathbb{E}_{X} \Bigg[ & \sum_{k=0}^{K_X-1}(p_{k+1}(X)-\overline{p}_k(X))\int_0^1 Q_{Y,k\mid X}(t)\,Q_{\omega,k\mid X}(1-t)\,\mathrm{d}t \\ & + \sum_{k=1}^{K_X}\int_{p_k(X)}^{\overline{p}_k(X)} \left( \frac{{\rm d} }{{\rm d} u} \mathbb{E}_{obs}[Y W \mid p(Z,X)=u, X] \right) \omega(X, u) \, \mathrm{d}u \\ & + \int_{\overline{p}_{K_X}(X)}^1 \Big( y_{\min} \max\{0, \omega(X, u)\} + y_{\max} \min\{0, \omega(X, u)\} \Big) \mathrm{d}u \Bigg]. \end{split} \end{equation} Similarly, the sharp upper bound is given by replacing $Q_{\omega,k\mid X}(1-t)$ with $Q_{\omega,k\mid X}(t)$ in the OT term and swapping $y_{\min}$ and $y_{\max}$ in the trivial tail term.

The proof of (ref) is given in (ref). As in the no-covariate case, (ref) subsumes both the finite-discrete and the connected-continuous settings: when every conditional LIV interval is degenerate ($\underline{p}_k(X) = \overline{p}_k(X)$ a.s.), the LIV sum vanishes and the expression reduces to the discrete covariate bound; when $K_X = 1$ with $\underline{p}_1(X) < \overline{p}_1(X)$, the gap sum collapses to a single term on $(0, \underline{p}_1(X))$ and we recover the standard LIV covariate bound.



Extension to General Treatment

In this section, we generalize our partial identification framework to accommodate continuous treatments. We allow the treatment domain $\mathcal{W}$ to be a compact subset of $\mathbb{R}$. We impose the following generalized structural assumptions.

assumption[Generalized Structural Assumptions] The potential outcomes and the treatment selection mechanism satisfy the following conditions: \begin{enumerate} • Consistency: $Y = Y(W)$. • Conditional Instrumental Exogeneity: $Z \perp\!\!\!\perp (U, \{Y(w)\}_{w \in \mathcal{W}}) \mid X$. • Selection Mechanism: The treatment $W$ is selected via the structural equation: \begin{equation} W = F_{W \mid Z,X}^{-1}(U \mid Z,X), \end{equation} where $U \mid X \sim \text{Unif}(0,1)$, and $F_{W \mid Z,X}^{-1}(\cdot \mid z,x)$ is the conditional quantile function of $W$ given $Z=z$ and $X=x$. • Common Instrument Support: The conditional support of $Z$ given $X$ does not depend on $X$: $\mathrm{Supp}(Z \mid X=x) = \mathcal{Z}$ for all $x \in \mathcal{X}$. \end{enumerate}

As in the binary case, the common instrument support condition ensures that the conditional distribution $F_{W \mid Z=z, X=x}$ is well-defined for every $z \in \mathcal{Z}$ and every $x \in \mathcal{X}$, which is necessary for the structural equation ((ref)) to be identified uniformly across covariates. If $W$ is a binary treatment, (ref) reduces to the threshold-crossing model utilized in our previous assumptions after the harmless reparameterization $U' = 1-U$.\footnote{Under the generalized ordered-choice convention, $W = F^{-1}_{W\mid Z,X}(U\mid Z,X)$ is increasing in $U$, whereas our earlier binary threshold-crossing convention writes treatment as $W=\indicator(U \le p(Z,X))$. Setting $U' = 1-U$ aligns the two models.} If $W$ is multi-valued, this model is equivalent to the discrete ordered choice model heckman2007econometric. In the continuous treatment regime, this assumption provides the structural foundation for the control function approach in nonseparable models imbens2009identification, blundell2003endogeneity, chesher2003identification. We discuss this connection in detail in (ref).

In this model, $U$ can be interpreted as a latent rank that orders individuals' treatment intensity within each $(Z,X)$ cell. Under the ordered-choice convention in ((ref)), larger values of $U$ correspond to larger realizations of $W$.

Following the binary treatment case, we consider the following general target parameter. Let $\pi_{w,x}$ be the structural joint distribution of $(U, Y(w)) \mid X=x$, which can be seen as a probability kernel. For any identifiable weight function $\omega(w,x,u)$, we define:

equation[equation omitted — 210 chars of source]

where $\lambda(\cdot)$ denotes the Lebesgue measure for continuous treatments and the counting measure for discrete treatments. This directly generalizes $\theta_\omega$ from ((ref)). To see this, consider the binary treatment case $\mathcal{W} = \{0,1\}$, where $\int_{\mathcal{W}} \cdot \,\mathrm{d}w$ reduces to the sum over $\{0,1\}$. Setting $\omega(1,x,u) = \omega(x,u)$ and $\omega(0,x,u) = -\omega(x,u)$, the expression ((ref)) becomes:

equation*[equation* omitted — 239 chars of source]

recovering ((ref)).

As an example, consider a counterfactual policy that alters two components while holding the covariate distribution fixed: (i) the treatment assignment mechanism changes from $F_{W \mid Z,X}$ to $\tilde{F}_{W \mid Z,X}$, and (ii) the instrument distribution shifts from $F_{Z \mid X}$ to $\tilde{F}_{\tilde{Z} \mid X}$. Under this new policy, the counterfactual treatment is $\tilde{W} = \tilde{F}_{W \mid Z,X}^{-1}(U \mid \tilde{Z},X)$, and the realized outcome is $\tilde{Y} = Y(\tilde{W})$. The policy effect $\theta = \mathbb{E}[\tilde{Y} - Y]$ can be expressed as $\theta_\omega$ with the specific weight:

equation[equation omitted — 165 chars of source]

where the right-hand side denotes the $w$-section of the signed measure on $\mathcal{W} \times [0,1]$. The policy weight ((ref)) reduces to ((ref)) in the binary case.

The observed data constrains the structural kernel $\pi_{w,x}$. Note that the conditional distribution $\mathbb{P}_{\text{obs}}(\mathrm{d}y \mid W=w, Z=z, X=x)$ is only well-defined when $w$ belongs to the conditional support of $W$ given $(Z,X)$. We therefore define the effective instrument set:

equation*[equation* omitted — 111 chars of source]

and impose the observational constraints only for $z \in \mathcal{Z}(x,w)$. Specifically, for any $w \in \mathcal{W}$ and $x \in \mathcal{X}$, integrating the conditional distribution of the potential outcome over the latent variable $U$ must perfectly reconstruct the observed conditional outcome measure:

equation*[equation* omitted — 203 chars of source]

Crucially, because both the objective functional and the observational constraints are additively separable across the support of $X$ and $W$, the global optimization over all joint distributions decomposes into a family of independent CCOT sub-problems. Formally, this is equivalent to solving the following CCOT problem for $\lambda$-a.e.\ $w \in \mathcal{W}$ and $\mathbb{P}_{\mathrm{obs}}$-a.e.\ $x \in \mathcal{X}$:

equation[equation omitted — 558 chars of source]

Let $\overline{\theta}(x,w)$ and $\underline{\theta}(x,w)$ denote the optimal values of the maximization and minimization in ((ref)), respectively. The tight bounds for the global parameter $\theta_\omega$ are then obtained by aggregating over all $(x,w)$ strata:

equation[equation omitted — 300 chars of source]

To formally guarantee that ((ref)) yields sharp bounds for the continuous treatment setting, we must establish that any family of measures satisfying the generalized observational and marginal constraints corresponds to a valid data-generating process. Let $\Gamma_{\text{gen}}(\mathbb{P}_{\text{obs}})$ denote the set of all measure families $\{\pi_{w,x}\}_{w \in \mathcal{W}, x \in \mathcal{X}}$ that satisfy the constraints in the CCOT problem above. The following proposition verifies this observational equivalence, extending the sharpness guarantee from the binary framework, (ref).

proposition[Generalized Observational Equivalence and Sharpness] Suppose the observed distribution $\mathbb{P}_{\text{obs}}$ is generated by a true structural model satisfying (ref). Then, for any candidate family of probability kernels $\{\tilde{\pi}_{w,x}\}_{w \in \mathcal{W}, x \in \mathcal{X}} \in \Gamma_{\text{gen}}(\mathbb{P}_{\text{obs}})$, there exists a probability space supporting a jointly measurable process $\{Y(w)\}_{w \in \mathcal{W}}$ together with random variables $(U, X, Z, W)$ such that: \begin{enumerate} • For $\lambda$-a.e.\ $w \in \mathcal{W}$ and $\mathbb{P}_{\text{obs}}$-a.e.\ $x \in \mathcal{X}$, the conditional distribution of $(Y(w), U) \mid X=x$ is exactly $\tilde{\pi}_{w,x}$. Moreover, conditional on $(U, X)$, the process $\{Y(w)\}_{w \in \mathcal{W}}$ is independent of $(Z, W)$ and has mutually independent coordinates. • (ref) holds for $(\{Y(w)\}_{w \in \mathcal{W}}, U, X, Z, W)$. • The induced distribution of the observable variables $(Z,X,W,Y)$ exactly matches the observed data distribution $\mathbb{P}_{\text{obs}}$. \end{enumerate}

This formal CCOT formulation highlights that variation in $F_{U \mid Z=z, X=x, W=w}$ across different instruments $Z$ imposes multiple marginal constraints on the structural distribution $\pi_{w,x}$. In general, the solution depends on the union of these supports, $$\mathcal{U}_{\text{id}}(x,w) \coloneqq \bigcup_{z \in \mathcal{Z}(x,w)} \text{supp}(F_{U \mid Z=z, X=x, W=w}).$$ Note that each $\text{supp}(F_{U \mid Z=z, X=x, W=w})$ is either a single point or a closed interval in $[0,1]$. Throughout this section, we assume that $\mathcal{U}_{\text{id}}(x,w)$ is measurable for each $(x,w)$. This measurability condition holds, for instance, if $\mathcal{Z}(x,w)$ is compact and the map $z \mapsto F_{W \mid Z,X}(w \mid z,x)$ is continuous, in which case $\mathcal{U}_{\text{id}}(x,w)$ is a compact subset of $[0,1]$.

In the following subsections, we demonstrate how to obtain closed-form solutions under specific structural settings. The general case is much more complicated and can be seen as a mix of these settings, which we will leave for future work.

Strictly Monotonic Treatment Selection

imbens2009identification considers the following assumption for the IV model.

assumption[Strict Monotonicity] The structural treatment function $W = h(Z,X,U)$ is strictly increasing in the latent variable $U$ almost surely.

Under (ref), it is observationally equivalent to the conditional quantile representation in ((ref)). Crucially, the unobserved confounder $U$ can be perfectly inverted and recovered from the observables as the conditional rank of the treatment, $U = F_{W \mid Z,X}(W \mid Z,X)$. This recovered latent rank serves directly as a control variable: conditioning on $U$ alongside the covariates $X$ effectively absorbs the endogeneity, allowing us to isolate the exogenous variation in $W$. This formal equivalence connects our generalized partial identification framework to classical control function approach blundell2003endogeneity, chesher2003identification.

Because $U$ is perfectly recoverable, the conditional distribution $F_{U \mid Z=z, X=x, W=w}$ degenerates to a point mass at $u = F_{W \mid Z,X}(w \mid z, x)$. Consequently, the integral constraint in ((ref)) simplifies to:

equation*[equation* omitted — 169 chars of source]

where $z_u \in \mathcal{Z}(x,w)$ satisfies $F_{W \mid Z,X}(w \mid z_u, x) = u$. When multiple instrument values map to the same $u$, the observational constraints ensure that $\mathbb{P}_{\text{obs}}(\mathrm{d}y \mid W=w, Z=z_u, X=x)$ is identical for all such $z_u$, so the choice is immaterial. That is, the structural kernel $\pi_{w,x}$ is point-identified on the identifiable region $\mathcal{U}_{\text{id}}(x,w) = \bigcup_{z \in \mathcal{Z}(x,w)} \{F_{W\mid Z,X}(w \mid z,x)\}$. Therefore, if $u \in \mathcal{U}_{\text{id}}(x,w)$, we have:

equation[equation omitted — 132 chars of source]

where $z_u \in \mathcal{Z}$ is the specific baseline instrument realization that satisfies $w = F_{W \mid Z,X}^{-1}(u \mid z_u, x)$.

Outside this identifiable support, there is no constraint, and the optimal solution trivially assigns probability mass to the global outcome bounds to minimize or maximize the objective. Substituting this into our separated CCOT formulation yields the closed-form bounds.

theorem[Closed-Form Bounds under Strictly Monotonic Treatment Selection] Suppose the generalized structural conditions in (ref), the strict monotonicity in (ref), and the outcome boundedness in (ref) hold. For any identifiable weight $\omega(w,x,u)$, the sharp lower bound for $\theta_\omega$ is: \begin{align*} \theta &= \int_{\mathcal{W}} \mathbb{E}_{X} \Bigg[ \int_0^1 \indicator\Big(u \in \mathcal{U}_{id}(x,w)\Big) \mathbb{E}_{obs}[Y \mid w, Z=z_u, X] \,\omega(w,X,u)\,\mathrm{d}u \Bigg] \mathrm{d}w \\ &\quad + \int_{\mathcal{W}} \mathbb{E}_{X} \Bigg[ \int_0^1 \indicator\Big(u \notin \mathcal{U}_{id}(x,w)\Big) \Big( y_{\min} \max\{0, \omega(w,X,u)\} + y_{\max} \min\{0, \omega(w,X,u)\} \Big) \mathrm{d}u \Bigg] \mathrm{d}w, \end{align*} where $z_u$ satisfies $F_{W \mid Z,X}(w \mid z_u, x) = u$. The sharp upper bound $\overline{\theta}$ is obtained symmetrically by swapping $y_{\min}$ and $y_{\max}$ in the trivial bound term.

Ordered Choice Model

Next, let us consider the multi-valued treatment setting, which is known as the ordered choice model heckman2007econometric. In this case, $\mathcal{W} $ is a finite set and $\text{supp}(F_{U \mid Z=z, X=x, W=w})$ is an interval for each $z$. Let us denote this interval as $I_{x,w}(z)$. Then, the first constraint of ((ref)) simplifies to:

equation[equation omitted — 237 chars of source]

In particular, if $|I_{x,w}(z)|=0$, the interval $I_{x,w}(z)$ degenerates to a single point $\{u_0\}$ and the constraint reduces to $\pi_{w,x}(\mathrm{d}y \mid U=u_0) = \mathbb{P}_{\text{obs}}(\mathrm{d}y \mid W=w, Z=z, X=x)$, i.e., the conditional distribution is point-identified at $u_0$. By convention, the left-hand side of ((ref)) is interpreted as $\pi_{w,x}(\mathrm{d}y \mid U=u_0)$ in this degenerate case.

The main difficulty of ((ref)) is that the intervals in $\{I_{x,w}(z)\}_{z \in \mathcal{Z}(x,w)}$ may overlap with each other, making it hard to disentangle the constraints like in the binary treatment setting. Surprisingly, we find that if $\{I_{x,w}(z)\}_{z \in \mathcal{Z}(x,w)}$ forms a $\pi$-system, closed-form solutions still exist.

assumption[$\pi$-System] Assume that the collection of intervals $\{I_{x,w}(z)\}_{z \in \mathcal{Z}(x,w)}$ forms a $\pi$-system for all $x \in \mathcal{X}$ and $w \in \mathcal{W}$. That is, for any $z, z' \in \mathcal{Z}(x,w)$, the intersection is also in the family: $I_{x,w}(z) \cap I_{x,w}(z') \in \{I_{x,w}(z'')\}_{z'' \in \mathcal{Z}(x,w)} \cup \{\emptyset\}$.

In particular, if the $\{I_{x,w}(z)\}_{z \in \mathcal{Z}(x,w)}$ share the same start point or end point, they naturally satisfy this algebraic property, which is exactly the case in the binary treatment setting. More generally, (ref) endows $\{I_{x,w}(z)\}$ with a hierarchical Directed Acyclic Graph (DAG) ordering under strict inclusion---$z$ is an ancestor of $z'$ whenever $I_{x,w}(z') \subsetneq I_{x,w}(z)$---which, as we show below, is exactly the structure needed to disentangle the overlapping constraints.

For each $z \in \mathcal{Z}(x,w)$, let $\text{children}(z)$ and $\text{Dec}(z)$ denote its direct children and all descendants in this DAG. Define the disjoint isolated sub-region

equation*[equation* omitted — 103 chars of source]

and, recursively from the leaves upward, the isolated measure

equation[equation omitted — 245 chars of source]

The measure $\mu_{w,z,x}$ is the general analogue of $\mu_{1,i}$ from the binary case: it isolates the conditional distribution of $Y(w)$ on the disjoint slice $J_{x,w}(z)$. As in the binary setting, $\mu_{w,z,x}$ is a valid (nonnegative) probability measure under correct model specification, and its nonnegativity serves as a testable implication of the structural model. We also let $J_{x,w}(\emptyset) = [0,1] \setminus \bigcup_{z \in \mathcal{Z}(x,w)} I_{x,w}(z)$ denote the unconstrained domain.

In (ref), we show that under (ref) the original constraint system ((ref)) is equivalent to the {disentangled} marginal constraints

equation[equation omitted — 211 chars of source]

each acting on a single disjoint region $J_{x,w}(z)$. Because the $\{J_{x,w}(z)\}_{z \in \mathcal{Z}(x,w)}$ partition the identified support $\mathcal{U}_{\text{id}}(x,w)$ and the objective in ((ref)) is additively separable across these regions, the global problem decomposes into independent 1D optimal transport sub-problems with product cost---one per $J_{x,w}(z)$, pairing $\mu_{w,z,x}$ with $\mathrm{Unif}(J_{x,w}(z))$---plus a pointwise trivial bound on the unconstrained domain $J_{x,w}(\emptyset)$. Applying (ref) to each sub-problem yields the following closed-form result.

theorem[Closed-Form Bounds for Multi-Valued Treatment] Suppose (ref), (ref), and (ref) hold, $\mathcal{Z}$ is a discrete set, $\mathcal{W}$ is finite, and for every $(x,w)$ and $z \in \mathcal{Z}(x,w)$ the support $\mathrm{supp}(F_{U \mid Z=z, X=x, W=w})$ is an interval $I_{x,w}(z) \subseteq [0,1]$. Let $Q_{Y, J_{x,w}(z)}(t)$ be the quantile function of the isolated measure $\mu_{w,z,x}$, and let $Q_{\omega, J_{x,w}(z)}(t)$ be the quantile function of $\omega(w,x,U)$ for $U \sim \text{Unif}(J_{x,w}(z))$. For a given $x \in \mathcal{X}$ and $w \in \mathcal{W}$, let $J_{x,w}(\emptyset) = [0,1] \setminus \bigcup_{z \in \mathcal{Z}(x,w)} I_{x,w}(z)$ be the unconstrained domain. The sharp lower bound $\underline{\theta}(x,w)$ is given by: \begin{align*} \theta(x,w) &= \sum_{z \in \mathcal{Z}(x,w)} |J_{x,w}(z)| \int_0^1 Q_{Y, J_{x,w}(z)}(t) Q_{\omega, J_{x,w}(z)}(1-t) \, \mathrm{d}t \\ &\quad + \int_{J_{x,w}(\emptyset)} \Big( y_{\min} \max\{0, \omega(w,x,u)\} + y_{\max} \min\{0, \omega(w,x,u)\} \Big) \mathrm{d}u. \end{align*} The global sharp lower bound is $\underline{\theta}_\omega = \sum_{w \in \mathcal{W}} \mathbb{E}_{X}[\underline{\theta}(X,w)]$. The sharp upper bound $\overline{\theta}$ is obtained symmetrically by taking the comonotonic coupling integral $\int_0^1 Q_{Y, J_{x,w}(z)}(t) Q_{\omega, J_{x,w}(z)}(t) \mathrm{d}t$ and swapping $y_{\min}$ and $y_{\max}$ in the trivial bound term.
remark[Mixed Continuous-Discrete Treatments] In general, the solution to the global CCOT problem in ((ref)) is a hybrid of the results in (ref) and (ref). The exact nature of the subproblem for a given stratum $(x,w)$ depends directly on the localized behavior of the treatment distribution. If the CDF $F_{W\mid Z,X}(w)$ is continuous at the evaluation point $w$, the structural quantile function is strictly increasing, collapsing the conditional distribution $F_{U \mid Z=z, X=x, W=w}$ to a point mass. In these regions, the subproblem is perfectly constrained, and the potential outcome expectation is point-identified as in ((ref)). Conversely, if the data exhibit a discrete probability mass at $w$, the latent variable $U$ maps to an interval. Over these discrete mass points, the subproblem natively transitions into the localized 1D optimal transport problem described in (ref). Thus, our unified CCOT framework naturally accommodates complex empirical settings with mixed continuous-discrete treatments.

Estimation and Inference

In this section, we develop estimation and inference results for the closed-form bounds derived in (ref). For the discrete instrument setting, we leverage DML chernozhukov2018double to accommodate high-dimensional covariates, constructing Neyman-orthogonal scores that yield $\sqrt{n}$-consistent and asymptotically normal estimators. In the continuous instrument setting, we characterize the corresponding nonparametric convergence rates.

We focus our exposition on the bounds for the PRTE, as the derivation for other causal quantities follows similarly. For simplicity, we first develop theory for the unscaled numerator $\mathbb{E}[Y^{q}-Y]$. In the empirical designs considered below, the PRTE denominator is the known policy-shift size $\alpha > 0$, so the reported bounds and confidence intervals for $\text{PRTE}_\alpha$ are obtained by dividing the corresponding numerator bounds by $\alpha$. More generally, when $\mathbb{E}[q(Z,X)-p(Z,X)]$ is not fixed by design, it is point-identified and can be estimated separately. Thus, our target weight function simplifies to:

align*[align* omitted — 138 chars of source]

We model the alternative policy as $q(z,x) = \phi(z,x,p(z,x))$ for a known function $\phi$. Common examples for $\phi$ include uniform propensity shifts $\phi(z,x,p) = p + \alpha$ or $p(1+\alpha)$ for a constant $\alpha$, or setting $\phi(z,x,p) = r(z,x)$ to evaluate a specific alternative targeting policy $r$.

Discrete Instrument Setting

For a fixed $x$, the weight function $\omega(x,u)$ is a step function that is constant and monotonic within each latent interval $u \in (p_i(x), p_{i+1}(x)]$. This piecewise structure allows us to decompose the closed-form lower bound integral into a finite weighted sum over the instrument level sets.

Define $q_{j,i}(x)$ as the $j$-th smallest value of the alternative policy propensity $q(z,x)$ that falls strictly inside the baseline interval $[p_i(x), p_{i+1}(x))$, such that:

equation*[equation* omitted — 87 chars of source]

with the convention $ q_{0,i}(x) = p_i(x)$. Let the baseline and alternative instrument level sets be:

align*[align* omitted — 125 chars of source]

Before presenting the rewriting of the target parameter, we impose two regularity conditions on the alternative policy and the propensity scores.

assumption[Smoothness of $\phi$ and Propensity Score] (i) $\phi(z,x,p)$ is differentiable with respect to $p$, and $\frac{\partial \phi}{\partial p}$ is bounded almost surely for all $z \in \mathcal{Z}$ and $x \in \mathcal{X}$. (ii) For each $z \in \mathcal{Z}$, the conditional propensity score $p(z,\cdot)$ is continuous on $\mathcal{X}$.

Part (i) will be used in the pathwise derivation of the orthogonal score and in the asymptotic analysis. Because $q(z,x) = \phi(z,x,p(z,x))$, part (ii) combined with (i) also implies that $q(z,\cdot)$ is continuous on $\mathcal{X}$ for each $z \in \mathcal{Z}$.

To ensure the level sets can be consistently estimated from the data without asymptotic ambiguity, we further impose the following gap assumption on the propensity scores.

assumption[Propensity Score Gap] There exists a constant $c_{\text{gap}} > 0$ such that for all $x \in \mathcal{X}$ and for all $v, v' \in R_x = \{p(z,x), q(z,x)\}_{z \in \mathcal{Z}}$, either $v' = v$ or $|v - v'| > c_{\text{gap}}$.

Fix any two pairs $(z,\tau)$ and $(z',\tau')$ with $\tau, \tau' \in \{p,q\}$. The map $x \mapsto \tau(z,x) - \tau'(z',x)$ is continuous on $\mathcal{X}$ by (ref), while (ref) requires its value at every $x$ to be either exactly zero or of absolute value strictly greater than $c_{\text{gap}}$. A continuous function whose range avoids the punctured neighborhood $(-c_{\text{gap}}, c_{\text{gap}}) \setminus \{0\}$ cannot change sign, so the relative ordering of the elements of $R_x$ is invariant across $\mathcal{X}$. We therefore suppress the dependence on $x$ in $S_k$, $T_{j,k}$, $K$, and $l_k$ from now on.

The bound in (ref) admits a representation as an expectation of a finite weighted sum over instrument level sets and their sub-intervals. Before stating this representation, we introduce the nuisance components that carry the observable content of the bound.

\paragraph{Nuisance components.} For each interval index $k \in \{0,\dots,K\}$ and sub-interval index $j \in \{0,\dots,l_k\}$, define:

itemize• Instrument-level-set weights. \begin{align*} \gamma_{full,k}(x) &\coloneqq \mathbb{P}_{obs}\!\left(Z \in \bigcup_{i=k+1}^K \bigcup_{l=0}^{l_i} T_{l,i} \mathrel{\Big|} X=x\right) - \mathbb{P}_{obs}\!\left(Z \in \bigcup_{i=k+1}^K S_i \mathrel{\Big|} X=x\right), \\ \gamma_{j,k}(x) &\coloneqq \mathbb{P}_{obs}\!\left(Z \in \bigcup_{l=j}^{l_k} T_{l,k} \mathrel{\Big|} X=x\right), \qquad \gamma_K(x) \coloneqq \mathbb{P}_{obs}\!\left(Z \in \bigcup_{j=1}^{l_K} T_{j,K} \mathrel{\Big|} X=x\right). \end{align*} These are the aggregate masses that the step weight $\omega$ assigns to instruments; each is estimable by a standard regression of an instrument-set indicator on $X$. • \emph{Scaled complier mean and sub-interval quantile integral.} With the relative threshold $\kappa_{j,k}(x) \coloneqq (q_{j,k}(x) - p_k(x))/(p_{k+1}(x) - p_k(x)) \in [0,1]$, \begin{equation*} J_{\text{full},k}(x) \coloneqq (p_{k+1}(x)-p_k(x))\,\mathbb{E}_{\mu_{1,k\mid x}}[Y], \qquad J_{j,k}(x) \coloneqq (p_{k+1}(x)-p_k(x))\int_{\kappa_{j,k}(x)}^{\kappa_{j+1,k}(x)} Q_{Y,k\mid x}(u)\,\mathrm{d}u, \end{equation*} where $\mu_{1,k\mid x}$ is the identified complier distribution of (ref) and $Q_{Y,k\mid x}$ is its quantile function. • \emph{Trivial-bound boundary term.} \begin{equation*} \Delta_K(x) \coloneqq y_{\min}\, \mathbb{E}_{\text{obs}}\!\left[ \indicator\!\Big(Z \in \bigcup_{j=1}^{l_K} T_{j,K}\Big) (q(Z,X) - p_K(X)) \mathrel{\Big|} X=x \right]. \end{equation*}
proposition[Discrete-Instrument Reduction to Nuisance Functional] Suppose (ref), (ref), (ref), and (ref) hold and $\mathcal{Z}$ is finite. Then the sharp lower bound in (ref) admits the representation \begin{align} \theta_{\omega,1} = \mathbb{E}\Bigg[\sum_{k=0}^{K-1}\bigg(\gamma_{full,k}(X)\, J_{full,k}(X) + \sum_{j=1}^{l_k} \gamma_{j,k}(X)\, J_{j-1,k}(X)\bigg) + \Delta_K(X) \Bigg]. \end{align} The sharp upper bound admits an analogous representation, obtained by replacing $Q_{Y,k\mid x}(u)$ by $Q_{Y,k\mid x}(1-u)$ inside $J_{j,k}$ and swapping $y_{\min}$ with $y_{\max}$ in $\Delta_K$.

(ref) is the key analytical step that converts the optimal-transport integral in (ref) into a Neyman-orthogonalizable functional of standard nuisance quantities---conditional probabilities, conditional means, and conditional quantile integrals. The proof, given in (ref), exploits the piecewise-constant, monotone structure of $\omega(X,\cdot)$ on each baseline interval and identifies the step-function value attached to each sub-interval $(q_{j,k}(X), q_{j+1,k}(X))$.

The sub-interval term $J_{j,k}$ still involves the conditional quantile function $Q_{Y,k\mid X}$, which is delicate to estimate directly. The next proposition replaces it with conditional expectations in the two outcome regimes of interest.

proposition[Specialization of the Sub-Interval Contribution] Fix $k \in \{0,\dots,K-1\}$ and let $\nu_{j,k}(x) \coloneqq Q_{Y,k\mid x}(\kappa_{j,k}(x))$ for $j \in \{0,\dots,l_k\}$. \begin{enumerate}[label=(\roman*)] • Continuous outcome. If $\mu_{1,k\mid X}$ is continuous at $\nu_{j,k}(X)$ and $\nu_{j+1,k}(X)$ almost surely, then, for $j \in \{0,\dots,l_k-1\}$, \begin{equation} J_{j,k}(X) = (p_{k+1}(X) - p_k(X))\,\mathbb{E}_{\mu_{1,k\mid X}}\!\big[ Y\,\indicator\!\big(\nu_{j,k}(X) < Y \leqslant \nu_{j+1,k}(X)\big) \big]. \end{equation} • Binary outcome. If $Y \in \{0,1\}$, then $J_{\text{full},k}(X) = P_{1,k+1}(X) - P_{1,k}(X)$ with $P_{1,k}(x) \coloneqq \mathbb{E}_{\text{obs}}[YW \mid Z \in S_k, X=x]$, and for $j \in \{1,\dots,l_k\}$, \begin{equation} J_{j-1,k}(X) = h_{j,k}^{-}(X) \coloneqq \max\!\Big(0,\; q_{j,k}(X) - \max\!\big(q_{j-1,k}(X),\; p_{k+1}(X) - J_{full,k}(X)\big)\Big). \end{equation} Consequently, (ref) collapses to \begin{align} \theta_{\omega,1} = \mathbb{E}\Bigg[\sum_{k=0}^{K-1}\bigg(\gamma_{full,k}(X)\, J_{full,k}(X) + \sum_{j=1}^{l_k} \gamma_{j,k}(X)\, h_{j,k}^{-}(X)\bigg) + \Delta_K(X)\Bigg]. \end{align} \end{enumerate}

Part (i) converts the sub-interval quantile integral into a plain conditional expectation of $Y$ on the quantile-defined band $(\nu_{j,k}, \nu_{j+1,k}]$, so only the $l_k+1$ boundary quantiles $\nu_{j,k}$---rather than the full quantile process---enter the estimation. Part (ii) leverages the step-function form of $Q_{Y,k\mid X}$ when $Y$ is binary: the complier mean $\theta_k(X) = \mathbb{E}_{\mu_{1,k\mid X}}[Y]$ reduces to a difference of observed joint probabilities, and the integral of the step function over $[\kappa_{j-1,k}, \kappa_{j,k}]$ evaluates to the closed-form $h_{j,k}^{-}$. Both parts are proved in (ref).

Together, (ref) expose (ref) as a functional of conditional expectations, boundary quantiles, and instrument-level-set probabilities---the form required for DML-style $\sqrt{n}$-inference. We construct the corresponding estimators and establish their asymptotic properties next, treating the two outcome regimes in turn.

Continuous Outcome

We first treat the case where the conditional complier distribution $\mu_{1,k\mid X}$ is continuous at the boundary quantiles $\nu_{j,k}(X)$, so (ref)(i) applies. The representation (ref) converts each sub-interval term $J_{j,k}(X)$ into a conditional expectation over the band $(\nu_{j,k}(X), \nu_{j+1,k}(X)]$, thereby reducing the nuisance list from the full quantile process to the boundary quantiles $\nu_{j,k}$ alone. The Neyman-orthogonal score construction additionally requires the conditional indicator expectations

align*[align* omitted — 229 chars of source]

which serve as Riesz representers for the pathwise derivatives with respect to $\nu_{j,k}$.

Given the functional form in ((ref)) and ((ref)), our DML estimator takes the aggregated Neyman-orthogonal form

equation[equation omitted — 339 chars of source]

where each $\hat{\psi}$ is a Neyman-orthogonal score that corrects for the estimation error of its associated nuisance (explicit formulas are given in (ref)). The rest of this subsection develops the nuisance estimators $\hat{\eta}$ trained on an auxiliary sample $I_1$; (ref) below then establishes $\sqrt{n}$-consistency and asymptotic normality of ((ref)).

\paragraph{Estimation Procedure.} Following the standard DML template, we randomly partition the sample into an auxiliary set $I_1$ and a main estimation set $I_2$, train all nuisance estimators on $I_1$, and average plug-in orthogonal scores on $I_2$. (ref) summarizes the full procedure; the explicit moment equation for the conditional quantile $\nu_{j,k}$, the conditional-expectation components $M_{j,k}^{\pm}$, $J_{j,k}^{\pm}$, $J_{\mathrm{full},k}^{\pm}$, and the Neyman-orthogonal scores $\psi^{\mathrm{prod}}$ are deferred to (ref) and (ref).

algorithm[algorithm omitted — 1,461 chars of source]

To ensure the debiased terms remain well-behaved and that inverse probability weights do not explode, we require the following strict overlap assumption.

assumption[Overlap] There exists a constant $c_\pi > 0$ such that $\pi_k(x) \geqslant c_\pi$ for all $x \in \mathcal{X}$ and for all interval indices $k \in \{0, \dots, K\}$.

The following theorem formally establishes the asymptotic normality of our DML estimator.

theorem[Asymptotic Normality] Suppose (ref), (ref), (ref), (ref) and (ref) hold, $\mathcal{Z}$ is finite and the estimators for the nuisance parameters $\eta$ trained on $I_1$ converge at a rate of at least $o_P(n^{-1/4})$ in the $L_2$ norm. Furthermore, assume that for each $j = 1, \ldots, l_k$, $k = 0, \ldots, K-1$ and $k' \in \{k, k+1\}$, the conditional density $f_{Y \mid X, Z \in S_{k'}, W=1}(y \mid X)$ exists and is continuous at $y = \nu_{j,k}(X)$ almost surely. Then, the sample-split estimator is $\sqrt{n}$-consistent and asymptotically normal: \begin{equation*} \sqrt{n/2} \big( \hat{\theta}_{\omega, 1} - \theta_{\omega, 1} \big) \xrightarrow{d} \mathcal{N}(0, \sigma^2), \end{equation*} where the asymptotic variance is the variance of the true orthogonal score, $\sigma^2 = \mathbb{E}\big[ ({\psi}(O; \eta) - \underline{\theta}_{\omega, 1})^2 \big]$. It can be consistently estimated by its sample analog $\hat{\sigma}^2 = \frac{1}{|I_2|} \sum_{i \in I_2} \big(\hat{\psi}(O_i; \hat{\eta}) - \hat{\underline{\theta}}_{\omega, 1}\big)^2$.

Discrete Outcome

When the outcome $Y$ takes finitely many values---in particular, when $Y \in \{0,1\}$ is binary---the conditional quantile function $Q_{Y,k\mid X}$ becomes a step function with atoms, and the continuity hypothesis of (ref)(i) fails. Instead, (ref)(ii) applies: the representation (ref) expresses $J_{\text{full},k}$ as a difference of observed joint probabilities $P_{1,k+1} - P_{1,k}$ and evaluates the sub-interval contribution in closed form through $h_{j,k}^{-}$, so the nuisance list collapses to conditional expectations alone and quantile estimation is avoided entirely. We present the binary case; the extension to general discrete outcomes is analogous.

Given the functional form in (ref), our DML estimator takes an aggregated Neyman-orthogonal form analogous to the continuous-outcome case:

equation[equation omitted — 326 chars of source]

where each $\hat{\psi}$ is a Neyman-orthogonal score (explicit formulas in (ref)). (ref) below establishes $\sqrt n$-consistency and asymptotic normality.

\paragraph{Estimation Procedure.} The procedure mirrors (ref), with the conditional-quantile and indicator-expectation steps removed. (ref) summarizes; full details and the construction of the orthogonal scores are in (ref) and (ref).

algorithm[algorithm omitted — 835 chars of source]

To ensure differentiability of the target functional with respect to the nuisance parameters, we impose the following gap condition, which replaces the conditional-density smoothness assumption of (ref).

assumption[Discrete Outcome Gap] For all $k = 0, \ldots, K-1$ and $x \in \mathcal{X}$, the integration boundaries $p_{k+1}(x) - J_{\text{full},k}(x)$ and $p_k(x) + J_{\text{full},k}(x)$ do not coincide with any alternative policy level $q_{j,k}(x)$ for $j = 1, \ldots, l_k$.

This condition guarantees that the target is pathwise-differentiable in the nuisances despite the $\max$ operators appearing in $h_{j,k}^{-}$. The following theorem establishes the asymptotic normality of our DML estimator in the discrete-outcome setting.

theorem[Asymptotic Normality, Discrete Outcome] Suppose (ref), (ref), (ref), (ref), (ref) and (ref) hold, $\mathcal{Z}$ is finite, outcome is binary and the nuisance estimators trained on $I_1$ converge at a rate of at least $o_P(n^{-1/4})$ in the $L_2$ norm. Then, the sample-split estimator $\hat{\underline{\theta}}_{\omega,1}$ defined in (ref) is $\sqrt{n}$-consistent and asymptotically normal: \begin{equation*} \sqrt{n/2}\big(\hat{\theta}_{\omega,1} - \theta_{\omega,1}\big) \xrightarrow{d} \mathcal{N}(0, \sigma^2), \end{equation*} where $\sigma^2 = \mathbb{E}\big[(\psi(O; \eta) - \underline{\theta}_{\omega,1})^2\big]$. It can be consistently estimated by $\hat{\sigma}^2 = \frac{1}{|I_2|}\sum_{i \in I_2}\big(\hat{\psi}(O_i; \hat{\eta}) - \hat{\underline{\theta}}_{\omega,1}\big)^2$.

Continuous Instrument Setting

We now proceed to estimation in the continuous instrument setting. For simplicity, we will assume that $\mathcal{Z}$ is connected throughout this subsection; the multiply connected component case can be derived similarly. When the instrumental variable $Z$ is continuous, the propensity score $p(Z,X)$ takes a continuous support, and the piecewise-constant structure of the weight function no longer holds. The general bound ((ref)) contains a derivative $\frac{\partial}{\partial u}\mathbb{E}_{\text{obs}}[YW \mid p(Z,X)=u, X]$ that is non-trivial to estimate directly. The following proposition leverages Fubini's theorem to collapse the derivative into a finite difference of conditional-expectation evaluations, yielding a representation amenable to plug-in ML estimation.

proposition[Continuous-Instrument Reduction via Fubini] Suppose (ref), (ref), (ref), and (ref) hold, and $\mathcal{Z}$ is connected. Let $g_1(u,X) \coloneqq \mathbb{E}_{\text{obs}}[Y W \mid p(Z,X)=u, X]$, and let $\underline{p}(X), \overline{p}(X)$ denote the conditional infimum and supremum of $p(Z,X)$ given $X$. Then the sharp lower bound in (ref) admits the representation \begin{equation} \begin{split} \theta_{\omega, 1} &= \mathbb{E}_{Z,X} \Bigg[ - \int_{\min\{p(X), q(Z,X)\}}^{p(X)} Q_{Y,p \mid X}\!\left({v}/{p(X)}\right) \, \mathrm{d}v \\ &\qquad + g_1\big(\min\{\overline{p}(X), q(Z,X)\}, X\big) - g_1\big(p(Z,X), X\big) \\ &\qquad + y_{\min} \max\{0, q(Z,X) - \overline{p}(X)\} \Bigg]. \end{split} \end{equation} The sharp upper bound admits an analogous representation, obtained by replacing $Q_{Y,\underline{p} \mid X}(v/\underline{p}(X))$ by $Q_{Y,\underline{p} \mid X}(1 - v/\underline{p}(X))$ in the first term and $y_{\min}$ by $y_{\max}$ in the third term.

(ref) is the continuous-instrument analog of (ref): it converts the sharp bound into an expectation over $(Z,X)$ that depends on estimable nuisances only. The first term is a boundary-quantile tail integral capturing the optimal-coupling contribution on $(0, \underline{p}(X))$ through the conditional quantile function of the identified complier distribution $\mu_{1,\underline{p}\mid X}$. The second term replaces the derivative of $g_1$ by a finite difference of $g_1$-values at the truncated limits---this is the key simplification enabled by Fubini, since $g_1$ itself is a standard conditional expectation estimable by flexible ML regressions. The third term is the closed-form trivial-bound contribution on $(\overline{p}(X),1]$, where $\omega(X,u) \geqslant 0$ for the PRTE weight. Crucially, the entire target requires only conditional expectations and tail quantiles, avoiding derivative estimation entirely. The proof is given in (ref).

(ref) motivates the following plug-in estimator, obtained by replacing each nuisance in ((ref)) with its machine-learning estimate trained on an auxiliary sample:

equation[equation omitted — 210 chars of source]

where $\hat{\psi}_1, \hat{\psi}_2, \hat{\psi}_3$ correspond respectively to the boundary-quantile integral, the $g_1$ difference, and the trivial-bound term in ((ref)). (ref) below establishes its nonparametric convergence rate. The rest of this subsection develops the nuisance estimators $\hat{p}, \hat{g}_1$, and $\widehat{Q}_{Y,\underline{p}\mid X}$ that appear in ((ref)).

\paragraph{Estimation Procedure.} Because the continuous instrument setting requires nonparametric estimation of conditional expectations and quantiles, we employ a three-fold sample split: a propensity-estimation sample $I_0$, a nuisance-estimation sample $I_1$, and a main evaluation sample $I_2$. Splitting off $I_0$ separately is what allows the localized boundary-quantile subsample $\mathcal{I}_\delta = \{i\in I_1 : \hat p(Z_i,X_i) \leqslant \widehat{\underline p}(X_i) + \delta_n\}$ to remain conditionally i.i.d.\ given $I_0$. The conditional expectation $g_1(u,x)$ is estimated by a kernel-weighted local $M$-estimator in $u$ combined with a flexible ML fit in $x$, yielding a piecewise-linear-in-$u$ estimator that is deterministically Lipschitz in $u$---a property no joint ML regression on $(u, X)$ is known to deliver. This construction, related to local-polynomial DR-learning kennedy2023towards and generalized random forests athey2019generalized, decouples the covariate-dimension ML rate $r_{X,n}$ from the smoothing bias $h_n^2$ in $u$. (ref) summarizes; the explicit kernel $M$-estimator, grid quantile regression, and target evaluation are in (ref).

algorithm[algorithm omitted — 1,084 chars of source]

To establish the formal convergence rate of our estimator in the continuous instrument setting, we must account for the estimation error of the generated regressor $\hat{p}$, the kernel-smoothed conditional expectation $\hat{g}_1$, and the localized boundary quantile $\widehat{Q}_{Y,\underline{p} \mid X}$. We impose the following regularity condition on the data generating process.

assumption[Regularity] The marginal density of the propensity score $p(Z,X)$ is bounded away from zero and bounded above on its support. Furthermore, the density at the lower boundary $\underline{p}(X)$ is strictly positive almost surely.
assumption[Nuisance Rates] Recall that $d_X$ is the dimension of the covariates $X$. We assume the nuisance estimators satisfy the following conditions: \begin{enumerate} • Rate for Propensity Score: The estimator $\hat{p}(z,x)$ converges to $p(z,x)$ at the nonparametric rate in the supremum norm, denoted by $\| \hat{p} - p \|_\infty = O_P(r_{p,n})$. • Regularity of $g_1$: The conditional expectation function $g_1(u, x)$ is twice continuously differentiable with respect to $u$. Furthermore, its first and second derivatives, $\frac{\partial g_1(u,x)}{\partial u}$ and $\frac{\partial^2 g_1(u,x)}{\partial u^2}$, are bounded uniformly over all $u \in [0,1]$ and $x \in \mathcal{X}$. • Localized Kernel Bandwidth: The 1D kernel function $K(\cdot)$ is continuously differentiable with bounded derivative. The bandwidth sequence is chosen such that $h_n \to 0, n h_n \to \infty$. • ML Rate and Realizability: Let $\mathcal{F}_n$ be the machine learning hypothesis class used for a sample of size $n$. We assume $\mathcal{F}_n$ is sufficiently rich such that $g_1(u,\cdot) \in \mathcal{F}, \forall u \in[0,1]$ and the statistical estimation error converges at a rate $r_{X,n}$ for an effective sample of size $n$. • $W_1$ rate for the conditional distribution of $Y$: Given $m$ i.i.d.\ samples $\{(Y_i, X_i)\}_{i=1}^m$ on $[y_{\min}, y_{\max}] \times \mathcal{X}$, the conditional distribution estimator satisfies $\sup_{x \in \mathcal{X}} W_1\bigl(\widehat{F}_{Y \mid X=x},\, F_{Y \mid X=x}\bigr) = O_P(r_Q(m))$ for some $r_Q(m) \to 0$. • $W_1$-Lipschitz boundary: Letting $F_{Y,\delta \mid x}$ denote the conditional distribution of $Y$ given $\{p(Z,X) \leqslant \underline{p}(x) + \delta, X = x\}$, this distribution is $W_1$-Lipschitz in $\delta$ at $\delta = 0$, uniformly in $x$: there exists $C_{\mathrm{Lip}} > 0$ such that $\sup_{x \in \mathcal{X}} W_1\!\bigl(F_{Y,\delta \mid x},\, F_{Y,\underline{p} \mid x}\bigr) \leqslant C_{\mathrm{Lip}}\,\delta$ for all sufficiently small $\delta > 0$. \end{enumerate}

Conditions 1--3 are standard regularity requirements: Condition 1 is satisfied by kernel, sieve, or $\ell_1$-penalized propensity estimators; Condition 2 ensures the piecewise-linear interpolation $\hat{g}_1$ incurs only an $h_n^2$ second-order bias; Condition 3 imposes standard kernel and bandwidth restrictions. Condition 4 is deliberately flexible: any ML procedure (random forests, neural networks, penalized linear methods) achieving kernel-weighted error rate $r_{X,n}$ is admissible. Conditions 5 and 6 ask only for conditional $W_1$-type rates---Condition 5 is strictly weaker than the uniform sup-norm rates established for standard conditional distribution and quantile regression estimators hall1999methods, guerre2012uniform, belloni2019conditional, while Condition 6 controls the bias from localizing to observations within $\delta_n$ of $\underline{p}(x)$.

The following theorem formally establishes the nonparametric convergence rate of the lower bound estimator.

theorem[Nonparametric Convergence Rate] Suppose (ref), (ref), (ref), (ref), (ref), and (ref) hold, the grid resolution satisfies $M \gtrsim \sqrt{n h_n}$, the localization bandwidth satisfies $\delta_n \to 0$ with $n\delta_n \to \infty$, and $r_{p,n} = o(\delta_n)$. Then, the sample-split estimator $\hat{\underline{\theta}}_{\omega, 1}$ converges to the true lower bound $\underline{\theta}_{\omega, 1}$ at the following rate: \begin{equation*} |\hat{\theta}_{\omega, 1} - \theta_{\omega, 1}| = O_P\!\left(\frac{r_{p,n}}{h_n} + r_{X,n} + h_n^2 + r_Q(n\delta_n) + \delta_n\right). \end{equation*}

Simulation

We validate our theoretical results through synthetic experiments and a real-data empirical application. All experiments compare our method (hereafter IVOT) bounds against the moment-relaxation bounds of magne2018using (hereafter IVMTE), which represent the state-of-the-art alternative. Detailed numerical tables and additional results are deferred to (ref).

Synthetic Experiment

We consider two synthetic settings, one with a continuous instrument and one with a discrete instrument, each using a distinct data-generating process (DGP). In both cases the target is $\theta_\alpha := \mathbb{E}[Y^{q_\alpha} - Y]$ under the policy $q_\alpha(Z) = \operatorname{clip}(p(Z)+\alpha, 0, 1)$ for $\alpha \in [-0.12, 0.12]$. For IVMTE, we specify both MTR functions $m_0,m_1$ as degree-$9$ $u$-splines on $[0,1]$ with nine interior knots at $\{0.1,0.2,\ldots,0.9\}$; this is a deliberately flexible, near-nonparametric sieve, chosen so that IVMTE is not disadvantaged by an overly restrictive parametric choice. Full DGP specifications are given in (ref).

\paragraph{Continuous instrument.} The continuous-instrument DGP uses a logistic propensity $p(Z) = \operatorname{logistic}(-1+2Z)$ with limited support $p(Z) \in [0.27, 0.73]$ and a constant marginal treatment effect $\text{MTE}(u) = 0.5$, so the ground truth is $\theta_\alpha \approx 0.5\alpha$ for small $|\alpha|$. This setting deliberately isolates the identification difficulty: even though the MTE is constant, the limited propensity support prevents moment-relaxation methods from recovering tight bounds. (ref)(a) displays the identified sets at $n = 5{,}000$. IVOT yields tighter bounds than IVMTE; at $\alpha = 0.05$, the IVOT interval width is $0.011$ versus the IVMTE width of $0.096$, a roughly $8\times$ reduction. The IVMTE bounds are notably asymmetric---the upper bound extends substantially above the truth while the lower bound crosses zero---illustrating how discarding distributional information leads to spuriously wide identified sets. Across the 25 reported $\alpha$ grid points in this design, the true $\theta_\alpha$ falls inside the IVOT interval at 22 grid points. The three near-misses occur at $|\alpha| \leqslant 0.02$ where the IVOT interval is extremely tight (width ${\approx}\,0.001$) and finite-sample error in propensity estimation pushes the truth just outside the bounds. With $n = 10{,}000$, this pointwise inclusion check holds at all 25 grid points ((ref)), confirming this is a finite-sample phenomenon.

\paragraph{Discrete instrument.} The discrete-instrument DGP uses $Z \sim \text{Bernoulli}(0.5)$ with two-point propensity score $\{0.25, 0.75\}$ and a heterogeneous $\text{MTE}(u) = 0.48 + 0.18u$, designed to test the closed-form bounds ((ref)) when the propensity takes finitely many values. (ref)(b) shows the results at $n = 5{,}000$. IVOT again yields tighter bounds, and the true $\theta_\alpha$ lies inside the IVOT interval at all reported $\alpha$ grid points. At $\alpha = 0.05$, the IVOT interval is approximately $2.4\times$ narrower than IVMTE (width $0.041$ versus $0.097$), and the IVOT 95% delta-method confidence interval $[-0.007,\; 0.046]$ is tighter than the IVMTE 95% backward CI $[-0.050,\; 0.050]$. Because the propensity support consists of only two points, the MTE is identified only on $(0.25, 0.75)$; IVOT tightens bounds by exploiting the full distributional information in the identified region when bounding the contribution from the unidentified regions $(0, 0.25)$ and $(0.75, 1)$.

\paragraph{Discussion.} The disparity between the two methods is consistent with the theoretical mechanism identified in (ref): IVMTE reduces the observed conditional distributions to first-moment constraints, whereas IVOT retains the full quantile structure. The advantage is largest when the propensity score support is contained to $[0,1]$ (so that the unidentified region is large) and when the outcome distribution is informative about its quantile structure in the identified region. Numerical details for selected $\alpha$ values are tabulated in (ref).

figure[figure omitted — 700 chars of source]

Effect of Price Subsidies for Bed Nets

\paragraph{Background and setup.} We apply our method to the bed net subsidy experiment of dupas2014short. A longstanding debate in development economics concerns the optimal pricing strategy for health-protective goods: while high prices screen out low-valuation users, subsidies expand take-up but may attract individuals who ultimately do not use the product. Understanding the causal effect of price reductions on product usage---not just on take-up---is therefore essential for welfare analysis. The dataset records insecticide-treated bed net (ITN) transactions and follow-up usage for $n = 1{,}078$ Kenyan households; the instrument $Z$ is the offered price (17 distinct values spanning 0--250 KSh, randomized across distribution points), the treatment $W \in \{0,1\}$ indicates purchase, and the outcome $Y \in \{0,1\}$ indicates one-year follow-up usage. The baseline policy is the reference price $z_0 = 150$ KSh. For each $\alpha \in [0.05, 0.62]$, we study the counterfactual policy target that raises the purchase propensity at that baseline price from $\hat p(z_0)$ to $q_\alpha := \min(\hat{p}(z_0) + \alpha, 1)$; thus the denominator of $\text{PRTE}_\alpha = \mathbb{E}[Y^{q_\alpha} - Y]/\alpha$ is the known shift size $\alpha$. The maximum $\alpha_{\max}\approx 0.621$ equals the propensity at zero price minus the baseline propensity. Because the instrument is discrete, the closed-form bounds of (ref) apply directly to the numerator, after which we divide by $\alpha$. For IVMTE, we report two sieve sizes (degree-$10$ and degree-$20$ $u$-splines on $[0,1]$) so that the shape-restriction trade-off is explicit: degree $20$ enlarges the MTR class and imposes weaker shape restrictions, so its identified set should widen relative to degree $10$. Full data, propensity-estimation, and policy-construction details are in (ref).

\paragraph{Results.} (ref) plots the IVOT and IVMTE identified sets for $\text{PRTE}_\alpha$ across subsidy levels. IVOT consistently yields a tighter identified set than IVMTE across the reported range of $\alpha$ and under both spline specifications. At the smallest subsidy ($\alpha = 0.05$), the IVOT interval is $[0.740, 1.000]$ (width $0.260$), while the IVMTE interval is $[0.129, 0.990]$ (width $0.861$) at degree $10$ and widens to $[-0.372, 1.002]$ (width $1.374$) at degree $20$---roughly a $3.3\times$ and $5.3\times$ reduction in width relative to IVOT, respectively. The IVMTE bound visibly widens as the spline degree increases from $10$ to $20$, empirically confirming that the higher-degree sieve is more nonparametric and imposes weaker shape restrictions on the MTR. IVOT's advantage is therefore robust across these two IVMTE specifications.

Both methods indicate a positive policy effect across most of the subsidy range---the IVOT lower bound is positive throughout, and the degree-$10$ IVMTE lower bound is positive for all $\alpha$, with the more conservative degree-$20$ lower bound turning positive once $\alpha \geqslant 0.14$. Moreover, IVOT achieves near-point identification (width $< 0.05$) for the majority of $\alpha$ values, exhibiting a periodic pattern in which the bounds tighten to near-zero width before widening slightly as additional complier groups enter the policy margin: for instance, the identified set is essentially a point for $\alpha \in \{0.08, 0.09, 0.11\}$ and again for $\alpha \in \{0.23, 0.24, 0.26\}$. At the maximum feasible subsidy ($\alpha \approx 0.621$), IVOT point-identifies $\text{PRTE}_\alpha \approx 0.597$. Full numerical results, including IVOT 95% delta-method confidence intervals, are reported in (ref).

figure[figure omitted — 693 chars of source]

Conclusion

We established a novel connection between the partial identification of PRTEs in IV models and OT. By formulating the problem as a CCOT problem over joint distributions compatible with the observed data, we showed that the multidimensional optimization analytically reduces to one-dimensional OT problems with product costs, yielding explicit closed-form expressions for the sharp bounds and bypassing the computationally intensive optimization required by existing approaches. The framework extends to general treatment settings (continuous and multi-valued), and we develop corresponding estimation results: Neyman-orthogonal DML scores deliver $\sqrt{n}$-consistency and asymptotic normality for discrete instruments, while for continuous instruments we characterize nonparametric convergence rates. A central insight, demonstrated theoretically and empirically, is that preserving full distributional information is both tractable and valuable for tight partial identification: moment-relaxation approaches that reduce the observed data to conditional expectations discard higher-order distributional features and can produce substantially wider bounds, whereas the CCOT formulation enforces compatibility with the entire observed distribution and makes the underlying optimization analytically transparent. We believe this insight extends beyond the IV setting studied here.

\paragraph{Acknowledgements. } Jose Blanchet gratefully acknowledges support from DoD through the grants Air Force Office of Scientific Research under award number FA9550-20-1-0397 and ONR 1398311, also support from NSF via grants 2229012, 2312204, 2403007 is gratefully acknowledged. Vasilis Syrgkanis gratefully acknowledges support from NSF Award IIS-2337916, an Amazon Research Award and a Google Research Award.