Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
142,003 characters · 24 sections · 77 citation commands
Partial Identification of Policy-Relevant Treatment Effects with Instrumental Variables via Optimal Transport
Instrumental variables (IVs) are a standard tool for identifying causal effects when treatment is endogenous in observational studies. Under the classical IV assumptions, angrist1995identification show that the Local Average Treatment Effect (LATE), defined as the average effect for the subpopulation of compliers, is point-identified. However, many policy questions are not about compliers themselves; they ask how outcomes would change under a counterfactual policy that shifts treatment take-up in the broader population. Such questions are naturally captured by the Policy-Relevant Treatment Effects (PRTEs), a central target in the generalized Roy and Marginal Treatment Effect (MTE) framework heckman1999local,heckman2005structural,heckman2006understanding, which encompass a wide range of causal estimands including the Average Treatment Effect (ATE). Empirical policy evaluations within this framework include carneiro2010evaluating,carneiro2011estimating. Unlike LATE, PRTEs are generally not point-identified when the instrument induces only limited support in the treatment propensity.
This limited-support problem is central in the generalized Roy model heckman2005structural, where treatment selection follows a threshold-crossing mechanism formally equivalent to the classical monotonicity condition angrist1995identification,vytlacil2002independence. In this framework, PRTEs can be written as weighted averages of the MTE, so identification depends on how much of the latent resistance distribution is reached by the propensity score induced by the instrument. When the propensity score lacks full support over the unit interval, PRTEs and related quantities such as the ATE remain only partially identified, and a long line of work has studied such identified sets under various assumptions manski1990nonparametric,manski1997monotone,manski_pepper2000monotone,heckman2001instrumental,magne2018using.
Despite its importance, deriving tight, closed-form bounds for general PRTEs remains a significant challenge. Existing strategies fall into three broad approaches: moment-relaxation over MTR functions magne2018using,mogstad2018identification, second-order stochastic dominance constraints on the MTE function solved by rearrangement marx2024sharp, and infinite-dimensional linear programming over the joint law of outcome and latent resistance solved by sieve approximation han_yang2024computational. Each faces limitations---discarded distributional information, restricted scope to weights on the complier interval under a binary instrument, lack of closed-form solutions or incompatibility with continuous instruments---that we discuss in detail in (ref).
In this paper, we develop a unified Optimal Transport (OT) framework for partial identification in the generalized Roy model, and use it to derive sharp, closed-form bounds for PRTEs. Specifically, we recast PRTE partial identification as a Constrained Conditional Optimal Transport (CCOT) problem, a structured instance within the broader OT framework that is equivalent to the infinite-dimensional LP of han_yang2024computational. Rather than summarizing the observed distributions through moment conditions or imposing distributional constraints on conditional mean (MTE) functions, we place the distributional constraints directly on the joint conditional law of the potential outcome and latent resistance. Then, we exploit the OT structure to prove that this heavily constrained, multidimensional CCOT problem analytically reduces to separable standard one-dimensional OT problems. This reduction completely bypasses computationally intensive linear programming and sieve approximation, yielding explicit, closed-form solutions for the sharp bounds. Because the framework constrains the full joint latent law rather than a conditional mean, it applies in principle to a broader class of partial identification problems in the generalized Roy model; here we focus on PRTEs.
Our main contributions can be summarized as follows.
\paragraph{Organization of this paper.} The remainder of the paper is organized as follows. (ref) introduces the setup and optimal transport background; (ref) presents the main identification theory and closed-form bounds; (ref) extends the framework to general treatments; (ref) develops estimation and inference; and (ref) reports simulation and empirical results.
\paragraph{Local IV and MTE.} The Marginal Treatment Effect (MTE), introduced by heckman1999local and systematically developed by heckman2005structural, is the foundational building block of our identification analysis. The MTE is defined as the treatment effect for individuals at a specific quantile of unobserved resistance to treatment, and serves as a unifying object from which all standard IV estimands--including LATE, ATE, ATT, and PRTE--can be recovered as weighted averages heckman2005structural, heckman2006understanding. The formal equivalence between the threshold-crossing selection model and the monotonicity assumption of angrist1995identification is established by vytlacil2002independence, who shows that the two frameworks are equivalent under a continuity condition on the unobservable. Point identification of the MTE is achieved via the Local Instrumental Variables (LIV) estimand heckman1999local, which requires full propensity score support over $[0,1]$---a condition that fails under limited instrument variation, the regime our paper addresses. Empirical MTE-based policy evaluations include carneiro2010evaluating,carneiro2011estimating.
\paragraph{IV with General Treatment.} angrist1995two extends the LATE framework to treatments with multiple discrete ordered levels. heckman2007econometric extends the MTE framework to the ordered choice model. kirkeboen2016field provides methodology for identifying treatment effects in settings with multiple unordered discrete choices--specifically fields of study--by accounting for how individuals self-select into alternatives based on their unobserved comparative advantages. lee2018identifying develop identification results for multivalued treatments. In the continuous treatment setting, florens2008identification use the control function method to identify the ATE and ATT, and imbens2009identification use a similar approach to identify structural equations in models with continuous endogenous variables. While these results establish point identification for certain causal quantities, our framework instead targets partial identification of PRTEs, and we extend our results to the continuous and multi-valued treatment settings in (ref).
\paragraph{Partial Identification.} manski1989anatomy,manski1990nonparametric,manski1997monotone,manski_pepper2000monotone,manskiInferenceRegressionsInterval2002 pioneered partial identification, providing closed-form bounds for the IV problem under a range of assumptions. In the econometric literature, partial identification is often cast as estimating the identified set defined by moment inequalities chernozhukov_hong_tamer2007,rosen2008confidence,romano2008inference,canay2010inference,bugni2010bootstrap; see tamer2010partial for a comprehensive survey. As a special case, our bound recovers the classical Manski bound for ATE. The computational and analytical approaches of han_yang2024computational and marx2024sharp to PRTE partial identification in the MTE framework are discussed in the PRTE Partial Identification paragraph below.
balke_pearl1997bounds develop the linear programming approach for partial identification. Recently, this optimization-based approach has been revisited and generalized using modern techniques balazadeh_meresht_syrgkanis_krishnan2022,guo_yin_wang_jordan2022,duarte2024automated,voronin2025linear,levis2025covariate. In particular, levis2025covariate use covariates to tighten balke_pearl1997bounds's classical bound and develop estimation theory in the IV setting. Several works establish connections between partial identification and robust optimization guo_yin_wang_jordan2022,balazadeh_meresht_syrgkanis_krishnan2022,gao_ge_qian2024,penn_gunderson_bravohermsdorff_silva_watson2025,fan_pass_shi2025,tan2024consistency. The partial identification problem can be formulated as an optimal transport problem with marginal constraints derived from the observational distribution ji_lei_spector2024,lin2025estimation,lin2025tightening. For a recent primer on optimal transport methods for causal inference, see gunsilius2025primer. In a related but distinct setting, fan_guerre_zhu2017partial study partial identification of functionals of the joint distribution of potential outcomes in settings where the conditional marginal distributions are identified---including the latent threshold-crossing model of heckman2005structural under full propensity-score support. oberreynolds2023estimating characterizes sharp bounds on functionals of the joint complier outcome distribution via optimal transport under the classical binary-instrument LATE framework of angrist1995identification, focusing on the identified complier subpopulation rather than the partially identified PRTEs we target here.
\paragraph{PRTE Partial Identification.} Three strategies have emerged for PRTE partial identification under limited support. magne2018using optimize over MTR functions subject to moment constraints mogstad2018identification, which discards distributional information beyond the first moment. han_yang2024computational preserve full distributional content by formulating an infinite-dimensional LP over the joint law of latent states, solved by sieve approximation for discrete instruments. marx2024sharp generalize magne2018using by imposing full distributional content on the MTE function via a second-order stochastic dominance (SSD) constraint, obtaining sharp closed-form bounds by rearrangement bounds; his explicit theorems cover weights on the complier interval under a binary instrument, with extensions beyond compliers requiring rank similarity.
Our work differs from all three in scope: we impose distributional constraints directly on the joint conditional law of the potential outcome and latent resistance rather than on a conditional mean, cast the resulting program as a CCOT problem equivalent to the LP of han_yang2024computational, and show it reduces analytically to separable one-dimensional OT problems with product costs. This delivers (i) sharp closed-form bounds for the full PRTE on all of $[0,1]$ under multi-valued instruments via an interval decomposition across $K\!+\!1$ compliance types, including boundary intervals where only one potential outcome marginal is identified, (ii) extensions to continuous and mixed instruments with sub-problems indexed by the propensity score continuum, and (iii) discrete-instrument inference together with continuous-instrument rate results, which are not available in either marx2024sharp or han_yang2024computational. On the complier interval under a binary instrument, the CCOT solution and Marx's SSD rearrangement coincide. Beyond PRTE, our sharpness result characterizes the identified set at the level of the full joint latent law and sheds light on the joint-identification question raised in marx2024sharp through its reduction to independent one-dimensional marginal constraints.
For most of this paper, we work with the generalized Roy model heckman2005structural. In the model, the observable variables are treatment $W\in \mathcal{W} = \{0,1\}$, instrument $ Z \in \mathcal{Z}$, covariates $X\in\mathcal{X}$ and the outcome $Y \in \mathcal{Y}$, where $\mathcal{X} \subset\mathbb{R}^{d_X}, \mathcal{Z}$ and $ \mathcal{Y} \subset \mathbb{R}$ are the domains of the corresponding variables. We consider both discrete and continuous instrument settings in this paper. Throughout the paper, we assume the domains of all random variables are compact. The unobservables are the potential outcomes $Y(0), Y(1)$, and the variable $U$ in the selection equation ((ref)).
It is without loss of generality to assume that $U\sim\text{Unif}(0,1)$ if $U$ is continuously distributed, because we can always normalize its marginal distribution.\footnote{By the probability integral transform, if $U \mid X$ has a continuous CDF $F_{U\mid X}(\cdot \mid x)$, then $\tilde{U} := F_{U\mid X}(U \mid X) \sim \mathrm{Unif}(0,1)$ conditional on $X$. Since $\tilde{U}$ is a strictly increasing function of $U$, the threshold-crossing model is preserved under this reparametrization, and we may work with $\tilde{U}$ in place of $U$ without loss of generality.} The common instrument support condition is essential for the identification of the propensity score $p(z,x)$: without $z \in \mathrm{Supp}(Z \mid X=x)$ for all $x$, the conditional probability $\mathbb{P}(W=1 \mid Z=z, X=x)$ is not well-defined and hence $p(z,x)$ is not identified.
We also require the following boundedness condition on the potential outcomes, which is needed to obtain informative bounds.
The fundamental building block for evaluating causal parameters in this framework is the MTE, originally introduced by heckman1999local,heckman2005structural. The MTE is defined as the average treatment effect for individuals with covariates $X=x$ and latent variable $U=u$: $$ \text{MTE}(x,u) = \mathbb{E}[Y(1) - Y(0) \mid X=x, U=u]. $$
In policy analysis, we are often interested in the effect of an intervention that shifts the distribution of the instrument $Z$, or more generally, alters the propensity score from a baseline status to a new regime. Let the baseline policy be characterized by the current propensity score $p(Z,X)$ and the corresponding treatment status $W$. Consider an alternative policy that induces a new propensity score $q(Z,X)$ and a new counterfactual treatment status $W^q = \indicator(U \leqslant q(Z,X))$, where we assume the latent variable $U$ is invariant to the policy change, as is standard in the MTE framework.
The PRTE evaluates the average per-person impact of shifting from the baseline policy to the alternative policy. Let $Y^q$ denote the outcome realized under the alternative policy, such that $Y^q = W^qY(1) + (1-W^q)Y(0)$. The PRTE is formally defined as: $$ \text{PRTE} = \frac{\mathbb{E}[Y^q - Y]}{\mathbb{E}[W^q - W]}. $$
Many common treatment effect parameters can be expressed as a weighted average of the MTE. Specifically, the PRTE can be rewritten as: $$ \text{PRTE} = \mathbb{E}_{X,U}[\text{MTE}(X, U) \omega_{\text{PRTE}}(X,U)], $$ where the policy-specific weight function is given by: $$ \omega_{\text{PRTE}}(x,u) = \frac{\mathbb{P}(q(Z,X) \geqslant u \mid X=x) - \mathbb{P}(p(Z,X) \geqslant u \mid X=x)}{\mathbb{E}[q(Z,X) - p(Z,X)]}, $$ provided $\mathbb{E}[q(Z,X) - p(Z,X)] \neq 0$. Following heckman2005structural,magne2018using, we consider the following more general target parameter.
where $\omega(x,u)$ is an identifiable function. We can obtain different PRTEs by taking
for given propensity functions $q_0$ and $q_1$, where $q_0 = p$ and $q_1 = q$ correspond to the baseline and alternative propensity score functions, respectively, in the PRTE case. The coefficient hidden inside the proportionality symbol is identifiable. (ref) shows a variety of common causal parameters under different choices of $\omega$.
While our bounds recover the more classical bound for the first 4 quantities in (ref), we will focus in particular on PRTEs. As a leading example, consider a uniform policy shift of the form $q(Z,X) = \min(p(Z,X) + \alpha, 1)$ for some $\alpha > 0$. This corresponds to a policy that uniformly raises the probability of treatment uptake by $\alpha$. In our empirical application ((ref)), the instrument is the offered price of an insecticide-treated bed net dupas2014subsidies, and $q = p + \alpha$ models a uniform price subsidy that lowers the cost of purchase and thereby increases take-up probability by $\alpha$.
heckman2001instrumental shows that if the support of $p(Z,x)$ spans over the entire interval $(0,1)$ for all $x$, then MTE is identifiable. However, because the common support of $p(Z,X)$ is often a strict subset of $(0,1)$, the MTE is generally not identifiable. For instance, when both $X$ and $Z$ are discrete, the supports of $p(Z,x)$ are discrete and this common support assumption is violated. Consequently, the integral defining the PRTE is not point-identified without imposing strong parametric extrapolations.
To build the framework for our partial identification results, we briefly introduce the relevant concepts from OT theory. OT provides a mathematical framework for coupling two probability distributions at minimum cost, and has emerged as a powerful toolkit for bounding joint distributions when only their marginals are known galichon2016optimal.
Let $\mu$ and $\nu$ be two probability measures on state spaces $\mathcal{A}$ and $\mathcal{B}$, respectively. In the context of partial identification, we often observe the marginal distributions of two variables but not their joint distribution. We define $\Pi(\mu, \nu)$ as the set of all couplings between $\mu$ and $\nu$. Formally, a coupling $\pi \in \Pi(\mu, \nu)$ is a joint probability measure on the product space $\mathcal{A} \times \mathcal{B}$ such that its marginals are exactly $\mu$ and $\nu$.
Given a cost function $c: \mathcal{A} \times \mathcal{B} \to \mathbb{R} \cup \{+\infty\}$ that represents the cost of associating state $a$ with state $b$, the standard Kantorovich optimal transport problem seeks the coupling that minimizes the expected cost:
Generally, solving ((ref)) in multidimensional spaces requires computationally intensive linear programming. However, a celebrated result in OT literature demonstrates that when the spaces are one-dimensional and the cost function satisfies specific structural conditions--such as supermodularity--the optimal coupling admits a sharp, closed-form solution.
We formalize this one-dimensional closed-form solution in the following theorem, which will be instrumental in deriving the analytic bounds for the PRTE.
(ref) provides a closed-form characterization of the optimal value of 1D OT problems. We will leverage this characterization to derive tight bounds for PRTEs later.
We use $\mathbb{P}$ and $\mathbb{E}$ for probability and expectation, and $\perp\!\!\!\perp$ for conditional independence. $\indicator(\cdot)$ denotes the indicator function. We write $\mathrm{d}$ for the differential or Lebesgue measure. For a probability measure $\mu$ on $\mathbb{R}$, $Q_\mu(t)$ denotes its quantile function. $\Pi(\mu, \nu)$ denotes the set of all couplings of two probability measures $\mu$ and $\nu$. Calligraphic letters $\mathcal{X}, \mathcal{Y}, \mathcal{Z}, \mathcal{W}$ denote domains of the corresponding variables. We write $O_P(\cdot)$ and $o_P(\cdot)$ for stochastic order notation: $X_n = O_P(a_n)$ means $X_n/a_n$ is bounded in probability, and $X_n = o_P(a_n)$ means $X_n/a_n \xrightarrow{P} 0$.
In this section, we demonstrate how to derive closed-form, tight bounds for the PRTE. We begin by formally mapping the partial identification problem into an optimal transport framework.
The core intuition is to search over all possible joint distributions of the latent and observable variables $(Y(0), Y(1), U, X, Z, W)$ that are compatible with both the observed data and the structural conditions in (ref). Because the potential outcomes are conditionally independent of the instrument given $X$, the identifying power of our model is entirely captured by the joint distributions of $(Y(w), U, X)$ for $w \in \{0, 1\}$.
Let $\pi_w$ denote the joint probability measure of $(Y(w), U, X)$ and $ \pi_w(\cdot\mid x,u) $ be the conditional distribution of $Y(w)$ given $x,u$. Recall that our target parameter is a linear functional of these measures: $$\theta_\omega(\pi_0, \pi_1) = \mathbb{E}_{(X,U,Y(1))\sim\pi_1} [Y(1) \omega(X,U)] - \mathbb{E}_{(X,U,Y(0))\sim\pi_0} [Y(0) \omega(X,U)]. $$ To achieve sharp bounds, we must constrain $\pi_w$ using all available observational information. Let $\mathbb{P}_{obs}$ denote the true observed joint distribution of $(Y, W, Z, X)$. Under the threshold-crossing model in (ref), while $U$ is generally unobserved, its conditional distribution is explicitly tied to the treatment status and the propensity score. Specifically, conditional on $X$ and $Z$, the event $W=1$ perfectly corresponds to $U \leqslant p(Z,X)$, and $W=0$ corresponds to $U > p(Z,X)$.
Consequently, the joint distributions $\pi_0$ and $\pi_1$ fully parameterize the observational distribution of $(X,Z,W,Y)$. To see this, we can link the unknown structural distributions directly to the observed conditional distribution of the data by exploiting the consistency equation $Y = W Y(1) + (1-W) Y(0)$ and the selection mechanism $W = \indicator(U \leqslant p(Z,X))$. For any measurable set $A \subseteq \mathcal{Y}$ and $W=1$, we have:
By the conditional independence assumption $Z \perp\!\!\!\perp (Y(1), U) \mid X$, we can omit the $Z$ in the second conditioning:
Since instruments sharing the same propensity score $p(z,x)$ yield the same right-hand side value, we can replace the conditioning on $Z$ with conditioning on $p(Z)$. Next, we disintegrate the joint measure $\pi_1$ into the conditional distribution of the potential outcome given the covariates and the latent variable. Because $U \mid X \sim \text{Unif}(0,1)$, integrating over the region $U \leqslant p(z,x)$ yields:
Following the exact same sequence of steps for $W=0$, where the observed outcome is $Y(0)$ and the selection mechanism dictates $U > p(z,x)$, we obtain the symmetric result:
Furthermore, by construction, the marginal distribution of $(U, X)$ under $\pi_w$ must equal the known population distribution $\mathbb{P}_{U,X}$.
This allows us to formulate the partial identification of the PRTE as the following infinite dimensional linear programming problem:
We call the optimization in (ref) the Constrained Conditional Optimal Transport (CCOT) problem. The feasible set is constrained by the observational compatibility conditions in (ref); the problem is conditional in that, for fixed covariates $x$ and latent resistance $u$, it seeks a conditional measure $\pi_w(\cdot \mid u, x)$ on $\mathcal{Y}$; and it is a transport problem because the objective is linear in $\pi_w(\cdot \mid u, x)$ with effective cost $c(y, u, x) = y \cdot \omega(x, u)$, so that mass is transported from the latent space $U$ to the outcome space $Y(w)$ conditional on $X$.
To formally guarantee that the CCOT formulation in ((ref)) yields sharp bounds for the PRTE, we must show that any pair of measures $(\pi_0, \pi_1)$ satisfying the CCOT constraints corresponds to a valid data-generating process that satisfies our structural assumptions and perfectly reproduces the observed data. Let $\Gamma(\mathbb{P}_{obs})$ be the set of all measure pairs $(\pi_0, \pi_1)$ satisfying the constraints in ((ref)). The following proposition characterizes the joint distribution of $(Y(0), Y(1), U, X)$, from which observational equivalence of the marginal pair $(\pi_0, \pi_1)$ follows immediately.
The content of (ref) is twofold. First, the CCOT constraints act only on the two marginals $\tilde{\pi}_0$ and $\tilde{\pi}_1$: any coupling of these marginals---including every joint law of $(Y(0), Y(1))$ given $(U, X)$---is observationally equivalent. Second, the converse makes the characterization sharp: no joint law outside this set can be generated by a structural model matching $\mathbb{P}_{obs}$. In particular, applying (ref) to the independent coupling $\tilde{\pi} := \tilde{\pi}_0 \otimes_{(U, X)} \tilde{\pi}_1$ shows that every marginal pair $(\tilde{\pi}_0, \tilde{\pi}_1) \in \Gamma(\mathbb{P}_{obs})$ is observationally equivalent to the true data-generating process. Consequently, the identified set for the PRTE target is $\{\theta_\omega(\pi_0, \pi_1) : (\pi_0, \pi_1) \in \Gamma(\mathbb{P}_{obs})\}$, and the maximum and minimum of $\theta_\omega(\pi_0, \pi_1)$ over $\Gamma(\mathbb{P}_{obs})$ are the sharp upper and lower bounds on the PRTE.
Next, we show how to obtain the closed-form solution of the CCOT problem ((ref)). Notice that in ((ref)), the objective is the difference of two linear functionals of $\pi_0$ and $\pi_1$, and the constraints on $\pi_0$ and $\pi_1$ are separate. Therefore, we only need to consider one side of the problem, as the other side follows identically. In what follows, we only consider the $W=1$ side minimization problem, i.e.,
The solution for the maximization problem can be derived similarly. Writing $\theta_\omega = \theta_{\omega,1} - \theta_{\omega,0}$, let $[\underline{\theta}_{\omega,w}, \overline{\theta}_{\omega,w}]$ denote the sharp identified interval for the $w$-th component. Because the feasible sets for $\pi_0$ and $\pi_1$ are Cartesian products, the sharp identified interval for the full target is \[ [\underline{\theta}_\omega, \overline{\theta}_\omega] = [\underline{\theta}_{\omega,1} - \overline{\theta}_{\omega,0},\, \overline{\theta}_{\omega,1} - \underline{\theta}_{\omega,0}]. \] Accordingly, we state the main derivations for $\theta_{\omega,1}$; the formulas for $\theta_{\omega,0}$ are obtained by the same arguments with the untreated analogues.
To better illustrate the idea, we first derive the results without covariates and generalize them in the next subsection. Therefore, in this subsection, $\omega$ and $p$ do not depend on $x$.
We first consider the setting where there are no covariates and the instrument $Z$ takes values in a finite discrete set $\mathcal{Z}$. Let the unique values of the propensity score be ordered as: \[ \mathrm{Range}(p(z)) := \{p(z) : z \in \mathcal{Z}\} = \{p_1, \dots, p_K\}, \] with the conventions $p_0 := 0$ and $p_{K+1} := 1$, such that $0 = p_0 \leqslant p_1 < p_2 < \dots < p_K \leqslant p_{K+1} = 1$. We allow different instruments to be mapped to the same propensity value. This discretizes the unit interval of the latent variable $U$ into $K+1$ disjoint sub-intervals, $I_i = (p_i, p_{i+1}]$ for $i=0, \dots, K$.
Because the objective function ((ref)) is an integral over $U$, we can additively decompose it across these sub-intervals:
The fundamental advantage of the discrete setting is that the global observational constraints difference out, perfectly isolating the conditional distribution of $Y(1)$ within each interval $I_i$. Recall that from ((ref)),
The key insight of this work is that ((ref)) can be decomposed into separate marginal constraints, which enables us to convert ((ref)) into a series of seperate standard one-dimensional OT problems. For $i = 1, \dots, K-1$, subtracting the $i$-th constraint from the $(i+1)$-th in ((ref)) isolates the integral over $(p_i, p_{i+1}]$; for $i=0$, the first constraint directly governs $(0, p_1]$. Since cumulative sums recover the original constraints, this invertible transformation shows that the system ((ref)) is equivalent to
where $\mu_{1,i}$ is given by
Note that for the final interval $I_K = (p_K, 1]$, the potential outcome $Y(1)$ is never observed because no individual in the population has a propensity score high enough to induce treatment when $U > p_K$. By ((ref)), since $\pi_1$ is nonnegative, $\mathbb{P}_{obs}(\mathrm{d}y, W=1 \mid p(Z)=p)$ is non-decreasing in $p$. Note that $\mu_{1,i}(\mathcal{Y}) = 1$ since $\mathbb{P}_{obs}(W=1 \mid p(Z)=p) = p$. Therefore, the $\mu_{1,i}$ are valid distributions. mourifie2017testing provide a simple method to test the monotonicity of $\mathbb{P}_{obs}(\mathrm{d}y, W=1 \mid p(Z)=p)$.
For $0<i<K$, notice that
for any measurable set $A$, where we use (ref).(3) and consistency in the second equality and (ref).(2) in the last equality. Therefore, $\mu_{1,i}$ can be interpreted as the distribution of the potential outcome under treatment for the $i$-th complier group. The $i$-th complier group consists of units that refuse treatment when $p(Z) \leqslant p_i$ and comply only when $p(Z) > p_i$. In particular, when the instrument is binary, this recovers the identification results of imbens1997estimating
Substituting the objective decomposition and the interval-wise constraints derived above, the original problem ((ref)) can be equivalently written as
subject to the $K$ independent marginal constraints
with the support constraint $\text{supp}(\pi_1(\cdot \mid u)) \subseteq \mathcal{Y}$ for all $u \in [0,1]$, and no further distributional constraint on $\pi_1(\cdot \mid u)$ for $u \in I_K = (p_K, 1]$. Crucially, each constraint in ((ref)) involves $\pi_1$ only on the single interval $I_i$, and the objective ((ref)) is additively separable across these intervals. Therefore, the global minimization problem decomposes into $K+1$ independent subproblems, one for each interval.
Specifically, for each interval $i = 0, \dots, K-1$, the data perfectly fixes the marginal distribution of $Y(1)$ as $\mu_{1, i}$ and the marginal distribution of $U$ as uniform on $(p_i, p_{i+1}]$. Let $\nu_i$ denote the uniform probability measure $\text{Unif}(p_i, p_{i+1})$. The set of valid joint distributions for $(Y(1), U)$ conditional on $U \in I_i$ is exactly the set of all couplings $\Pi(\mu_{1, i}, \nu_i)$. Therefore, for each of these $K$ intervals, we solve a standard 1D optimal transport problem:
Conversely, for the final interval $I_K = (p_K, 1]$, the propensity score never reaches a value high enough to induce treatment, meaning the data provides absolutely no constraints on the distribution of $Y(1)$ for this subpopulation. Without distributional constraints, the optimal transport problem degenerates into a trivial pointwise minimization. By (ref), for each $u \in I_K$, the optimal solution assigns all probability mass to $y_{\min}$ if $\omega(u) \geqslant 0$, or to $y_{\max}$ if $\omega(u) < 0$:
Summing the optimal values of the $K$ constrained subproblems in ((ref)) scaled by their respective interval lengths $(p_{i+1} - p_i)$ and adding the trivial unconstrained bound from ((ref)), we obtain the global minimum. Applying (ref) to each subproblem in ((ref)) yields the countermonotonic coupling, which leads directly to the following closed-form sharp bounds.
We defer the formal proof of this theorem to (ref), since the present theorem is a special case ($\mathcal{X}=\emptyset$, every LIV interval degenerate). (ref) provides a closed-form, analytically tractable expression for the sharp bounds, bypassing any need for linear programming and requiring only estimation of the conditional quantile function from the observed data.
As a special case, if we set $ \omega(u) = 1 $, (ref) recovers the classical bound by manski1990nonparametric for $\mathbb{E}[Y(1)]$, which is known to be tight heckman2001instrumental:
(ref) illustrates the three identification regions that compose the bound in (ref); (ref) in the appendix works through a concrete binary-instrument example with a constant propensity score, deriving the explicit three-part decomposition.
By (ref), the length of the bound is
which depends on the range of the propensity function and distribution $\mu_{1,i}$. If $p_K$ is close to $1$ and the support of $\mu_{1,i}$ is small, the length of the bound is also small. In particular, if each $Q_{Y,i}$ is constant and $p_K = 1,p_0 = 0$, $\theta_{\omega,1}$ is identifiable from the data.
We now unify consider a general instrument setting under a single regularity condition on the instrument domain. Rather than assuming $\mathcal{Z}$ is finite, we only require it to be compact with finitely many connected components. This framework nests the finite discrete case (each component a singleton), the connected continuous case of the LIV literature heckman1999local, and any intermixture of the two.
Under (ref), the image of each connected component of $\mathcal{Z}$ under $p(\cdot)$ is a closed interval (possibly a singleton if $p$ is constant on the component or if the component is a single point). Merging overlaps and ordering from left to right, the range of the propensity score admits a maximal-interval decomposition
where we explicitly allow degenerate intervals $\underline{p}_k = \overline{p}_k$ corresponding to isolated propensity values, as in the finite discrete case. Adopting the boundary conventions $\overline{p}_0 := 0$ and $\underline{p}_{K+1} := 1$, this decomposition partitions the unit interval of the latent variable $U$ into three types of regions: the LIV intervals $L_k := [\underline{p}_k,\overline{p}_k]$ for $k=1,\dots,K$, the OT gaps $G_k := (\overline{p}_k, \underline{p}_{k+1})$ for $k=0,\dots,K-1$, and the trivial tail $(\overline{p}_K,1]$.
The CCOT problem for the treated ($W=1$) component becomes
Following the same strategy as in (ref), we rewrite ((ref)) into an obviously separable form across the three region types. The objective splits additively as
On each gap $G_k$, the same telescoping argument that produced ((ref)) --- differencing the observational constraint in ((ref)) at $p=\overline{p}_k$ and $p=\underline{p}_{k+1}$ --- fixes the marginal of $Y(1)$ on $G_k$ to be
with the convention $\mathbb{P}_{obs}(\mathrm{d}y, W=1 \mid p(Z)=0) \equiv 0$ so that $k=0$ reduces to $\mu_{1,0}(\mathrm{d}y) = \mathbb{P}_{obs}(\mathrm{d}y, W=1 \mid p(Z)=\underline{p}_1)/\underline{p}_1$. On each non-degenerate LIV interval $L_k$, the constraint in ((ref)) holds for every $p$ in the continuum $[\underline{p}_k,\overline{p}_k]$; differentiating in $p$ pins $\pi_1(\mathrm{d}y\mid u)$ down pointwise via the LIV identity
whose absolute continuity in $p$ follows from the integral representation in ((ref)). The trivial tail $(\overline{p}_K,1]$ carries no constraint. Because cumulative sums of the gap marginals plus integration of the LIV densities recover the original constraint at every $p\in p(\mathcal{Z})$, this transformation is invertible and the two constraint systems are equivalent. Each rewritten constraint involves $\pi_1$ only on its own region and the objective is additively separable, so the global problem decomposes into three types of sub-problems, one for each region type.
Combining the solutions across these regions yields the closed-form sharp bounds for the general instrument setting, which are illustrated in (ref).
The proof combines the subtraction-of-constraints decomposition from (ref) (applied region-by-region to identify $\mu_{1,k}$ on each gap) with the LIV-derivative argument on each non-degenerate interval $L_k$, followed by the 1D OT coupling of (ref). We defer the formal proof to (ref), which restates the result in the unified form in the covariate subsection below. (ref) subsumes both classical settings as limit cases. When every LIV interval is degenerate ($\underline{p}_k = \overline{p}_k$ for all $k$), the LIV sum vanishes identically, each $\mu_{1,k}$ coincides with the discrete complier distribution in ((ref)), and ((ref)) reduces exactly to (ref). Conversely, when $K=1$ with $\underline{p}_1 < \overline{p}_1$, the gap sum contains a single OT term on $(0,\underline{p}_1)$, the LIV sum contributes a single integral on $[\underline{p}_1,\overline{p}_1]$, and we recover the connected-continuous case; in particular, we match the identification results of heckman1999local when additionally $\underline{p}_1 = 0$ and $\overline{p}_1 = 1$, so the propensity score has full support with no gaps. The length of the bound depends on the gap widths, the identified gap distributions $\mu_{1,k}$, and the tail beyond $\overline{p}_K$:
Generalizing the previous results to accommodate covariates $X$ is mathematically straightforward. Because our structural assumption (ref) imposes conditional independence $Z \perp\!\!\!\perp (Y(w), U) \mid X$, the global CCOT problem naturally disintegrates into a family of conditional one-dimensional optimal transport problems, indexed by $x \in \mathcal{X}$. We can therefore solve the optimization conditional on $X=x$ and aggregate the resulting conditional bounds over the marginal distribution of $X$ via the law of iterated expectations.
Under (ref), for $\mathbb{P}_{obs}$-a.e.\ $x \in \mathcal{X}$ the conditional range of the propensity score admits a maximal-interval decomposition
with degenerate intervals $\underline{p}_k(x) = \overline{p}_k(x)$ corresponding to isolated conditional propensity values and boundary conventions $\overline{p}_0(x) := 0$ and $\underline{p}_{K_x+1}(x) := 1$. For each $x$, this partitions $[0,1]$ into the conditional LIV intervals $L_k(x) := [\underline{p}_k(x), \overline{p}_k(x)]$, the OT gaps $G_k(x) := (\overline{p}_k(x), \underline{p}_{k+1}(x))$, and the trivial tail $(\overline{p}_{K_x}(x), 1]$. The conditional gap-identified measure of $Y(1)$ on $G_k(x)$ is
for $k=0,\dots,K_x-1$, with the convention $\mathbb{P}_{obs}(\mathrm{d}y, W=1 \mid p(Z,X)=0, X=x) \equiv 0$.
The proof of (ref) is given in (ref). As in the no-covariate case, (ref) subsumes both the finite-discrete and the connected-continuous settings: when every conditional LIV interval is degenerate ($\underline{p}_k(X) = \overline{p}_k(X)$ a.s.), the LIV sum vanishes and the expression reduces to the discrete covariate bound; when $K_X = 1$ with $\underline{p}_1(X) < \overline{p}_1(X)$, the gap sum collapses to a single term on $(0, \underline{p}_1(X))$ and we recover the standard LIV covariate bound.
In this section, we generalize our partial identification framework to accommodate continuous treatments. We allow the treatment domain $\mathcal{W}$ to be a compact subset of $\mathbb{R}$. We impose the following generalized structural assumptions.
As in the binary case, the common instrument support condition ensures that the conditional distribution $F_{W \mid Z=z, X=x}$ is well-defined for every $z \in \mathcal{Z}$ and every $x \in \mathcal{X}$, which is necessary for the structural equation ((ref)) to be identified uniformly across covariates. If $W$ is a binary treatment, (ref) reduces to the threshold-crossing model utilized in our previous assumptions after the harmless reparameterization $U' = 1-U$.\footnote{Under the generalized ordered-choice convention, $W = F^{-1}_{W\mid Z,X}(U\mid Z,X)$ is increasing in $U$, whereas our earlier binary threshold-crossing convention writes treatment as $W=\indicator(U \le p(Z,X))$. Setting $U' = 1-U$ aligns the two models.} If $W$ is multi-valued, this model is equivalent to the discrete ordered choice model heckman2007econometric. In the continuous treatment regime, this assumption provides the structural foundation for the control function approach in nonseparable models imbens2009identification, blundell2003endogeneity, chesher2003identification. We discuss this connection in detail in (ref).
In this model, $U$ can be interpreted as a latent rank that orders individuals' treatment intensity within each $(Z,X)$ cell. Under the ordered-choice convention in ((ref)), larger values of $U$ correspond to larger realizations of $W$.
Following the binary treatment case, we consider the following general target parameter. Let $\pi_{w,x}$ be the structural joint distribution of $(U, Y(w)) \mid X=x$, which can be seen as a probability kernel. For any identifiable weight function $\omega(w,x,u)$, we define:
where $\lambda(\cdot)$ denotes the Lebesgue measure for continuous treatments and the counting measure for discrete treatments. This directly generalizes $\theta_\omega$ from ((ref)). To see this, consider the binary treatment case $\mathcal{W} = \{0,1\}$, where $\int_{\mathcal{W}} \cdot \,\mathrm{d}w$ reduces to the sum over $\{0,1\}$. Setting $\omega(1,x,u) = \omega(x,u)$ and $\omega(0,x,u) = -\omega(x,u)$, the expression ((ref)) becomes:
recovering ((ref)).
As an example, consider a counterfactual policy that alters two components while holding the covariate distribution fixed: (i) the treatment assignment mechanism changes from $F_{W \mid Z,X}$ to $\tilde{F}_{W \mid Z,X}$, and (ii) the instrument distribution shifts from $F_{Z \mid X}$ to $\tilde{F}_{\tilde{Z} \mid X}$. Under this new policy, the counterfactual treatment is $\tilde{W} = \tilde{F}_{W \mid Z,X}^{-1}(U \mid \tilde{Z},X)$, and the realized outcome is $\tilde{Y} = Y(\tilde{W})$. The policy effect $\theta = \mathbb{E}[\tilde{Y} - Y]$ can be expressed as $\theta_\omega$ with the specific weight:
where the right-hand side denotes the $w$-section of the signed measure on $\mathcal{W} \times [0,1]$. The policy weight ((ref)) reduces to ((ref)) in the binary case.
The observed data constrains the structural kernel $\pi_{w,x}$. Note that the conditional distribution $\mathbb{P}_{\text{obs}}(\mathrm{d}y \mid W=w, Z=z, X=x)$ is only well-defined when $w$ belongs to the conditional support of $W$ given $(Z,X)$. We therefore define the effective instrument set:
and impose the observational constraints only for $z \in \mathcal{Z}(x,w)$. Specifically, for any $w \in \mathcal{W}$ and $x \in \mathcal{X}$, integrating the conditional distribution of the potential outcome over the latent variable $U$ must perfectly reconstruct the observed conditional outcome measure:
Crucially, because both the objective functional and the observational constraints are additively separable across the support of $X$ and $W$, the global optimization over all joint distributions decomposes into a family of independent CCOT sub-problems. Formally, this is equivalent to solving the following CCOT problem for $\lambda$-a.e.\ $w \in \mathcal{W}$ and $\mathbb{P}_{\mathrm{obs}}$-a.e.\ $x \in \mathcal{X}$:
Let $\overline{\theta}(x,w)$ and $\underline{\theta}(x,w)$ denote the optimal values of the maximization and minimization in ((ref)), respectively. The tight bounds for the global parameter $\theta_\omega$ are then obtained by aggregating over all $(x,w)$ strata:
To formally guarantee that ((ref)) yields sharp bounds for the continuous treatment setting, we must establish that any family of measures satisfying the generalized observational and marginal constraints corresponds to a valid data-generating process. Let $\Gamma_{\text{gen}}(\mathbb{P}_{\text{obs}})$ denote the set of all measure families $\{\pi_{w,x}\}_{w \in \mathcal{W}, x \in \mathcal{X}}$ that satisfy the constraints in the CCOT problem above. The following proposition verifies this observational equivalence, extending the sharpness guarantee from the binary framework, (ref).
This formal CCOT formulation highlights that variation in $F_{U \mid Z=z, X=x, W=w}$ across different instruments $Z$ imposes multiple marginal constraints on the structural distribution $\pi_{w,x}$. In general, the solution depends on the union of these supports, $$\mathcal{U}_{\text{id}}(x,w) \coloneqq \bigcup_{z \in \mathcal{Z}(x,w)} \text{supp}(F_{U \mid Z=z, X=x, W=w}).$$ Note that each $\text{supp}(F_{U \mid Z=z, X=x, W=w})$ is either a single point or a closed interval in $[0,1]$. Throughout this section, we assume that $\mathcal{U}_{\text{id}}(x,w)$ is measurable for each $(x,w)$. This measurability condition holds, for instance, if $\mathcal{Z}(x,w)$ is compact and the map $z \mapsto F_{W \mid Z,X}(w \mid z,x)$ is continuous, in which case $\mathcal{U}_{\text{id}}(x,w)$ is a compact subset of $[0,1]$.
In the following subsections, we demonstrate how to obtain closed-form solutions under specific structural settings. The general case is much more complicated and can be seen as a mix of these settings, which we will leave for future work.
imbens2009identification considers the following assumption for the IV model.
Under (ref), it is observationally equivalent to the conditional quantile representation in ((ref)). Crucially, the unobserved confounder $U$ can be perfectly inverted and recovered from the observables as the conditional rank of the treatment, $U = F_{W \mid Z,X}(W \mid Z,X)$. This recovered latent rank serves directly as a control variable: conditioning on $U$ alongside the covariates $X$ effectively absorbs the endogeneity, allowing us to isolate the exogenous variation in $W$. This formal equivalence connects our generalized partial identification framework to classical control function approach blundell2003endogeneity, chesher2003identification.
Because $U$ is perfectly recoverable, the conditional distribution $F_{U \mid Z=z, X=x, W=w}$ degenerates to a point mass at $u = F_{W \mid Z,X}(w \mid z, x)$. Consequently, the integral constraint in ((ref)) simplifies to:
where $z_u \in \mathcal{Z}(x,w)$ satisfies $F_{W \mid Z,X}(w \mid z_u, x) = u$. When multiple instrument values map to the same $u$, the observational constraints ensure that $\mathbb{P}_{\text{obs}}(\mathrm{d}y \mid W=w, Z=z_u, X=x)$ is identical for all such $z_u$, so the choice is immaterial. That is, the structural kernel $\pi_{w,x}$ is point-identified on the identifiable region $\mathcal{U}_{\text{id}}(x,w) = \bigcup_{z \in \mathcal{Z}(x,w)} \{F_{W\mid Z,X}(w \mid z,x)\}$. Therefore, if $u \in \mathcal{U}_{\text{id}}(x,w)$, we have:
where $z_u \in \mathcal{Z}$ is the specific baseline instrument realization that satisfies $w = F_{W \mid Z,X}^{-1}(u \mid z_u, x)$.
Outside this identifiable support, there is no constraint, and the optimal solution trivially assigns probability mass to the global outcome bounds to minimize or maximize the objective. Substituting this into our separated CCOT formulation yields the closed-form bounds.
Next, let us consider the multi-valued treatment setting, which is known as the ordered choice model heckman2007econometric. In this case, $\mathcal{W} $ is a finite set and $\text{supp}(F_{U \mid Z=z, X=x, W=w})$ is an interval for each $z$. Let us denote this interval as $I_{x,w}(z)$. Then, the first constraint of ((ref)) simplifies to:
In particular, if $|I_{x,w}(z)|=0$, the interval $I_{x,w}(z)$ degenerates to a single point $\{u_0\}$ and the constraint reduces to $\pi_{w,x}(\mathrm{d}y \mid U=u_0) = \mathbb{P}_{\text{obs}}(\mathrm{d}y \mid W=w, Z=z, X=x)$, i.e., the conditional distribution is point-identified at $u_0$. By convention, the left-hand side of ((ref)) is interpreted as $\pi_{w,x}(\mathrm{d}y \mid U=u_0)$ in this degenerate case.
The main difficulty of ((ref)) is that the intervals in $\{I_{x,w}(z)\}_{z \in \mathcal{Z}(x,w)}$ may overlap with each other, making it hard to disentangle the constraints like in the binary treatment setting. Surprisingly, we find that if $\{I_{x,w}(z)\}_{z \in \mathcal{Z}(x,w)}$ forms a $\pi$-system, closed-form solutions still exist.
In particular, if the $\{I_{x,w}(z)\}_{z \in \mathcal{Z}(x,w)}$ share the same start point or end point, they naturally satisfy this algebraic property, which is exactly the case in the binary treatment setting. More generally, (ref) endows $\{I_{x,w}(z)\}$ with a hierarchical Directed Acyclic Graph (DAG) ordering under strict inclusion---$z$ is an ancestor of $z'$ whenever $I_{x,w}(z') \subsetneq I_{x,w}(z)$---which, as we show below, is exactly the structure needed to disentangle the overlapping constraints.
For each $z \in \mathcal{Z}(x,w)$, let $\text{children}(z)$ and $\text{Dec}(z)$ denote its direct children and all descendants in this DAG. Define the disjoint isolated sub-region
and, recursively from the leaves upward, the isolated measure
The measure $\mu_{w,z,x}$ is the general analogue of $\mu_{1,i}$ from the binary case: it isolates the conditional distribution of $Y(w)$ on the disjoint slice $J_{x,w}(z)$. As in the binary setting, $\mu_{w,z,x}$ is a valid (nonnegative) probability measure under correct model specification, and its nonnegativity serves as a testable implication of the structural model. We also let $J_{x,w}(\emptyset) = [0,1] \setminus \bigcup_{z \in \mathcal{Z}(x,w)} I_{x,w}(z)$ denote the unconstrained domain.
In (ref), we show that under (ref) the original constraint system ((ref)) is equivalent to the {disentangled} marginal constraints
each acting on a single disjoint region $J_{x,w}(z)$. Because the $\{J_{x,w}(z)\}_{z \in \mathcal{Z}(x,w)}$ partition the identified support $\mathcal{U}_{\text{id}}(x,w)$ and the objective in ((ref)) is additively separable across these regions, the global problem decomposes into independent 1D optimal transport sub-problems with product cost---one per $J_{x,w}(z)$, pairing $\mu_{w,z,x}$ with $\mathrm{Unif}(J_{x,w}(z))$---plus a pointwise trivial bound on the unconstrained domain $J_{x,w}(\emptyset)$. Applying (ref) to each sub-problem yields the following closed-form result.
In this section, we develop estimation and inference results for the closed-form bounds derived in (ref). For the discrete instrument setting, we leverage DML chernozhukov2018double to accommodate high-dimensional covariates, constructing Neyman-orthogonal scores that yield $\sqrt{n}$-consistent and asymptotically normal estimators. In the continuous instrument setting, we characterize the corresponding nonparametric convergence rates.
We focus our exposition on the bounds for the PRTE, as the derivation for other causal quantities follows similarly. For simplicity, we first develop theory for the unscaled numerator $\mathbb{E}[Y^{q}-Y]$. In the empirical designs considered below, the PRTE denominator is the known policy-shift size $\alpha > 0$, so the reported bounds and confidence intervals for $\text{PRTE}_\alpha$ are obtained by dividing the corresponding numerator bounds by $\alpha$. More generally, when $\mathbb{E}[q(Z,X)-p(Z,X)]$ is not fixed by design, it is point-identified and can be estimated separately. Thus, our target weight function simplifies to:
We model the alternative policy as $q(z,x) = \phi(z,x,p(z,x))$ for a known function $\phi$. Common examples for $\phi$ include uniform propensity shifts $\phi(z,x,p) = p + \alpha$ or $p(1+\alpha)$ for a constant $\alpha$, or setting $\phi(z,x,p) = r(z,x)$ to evaluate a specific alternative targeting policy $r$.
For a fixed $x$, the weight function $\omega(x,u)$ is a step function that is constant and monotonic within each latent interval $u \in (p_i(x), p_{i+1}(x)]$. This piecewise structure allows us to decompose the closed-form lower bound integral into a finite weighted sum over the instrument level sets.
Define $q_{j,i}(x)$ as the $j$-th smallest value of the alternative policy propensity $q(z,x)$ that falls strictly inside the baseline interval $[p_i(x), p_{i+1}(x))$, such that:
with the convention $ q_{0,i}(x) = p_i(x)$. Let the baseline and alternative instrument level sets be:
Before presenting the rewriting of the target parameter, we impose two regularity conditions on the alternative policy and the propensity scores.
Part (i) will be used in the pathwise derivation of the orthogonal score and in the asymptotic analysis. Because $q(z,x) = \phi(z,x,p(z,x))$, part (ii) combined with (i) also implies that $q(z,\cdot)$ is continuous on $\mathcal{X}$ for each $z \in \mathcal{Z}$.
To ensure the level sets can be consistently estimated from the data without asymptotic ambiguity, we further impose the following gap assumption on the propensity scores.
Fix any two pairs $(z,\tau)$ and $(z',\tau')$ with $\tau, \tau' \in \{p,q\}$. The map $x \mapsto \tau(z,x) - \tau'(z',x)$ is continuous on $\mathcal{X}$ by (ref), while (ref) requires its value at every $x$ to be either exactly zero or of absolute value strictly greater than $c_{\text{gap}}$. A continuous function whose range avoids the punctured neighborhood $(-c_{\text{gap}}, c_{\text{gap}}) \setminus \{0\}$ cannot change sign, so the relative ordering of the elements of $R_x$ is invariant across $\mathcal{X}$. We therefore suppress the dependence on $x$ in $S_k$, $T_{j,k}$, $K$, and $l_k$ from now on.
The bound in (ref) admits a representation as an expectation of a finite weighted sum over instrument level sets and their sub-intervals. Before stating this representation, we introduce the nuisance components that carry the observable content of the bound.
\paragraph{Nuisance components.} For each interval index $k \in \{0,\dots,K\}$ and sub-interval index $j \in \{0,\dots,l_k\}$, define:
(ref) is the key analytical step that converts the optimal-transport integral in (ref) into a Neyman-orthogonalizable functional of standard nuisance quantities---conditional probabilities, conditional means, and conditional quantile integrals. The proof, given in (ref), exploits the piecewise-constant, monotone structure of $\omega(X,\cdot)$ on each baseline interval and identifies the step-function value attached to each sub-interval $(q_{j,k}(X), q_{j+1,k}(X))$.
The sub-interval term $J_{j,k}$ still involves the conditional quantile function $Q_{Y,k\mid X}$, which is delicate to estimate directly. The next proposition replaces it with conditional expectations in the two outcome regimes of interest.
Part (i) converts the sub-interval quantile integral into a plain conditional expectation of $Y$ on the quantile-defined band $(\nu_{j,k}, \nu_{j+1,k}]$, so only the $l_k+1$ boundary quantiles $\nu_{j,k}$---rather than the full quantile process---enter the estimation. Part (ii) leverages the step-function form of $Q_{Y,k\mid X}$ when $Y$ is binary: the complier mean $\theta_k(X) = \mathbb{E}_{\mu_{1,k\mid X}}[Y]$ reduces to a difference of observed joint probabilities, and the integral of the step function over $[\kappa_{j-1,k}, \kappa_{j,k}]$ evaluates to the closed-form $h_{j,k}^{-}$. Both parts are proved in (ref).
Together, (ref) expose (ref) as a functional of conditional expectations, boundary quantiles, and instrument-level-set probabilities---the form required for DML-style $\sqrt{n}$-inference. We construct the corresponding estimators and establish their asymptotic properties next, treating the two outcome regimes in turn.
We first treat the case where the conditional complier distribution $\mu_{1,k\mid X}$ is continuous at the boundary quantiles $\nu_{j,k}(X)$, so (ref)(i) applies. The representation (ref) converts each sub-interval term $J_{j,k}(X)$ into a conditional expectation over the band $(\nu_{j,k}(X), \nu_{j+1,k}(X)]$, thereby reducing the nuisance list from the full quantile process to the boundary quantiles $\nu_{j,k}$ alone. The Neyman-orthogonal score construction additionally requires the conditional indicator expectations
which serve as Riesz representers for the pathwise derivatives with respect to $\nu_{j,k}$.
Given the functional form in ((ref)) and ((ref)), our DML estimator takes the aggregated Neyman-orthogonal form
where each $\hat{\psi}$ is a Neyman-orthogonal score that corrects for the estimation error of its associated nuisance (explicit formulas are given in (ref)). The rest of this subsection develops the nuisance estimators $\hat{\eta}$ trained on an auxiliary sample $I_1$; (ref) below then establishes $\sqrt{n}$-consistency and asymptotic normality of ((ref)).
\paragraph{Estimation Procedure.} Following the standard DML template, we randomly partition the sample into an auxiliary set $I_1$ and a main estimation set $I_2$, train all nuisance estimators on $I_1$, and average plug-in orthogonal scores on $I_2$. (ref) summarizes the full procedure; the explicit moment equation for the conditional quantile $\nu_{j,k}$, the conditional-expectation components $M_{j,k}^{\pm}$, $J_{j,k}^{\pm}$, $J_{\mathrm{full},k}^{\pm}$, and the Neyman-orthogonal scores $\psi^{\mathrm{prod}}$ are deferred to (ref) and (ref).
To ensure the debiased terms remain well-behaved and that inverse probability weights do not explode, we require the following strict overlap assumption.
The following theorem formally establishes the asymptotic normality of our DML estimator.
When the outcome $Y$ takes finitely many values---in particular, when $Y \in \{0,1\}$ is binary---the conditional quantile function $Q_{Y,k\mid X}$ becomes a step function with atoms, and the continuity hypothesis of (ref)(i) fails. Instead, (ref)(ii) applies: the representation (ref) expresses $J_{\text{full},k}$ as a difference of observed joint probabilities $P_{1,k+1} - P_{1,k}$ and evaluates the sub-interval contribution in closed form through $h_{j,k}^{-}$, so the nuisance list collapses to conditional expectations alone and quantile estimation is avoided entirely. We present the binary case; the extension to general discrete outcomes is analogous.
Given the functional form in (ref), our DML estimator takes an aggregated Neyman-orthogonal form analogous to the continuous-outcome case:
where each $\hat{\psi}$ is a Neyman-orthogonal score (explicit formulas in (ref)). (ref) below establishes $\sqrt n$-consistency and asymptotic normality.
\paragraph{Estimation Procedure.} The procedure mirrors (ref), with the conditional-quantile and indicator-expectation steps removed. (ref) summarizes; full details and the construction of the orthogonal scores are in (ref) and (ref).
To ensure differentiability of the target functional with respect to the nuisance parameters, we impose the following gap condition, which replaces the conditional-density smoothness assumption of (ref).
This condition guarantees that the target is pathwise-differentiable in the nuisances despite the $\max$ operators appearing in $h_{j,k}^{-}$. The following theorem establishes the asymptotic normality of our DML estimator in the discrete-outcome setting.
We now proceed to estimation in the continuous instrument setting. For simplicity, we will assume that $\mathcal{Z}$ is connected throughout this subsection; the multiply connected component case can be derived similarly. When the instrumental variable $Z$ is continuous, the propensity score $p(Z,X)$ takes a continuous support, and the piecewise-constant structure of the weight function no longer holds. The general bound ((ref)) contains a derivative $\frac{\partial}{\partial u}\mathbb{E}_{\text{obs}}[YW \mid p(Z,X)=u, X]$ that is non-trivial to estimate directly. The following proposition leverages Fubini's theorem to collapse the derivative into a finite difference of conditional-expectation evaluations, yielding a representation amenable to plug-in ML estimation.
(ref) is the continuous-instrument analog of (ref): it converts the sharp bound into an expectation over $(Z,X)$ that depends on estimable nuisances only. The first term is a boundary-quantile tail integral capturing the optimal-coupling contribution on $(0, \underline{p}(X))$ through the conditional quantile function of the identified complier distribution $\mu_{1,\underline{p}\mid X}$. The second term replaces the derivative of $g_1$ by a finite difference of $g_1$-values at the truncated limits---this is the key simplification enabled by Fubini, since $g_1$ itself is a standard conditional expectation estimable by flexible ML regressions. The third term is the closed-form trivial-bound contribution on $(\overline{p}(X),1]$, where $\omega(X,u) \geqslant 0$ for the PRTE weight. Crucially, the entire target requires only conditional expectations and tail quantiles, avoiding derivative estimation entirely. The proof is given in (ref).
(ref) motivates the following plug-in estimator, obtained by replacing each nuisance in ((ref)) with its machine-learning estimate trained on an auxiliary sample:
where $\hat{\psi}_1, \hat{\psi}_2, \hat{\psi}_3$ correspond respectively to the boundary-quantile integral, the $g_1$ difference, and the trivial-bound term in ((ref)). (ref) below establishes its nonparametric convergence rate. The rest of this subsection develops the nuisance estimators $\hat{p}, \hat{g}_1$, and $\widehat{Q}_{Y,\underline{p}\mid X}$ that appear in ((ref)).
\paragraph{Estimation Procedure.} Because the continuous instrument setting requires nonparametric estimation of conditional expectations and quantiles, we employ a three-fold sample split: a propensity-estimation sample $I_0$, a nuisance-estimation sample $I_1$, and a main evaluation sample $I_2$. Splitting off $I_0$ separately is what allows the localized boundary-quantile subsample $\mathcal{I}_\delta = \{i\in I_1 : \hat p(Z_i,X_i) \leqslant \widehat{\underline p}(X_i) + \delta_n\}$ to remain conditionally i.i.d.\ given $I_0$. The conditional expectation $g_1(u,x)$ is estimated by a kernel-weighted local $M$-estimator in $u$ combined with a flexible ML fit in $x$, yielding a piecewise-linear-in-$u$ estimator that is deterministically Lipschitz in $u$---a property no joint ML regression on $(u, X)$ is known to deliver. This construction, related to local-polynomial DR-learning kennedy2023towards and generalized random forests athey2019generalized, decouples the covariate-dimension ML rate $r_{X,n}$ from the smoothing bias $h_n^2$ in $u$. (ref) summarizes; the explicit kernel $M$-estimator, grid quantile regression, and target evaluation are in (ref).
To establish the formal convergence rate of our estimator in the continuous instrument setting, we must account for the estimation error of the generated regressor $\hat{p}$, the kernel-smoothed conditional expectation $\hat{g}_1$, and the localized boundary quantile $\widehat{Q}_{Y,\underline{p} \mid X}$. We impose the following regularity condition on the data generating process.
Conditions 1--3 are standard regularity requirements: Condition 1 is satisfied by kernel, sieve, or $\ell_1$-penalized propensity estimators; Condition 2 ensures the piecewise-linear interpolation $\hat{g}_1$ incurs only an $h_n^2$ second-order bias; Condition 3 imposes standard kernel and bandwidth restrictions. Condition 4 is deliberately flexible: any ML procedure (random forests, neural networks, penalized linear methods) achieving kernel-weighted error rate $r_{X,n}$ is admissible. Conditions 5 and 6 ask only for conditional $W_1$-type rates---Condition 5 is strictly weaker than the uniform sup-norm rates established for standard conditional distribution and quantile regression estimators hall1999methods, guerre2012uniform, belloni2019conditional, while Condition 6 controls the bias from localizing to observations within $\delta_n$ of $\underline{p}(x)$.
The following theorem formally establishes the nonparametric convergence rate of the lower bound estimator.
We validate our theoretical results through synthetic experiments and a real-data empirical application. All experiments compare our method (hereafter IVOT) bounds against the moment-relaxation bounds of magne2018using (hereafter IVMTE), which represent the state-of-the-art alternative. Detailed numerical tables and additional results are deferred to (ref).
We consider two synthetic settings, one with a continuous instrument and one with a discrete instrument, each using a distinct data-generating process (DGP). In both cases the target is $\theta_\alpha := \mathbb{E}[Y^{q_\alpha} - Y]$ under the policy $q_\alpha(Z) = \operatorname{clip}(p(Z)+\alpha, 0, 1)$ for $\alpha \in [-0.12, 0.12]$. For IVMTE, we specify both MTR functions $m_0,m_1$ as degree-$9$ $u$-splines on $[0,1]$ with nine interior knots at $\{0.1,0.2,\ldots,0.9\}$; this is a deliberately flexible, near-nonparametric sieve, chosen so that IVMTE is not disadvantaged by an overly restrictive parametric choice. Full DGP specifications are given in (ref).
\paragraph{Continuous instrument.} The continuous-instrument DGP uses a logistic propensity $p(Z) = \operatorname{logistic}(-1+2Z)$ with limited support $p(Z) \in [0.27, 0.73]$ and a constant marginal treatment effect $\text{MTE}(u) = 0.5$, so the ground truth is $\theta_\alpha \approx 0.5\alpha$ for small $|\alpha|$. This setting deliberately isolates the identification difficulty: even though the MTE is constant, the limited propensity support prevents moment-relaxation methods from recovering tight bounds. (ref)(a) displays the identified sets at $n = 5{,}000$. IVOT yields tighter bounds than IVMTE; at $\alpha = 0.05$, the IVOT interval width is $0.011$ versus the IVMTE width of $0.096$, a roughly $8\times$ reduction. The IVMTE bounds are notably asymmetric---the upper bound extends substantially above the truth while the lower bound crosses zero---illustrating how discarding distributional information leads to spuriously wide identified sets. Across the 25 reported $\alpha$ grid points in this design, the true $\theta_\alpha$ falls inside the IVOT interval at 22 grid points. The three near-misses occur at $|\alpha| \leqslant 0.02$ where the IVOT interval is extremely tight (width ${\approx}\,0.001$) and finite-sample error in propensity estimation pushes the truth just outside the bounds. With $n = 10{,}000$, this pointwise inclusion check holds at all 25 grid points ((ref)), confirming this is a finite-sample phenomenon.
\paragraph{Discrete instrument.} The discrete-instrument DGP uses $Z \sim \text{Bernoulli}(0.5)$ with two-point propensity score $\{0.25, 0.75\}$ and a heterogeneous $\text{MTE}(u) = 0.48 + 0.18u$, designed to test the closed-form bounds ((ref)) when the propensity takes finitely many values. (ref)(b) shows the results at $n = 5{,}000$. IVOT again yields tighter bounds, and the true $\theta_\alpha$ lies inside the IVOT interval at all reported $\alpha$ grid points. At $\alpha = 0.05$, the IVOT interval is approximately $2.4\times$ narrower than IVMTE (width $0.041$ versus $0.097$), and the IVOT 95% delta-method confidence interval $[-0.007,\; 0.046]$ is tighter than the IVMTE 95% backward CI $[-0.050,\; 0.050]$. Because the propensity support consists of only two points, the MTE is identified only on $(0.25, 0.75)$; IVOT tightens bounds by exploiting the full distributional information in the identified region when bounding the contribution from the unidentified regions $(0, 0.25)$ and $(0.75, 1)$.
\paragraph{Discussion.} The disparity between the two methods is consistent with the theoretical mechanism identified in (ref): IVMTE reduces the observed conditional distributions to first-moment constraints, whereas IVOT retains the full quantile structure. The advantage is largest when the propensity score support is contained to $[0,1]$ (so that the unidentified region is large) and when the outcome distribution is informative about its quantile structure in the identified region. Numerical details for selected $\alpha$ values are tabulated in (ref).
\paragraph{Background and setup.} We apply our method to the bed net subsidy experiment of dupas2014short. A longstanding debate in development economics concerns the optimal pricing strategy for health-protective goods: while high prices screen out low-valuation users, subsidies expand take-up but may attract individuals who ultimately do not use the product. Understanding the causal effect of price reductions on product usage---not just on take-up---is therefore essential for welfare analysis. The dataset records insecticide-treated bed net (ITN) transactions and follow-up usage for $n = 1{,}078$ Kenyan households; the instrument $Z$ is the offered price (17 distinct values spanning 0--250 KSh, randomized across distribution points), the treatment $W \in \{0,1\}$ indicates purchase, and the outcome $Y \in \{0,1\}$ indicates one-year follow-up usage. The baseline policy is the reference price $z_0 = 150$ KSh. For each $\alpha \in [0.05, 0.62]$, we study the counterfactual policy target that raises the purchase propensity at that baseline price from $\hat p(z_0)$ to $q_\alpha := \min(\hat{p}(z_0) + \alpha, 1)$; thus the denominator of $\text{PRTE}_\alpha = \mathbb{E}[Y^{q_\alpha} - Y]/\alpha$ is the known shift size $\alpha$. The maximum $\alpha_{\max}\approx 0.621$ equals the propensity at zero price minus the baseline propensity. Because the instrument is discrete, the closed-form bounds of (ref) apply directly to the numerator, after which we divide by $\alpha$. For IVMTE, we report two sieve sizes (degree-$10$ and degree-$20$ $u$-splines on $[0,1]$) so that the shape-restriction trade-off is explicit: degree $20$ enlarges the MTR class and imposes weaker shape restrictions, so its identified set should widen relative to degree $10$. Full data, propensity-estimation, and policy-construction details are in (ref).
\paragraph{Results.} (ref) plots the IVOT and IVMTE identified sets for $\text{PRTE}_\alpha$ across subsidy levels. IVOT consistently yields a tighter identified set than IVMTE across the reported range of $\alpha$ and under both spline specifications. At the smallest subsidy ($\alpha = 0.05$), the IVOT interval is $[0.740, 1.000]$ (width $0.260$), while the IVMTE interval is $[0.129, 0.990]$ (width $0.861$) at degree $10$ and widens to $[-0.372, 1.002]$ (width $1.374$) at degree $20$---roughly a $3.3\times$ and $5.3\times$ reduction in width relative to IVOT, respectively. The IVMTE bound visibly widens as the spline degree increases from $10$ to $20$, empirically confirming that the higher-degree sieve is more nonparametric and imposes weaker shape restrictions on the MTR. IVOT's advantage is therefore robust across these two IVMTE specifications.
Both methods indicate a positive policy effect across most of the subsidy range---the IVOT lower bound is positive throughout, and the degree-$10$ IVMTE lower bound is positive for all $\alpha$, with the more conservative degree-$20$ lower bound turning positive once $\alpha \geqslant 0.14$. Moreover, IVOT achieves near-point identification (width $< 0.05$) for the majority of $\alpha$ values, exhibiting a periodic pattern in which the bounds tighten to near-zero width before widening slightly as additional complier groups enter the policy margin: for instance, the identified set is essentially a point for $\alpha \in \{0.08, 0.09, 0.11\}$ and again for $\alpha \in \{0.23, 0.24, 0.26\}$. At the maximum feasible subsidy ($\alpha \approx 0.621$), IVOT point-identifies $\text{PRTE}_\alpha \approx 0.597$. Full numerical results, including IVOT 95% delta-method confidence intervals, are reported in (ref).
We established a novel connection between the partial identification of PRTEs in IV models and OT. By formulating the problem as a CCOT problem over joint distributions compatible with the observed data, we showed that the multidimensional optimization analytically reduces to one-dimensional OT problems with product costs, yielding explicit closed-form expressions for the sharp bounds and bypassing the computationally intensive optimization required by existing approaches. The framework extends to general treatment settings (continuous and multi-valued), and we develop corresponding estimation results: Neyman-orthogonal DML scores deliver $\sqrt{n}$-consistency and asymptotic normality for discrete instruments, while for continuous instruments we characterize nonparametric convergence rates. A central insight, demonstrated theoretically and empirically, is that preserving full distributional information is both tractable and valuable for tight partial identification: moment-relaxation approaches that reduce the observed data to conditional expectations discard higher-order distributional features and can produce substantially wider bounds, whereas the CCOT formulation enforces compatibility with the entire observed distribution and makes the underlying optimization analytically transparent. We believe this insight extends beyond the IV setting studied here.
\paragraph{Acknowledgements. } Jose Blanchet gratefully acknowledges support from DoD through the grants Air Force Office of Scientific Research under award number FA9550-20-1-0397 and ONR 1398311, also support from NSF via grants 2229012, 2312204, 2403007 is gratefully acknowledged. Vasilis Syrgkanis gratefully acknowledges support from NSF Award IIS-2337916, an Amazon Research Award and a Google Research Award.