Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
86,597 characters · 37 sections · 42 citation commands
Conformal Prediction for Nonparametric Instrumental Regression
Instrumental variables (IVs) are widely used for causal and policy evaluation when regressors are endogenous. We consider structural prediction problems characterized by moment conditions conditional on IVs, for which a representative example is nonparametric IV regression Newey2003instrumentalvariable,Ai2003efficientestimation,Darolles2011nonparametricinstrumental. A large literature studies how to estimate $h_0$ under ill-posedness, using series, minimum-distance, reproducing kernel Hilbert space (RKHS), neural network, and minimax procedures Newey2003instrumentalvariable,Hartford2017deepiv,Singh2019kernelinstrumental,Dikkala2020minimaxestimation.
Here, we briefly introduce the NPIV setup. For details, see Section (ref). In NPIV, we typically consider the following data-generating process:
where $X$ is an endogenous variable, $\varepsilon$ is the error term, $Z$ is an IV, and $h_0(\cdot)$ is called the structural function. While the structural function $h_0(\cdot)$ relates the endogenous variable $X$ to the outcome $Y$, it cannot be estimated by regressing $Y$ on $X$ because of endogeneity, that is, the correlation between $X$ and $\varepsilon$.
In this paper, we study prediction intervals for future outcomes in NPIV. Existing inference methods in NPIV typically target the structural function $h_0$ or its functionals through asymptotic approximations to regularized estimators. Our goal is different. We aim to construct prediction intervals for $Y$ itself, with validity formulated relative to the IV. For example, consider predicting future demand for a good. In conventional approaches, one first estimates the demand curve $h_0$ and then predicts demand $Y$. In that case, one constructs confidence or prediction intervals for $h_0$. In this study, by contrast, we target $Y$ itself, that is, we construct prediction intervals for $Y$, not $h_0$. Our prediction intervals for $Y$ are justified by the IV $Z$.
Conformal prediction provides a natural starting point for this purpose, since it yields distribution-free finite-sample coverage under exchangeability Vovk2005algorithmiclearning. While conformal prediction is a convenient approach in many tasks, it is not immediately applicable to NPIV, because exact conditional coverage is impossible Lei2013distributionfree,Barber2020thelimits. That is, exact conditional coverage given the IV is impossible in a fully distribution-free, finite-sample sense. To address this issue, we build on the results in Tibshirani2019conformalprediction and Gibbs2025conformalprediction, where the former considers conformal prediction under covariate shift, and the latter considers conditional guarantees in conformal prediction.
A basic modeling question then arises. In IV problems, what should the final interval depend on? The structural regression itself is a function of $X$ alone, which suggests an $X$-indexed prediction interval $C(X)$ as the most natural predictive object. At the same time, one may also allow the radius to vary jointly with $(X,Z)$ or only with $Z$. These three possibilities have different interpretations and different algorithmic properties. We study conformal procedures for all three cases.
Within that common target, the three choices of radius lead to different methods. When the radius depends jointly on $(X,Z)$, the finite-dimensional conditional conformal machinery of Gibbs2025conformalprediction applies directly with the covariate $(X,Z)$, yielding exact finite-sample coverage over joint shifts in $(X,Z)$. When the radius depends only on $Z$, the same machinery specializes to an exact IV-specific procedure, which we call IV-conditional conformal prediction, with coverage over IV shifts. When the radius depends only on $X$, exact finite-sample family-wise calibration is not currently available. To handle that case, we use the importance-weighting logic of Kato2022learningcausal to convert the conditional coverage moment into weighted unconditional moments, and then combine that idea with weighted conformal recalibration under a fixed target shift in the spirit of Tibshirani2019conformalprediction.
We propose methods for prediction intervals in NPIV with distribution-free finite-sample coverage guarantees. In particular, we refer to the IV-specific conformal construction as IV-CCP. IV-CCP is based on conformal prediction, in particular, on conditional guarantees through conditional calibration.
The methodological contribution is to transform the NPIV problem with conditional moment restrictions into one under marginal moment restrictions so that we can apply conformal prediction. We recast exact conditional coverage as a moment condition against all measurable reweightings of the IV and relax that impossible target to robust coverage over a user-specified class of IV shifts ${\mathcal{F}}$. This approach is closely related to covariate shift Shimodaira2000improvingpredictive.
In implementation, we use an estimate of $h_0(x)$ to tighten the prediction intervals. As a basic construction, we first estimate the structural function $h_0(x)$ and then construct prediction intervals by adding a radius to the estimator $\widehat h$. We study three types of radii. The first depends on both $X$ and $Z$, the second depends only on $Z$, and the third depends only on $X$. For each approach, we present a concrete implementation and theoretical guarantees.
Our work lies at the intersection of NPIV estimation and conformal prediction. Estimation of causal parameters under conditional moment restrictions is a core problem in causal inference, and NPIV is a representative example. In this problem, the technical difficulties come from identification and stable estimation, including ill-posedness and regularization for $h_0$ Newey2003instrumentalvariable,Ai2003efficientestimation,Darolles2011nonparametricinstrumental. In the machine learning literature, several studies use flexible models for the structural function based on neural networks and RKHS, or employ a minimax approach to estimation Hartford2017deepiv,Singh2019kernelinstrumental,Dikkala2020minimaxestimation. Our method is agnostic to the particular NPIV models and estimation methods and therefore acts as a predictive wrapper around these estimators rather than a replacement for them.
On the conformal side, the starting point is distribution-free finite-sample validity under exchangeability Vovk2005algorithmiclearning. Exact distribution-free conditional coverage is impossible in general Lei2013distributionfree,Barber2020thelimits, so recent work has studied structured relaxations, including weighted conformal prediction under a fixed target shift Tibshirani2019conformalprediction, localized conformal prediction Guan2022localizedconformal, distributional conformal prediction Chernozhukov2021distributionalconformal, and conditional guarantees over a family of shifts Gibbs2025conformalprediction. The present paper specializes these ideas to IV settings and combines the conformal perspective with the importance-weighted conditional-moment formulation of Kato2022learningcausal. Appendix (ref) provides a fuller discussion.
In this section, we formulate our problem.
\paragraph{Variables} Let $Y \in {\mathcal{Y}} \subseteq {\mathbb{R}}$ be a scalar outcome, $X \in {\mathcal{X}} \subseteq {\mathbb{R}}^{k_X}$ be a $k_X$-dimensional endogenous variable, and $Z \in {\mathcal{Z}} \subseteq {\mathbb{R}}^{k_Z}$ be $k_Z$-dimensional IVs, where ${\mathcal{Y}}$, ${\mathcal{X}}$, and ${\mathcal{Z}}$ denote the outcome, covariate, and IV spaces, respectively. Assume that $(Y, X, Z)$ jointly follows a distribution $P$.
\paragraph{Observations} Let $n \in {\mathbb{N}}$ be the sample size. Assume that we observe a dataset $\left\{(Y_i,X_i,Z_i)\right\}^n_{i=1}$, where each $(Y_i, X_i, Z_i)$ is an i.i.d. copy of $(Y, X, Z)$ following the distribution $P$.
This study considers prediction of $Y$ under endogeneity of $X$. We clarify the meaning of endogeneity below.
\paragraph{Recap of identification through conditional moment restrictions} The conventional approach defines the structural function through identification under conditional moment restrictions Ai2003efficientestimation,Darolles2011nonparametricinstrumental. In this approach, we define the causal relationship between $Y$ and $X$ as
where $h_0:{\mathcal{X}}\to{\mathcal{Y}}$ is the structural function and $\varepsilon$ is a sub-Gaussian error term with mean zero. To estimate $h_0$, suppose that the IV $Z$ satisfies the following conditional moment restriction:
Then we assume that $h_0$ is uniquely identified under ((ref)). The goal of the conventional setup is to estimate $h_0$ from the conditional moment restriction in ((ref)). If $Z = X$, this problem reduces to estimation of the regression function ${\mathbb{E}}\left[Y \mid X\right]$. However, when ${\mathbb{E}}\left[\varepsilon\mid X\right] \neq 0$, ${\mathbb{E}}\left[Y\mid X\right]$ is not equal to $h_0(X)$, so standard regression methods such as least squares may fail to recover the structural function. In such cases, we use NPIV methods to estimate $h_0$ using IVs $Z$.
\paragraph{Reformulation by conditional prediction given IVs} The conventional approach relies on identification of the structural function itself. In this study, we do not focus on identification. Instead, we consider prediction of $Y$ using $X$, with validity indexed by the IV $Z$. In this approach, we define the target as a prediction interval \[\widehat C\colon {\mathcal{X}}\times{\mathcal{Z}}\to \left\{\text{intervals in }{\mathbb{R}}\right\},\] which satisfies
We first define a prediction interval $\widehat C$ that depends on both $X$ and $Z$. We then introduce prediction intervals that depend only on $Z$ or only on $X$. Our main focus is the prediction interval that depends only on $X$. Note that in our proposed method, we also use an estimate of $h_0$ to tighten the prediction intervals, but this is not necessary.
In summary, our ideal goal is to predict $Y$ given $X$ while accounting for endogeneity and to obtain a prediction interval $\widehat C\colon {\mathcal{X}}\times{\mathcal{Z}}\to \left\{\text{intervals in }{\mathbb{R}}\right\}$ satisfying the conditional coverage guarantee ((ref)).
For the construction of $\widehat C(X_{n+1},Z_{n+1})$, we allow the use of an estimator $\widehat h\colon {\mathcal{X}}\to{\mathcal{Y}}$ of the structural function $h_0$. This reflects the common application setting in which $X$ is the argument of the prediction rule and $Z$ is used to calibrate validity, not necessarily to serve as a direct input to the point predictor.
In this study, we aim to guarantee distribution-free and finite-sample coverage for the prediction interval. However, distribution-free procedures cannot satisfy ((ref)) with nontrivial finite-length intervals in general Lei2013distributionfree,Barber2020thelimits. We therefore pursue a relaxation in the next section.
Let ${\mathbb{P}}$ denote the probability law induced by $P$, and let ${\mathbb{E}}$ denote expectation under ${\mathbb{P}}$. We refer to the environment in which we predict $Y_{n+1}$ as the test environment, and to the corresponding probability law as the test law. Let $\mathrm{len}(C)$ be the length of a prediction interval $C \subseteq {\mathbb{R}}$.
This section defines the target prediction intervals. We begin with prediction intervals with exact conditional coverage, rewrite this target in moment form, and then relax it to coverage over a class of IV shifts. The resulting target applies to all radius classes considered in the paper.
For a generic interval $\widehat C_\tau$, exact conditional coverage in the IV setting takes the form
The next proposition rewrites ((ref)) in moment form.
Proposition (ref) shows that exact conditional coverage is equivalent to marginal coverage weighted by every measurable function of the IV. This equivalence also clarifies the impossibility result. Exact conditional coverage asks for validity after every possible reweighting of $Z$, which is too demanding in a distribution-free finite-sample setting. Lei2013distributionfree shows that conditional guarantees of this type cannot generally be achieved by nontrivial finite-length prediction sets, and Barber2020thelimits sharpens this message by showing that even approximate conditional coverage requires intervals that are essentially as conservative as those built for much larger subsets. The conditioning variable is the IV rather than a generic predictor, but the impossibility itself is unchanged.
We therefore replace the full measurable class ${\mathcal{M}}$ in ((ref)) with a user-specified class ${\mathcal{F}}\subset {\mathcal{M}}$. This yields the relaxed requirement
Following Gibbs2025conformalprediction, ((ref)) can be interpreted as robust marginal coverage over a family of IV shifts. For any nonnegative $f\in{\mathcal{F}}$ with ${\mathbb{E}}\left[f(Z)\right]>0$, define a tilted distribution
Then ((ref)) implies \[ {\mathbb{P}}_f\left(Y_{n+1}\in \widehat C_\tau(X_{n+1},Z_{n+1})\right)\ge 1-\alpha \] for every such $f$. In applications with IVs, this interpretation is natural. Policies often change the marginal distribution of the IV while leaving the structural relationship ((ref)) intact. The one-sided inequality in ((ref)) is enough for lower coverage guarantees, which is the finite-sample target throughout the paper.
In constructing prediction intervals $\widehat C_\tau(x,z)$, we allow the use of an estimator $\widehat h$ of the structural function $h_0$. Let $\widehat h\colon {\mathcal{X}}\to{\mathcal{Y}}$ denote a predictor of $h_0$ obtained from some NPIV method. Then we define the general family of intervals as
where $\widehat\tau\colon {\mathcal{X}}\times{\mathcal{Z}}\to{\mathbb{R}}_+$ is a nonnegative radius function.
Based on their dependence on the variables, we distinguish the following three radius classes:
The broad class ${\mathcal{T}}_{XZ}$ lets the radius vary jointly with $(X,Z)$ and serves as the common parent class for the later specializations. The IV-indexed class ${\mathcal{T}}_Z$ keeps the center equal to $\widehat h(X)$ while allowing the radius to adapt to the IV. The $X$-indexed class ${\mathcal{T}}_X$ yields a prediction interval of the form \[ \widehat C_X(x)=\left[\widehat h(x)-\widehat\tau_X(x),\widehat h(x)+\widehat\tau_X(x)\right]. \] The common special case ${\mathcal{T}}_X\cap{\mathcal{T}}_Z$ is the constant-radius split conformal rule.
The distinction among the three classes is substantive. The class ${\mathcal{T}}_{XZ}$ is the most expressive benchmark. The class ${\mathcal{T}}_Z$ is IV-specific, because the IV indexes predictive uncertainty while the center remains a function of $X$ alone. The class ${\mathcal{T}}_X$ yields prediction intervals $C(X)$ that depend only on $X$, which is the most natural class from the NPIV motivation. This paper studies all three classes, with the main methodological focus on ${\mathcal{T}}_X$.
The relaxation ((ref)) is common across the three radius classes, but their feasible sets differ. Let ${\mathcal{T}}\subset {\mathcal{T}}_{XZ}$ be any radius class, and define the oracle problem
Let $L({\mathcal{T}},{\mathcal{F}})$ denote the infimum in ((ref)) under the constraint ((ref)), that is, $L({\mathcal{T}},{\mathcal{F}})=\inf\left\{{\mathbb{E}}\left[\tau(X,Z)\right]:\tau\in{\mathcal{T}} \text{ satisfies }(\ref{eq:oracle_radius_constraint})\right\}$. Here $L({\mathcal{T}},{\mathcal{F}})$ is the oracle expected half-length within the radius class ${\mathcal{T}}$ under the shift-robust target generated by ${\mathcal{F}}$.
Proposition (ref) clarifies the trade-off. The class ${\mathcal{T}}_{XZ}$ is the broadest feasible class and therefore the easiest one in which to reduce oracle length. The classes ${\mathcal{T}}_Z$ and ${\mathcal{T}}_X$ are more restrictive, but they have sharper substantive interpretations. The class ${\mathcal{T}}_Z$ treats the IV as an index of residual uncertainty. The class ${\mathcal{T}}_X$ yields prediction intervals $C(X)$ depending only on $X$, which is the most natural class from the NPIV motivation.
The constrained problem defined by ((ref)) and ((ref)) has the same minimax flavor as conditional-moment formulations in NPIV. In the continuum-of-moments and linear-inverse viewpoints of Carrasco2000generalizationof and Carrasco2007linearinverse, conditional moments are converted into a family of unconditional moments indexed by a function class. More recent adversarial procedures, such as DeepGMM and the minimax-optimal estimators of Bennett2019deepgeneralized and Dikkala2020minimaxestimation, estimate $h_0$ by searching for moment violations over a critic class.
Our use of the class ${\mathcal{F}}$ is different. We are not trying to identify $h_0$ by finding moments that are most informative about the structural equation. Instead, $h_0$ or its estimator $\widehat h$ is treated as fixed, and the radius $\tau$ is chosen so that the coverage residual $\mathbbm{1}\left[Y\in \widehat C_\tau(X,Z)\right]-(1-\alpha)$ has nonnegative moments against every $f\in{\mathcal{F}}$. Thus, ${\mathcal{F}}$ is not a critic class for structural estimation.
This viewpoint also connects naturally to importance weighting. Equation ((ref)) rewrites the relaxed coverage target as a family of ordinary marginal coverage statements under weighted laws. This is the same algebraic mechanism that underlies covariate-shift correction in Shimodaira2000improvingpredictive and importance-weighting formulations for NPIV such as Kato2022learningcausal. In our setting, however, the weights define the probability laws over which coverage is required.
In this section, we propose IV-conditional conformal prediction (IV-CCP). The details differ across the three target radius classes, ${\mathcal{T}}_{XZ}$, ${\mathcal{T}}_Z$, and ${\mathcal{T}}_X$.
In all cases, we construct prediction intervals using a split-type implementation of conformal prediction. We first split the sample into a training set $\mathcal{I}_{\mathrm{tr}}$ and a calibration set $\mathcal{I}_{\mathrm{cal}}$ with $|\mathcal{I}_{\mathrm{cal}}|=m$. We then fit any NPIV estimator $\widehat h$ on $\mathcal{I}_{\mathrm{tr}}$. On the calibration set, we compute the following nonconformity scores:
Conditional on the training split, $\widehat h$ is fixed, so the calibration scores and the test score are exchangeable.
The first method studies the broad benchmark class ${\mathcal{T}}_{XZ}$. This class is not the main target of the paper, but it is mathematically convenient because the same variable $(X,Z)$ indexes both the radius and the shift family.
\paragraph{Class of the joint shift.} Let $W=(X,Z)$ and let $\phi_W\colon {\mathcal{X}}\times{\mathcal{Z}}\to{\mathbb{R}}^d$ be a feature map. Define \[ {\mathcal{G}}_W=\left\{w\mapsto \beta^\top\phi_W(w): \beta\in{\mathbb{R}}^d\right\},\qquad {\mathcal{F}}_W={\mathcal{G}}_W\cap\left\{f\ge 0\right\}. \] Throughout this section, we take the first component of $\phi_W$ to be the constant $1$, so ${\mathcal{G}}_W$ contains the constant functions. The class ${\mathcal{G}}_W$ determines how the radius may vary with $(X,Z)$, and ${\mathcal{F}}_W$ determines the family of joint distribution shifts.
\paragraph{Calibration.} Fix a test pair $w=(x,z)$ and a candidate radius $s\ge 0$. Denote the pinball loss by \[\rho_\alpha(u)=\alpha\max\left\{u,0\right\}+(1-\alpha)\max\left\{-u,0\right\}.\] Using the calibration pairs $(W_i,S_i)$ and the augmented point $(w,s)$, consider
This is the finite-dimensional conditional conformal machinery of Gibbs2025conformalprediction with covariate $W=(X,Z)$. Form the LP dual corresponding to this augmented quantile regression problem, and let $\eta_{m+1}(s)$ denote the dual coordinate associated with the augmented point. As in Gibbs2025conformalprediction, $\eta_{m+1}(s)$ is nondecreasing in $s$, and $s$ belongs to the conformal upper set if and only if $\eta_{m+1}(s)<1-\alpha$. We therefore define $\widehat\tau_{XZ}(x,z)$ as the largest accepted value of $s$ and output
For any nonnegative $f\in{\mathcal{F}}_W$ with ${\mathbb{E}}\left[f(W)\right]>0$, define the joint tilted law
The finite-sample guarantee for this class is stated in Theorem (ref).
The second method studies the IV-indexed class ${\mathcal{T}}_Z$. This is the exact IV-specific conformal class, and it is the first main focus of the paper. The interval center remains $\widehat h(X)$, while the radius depends on $Z$ alone.
\paragraph{Class of the IV shift.} Let $\phi\colon {\mathcal{Z}}\to{\mathbb{R}}^d$ be a feature map, and define \[ {\mathcal{G}}=\left\{z'\mapsto \beta^\top\phi(z'): \beta\in{\mathbb{R}}^d\right\},\qquad {\mathcal{F}}={\mathcal{G}}\cap\left\{f\ge 0\right\}. \] Throughout the sequel, we take the first component of $\phi$ to be the constant $1$, so ${\mathcal{G}}$ contains the constant functions. Fix a test IV $z$ and a candidate radius $s\ge 0$.
\paragraph{Calibration.} Using the same scores ((ref)), consider
This is the augmented quantile-regression problem of Gibbs2025conformalprediction with covariate $Z$. Let $\eta_{m+1}(s)$ denote the dual variable associated with the augmented point. As before, $\eta_{m+1}(s)$ is nondecreasing in $s$, and the largest accepted value of $s$ defines the radius $\widehat\tau_Z(z)$. The resulting IV-indexed interval is
\paragraph{Interpretation} The class ${\mathcal{T}}_Z$ is less flexible than ${\mathcal{T}}_{XZ}$, but it restores the IV-specific interpretation that the relevant change occurs in the IV distribution. The interval center remains a function of $X$ alone. What changes with $Z$ is the residual uncertainty around that center. This is compatible with the exclusion restriction in ((ref)), because the IV does not enter the structural center directly.
\paragraph{Relationship to weighted conformal prediction.} Weighted conformal prediction under covariate shift Tibshirani2019conformalprediction protects against one fixed target distribution, provided the corresponding density ratio is known or can be estimated. IV-CCP is different. It protects simultaneously against every IV tilt in the class ${\mathcal{F}}$. In other words, it is a guarantee over a family of shifts rather than a guarantee for a single shift. This distinction is central in IV applications, where the relevant uncertainty is often not one fully specified probability law in a setting of interest but a set of plausible changes in the IV distribution.
The third method studies the class ${\mathcal{T}}_X$, which produces prediction intervals $C(X)$ depending only on $X$. This is the most natural target in NPIV, because the final prediction rule depends on the structural regressor alone. It is also the second main focus of the paper. The difficulty is that the interval is indexed by $X$, while the robustness target is indexed by shifts in $Z$.
\paragraph{IV-CCP for the $X$-indexed interval} In this case, we consider a prediction interval of the following form:
Let us define the nonconformity score as \[ S=\left|Y-\widehat h(X)\right|. \] For a candidate radius function $\tau_X\colon {\mathcal{X}}\to{\mathbb{R}}_+$, define the coverage residual
Then exact conditional coverage for $\widehat C_X$ is equivalent to
The relaxed family-of-shifts target becomes
The challenge is to construct an $X$-indexed radius while the protected shift family remains a class of functions of $Z$.
\paragraph{Importance-weighted representation} The importance-weighting approach of Kato2022learningcausal provides a natural route. Suppose the conditional density ratio
is well defined. Then conditional moments of $(S,X)$ given $Z$ can be rewritten as weighted unconditional moments.
Proposition (ref) is the key bridge from the conditional coverage target to unconditional weighted moments. It is exactly the same algebraic step that Kato2022learningcausal uses for general conditional moment learning, now applied to the interval residual $m_{\tau_X}(S,X)$.
To turn this identity into a learning criterion, let ${\mathcal{R}}_X\subset \left\{\tau_X\colon {\mathcal{X}}\to{\mathbb{R}}_+\right\}$ denote a chosen model class for the radius function, for example a sieve span, a neural network class, or a shape-restricted class. Let $\widehat r(s,x\mid z)$ be an estimated version of ((ref)), obtained on the training split or by cross-fitting. Let $\widetilde Z_1,\dots,\widetilde Z_M$ be evaluation points, for example the calibration IVs themselves or a deterministic grid. Because the indicator in ((ref)) is discontinuous, it is convenient in optimization to replace it with a smooth surrogate. Let $\psi_\kappa\colon {\mathbb{R}}\to{\mathbb{R}}$ be a smooth approximation to $u\mapsto \mathbbm{1}\left[u\ge 0\right]-(1-\alpha)$, indexed by a smoothing parameter $\kappa>0$.
We then consider the criterion
and the estimator
where $\lambda>0$ balances average interval length and weighted moment fit. The resulting prediction interval is ((ref)). This procedure is not distribution-free in finite samples, because it relies on density-ratio estimation and surrogate optimization. Its role is different. It provides a principled way to construct an $X$-indexed radius under the same conditional-moment logic that underlies NPIV.
\paragraph{Single-shift weighted conformal recalibration} The importance-weighted procedure above learns the shape of the $X$-indexed radius. To obtain a finite-sample guarantee for a fixed target distribution shift while keeping the final interval $X$-indexed, one can add a final scalar conformal recalibration step in the spirit of Tibshirani2019conformalprediction. Let $q\colon {\mathcal{X}}\to{\mathbb{R}}_+$ be a positive function learned on a sample independent of the final recalibration split, for example by ((ref)), or by any other method. Let $\mathcal I_{\mathrm{rcal}}$ denote the final recalibration split. Consider intervals of the form
where $f_0\in{\mathcal{F}}$ is a fixed target shift. Form the normalized scores as
The weights are the density ratios induced by the target shift, which are defined as
Assume furthermore that there is a known finite constant $B_{f_0}$ such that \[ w_{f_0}(z)\le B_{f_0}\qquad \text{for all $z$ in the support of $Z$.} \] Define $\widehat t_{f_0}$ as the conservative weighted split conformal cutoff obtained from the calibration weights $w_{f_0}(Z_i)$ and the worst-case test weight $B_{f_0}$, that is, \[ \widehat t_{f_0}=\inf\left\{t\in{\mathbb{R}}_+: \sum_{i\in\mathcal I_{\mathrm{rcal}}}\frac{w_{f_0}(Z_i)}{\sum_{j\in\mathcal I_{\mathrm{rcal}}}w_{f_0}(Z_j)+B_{f_0}}\mathbbm{1}\left[R_i\le t\right]\ge 1-\alpha\right\}, \] with the convention $\inf\emptyset=\infty$. By construction, the cutoff is scalar and the final interval remains a function of $X$ only. Because it upper bounds the test-point-specific weighted split conformal cutoff of Tibshirani2019conformalprediction, it yields the following finite-sample guarantee under the single target law ${\mathbb{P}}_{f_0}$: \[ {\mathbb{P}}_{f_0}\left(Y_{n+1}\in \widehat C_{X,f_0}(X_{n+1})\right)\ge 1-\alpha, \] where the detailed statement is shown in Theorem (ref).
This coverage guarantee is intentionally narrower than the guarantees for ${\mathcal{T}}_{XZ}$ and ${\mathcal{T}}_Z$. It holds for one specified distribution shift rather than uniformly over a whole class. This reflects the current methodological gap. Family-wise exact finite-sample coverage for the mixed-index target $C(X)$ is not presently available, whereas a conservative single-shift weighted conformal recalibration is available.
The three classes now fall into place. The benchmark class ${\mathcal{T}}_{XZ}$ is the broad exact conformal class. The class ${\mathcal{T}}_Z$ is the exact IV-specific class that delivers simultaneous finite-sample coverage over a family of IV tilts. The class ${\mathcal{T}}_X$ is the primary target of interest, but it currently requires a different strategy, namely, importance-weighted conditional-moment learning and, when needed, a final single-shift weighted conformal recalibration.
The choice of ${\mathcal{F}}$ is the main modeling decision in the paper. Through ((ref)), it determines the ambiguity set of shifted distributions under which coverage is required. Accordingly, ${\mathcal{F}}$ should encode how the marginal distribution of the IV may change, not how $Y$ depends on $Z$, and not how $\widehat h$ is estimated.
For the exact classes ${\mathcal{T}}_{XZ}$ and ${\mathcal{T}}_Z$, the shift class is induced by a finite-dimensional feature map. For ${\mathcal{T}}_Z$, with feature map $\phi\colon {\mathcal{Z}}\to{\mathbb{R}}^d$, define \[ {\mathcal{G}}_\phi=\left\{z\mapsto \beta^\top\phi(z):\beta\in{\mathbb{R}}^d\right\},\qquad {\mathcal{F}}_\phi={\mathcal{G}}_\phi\cap\left\{f\ge 0\right\}. \] The corresponding ambiguity set is \[ {\mathcal{Q}}({\mathcal{F}}_\phi)=\left\{{\mathbb{P}}_f: \frac{d{\mathbb{P}}_f(y,x,z)}{d{\mathbb{P}}(y,x,z)}=\frac{f(z)}{{\mathbb{E}}\left[f(Z)\right]},\ f\in{\mathcal{F}}_\phi,\ {\mathbb{E}}\left[f(Z)\right]>0\right\}. \] If ${\mathcal{F}}_1\subset {\mathcal{F}}_2$, then ${\mathcal{Q}}({\mathcal{F}}_1)\subset {\mathcal{Q}}({\mathcal{F}}_2)$. Enlarging ${\mathcal{F}}$ therefore protects against a broader family of shifted distributions. If the constant function belongs to ${\mathcal{F}}$, then the training distribution itself belongs to the ambiguity set. In practice, this is enforced by taking the first component of $\phi$ to be $1$.
The gain in robustness is not free. In the exact IV-indexed class, Theorem (ref) shows that the calibration contribution to interval length scales with $d/(m+1)$. Thus, the design problem is to choose the smallest class ${\mathcal{F}}$ whose ambiguity set still contains the distribution shifts that matter scientifically or empirically.
\paragraph{Indicator and partition classes} If $Z$ is discrete, or if the user is willing to discretize a continuous IV, the simplest construction is an indicator basis. For a partition $\Pi=\left\{{\mathcal{Z}}_1,\dots,{\mathcal{Z}}_G\right\}$ of the IV space, take \[ \phi(z)=\left(1,\mathbbm{1}\left[z\in {\mathcal{Z}}_1\right],\dots,\mathbbm{1}\left[z\in {\mathcal{Z}}_G\right]\right). \] Then the induced tilts are constant within cells, and the fitted radius is piecewise constant. This is the natural construction for binary encouragement designs, judge identifiers, hospital identifiers, or other categorical IVs.
The same idea extends to overlapping groups. If $G_1,\dots,G_L\subset {\mathcal{Z}}$ are policy-relevant subsets, one may take \[ \phi(z)=\left(1,\mathbbm{1}\left[z\in G_1\right],\dots,\mathbbm{1}\left[z\in G_L\right]\right). \] The resulting guarantee holds for every group indicator and every nonnegative linear combination of them. When the groups overlap, the fitted radius is constant on the atoms generated by the overlap pattern.
\paragraph{Multiscale discretization} For a continuous scalar IV, one may include indicators for a coarse partition together with indicators for a finer nested partition. This allows the ambiguity set to contain both global reweightings and more local perturbations of the IV law. When $Z$ is irregular or mixed discrete-continuous, tree-based or forest-based partitions play the same role. A tree learned on $\mathcal I_{\mathrm{tr}}$ induces leaves $L_1,\dots,L_G$, and one can use leaf indicators $\mathbbm{1}\left[z\in L_g\right]$ as features. This yields an adaptive discretization of the IV space.
\paragraph{Series and sieve classes} When $Z$ is continuous and low-dimensional, a natural alternative is a smooth series basis defined as \[ \phi(z)=\left(1,\psi_1(z),\dots,\psi_d(z)\right), \] where $\psi_1,\dots,\psi_d$ may be spline basis functions, orthogonal polynomials, wavelets, or trigonometric terms. The resulting class protects against smooth density-ratio perturbations that can be represented, or well approximated, by the sieve span. This is the same approximation logic that underlies sieve NPIV estimation Chen2012estimationof,Chen2015sievewald.
For a vector IV $Z=(Z_1,\dots,Z_{k_Z})$, one may use additive or tensor-product sieves. An additive construction takes the form \[ \phi(z)=\Big(1,\psi_{1,1}(z_1),\dots,\psi_{1,d_1}(z_1),\dots,\psi_{k_Z,1}(z_{k_Z}),\dots,\psi_{k_Z,d_{k_Z}}(z_{k_Z})\Big), \] while a tensor-product construction adds selected interaction terms of the form \[ \phi_{j_1,\dots,j_q}(z)=\prod_{\ell=1}^q \psi_{\ell,j_\ell}(z_\ell). \] The additive version is often preferable when $k_Z$ is moderate, because it keeps $d$ small. Tensor products are appropriate only when domain knowledge points to specific interactions that may change in the setting of interest.
\paragraph{RKHS and kernel dictionaries} A convenient idealization for smooth, local, and nonlinear shifts is an RKHS ball. Let $K\colon {\mathcal{Z}}\times{\mathcal{Z}}\to{\mathbb{R}}$ be a positive-definite kernel with RKHS ${\mathcal{H}}_K$. A natural infinite-dimensional candidate is \[ {\mathcal{F}}^{\infty}_{K,B}=\left\{f\in {\mathcal{H}}_K: \left\|f\right\|_{{\mathcal{H}}_K}\le B,\ f\ge 0\right\}. \] The exact finite-sample result in this paper is developed for finite-dimensional classes, so the practical route is to approximate ${\mathcal{H}}_K$ by a finite dictionary, \[ \phi_r(z)=\left(1,K(z,u_1),\dots,K(z,u_r)\right), \] where $u_1,\dots,u_r\in {\mathcal{Z}}$ are landmark points. The induced class ${\mathcal{F}}_{\phi_r}$ is then a finite-dimensional proxy for the ideal RKHS family, and the exact theorem applies to this proxy.
The kernel determines the geometry of the protected shifts. Gaussian kernels generate local tilts centered around the landmarks. Polynomial kernels generate global low-order smooth reweightings. If one expects the shift to be smooth after an appropriate transformation of the IV, one may replace $z$ by a representation $g(z)$ and use features of the form $K\left(g(z),g(u_j)\right)$. Low-rank approximations such as Nystr\"om dictionaries and random features provide practical ways to keep $d$ moderate Rahimi2007randomfeatures,Musco2017recursivesampling.
\paragraph{Sparse linear and representation-based classes} When the IV is high-dimensional, a direct sieve or kernel expansion in the raw coordinates can make $d$ too large relative to the calibration size. A natural compromise is to construct a low-dimensional summary map $\psi\colon {\mathcal{Z}}\to{\mathbb{R}}^q$ on the training split and then set \[ \phi(z)=\left(1,\psi_1(z),\dots,\psi_q(z)\right). \] The summary may consist of principal components, selected coordinates, geographic summaries, judge-share summaries, or a learned representation of a text, image, or network-valued IV. The point is not that the representation must be rich enough to estimate $h_0$. It only needs to span the directions along which the IV distribution is likely to move in the setting of interest.
\paragraph{Shape-restricted classes} For scalar IVs, or for low-dimensional summaries of vector IVs, it can be natural to focus on monotone, convex, or bounded-variation shifts. Examples include interventions that only increase encouragement among larger values of $Z$, or geographic changes that monotonically reweight distance. In the exact conformal setting, the cleanest implementation is to approximate a shape-restricted family by a finite grid and a basis of piecewise-constant or piecewise-linear functions. One may then use the induced low-dimensional span as a proxy for the desired cone. A stronger approach would impose linear inequality constraints on the coefficients. Such constrained calibration remains computationally natural, because it leads to linear or convex programs, but it goes beyond the unconstrained span class analyzed here.
For the $X$-indexed class ${\mathcal{T}}_X$, there are two distinct modeling choices. The first is the shift class ${\mathcal{F}}$, which still describes how the IV distribution may move in the setting of interest. The second is the radius class for $\tau_X$, which determines how the interval width may vary with $X$. The two should not be conflated.
In practice, one may choose the radius class on $X$ from the same broad families used in nonparametric regression, for example bins in $X$, additive sieves, tensor-product bases, shape-restricted classes, RKHS models, or neural-network classes. The role of that class is completely different from the role of ${\mathcal{F}}$. The radius class controls how much heterogeneity the final prediction interval may have across values of $X$. The shift class controls which settings of interest must be protected. In the exact classes ${\mathcal{T}}_{XZ}$ and ${\mathcal{T}}_Z$, these two roles are partially tied together by the conditional conformal algorithm. In the $X$-indexed method, they are separate.
This section collects the main theoretical statements. Proofs appear in Appendix (ref).
We first show that the coverage ratio with an $(X, Z)$-indexed radius works as intended.
We next show that the coverage ratio with a $Z$-indexed radius works as intended.
\paragraph{Oracle comparison for the IV-indexed class} We now relate interval length in the exact IV-indexed class to NPIV estimation error and calibration complexity.
Let the oracle structural function be $h_0$ and define oracle scores $S_i^\star=\left|Y_i-h_0(X_i)\right|$. Let $\widehat\tau_Z$ be computed from $\left\{S_i\right\}$ and let $\widehat\tau_Z^\star$ be computed from $\left\{S_i^\star\right\}$ using the same IV-indexed calibration algorithm and features.
Next, compare this with the oracle IV-indexed radius \[ \tau_0(z)=\inf\left\{t: {\mathbb{P}}\left(\left|\varepsilon\right|\le t\mid Z=z\right)\ge 1-\alpha\right\}. \] We require two regularity conditions.
The next proposition records the population criterion underlying the $X$-indexed importance-weighted method.
Then, the following theorem provides a complementary finite-sample statement for the $X$-indexed class. It applies to one fixed target shift rather than to a whole class, and it requires an independent final recalibration split together with a known upper bound on the target weight. This is the exact finite-sample guarantee currently available for prediction intervals $C(X)$ under a fixed target shift.
We consider three synthetic NPIV designs of increasing difficulty. In all designs, the instruments are sampled from a uniform distribution on $\left[-1,1\right]^{d_Z}$, the endogenous regressors are nonlinear functions of the IVs and latent variables, and the outcome is generated as \[ Y=h_0(X)+\text{confounding}+\sigma(X,Z)V, \] where the latent variables shared by $X$ and $Y$ induce endogeneity and $\sigma(X,Z)$ yields heteroskedastic predictive uncertainty.
\paragraph{Dataset 1 (\texorpdfstring{$d_X=d_Z=1$}{dX=dZ=1}).} Let $Z\sim \mathrm{Unif}[-1,1]$ and let $U,E,V$ be independent standard normal variables. We generate \[ X=1.05\sin(\pi Z)+0.55U+0.12E, \qquad h_0(X)=\frac{\sin(\pi X)}{1+0.45X^2}, \] and \[ Y=h_0(X)+0.55U+\sigma(X,Z)V, \] with \[ \sigma(X,Z)=0.35+0.10|X|+0.12\frac{Z+1}{2}+0.08\left(1+\exp(-XZ)\right)^{-1}. \] This one-dimensional design is the basis for the visualization in Figure (ref).
\paragraph{Dataset 2 (\texorpdfstring{$d_X=3,d_Z=1$}{dX=3,dZ=1}).} Let $Z\sim \mathrm{Unif}[-1,1]$ and let $U_1,U_2,U_3,E_1,E_2,E_3,V$ be mutually independent standard normal variables. We set \[
\] \[ h_0(X)=\sin(X_1)+0.45X_2^2-0.35X_1X_3+0.40\cos(X_3), \] and \[ Y=h_0(X)+(0.40U_1-0.30U_2+0.22U_3)+\sigma(X,Z)V, \] where \[ \sigma(X,Z)=0.32+0.08|X_1|+0.07X_2^2+0.10\frac{Z+1}{2}+0.05\max\left\{X_3Z,0\right\}. \]
\paragraph{Dataset 3 (\texorpdfstring{$d_X=d_Z=3$}{dX=dZ=3}).} Let $Z=(Z_1,Z_2,Z_3)$ have independent coordinates distributed as $\mathrm{Unif}[-1,1]$, and let $U_1,U_2,U_3,E_1,E_2,E_3,V$ be mutually independent standard normal variables. We generate \[
\] \[ h_0(X)=0.60\sin(X_1)+0.38X_2^2-0.25X_1X_3+0.48\cos(X_3)+0.18X_1X_2, \] and \[ Y=h_0(X)+(0.36U_1-0.24U_2+0.18U_3)+\sigma(X,Z)V, \] with \[ \sigma(X,Z)=0.30+0.03(X_1^2+X_2^2)+0.07\max\{X_3,0\}+0.07\frac{Z_1+1}{2}+0.05\frac{Z_2Z_3+1}{2}. \] This is the most difficult design, because both the structural function and the heteroskedastic scale depend on higher-dimensional nonlinear interactions.
We set the nominal coverage level to $1-\alpha=0.9$ and run $100$ Monte Carlo replications for each synthetic design. Each replication uses $n_{\mathrm{train}}=1000$ observations to fit the NPIV base learner, $n_{\mathrm{cal}}=200$ observations for conformal calibration, and $n_{\mathrm{test}}=1000$ observations for evaluation. Across all methods, the center $\widehat h$ is a common series-2SLS NPIV estimator with cubic polynomial bases in both $X$ and $Z$.
For the exact classes ${\mathcal{T}}_{XZ}$ and ${\mathcal{T}}_Z$, we consider three finite-dimensional shift families. The Bins family uses a four-bin basis defined on a one-dimensional projection. For scalar IVs, this projection is $Z$ itself. For vector IVs, it is the standardized first principal component of the pooled training and calibration IVs. The Linear family uses an affine basis with an intercept, and the RKHS family uses a Gaussian-kernel dictionary with four landmarks and kernel parameter $\gamma=0.2$. For the $X$-indexed class ${\mathcal{T}}_X$, we consider four radius models: a linear basis, a six-bin basis, an RKHS basis with four landmarks and $\gamma=0.2$, and an MLP radius model. The conditional density ratio in the $X$-indexed learner is estimated by a neural network with hidden layers $(32,32)$ using two-fold cross-fitted training scores. We then split the calibration sample evenly into a shape-learning half and an independent final recalibration half. The $X$-indexed learner uses $\lambda=50$, $\kappa=0.05$, and, for the neural radius model, hidden layers $(32,32)$.
To evaluate robustness under IV shifts, we report four test laws. Let $u(z)$ be the standardized one-dimensional projection of $z$ computed from the pooled training and calibration IVs, and when $Z$ is multivariate, $u(z)$ is the first principal component after standardization. The raw weights are \[ w_{\mathrm{obs}}(u)=1,\qquad w_{\mathrm{lin}}(u)=\max\left\{1+0.95u,0.05\right\}, \] \[ w_{\mathrm{step}}(u)=
\qquad w_{\mathrm{loc}}(u)=0.20+1.60\exp\!\left[-\frac12\left(\frac{u-0.75}{0.35}\right)^2\right]. \] For each method and scenario, the reported coverage ratio and interval length are weighted test-sample averages computed with the normalized versions of these weights. An entry of inf in the length column means that at least one replication produced an unbounded interval, which makes the across-replication mean infinite.
Tables (ref)--(ref) summarize the synthetic experiments. Overall, the exact classes ${\mathcal{T}}_{XZ}$ and ${\mathcal{T}}_Z$ behave as intended for the Bins and Linear shift families: their coverage ratios stay close to the nominal level across the observed law and the three shifted IV laws. In Dataset 1, exact interval lengths are around $4$ for both classes, with the Linear specification slightly shorter than Bins and RKHS. In Dataset 2, the same pattern remains, and the exact finite-length intervals are even shorter, around $2.7$ to $3.1$. In Dataset 3, the exact procedures remain finite for Bins and Linear, but the intervals become much longer, reflecting the more difficult three-dimensional nonlinear design.
The RKHS proxy is more fragile in the exact classes. In Dataset 1, both ${\mathcal{T}}_{XZ}$ and ${\mathcal{T}}_Z$ with RKHS remain finite and are mildly conservative. In Dataset 2, the $Z$-indexed RKHS procedure is still well behaved, but the $(X,Z)$-indexed RKHS procedure yields unbounded intervals in at least one replication. In Dataset 3, the same contrast becomes sharper: the $Z$-indexed RKHS method remains finite, whereas the $(X,Z)$-indexed RKHS method again produces inf. This pattern is consistent with the fact that richer exact shift classes are harder to calibrate stably as the dimension of the conditioning variable grows.
The $X$-indexed results illustrate the trade-off between interpretability and difficulty. In Dataset 1, the Linear $X$-indexed model is the most stable $C(X)$ rule, with coverage near the nominal level and only a moderate increase in length relative to the exact methods. The RKHS $X$-indexed model is workable but more variable, while the Bins and especially the MLP radius models are unstable. In Dataset 2, the Linear, Bins, and RKHS $X$-indexed models all remain finite and near nominal, showing that fixed-shift recalibration can work well in a moderate-dimensional design. In Dataset 3, however, all $X$-indexed procedures become much longer and more variable, especially the Linear and MLP models. This is exactly the difficult regime for the target class ${\mathcal{T}}_X$: the final interval is required to depend only on $X$, even though the predictive uncertainty in the data-generating process varies with both $X$ and $Z$.
Figure (ref) visualizes the fitted intervals for Dataset 1 under the observed IV distribution using a single replication with seed $2026$. For the exact procedures, we use the Linear shift family, and for the $X$-indexed procedure we use the Linear radius model. The top row plots the lower and upper interval surfaces over a $41\times 41$ grid in $(x,z)$. The middle row fixes $x=0$ and plots the $Y$--$Z$ slices, while the bottom row fixes $z=0$ and plots the $Y$--$X$ slices.
Note that large or infinite-length prediction intervals do not necessarily imply poor performance. If one of the upper or lower bounds is large or infinite, the interval length becomes infinite. However, the opposite bound can still be finite, and if so, we can still use such prediction intervals in practice. A more important metric is the coverage ratio.
The figure makes the distinction among the three radius classes visually transparent. In the middle row, the $(X,Z)$-indexed and $Z$-indexed intervals vary with $z$, whereas the $X$-indexed interval is flat in $z$ by construction. In the bottom row, all three methods vary with $x$, but the width profiles differ: the exact methods adapt to IV-indexed uncertainty, while the $X$-indexed interval absorbs that uncertainty into a single $X$-only radius. In Dataset 1, the $(X,Z)$-indexed and $Z$-indexed surfaces are quite similar, which suggests that much of the relevant heteroskedasticity can already be captured by allowing the radius to depend on the IV alone.
The simulation evidence supports the main conceptual message of the paper. The exact $Z$-indexed class ${\mathcal{T}}_Z$ is the most reliable practical exact procedure: it consistently delivers near-nominal coverage under the shifted IV laws while keeping the interval lengths comparable to, and often slightly shorter than, those of the broader exact class ${\mathcal{T}}_{XZ}$. The $X$-indexed target ${\mathcal{T}}_X$ remains substantively appealing, because it yields intervals of the form $C(X)$, but its finite-sample performance depends much more strongly on the radius model and on the complexity of the design. In low- and moderate-dimensional settings, simple $X$-indexed models such as the Linear basis can work well. In harder settings, the price of insisting on $Z$-free final intervals can be substantial.
These results also clarify how to interpret richer model classes in practice. RKHS-based exact classes can become unstable in more difficult designs, and overly flexible $X$-indexed radius models can produce very long intervals. Accordingly, the finite-dimensional Bins and Linear exact classes are the most robust baselines in the current experiments, while richer classes are better viewed as sensitivity analyses rather than default choices.
We next study two empirical IV datasets using the same nominal level, shift scenarios, and evaluation metrics as in Section 7. In the CigarettesSW data Stock2007introductionto, we take \[ Y=\log\left(\texttt{packs}\right),\qquad X=\log\left(\texttt{price}/\texttt{cpi}\right),\qquad Z=\log\left(\texttt{taxs}/\texttt{cpi}\right), \] and include $\log\left(\texttt{income}/\texttt{population}/\texttt{cpi}\right)$ and a 1995 year indicator as exogenous controls. Each repetition uses a year-stratified split with $24$ training observations, $40$ calibration observations, and $32$ test observations. In the CollegeDistance data Stock2007introductionto,Rouse1995democratizationor, we take \[ Y=\log\left(\max\left\{\texttt{wage},10^{-2}\right\}\right),\qquad X=\texttt{education},\qquad Z=\texttt{distance}, \] and use score, tuition, unemp, and demographic dummy variables as controls. To keep the exact conditional procedures computationally manageable, each repetition first subsamples $1800$ observations and then uses a $600/600/600$ training, calibration, and test split. All empirical results are averaged over $100$ random replications.
Tables (ref) and (ref) show a sharper contrast between stable and unstable specifications than in the synthetic designs. In CigarettesSW, the exact Bins procedures and the exact $Z$-indexed Linear procedure remain finite and close to the nominal coverage level, whereas the richer exact classes frequently produce unbounded intervals. For the $X$-indexed class, the Observed and Step-Tilt scenarios are still workable for the Linear and RKHS radius models, but the Linear-Tilt and Local-Tilt scenarios often lead to inf. This is the small-sample regime in which aggressive reweightings of the IV distribution are hardest to support, so coarse exact shift classes are the most reliable practical choices.
In CollegeDistance, the exact Bins and Linear procedures are much more stable. Their coverage ratios stay close to the nominal level under all four scenarios, and their interval lengths are around $2.5$ to $2.7$. By contrast, the exact RKHS procedures again produce inf, which indicates that the richer exact shift class is still too unstable in this empirical design. The most striking pattern appears in the $X$-indexed class: the Linear, Bins, and RKHS radius models produce very long intervals, whereas the MLP radius model remains stable, with coverage around $0.895$ to $0.905$ and interval length around $2.54$ to $2.56$. Thus, in the larger empirical dataset, a flexible nonlinear $q(x)$ can be essential for obtaining a practical $C(X)$ interval.
Taken together, the empirical results complement the simulations. Exact IV-specific procedures based on Bins or Linear shift classes are the most robust default options across datasets. The $X$-indexed target can be competitive, but only when the radius model matches the complexity of the application and the final fixed-shift recalibration remains numerically stable.
We proposed IV-CCP for constructing prediction intervals in NPIV with a distribution-free finite-sample coverage guarantee. Our proposed prediction interval construction allows three types of radii, $(X, Z)$-indexed, $Z$-indexed, and $X$-indexed, whose classes are denoted by ${\mathcal{T}}_{XZ}$, ${\mathcal{T}}_Z$, and ${\mathcal{T}}_X$, respectively. The broad benchmark class is ${\mathcal{T}}_{XZ}$, which produces intervals whose radius may vary jointly with $(X,Z)$ and admits exact finite-sample coverage over joint shifts. The exact IV-specific class is ${\mathcal{T}}_Z$, which keeps the center equal to $\widehat h(X)$ and yields IV-CCP, an exact finite-sample procedure over policy-relevant IV shifts. The class ${\mathcal{T}}_X$ produces intervals $C(X)$, which may be the most natural operational target when the final prediction rule should depend on the structural regressor alone. For these approaches, we provide theoretical guarantees and confirm their validity through simulation studies and empirical analyses.
\onecolumn