The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
57,453 characters
Debiased Inference for Bounding Wage Inequality with Many Controls
\begin{abstract}
We study estimation and inference for a partially identified parameter whose identified set depends on a first-stage nuisance parameter that must itself be estimated.
Combining the criterion-function approach with the theory of Neyman-orthogonal moments that underlies double/debiased machine learning, we propose a two-step procedure: the point-identified nuisance is estimated by flexible machine-learning methods, and the set-identified target is recovered as a level set of a sample criterion built from orthogonal moment inequalities with cross-fitting.
When the contour level is bounded, we show that the resulting set estimator converges in Hausdorff distance at the parametric rate of the infeasible criterion built on the true nuisance. We further develop a subsampling procedure that delivers asymptotically valid coverage, provided the product of the first-stage estimation errors is $o(N^{-1/2})$. We illustrate the method on bounds for the wage distribution and the interquantile range under selection into employment and on the gender wage gap with an interval-censored wage. The empirical application studies the gender wage gap using the March supplement of the 2015 Current Population Survey.
\keyword{Bounds on wage distributions}
\keyword{bracketed (interval-valued) data}
\keyword{sample selection}
\keyword{double/debiased machine learning}
\keyword{subsampling}
\keyword{partial identification}
\keyword{moment inequalities.}
\end{abstract}
\section{Introduction}
\label{sec:intro}
Sample selection is a classic problem in econometrics \citep{Heckman74}. Economically important objects such as population wage distributions, gender and education wage differentials, and measures of wage inequality depend on the wages of individuals who do not work, and these wages are not observed. One approach imposes enough structure on the joint distribution of wages and employment to recover the latent wage distribution, for instance through a control function. For example, \citet{CFL} take this route with distribution regression to decompose British wages. An alternative approach leaves the selection process incompletely specified and asks what can be learned about the latent wage distribution under weaker restrictions, as in the partial identification analysis of \citet{Manski03}. In this paper, we focus on \citet{BGIM}, who take the second route to bound changes in the distribution of British wages. Their positive-selection restriction, under which the wage distribution of workers first-order stochastically dominates that of nonworkers conditional on observed characteristics, substantially tightens the worst-case bounds without point-identifying the missing wage distribution. The restriction is more plausible when the compared individuals are observationally similar, because it rules out selection only within narrow groups. Conditioning on many characteristics, however, creates a dimensionality problem.
This paper develops estimation and inference for partially identified models whose identifying inequalities depend on a high-dimensional first stage. The parameter of interest $\theta$ is set-identified by a finite collection of unconditional moment inequalities,
\begin{equation*}
\mathrm{E}\,g(Z;\theta,\xi_0)\,\bm{\leq}\,0,
\end{equation*}
where the nuisance parameter $\xi_0$ is point-identified but may be infinite-dimensional. We estimate $\xi_0$ by machine-learning methods with cross-fitting, construct Neyman-orthogonal versions of the inequalities, and estimate the identified set as a level set of the criterion function of \citet{CHT}. We illustrate it with the wage-distribution bounds of \citet{BGIM} and a regression with an interval-valued outcome. The object of inference is the identified set with the confidence region covering $\Theta_I$ with probability approaching $1-\tau$, where $1-\tau$ is the nominal confidence level.
The main technical contribution is subsampling inference that remains valid when the criterion is built from a cross-fitted first stage. Cross-fitting breaks the independence across subsamples that subsampling requires, because every block statistic uses a nuisance estimated on the whole sample and is maximized over a preliminary set built from the whole sample. We restore it using a carefully crafted balanced, disjoint-block construction together with a bracketing argument. Every block contains the same number of observations from each cross-fitting fold, so the observations it draws from fold $k$ are independent of the nuisance estimated without that fold. Each feasible block statistic is then bounded by statistics evaluated at the true nuisance on a small enlargement or contraction of the identified set. These bounding statistics depend only on the observations of their own block and can therefore be treated with ordinary block-level asymptotics. This yields consistency of the subsampling critical value and pointwise asymptotic coverage,
\begin{equation*}
{\mathrm{P}} \left(\Theta_I \subseteq C_N (\hat{c}_\tau; \hat{\xi})\right) = 1 - \tau + o(1),
\end{equation*}
under $b \to \infty$ and $b c_N / N \to 0$, with no strengthening of the first-stage rate condition used for set estimation. Combining cross-fitting with dependent data raises a related difficulty, resolved by a cross-fitting scheme tailored to the dependence, as in the weakly dependent panels of \citet{CGST}, or by an additional stability condition, as in the network and spatial designs of \citet{CaoLeung}.
A second contribution is a general theory that does not require the moment inequalities to be affine in $\theta$. Bounds on the interquantile range of the wage distribution under selection provide such an example.
We demonstrate our method on the partially identified model studied in \citet{CCMS}, with wage data from the March supplement of the 2015 Current Population Survey. In particular, we document that, unlike an approach that reports only separate bounds on individual coefficients, the criterion-function formulation directly delivers a joint confidence region for the two-dimensional coefficient set. Across the designs we consider, the estimated joint identified set occupies only about $50\%$ of the box formed by its coordinate projections. We interpret this as evidence that separate coefficient bounds contain many combinations that are inconsistent with the joint directional restrictions.
This paper contributes to the literature on partial identification by incorporating a function-valued first stage into the two-stage framework of \citet{KaidoWhite}. In doing so, we benefit from the Neyman-orthogonality and cross-fitting machinery of \citet{chernozhukov2016double}, \citet{LRSP} and \citet{chernozhukov2021automatic}, so that a high-dimensional first stage can be estimated by regularized methods without affecting first-order inference. In partial identification, earlier applications of these techniques are mostly confined to bounds on one-dimensional parameters \citep{Heiler2024,Semenova2025} and to support functions of convex identified sets \citep{Semenova2023,LiuMolinari}, building on \citet{BeresteanuMolinari2008}, \citet{BMM2} and \citet{BontempsMagnacMaurin2012}. \citet{LiuMolinari}, for instance, propose debiased estimation and inference for an algorithmic fairness-accuracy frontier, and \citet{LiuSPT2025} develops a partial identification framework for panel-data policy evaluation that is robust to violations of the identifying assumptions of difference-in-differences and synthetic control. In contrast, our moment inequalities may be nonlinear in $\theta$ and describe potentially non-convex identified sets. Last but not least, the paper's emphasis on selection problems contributes to the sample selection literature, see \citet{FVVW2019,FVV2018,FVHong2025}.
Since the first version of this paper\footnote{\Cref{app:theory} and its proofs, except the extension to moments that are discontinuous in $\theta$, first appeared in Chapter~4 of arXiv:1808.02569, formerly titled ``Machine Learning for Dynamic Discrete Choice.''} (arXiv:1808.02569, August 2018), \citet{Tian2025} has studied a closely related two-stage problem. He combines the refined conditional chi-squared test of \citet{CoxShi2023} with a correction that propagates first-stage estimation uncertainty into the covariance matrix and inverts the test point by point, so that his confidence set covers each $\theta$ in the identified set with asymptotic probability at least the nominal level. The correction relies on an asymptotically linear first stage, such as GMM or maximum likelihood, and on a high-level condition that the plug-in moment vector be $\sqrt{N}$-asymptotically normal, which a machine-learning first stage that converges more slowly does not deliver without orthogonalization. We instead remove the first-order contribution of the estimated nuisance by orthogonalization, allow a first stage that converges more slowly than $N^{-1/2}$, and cover $\Theta_I$ as a whole. The two procedures thus treat the first stage differently: \citet{Tian2025} carries its first-order sampling variation into inference, while our construction makes that variation second order under suitable rate conditions.
\paragraph{Structure of the paper.} \Cref{sec:setup} introduces the two-stage framework, the orthogonal criterion, and the estimation and inference procedure. \Cref{sec:example} presents the two examples: the wage distribution under selection of \citet{BGIM}, and the gender wage gap with an interval-valued outcome. \Cref{sec:examples-theory} gives primitive conditions and asymptotic guarantees for both examples, and \Cref{sec:empirical} reports the empirical application. Replication code is available at \texttt{https://github.com/korobkay/orthoset}. The appendix collects the general asymptotic theory, the proofs, and the verification of the high-level conditions for the two examples.
\section{Setup}
\label{sec:setup}
\subsection{Framework}
\label{sec:framework}
Let $Z_1,\ldots,Z_N$ be an i.i.d.\ sample from a distribution $P_0$ on a measurable space $\mathcal{Z}$. The \emph{target} parameter $\theta_0$ lies in a compact set $\Theta\subset\mathbb{R}^{d_\theta}$ and is the object of economic interest. The \emph{nuisance} parameter $\xi$ lies in a convex subset $\Xi$ of a normed space of functions of $Z$, with the $L^{2}(P_0)$ norm $\|\cdot\|_{P,2}$ unless an example specifies another norm. The nuisance has a true value $\xi_0\in\Xi$ that can be estimated without reference to $\theta_0$. Following \citet{chernozhukov2016double}, we assume that all objects are measurable and focus on estimation and inference. The pair $(\theta_0,\xi_0)$ has the two-stage structure of \citet{KaidoWhite}, with a point-identified first stage. Unlike their setting, the first stage here is an (infinite-dimensional) nuisance function rather than a Euclidean sub-vector.
A moment function $g:\mathcal{Z}\times\Theta\times\Xi\to\mathbb{R}^{L}$ defines the identified set through $L\geq 1$ inequalities evaluated at the true nuisance,
\begin{equation}\label{eq:idset}
\Theta_I := \big\{\theta\in\Theta:\ \mathrm{E}\,g(Z;\theta,\xi_0)\,\bm{\leq}\,0\big\},
\end{equation}
where $v\,\bm{\leq}\,0$ for a vector $v\in\mathbb{R}^{L}$ means that every coordinate of $v$ is non-positive. Write $\|v\|_{+}:=\|\max\{v,0\}\|$ for the Euclidean norm of the positive part of $v$, the maximum being taken coordinatewise, so that $\|v\|_{+}=0$ if and only if $v\,\bm{\leq}\,0$. Define the \emph{population criterion}
\begin{equation}\label{eq:Qpop}
Q(\theta,\xi) := \big\|\mathrm{E}\,g(Z;\theta,\xi)\big\|_{+}^{2}
\end{equation}
This criterion is non-negative for every $\theta\in\Theta$ and $\xi\in\Xi$. At the true value of the nuisance, it vanishes exactly on the identified set: $$\Theta_I=\{\theta\in\Theta: Q(\theta,\xi_0)=0\}.$$
We construct the sample criterion function by $K$-fold cross-fitting, with the number of folds $K\geq 2$ fixed. The indices $\{1,\ldots,N\}$ are partitioned into folds $J_1,\ldots,J_K$ of equal size $n=N/K$, and for $i\in J_k$ the nuisance $\widehat{\xi}_i:=\widehat{\xi}^{(k)}$ is estimated from the observations outside $J_k$ only. We write $\widehat{\xi}$ for the collection $(\widehat{\xi}^{(1)},\ldots,\widehat{\xi}^{(K)})$. The \emph{sample criterion} is
\begin{equation}\label{eq:Qn}
Q_N(\theta,\widehat{\xi}) := \bigg\|\frac{1}{N}\sum_{i=1}^{N} g(Z_i;\theta,\widehat{\xi}_i)\bigg\|_{+}^{2},
\end{equation}
and for a deterministic $\xi\in\Xi$ we write $Q_N(\theta,\xi)$ for the same expression with $\xi$ in place of every $\widehat{\xi}_i$. At $\xi=\xi_0$ it is the criterion of \citet{CHT} for these moment inequalities. For a level $c\geq 0$ and a nuisance $\xi$, the contour set of the sample criterion is
\begin{equation}\label{eq:contour}
C_N(c;\xi) := \big\{\theta\in\Theta:\ N\,Q_N(\theta,\xi)\leq c\big\}.
\end{equation}
\begin{definition}[Two-stage set estimator]
\label{def:setestim}
For a measurable level $\widehat{c}\geq 0$, the two-stage set estimator is
\begin{equation}\label{eq:setestim}
\widehat{\Theta}_I := C_N(\widehat{c};\widehat{\xi}).
\end{equation}
\end{definition}
The level $\widehat{c}$ is chosen in one of two ways. A deterministic level $c_N\to\infty$ with $c_N/N\to 0$, for instance $c_N=\log N$, makes $\widehat{\Theta}_I$ a consistent set estimator. A data-driven level $\widehat{c}_\tau$, computed by subsampling, makes it a confidence region for $\Theta_I$ with asymptotic coverage $1-\tau$.
\begin{definition}[Subsampling critical value]
\label{def:subsampling}
Fix a deterministic level $c_N\to\infty$ with $c_N/N\to 0$ and a block size $b=b_N\to\infty$, a multiple of $K$, with $b\,c_N/N\to 0$. Partition the sample into $B_N:=\lfloor N/b\rfloor$ disjoint blocks $\mathcal{I}_1,\ldots,\mathcal{I}_{B_N}$ of size $b$, each containing $b/K$ observations from every fold, so that the observations of a block in fold $k$ are independent of $\widehat{\xi}^{(k)}$. For each block $j=1,\dots,B_N$ compute
\begin{equation*}
\widehat{\mathcal{T}}_{j,b} \;:=\; \sup_{\theta\in C_N(c_N;\widehat{\xi})} b\,Q_{j,b}(\theta,\widehat{\xi}),
\qquad
Q_{j,b}(\theta,\widehat{\xi}) := \bigg\|\frac{1}{b}\sum_{i\in\mathcal{I}_j} g(Z_i;\theta,\widehat{\xi}_i)\bigg\|_{+}^{2},
\end{equation*}
so that $\widehat{\mathcal{T}}_{j,b}$ is the criterion of block $\mathcal{I}_j$, scaled by $b$ and maximized over the preliminary set $C_N(c_N;\widehat{\xi})$. Let $\widehat{G}$ be the empirical distribution function of $(\widehat{\mathcal{T}}_{j,b})_{j=1}^{B_N}$. For $\tau\in(0,1)$, the subsampling critical value is the $(1-\tau)$-quantile
\begin{equation*}
\widehat{c}_{\tau} \;:=\; \inf\{x\in\mathbb{R}:\widehat{G}(x)\geq 1-\tau\}.
\end{equation*}
\end{definition}
Rates for the set estimator are stated in the Hausdorff metric,
\begin{equation}\label{eq:hausdorff}
d_H(A,B) := \max\Big\{\sup_{x\in A}d(x,B),\ \sup_{y\in B}d(y,A)\Big\}, \qquad d(x,B):=\inf_{y\in B}\|x-y\|,
\end{equation}
and $\Theta_I^{\epsilon}:=\{\theta\in\Theta: d(\theta,\Theta_I)\leq\epsilon\}$ denotes the $\epsilon$-expansion of the identified set.
The first stage enters \eqref{eq:Qn} through $\widehat{\xi}$, whose estimation error is of larger order than $N^{-1/2}$ when $\xi_0$ is estimated by regularized methods. We require the moment function to be insensitive to that error to first order. For $\xi\in\Xi$ and $r\in[0,1]$ write $\xi_r:=\xi_0+r(\xi-\xi_0)$, which lies in $\Xi$.
\begin{definition}[Neyman orthogonality]
\label{def:ortho}
The moment function $g$ is \emph{Neyman-orthogonal} at $\xi_0$ on a set $\mathcal{N}\subset\Xi$ if, for every $\theta\in\Theta$ and $\xi\in\mathcal{N}$, the map $r\mapsto\mathrm{E}\,g(Z;\theta,\xi_r)$ is differentiable at $r=0$ with
\begin{equation}\label{eq:ortho}
\partial_r\,\mathrm{E}\,g(Z;\theta,\xi_r)\big|_{r=0} = 0 .
\end{equation}
\end{definition}
Orthogonality says that, at the truth, the expected moment is flat in the direction of any first-stage error, so an error in $\widehat{\xi}$ moves $\mathrm{E}\,g$ only to second order.
\subsection{Proposed procedure}
\label{sec:overview}
The following algorithm combines the cross-fitted first stage, the criterion \eqref{eq:Qn} and the contour sets \eqref{eq:contour} of \Cref{sec:framework} with a subsampling choice of the level. The orthogonal moment function $g$ is an input, supplied case by case. Step~1 is the sample splitting of \citet{schick1986asymptotically}, in the cross-fitted form of \citet{chernozhukov2016double}, and Steps~2--4 are the contour-set estimator and the subsampling procedure of \citet{CHT} applied to the cross-fitted criterion. Step~3 incorporates the proposed delicate combination of subsampling with cross-fitting.
\begin{algorithm}[Two-stage set estimator and confidence region]
\label{alg:main}
Inputs: an orthogonal moment function $g$ with nuisance $\xi$, the number of folds $K\geq 2$, a preliminary level $c_N\to\infty$ with $c_N/N\to 0$, a block size $b\to\infty$, a multiple of $K$, with $b\,c_N/N\to 0$, and a nominal confidence level $1 - \tau$.
\begin{enumerate}
\item \emph{Cross-fitted first stage.} Partition $\{1,\ldots,N\}$ into folds $J_1,\ldots,J_K$ of equal size $n=N/K$. For each $k$, compute an estimate $\widehat{\xi}^{(k)}$ of $\xi_0$ from the observations outside $J_k$, and for $i\in J_k$ set $\widehat{\xi}_i:=\widehat{\xi}^{(k)}$.
\item \emph{Criterion and contour sets.} Form the cross-fitted criterion \eqref{eq:Qn} and the contour sets \eqref{eq:contour},
\begin{equation*}
Q_N(\theta,\widehat{\xi}) = \bigg\|\frac{1}{N}\sum_{i=1}^{N} g(Z_i;\theta,\widehat{\xi}_i)\bigg\|_{+}^{2},
\qquad
C_N(c;\widehat{\xi}) = \big\{\theta\in\Theta:\ N\,Q_N(\theta,\widehat{\xi})\leq c\big\}.
\end{equation*}
\item \emph{Preliminary set and block statistics.} Form the preliminary set $C_N(c_N;\widehat{\xi})$ at the deterministic level $c_N$, for instance $c_N=\log N$. Partition the sample into $B_N=\lfloor N/b\rfloor$ disjoint blocks $\mathcal{I}_1,\ldots,\mathcal{I}_{B_N}$ of size $b$, each carrying $b/K$ observations from every fold. On block $j$ compute
\begin{equation*}
Q_{j,b}(\theta,\widehat{\xi}) = \bigg\|\frac{1}{b}\sum_{i\in\mathcal{I}_j} g(Z_i;\theta,\widehat{\xi}_i)\bigg\|_{+}^{2},
\qquad
\widehat{\mathcal{T}}_{j,b} = \sup_{\theta\in C_N(c_N;\widehat{\xi})} b\,Q_{j,b}(\theta,\widehat{\xi}),
\end{equation*}
\item \emph{Level and region.} Let $\widehat{G}$ be the empirical distribution function of $(\widehat{\mathcal{T}}_{j,b})_{j\leq B_N}$ and let
\begin{equation*}
\widehat{c}_{\tau} := \inf\big\{x:\ \widehat{G}(x)\geq 1-\tau\big\}
\end{equation*}
be its $(1-\tau)$-quantile, as in \Cref{def:subsampling}. Report the set estimator $\widehat{\Theta}_I=C_N(\widehat{c};\widehat{\xi})$ of \Cref{def:setestim} with $\widehat{c}=c_N$, or with $\widehat{c}=0$ when the sample criterion vanishes on a set with non-empty interior, together with the confidence region $C_N(\widehat{c}_{\tau};\widehat{\xi})$.
\end{enumerate}
\end{algorithm}
In applications, the supremum in Step~3 can be computed on a finite grid. The preliminary set is represented by the grid points at which $N\,Q_N(\theta,\widehat\xi)\leq c_N$, with spacing small relative to $\sqrt{c_N/N}$, and each block statistic is subsequently maximized over these points. When $g$ is affine in $\theta$, the block criterion and the preliminary set are convex, so only boundary grid points need to be evaluated.
\section{Examples}
\label{sec:example}
\subsection{Wage distribution under selection into employment}
\label{sec:example1}
Our first example is the wage-distribution problem of \citet{BGIM}. They study how the distribution of UK log hourly wages changed between 1978 and 2000 using the Family Expenditure Survey, and they report the interquartile range as a measure of within-group inequality together with educational, gender and cohort differentials in the median. Because wages are recorded only for active workers, these objects are only partially identified, and \citet{BGIM} bound them under restrictions on selection into work.
\subsubsection{Linear case: the distribution function on a grid}
\label{sec:example1-grid}
For each individual $i=1,\ldots,N$, let $W_i$ denote the log wage, $D_i\in\{0,1\}$ an employment indicator, and $X_i\in \mathbb{R}^{p}$ a vector of conditioning covariates. Selection into employment is non-ignorable, $W \not\perp D \mid X$, and we impose no exclusion restriction. Fix a grid of wage thresholds $w_1 < w_2 < \cdots < w_S$. The target parameter is the vector of population distribution-function values on the grid,
\begin{equation}\label{eq:gridtarget}
\theta_0 := \big(F_0(w_1),\ldots,F_0(w_S)\big)', \quad F_0(w_k) := {\mathrm{P}}(W\leq w_k), \quad k = 1, \ldots, S.
\end{equation}
Write $F_0(w\mid x) := {\mathrm{P}}(W\leq w\mid X=x)$ for the conditional distribution function, $F_0 ^1(w\mid x)$ and $F_0 ^0(w\mid x)$ for its counterparts among the employed and the unemployed, and $p_0(x) := {\mathrm{P}}(D=1\mid X=x)$ for the \emph{employment propensity}. The law of iterated expectations gives
\begin{equation}\label{eq:lie}
F_0(w\mid x) = F_0 ^1(w\mid x)\,p_0(x) + F_0 ^0(w\mid x)\,\big(1-p_0(x)\big).
\end{equation}
The propensity $p_0(x)$ and the employed-wage distribution $F_0 ^1 (w \mid x)$ are identified from the data (provided that $p_0(x) > 0$). On the contrary, the unemployed-wage distribution $F_0 ^0(w \mid x)$ is not identified, and it is restricted only by $F_0 ^0(w\mid x)\in[0,1]$. Substituting into \eqref{eq:lie} gives the worst-case bounds of \citet{Manski1994}, $$F_0 ^1(w\mid x)p_0(x)\leq F_0(w\mid x)\leq F_0 ^1(w\mid x)p_0(x)+1-p_0(x),$$ whose width $1-p_0(x)$ is the non-employment rate for a given value of $x$. \citet{BGIM} tighten the lower bound under positive selection into employment, that is, first-order stochastic dominance of the employed wage distribution, $F_0 ^1(w\mid x)\leq F_0 ^0(w\mid x)$. Under that restriction it holds that
\begin{equation}\label{eq:better_bounds}
F_0 ^1(w \mid x) \leq F_0(w \mid x) \leq F_0 ^1 (w \mid x)\,p_0(x) + \big(1 - p_0(x)\big).
\end{equation}
To convert the problem to unconditional moment inequalities framework of \cref{sec:setup}, we integrate the bounds in \eqref{eq:better_bounds} over the distribution of $X$, resulting in the bounds for each coordinate of $\theta_0$,
\begin{equation}\label{eq:coordbounds}
\ell_{0,k} := \mathrm{E}\big[F_0 ^1(w_k\mid X)\big], \quad
u_{0,k} := \mathrm{E}\big[F_0 ^1(w_k\mid X)p_0(X) + 1 - p_0(X)\big],
\end{equation}
which gives
$$
\ell_{0,k} \leq \theta_{0,k} \leq u_{0,k}.
$$
A second restriction, that $F_0 ^0(\cdot\mid x)$ is a distribution function, ties the coordinates together. Applying \eqref{eq:lie} at $w_j<w_k$ and subtracting gives $F_0(w_k\mid x) - F_0(w_j\mid x) \geq p_0(x)(F_0 ^1 (w_k\mid x) - F_0 ^1(w_j\mid x))$, since $F_0^0(\cdot\mid x)$ is non-decreasing. Integrating over $X$ and using $$\mathrm{E}[p_0(X)(F_0 ^1(w_k\mid X)-F_0 ^1(w_j\mid X))] = \mathrm{E}[D\,\mathbf{1}\{w_j<W\leq w_k\}]$$ turns this into a restriction on the coordinates of $\theta_0$,
\begin{equation}\label{eq:increment}
\theta_{0,k} - \theta_{0,j} \;\geq\; \mathrm{E}\big[D\,\mathbf{1}\{w_j<W\leq w_k\}\big] \;=\; u_{0,k} - u_{0,j}, \qquad j<k.
\end{equation}
Combining \eqref{eq:coordbounds} and \eqref{eq:increment} gives the identified set
\begin{equation}\label{eq:idsetgrid}
\Theta_I = \Big\{\theta\in\Theta:\; \ell_{0, k}\leq\theta_k\leq u_{0, k} \;\;\forall k, \quad \theta_{k}-\theta_{k-1}\geq u_{0, k}-u_{0, k - 1}\;\;\forall k\geq 2\Big\},
\end{equation}
a convex polytope cut out by \(3S-1\) inequalities.
In the notation of \Cref{sec:setup}, the observation is $Z_i := (X_i, D_i, D_i W_i)$, $i = 1, \ldots, N$, and the moment function has $3S-1$ coordinates. For $k=1,\ldots,S$ the two coordinate moments are $g_{k}^{\mathrm{lo}}(Z;\theta,\xi) := \phi_L(Z;w_k,\xi) - \theta_k$ and $g_{k}^{\mathrm{up}}(Z;\theta) := \theta_k - \phi_U(Z;w_k)$, where
\begin{equation}\label{eq:scores}
\begin{aligned}
\phi_L(Z;w,\xi) &:= F^1(w \mid X) + \frac{D\,\big(\mathbf{1}\{W \leq w\} - F^1(w \mid X)\big)}{p(X)}, \\
\phi_U(Z;w) &:= D\,\mathbf{1}\{W \leq w\} + 1 - D,
\end{aligned}
\end{equation}
and for $k\geq 2$ the increment moments are
\begin{equation}\label{eq:incmoment}
g_{k}^{\mathrm{inc}}(Z;\theta) := \mathbf{1}\{w_{k-1}<W\leq w_k\}D - (\theta_k-\theta_{k-1}).
\end{equation}
Collecting $\{g_k ^{\text{lo}}, g_k ^{\text{up}}\}_{k = 1} ^S$ and $\{g_k ^{\text{inc}}\}_{k = 2} ^{S}$ gives a moment function $g(Z;\theta,\xi)$ of length $L=3S-1$, and the identified set \eqref{eq:idset} for $\theta_0$ is \eqref{eq:idsetgrid}. The first-stage nuisance parameter is $\xi_0 := \big(p_0, F_0 ^1(w_1\mid \cdot),\ldots,F_0 ^1(w_S\mid \cdot)\big)$, estimated by cross-fitting.
Note that even in the linear case, it is necessary construct the orthogonal score $\phi_L$ as in \eqref{eq:scores}. The two plug-in estimators of the lower bound $\ell_{0,k}$, the regression form $N^{-1}\sum_{i}\widehat F^1(w_k\mid X_i)$ and the weighting form $N^{-1}\sum_{i}D_i\mathbf{1}\{W_i\leq w_k\}/\widehat p(X_i)$, inherit the first-stage error to first order. For a nuisance value $\xi=(p,F^1)$,
\begin{equation}\label{eq:pluginbias}
\begin{aligned}
\mathrm{E}\big[F^1(w_k\mid X)\big]-\ell_{0,k} &= \mathrm{E}\big[F^1(w_k\mid X)-F_0^1(w_k\mid X)\big],\\
\mathrm{E}\bigg[\frac{D\,\mathbf{1}\{W\leq w_k\}}{p(X)}\bigg]-\ell_{0,k} &= \mathrm{E}\bigg[F_0^1(w_k\mid X)\,\frac{p_0(X)-p(X)}{p(X)}\bigg],\\
\mathrm{E}\,\phi_L(Z;w_k,\xi)-\ell_{0,k} &= \mathrm{E}\bigg[\big(F^1(w_k\mid X)-F_0^1(w_k\mid X)\big)\,\frac{p(X)-p_0(X)}{p(X)}\bigg].
\end{aligned}
\end{equation}
The first two errors are linear in a single first-stage error, typically of larger order than $N^{-1/2}$, so coverage of $\Theta_I$ is no longer guaranteed. The third error is the product of the two first-stage errors, which is $o(N^{-1/2})$ when each converges faster than $N^{-1/4}$. Every inequality in \eqref{eq:idsetgrid} is affine in $\theta$, so the orthogonal estimates of $\ell_{0,k}$ and $u_{0,k}$ could also be potentially used in the standard moment-inequality inference on the estimated system \citep{andrews2010inference,CoxShi2023}.
\subsubsection{Nonlinear case: the interquantile range}
\label{sec:example1-iqr}
The interquartile range that \citet{BGIM} report is a difference of two quantiles resulting in moment inequalities that are not affine in the parameter. Fix $0<\alpha_1<\alpha_2<1$, and suppose that the distribution of $W$ is continuous. The target parameter is the pair of population quantiles,
\begin{equation}\label{eq:iqrtarget}
\theta_0 := \big(F_0^{-1}(\alpha_1),\,F_0^{-1}(\alpha_2)\big)', \qquad F_0(\theta_{0,j})=\alpha_j, \quad j=1,2,
\end{equation}
and the interquantile range is $\theta_{0,2}-\theta_{0,1}$. The parameter space is $\Theta:=\{\theta\in[\underline w,\bar w]^2:\theta_1\leq\theta_2\}$ for an interval $[\underline w,\bar w]$ that contains the support of $W$.
Write
\begin{equation*}
F_0^{\mathrm{lo}}(w) := \mathrm{E}\big[F_0^1(w\mid X)\big], \qquad F_0^{\mathrm{up}}(w) := \mathrm{E}\big[F_0^1(w\mid X)\,p_0(X)+1-p_0(X)\big],
\end{equation*}
so that $\ell_{0,k}=F_0^{\mathrm{lo}}(w_k)$, $u_{0,k}=F_0^{\mathrm{up}}(w_k)$, and $F_0^{\mathrm{lo}}(w)\leq F_0(w)\leq F_0^{\mathrm{up}}(w)$ for every $w$. Evaluating these bounds at $w=\theta_{0,j}$, and the increment restriction \eqref{eq:increment} at the thresholds $\theta_{0,1}<\theta_{0,2}$, gives the identified set
\begin{equation}\label{eq:iqrset}
\Theta_I=\Big\{\theta\in\Theta:\; F_0^{\mathrm{lo}}(\theta_j)\leq\alpha_j\leq F_0^{\mathrm{up}}(\theta_j),\;\; j=1,2, \quad F_0^{\mathrm{up}}(\theta_2)-F_0^{\mathrm{up}}(\theta_1)\leq\alpha_2-\alpha_1\Big\}.
\end{equation}
Suppose in addition that $\min\{\alpha_1,\alpha_2-\alpha_1\}>1-{\mathrm{P}}(D=1)$, $F_0^{\mathrm{lo}}$ and $F_0^{\mathrm{up}}$ are continuous, and strictly increasing near the points at which they equal $\alpha_1$ or $\alpha_2$. The coordinate projections of $\Theta_I$ are then the quantile bounds $[\underline\theta_j,\bar\theta_j]$, where $F_0^{\mathrm{up}}(\underline\theta_j)=\alpha_j$ and $F_0^{\mathrm{lo}}(\bar\theta_j)=\alpha_j$.
The increment restriction can make the joint bound on the interquantile range strictly tighter than the difference of the two coordinate bounds. Since
\[
F_0^{\mathrm{up}}(w)-F_0^{\mathrm{lo}}(w)
=
\mathrm{E}[(1-p_0(X))(1-F_0^1(w\mid X))],
\]the box corner \((\underline\theta_1,\bar\theta_2)\), which uniquely maximizes \(\theta_2-\theta_1\) over the Cartesian product of the coordinate bounds, satisfies
\[
F_0^{\mathrm{up}}(\bar\theta_2)
-
F_0^{\mathrm{up}}(\underline\theta_1)
=
\alpha_2-\alpha_1+
\mathrm{E}[(1-p_0(X))(1-F_0^1(\bar\theta_2\mid X))].
\]Hence this corner is excluded whenever the last expectation is positive. Since \(\Theta_I\) is compact, the upper bound on the interquantile range is then strictly smaller than \(\bar\theta_2-\underline\theta_1\). This is the unconditional analogue of the cdf coherence argument used by \citet[Section~2.1.3]{BGIM} to tighten bounds on conditional within-group inequality.
In the notation of \Cref{sec:setup}, the observation is $Z_i:=(X_i,D_i,D_iW_i)$, and the moment function has five coordinates, obtained by evaluating the scores \eqref{eq:scores} at the parameter,
\begin{equation}\label{eq:iqrmoments}
\begin{aligned}
g_j^{\mathrm{lo}}(Z;\theta,\xi) &:= \phi_L(Z;\theta_j,\xi)-\alpha_j, \qquad g_j^{\mathrm{up}}(Z;\theta):=\alpha_j-\phi_U(Z;\theta_j), \qquad j=1,2,\\
g^{\mathrm{inc}}(Z;\theta) &:= D\,\mathbf{1}\{\theta_1<W\leq\theta_2\}-(\alpha_2-\alpha_1).
\end{aligned}
\end{equation}
Because
\[
\mathrm{E}\,\phi_L(Z;w,\xi_0)=F_0^{\mathrm{lo}}(w),\qquad
\mathrm{E}\,\phi_U(Z;w)=F_0^{\mathrm{up}}(w),
\]and, for \(\theta_1\leq\theta_2\),
\[
\mathrm{E}[D\,\mathbf{1}\{\theta_1<W\leq\theta_2\}]
=
F_0^{\mathrm{up}}(\theta_2)-F_0^{\mathrm{up}}(\theta_1),
\]the zero set of the population criterion generated by \eqref{eq:iqrmoments} is \eqref{eq:iqrset}. The nuisance parameter is \(\xi_0=(p_0,F_0^1)\). The bias identity in \eqref{eq:pluginbias} holds at every threshold \(w\), so the lower-bound moments are Neyman-orthogonal pointwise in \(\theta\); \Cref{ass:iqrrates} supplies the uniform control over \(\theta\) required by the asymptotic theory.
Three features of \eqref{eq:iqrmoments} distinguish the interquantile case from the linear case. First, the moments are non-affine in \(\theta\), because the quantile locations enter through both the indicator functions and \(F^1(\theta_j\mid X)\). Second, the nuisance is evaluated at the parameter, so \(F_0^1(w\mid x)\) must be estimated uniformly over \(w\). Third, \(\Theta_I\) need not be convex: the increment restriction defines a nonlinear upper boundary for \(\theta_2\) as a function of \(\theta_1\), which need not be concave. Recovering quantiles from the fixed-threshold system of \Cref{sec:example1-grid} would instead require a threshold grid whose mesh tends to zero, and therefore a growing number of inequalities. Alternatively, one could estimate \(F_0^{\mathrm{lo}}\) and \(F_0^{\mathrm{up}}\) uniformly and invert the estimated bound functions. The criterion approach avoids this separate inversion step: the quantile locations enter directly as parameters, and the zero set of the criterion yields \(\Theta_I\) even when that set is nonconvex.
\subsection{Wage gap with an interval-censored outcome}
\label{sec:example2}
Our second example is the gender wage gap when the wage is recorded only as a bracket, a classic source of partial identification \citep{ManskiTamer02}. We describe the set of coefficients consistent with the brackets by a family of moment inequalities.
The target is the coefficient on the treatment in a partially linear projection of the log wage, of which only a bracket is observed. The projection is motivated by the model
\begin{equation}\label{eq:plm-emp}
Y = D'\theta_0 + f_0(X) + U, \qquad \mathrm{E}[U\mid X, D] = 0,
\end{equation}
where $D\in\mathbb{R}^{d}$ is a low-dimensional treatment vector, $f_0$ is an integrable function, and $X \in \mathbb{R}^p$ is a vector of controls whose dimension may be large ($p \gg N$). The target parameter is the partial-linear projection coefficient $\theta_0:=\Sigma_0^{-1}\mathrm{E}[V(\eta_0)Y]$, with $V(\eta_0)$ and $\Sigma_0$ as in \eqref{eq:idset-emp}. It equals the coefficient in \eqref{eq:plm-emp} when the model holds and remains well defined when effects are heterogeneous. The latent value of $Y$ is bracketed by the observable bounds,
\begin{equation}\label{eq:bracket}
Y_L \;\leq\; Y \;\leq\; Y_U, \qquad \Delta := Y_U-Y_L,
\end{equation}
where the width $\Delta$ can be random.
The identified set is the set of partial-linear projection coefficients generated by outcomes consistent with the brackets,
\begin{equation}\label{eq:idset-emp}
\Theta_I^{\ast} \;:=\; \big\{\Sigma_0^{-1}\mathrm{E}[V(\eta_0) Y]\;:\; Y_L\leq Y\leq Y_U \ \text{a.s.}\big\},
\end{equation}
where $\eta_0(X) := \mathrm{E}[D \mid X]$, $V(\eta) := D - \eta(X)$ is the residualized treatment and $\Sigma_0 := \mathrm{E} [ V(\eta_0)V(\eta_0)']$. Let $\ell_0(X):=\mathrm{E}[Y_L\mid X]$ and $U_L := Y_L-\ell_0(X)$ be the residualized lower bracket. By definition, $\Sigma_0\theta_0=\mathrm{E}[V(\eta_0) Y]$, and decomposing $Y=Y_L+R$ with $R:=Y-Y_L$ splits this into an observed and an unobserved part,
\begin{equation}\label{eq:split}
\Sigma_0\theta_0 \;=\; \mathrm{E}[V(\eta_0) U_L] \;+\; \mathrm{E}[V(\eta_0) R],
\end{equation}
where the first term uses $\mathrm{E}[V(\eta_0)\mid X]=0$.
Bounding the unobserved term produces one supporting half-space of $\Theta_I^{\ast}$ for every direction $q$. Since $0\leq R\leq\Delta$, the bound $(q'V(\eta_0))R\leq\Delta\,(q'V(\eta_0))_{+}$ holds pointwise for every $q$, where $a_+ := \max\{0, a\}$, and taking expectations gives
\begin{equation}\label{eq:halfspace}
q'\Sigma_0\theta_0 \;\leq\; \mathrm{E}[(q'V(\eta_0))U_L] + \mathrm{E}\big[\Delta\,(q'V(\eta_0))_{+}\big].
\end{equation}
The set $\Theta_I ^*$ is convex, so it is the intersection of its supporting half-spaces: \eqref{eq:halfspace} is the supporting half-space in direction $q$, and intersecting over all $q$ on the unit sphere recovers $\Theta_I ^*$. For implementation we use a finite grid $q_1, \ldots, q_L$ and let $\Theta_I$ denote the resulting polyhedral outer approximation,
\begin{equation}\label{eq:ex2finite_set}
\Theta_I = \left\{\theta \in \Theta: q_l' \Sigma_0 \theta \leq \mathrm{E}[(q_l'V(\eta_0))U_L] + \mathrm{E}\big[\Delta\,(q_l'V(\eta_0))_{+}\big], \quad l = 1, \ldots, L \right\}.
\end{equation}
Each direction gives one moment inequality, with its own orthogonalization because its nuisance enters through $q'V(\eta)$.
In the notation of \Cref{sec:setup}, the observation is $Z_i :=(X_i, D_i, Y_{L, i}, Y_{U, i})$, $i = 1, \ldots, N$. Consider recentering the positive-part term in \eqref{eq:halfspace} by $c(X) (q' V(\eta))$ for a measurable function $c$, which leaves the expectation at $\eta_0$ unchanged because $\mathrm{E}[V(\eta_0)\mid X]=0$. Whenever the positive-part expectation is pathwise differentiable at $\eta_0$, perturbing $\eta$ in the direction $\delta$ gives the derivative at $\eta_0$,
\begin{equation}\label{eq:piq}
-\mathrm{E}[(q'\delta(X))\{\pi_q(\eta_0, X)-c(X)\}], \quad \pi_q(\eta, X) \;:=\; \mathrm{E}\big[\Delta\,\mathbf{1}\{q'V(\eta)>0\}\mid X\big],
\end{equation}
which vanishes along every direction $\delta$ at $c(X) = \pi_q(\eta_0, X)$. The orthogonal moment function is therefore
\begin{equation}
\label{eq:emp-moment}
g_q(Z;\theta,\xi) := (q'V(\eta))(V(\eta)'\theta) - (q'V(\eta))U_L - \Big[\Delta\,(q'V(\eta))_{+} - (q'V(\eta))\,\pi_q(\eta, X)\Big].
\end{equation}
Collecting $g_1, \ldots, g_L$ gives a moment function $g(Z; \theta, \xi)$ of length $L$, and the identified set \eqref{eq:idset} for $\theta_0$ is \eqref{eq:ex2finite_set}. The first-stage nuisance parameter is $\xi_0 := (\eta_0, \ell_0, \pi_{0,1}, \ldots, \pi_{0,L})$, estimated by cross-fitting.
\section{Formal results}
\label{sec:examples-theory}
The three corollaries of this section apply \Cref{alg:main} to the moment functions of \Cref{sec:example} and state, under primitive conditions, the rate of the set estimator and the coverage of the confidence region.
\subsection{Wage distribution under selection into employment}\label{sec:formal_bgim}
\Cref{ass:ex1id} collects the restrictions that define the identified set \eqref{eq:idsetgrid}. \Cref{ass:ex1ovl} is the strict overlap that the inverse propensity in \eqref{eq:scores} requires. \Cref{ass:ex1rates} places standard restrictions on the convergence rates of the nuisance parameter estimators.
\begin{assumption}[Identification]
\label{ass:ex1id}
(1) For $j=1,\ldots,S$, the employed-wage distribution stochastically dominates the unemployed-wage distribution, that is, $F_0 ^1(w_j\mid X)\leq F_0 ^0(w_j\mid X)$ almost surely. (2) The parameter space is $\Theta=[0,1]^S$; under (1) the polytope \eqref{eq:idsetgrid} contains $\theta_0$ and is non-empty.
\end{assumption}
\begin{assumption}[Overlap]
\label{ass:ex1ovl}
There exists $\underline p>0$ such that $p_0(X)\geq\underline p$ almost surely and every fold estimate satisfies $\widehat p^{(k)}(X)\geq\underline p$ almost surely, $k=1,\ldots,K$.
\end{assumption}
\begin{assumption}[Convergence Rates]
\label{ass:ex1rates}
There exist sequences $g_N^{p}$ and $g_N^{F}$ such that, with probability approaching one, $\|\widehat p^{(k)}-p_0\|_{P,2}\leq g_N^{p}$ and $\max_{j\leq S}\|\widehat F^{1(k)}(w_j\mid\cdot)-F_0 ^1(w_j\mid\cdot)\|_{P,2}\leq g_N^{F}$ for every fold $k$, and
\begin{equation*}
g_N^{p}\,g_N^{F} = o(N^{-1/2}), \qquad g_N^{p}\vee g_N^{F} = o\big((\log N)^{-1/2}\big).
\end{equation*}
\end{assumption}
\begin{corollary}[Wage Distribution under Selection]
\label{cor:ex1}
Suppose \Cref{ass:ex1id,ass:ex1ovl,ass:ex1rates} hold, and let $g$ collect the $2S$ coordinate moments built from \eqref{eq:scores} and the $S-1$ increment moments \eqref{eq:incmoment}. Let $\widehat\Theta_I=C_N(c_N;\widehat\xi)$ be the set estimator of \Cref{def:setestim} at the level $c_N=\log N$. Then the following hold:
\begin{enumerate}
\item[(1)] ${\mathrm{P}}(\Theta_I\subseteq\widehat\Theta_I)\rightarrow 1$ and $d_H(\widehat\Theta_I,\Theta_I)=O_P\big(\sqrt{\log N/N}\big)$.
\item[(2)] Let $\tau\in(0,1)$ be such that the limit law of $\sup_{\theta\in\Theta_I}N\,Q_N(\theta,\xi_0)$ puts mass less than $1-\tau$ at zero and has a distribution function that is strictly increasing in a neighbourhood of its $(1-\tau)$-quantile, and let $b\rightarrow\infty$ with $b\log N/N\rightarrow 0$. The critical value $\widehat c_\tau$ of \Cref{alg:main} then satisfies $${\mathrm{P}}\big(\Theta_I\subseteq C_N(\widehat c_\tau;\widehat\xi)\big)= 1-\tau+o(1).$$
\item[(3)] Suppose $\Theta_I$ has non-empty interior. Then $d_H\big(C_N(\widehat c';\widehat\xi),\Theta_I\big)=O_P(N^{-1/2})$ for every level $\widehat c'=O_P(1)$ that is at least $N\min_{\theta\in\Theta}Q_N(\theta,\widehat\xi)$ with probability approaching one. The level $\widehat c'=0$ is admissible whenever the sample criterion attains zero with probability approaching one.
\end{enumerate}
\end{corollary}
\Cref{cor:ex1} gives the asymptotic theory for the set estimator and the confidence region in the wage-distribution example. The rate in (1) and the coverage in (2) match the infeasible criterion $Q_N(\theta,\xi_0)$ of \citet{CHT}, and (3) is their degenerate case. The first stage converges more slowly than $N^{-1/2}$ but does not affect the first-order asymptotics: \Cref{ass:ex1rates} restricts only the product of the two first-stage errors, so an $o(N^{-1/4})$ rate for each suffices.
\subsection{Interquantile range under selection into employment}\label{sec:formal_iqr}
\Cref{ass:iqrid} extends the stochastic-dominance restriction of \Cref{ass:ex1id} to every threshold and requires the upper bound function $F_0^{\mathrm{up}}$ of \Cref{sec:example1-iqr} to increase at a positive rate near the quantile bounds. \Cref{ass:iqrrates} restates \Cref{ass:ex1rates} uniformly over thresholds. Distribution regression, which estimates $F_0^1(w\mid x)$ at each threshold $w$ by an $\ell_1$-penalized logistic regression of $\mathbf{1}\{W\leq w\}$ on $X$ among the employed, attains rates of this type uniformly over thresholds \citep{Program}. Its fitted values need not be monotone in $w$. Rearranging them in $w$ restores monotonicity and does not increase $\sup_w|\widehat F^{1(k)}(w\mid x)-F_0^1(w\mid x)|$ at any $x$ \citep{CFG2010}, so it preserves the rate condition whenever that condition holds in the stronger norm $\|\sup_w|\cdot|\|_{P,2}$.
\begin{assumption}[Identification]
\label{ass:iqrid}
(1) For every $w$, the employed-wage distribution stochastically dominates the unemployed-wage distribution, that is, $F_0^1(w\mid X)\leq F_0^0(w\mid X)$ almost surely. (2) The distribution of $W$ is continuous, its support lies in $[\underline w,\bar w]$, and given $D=1$ and $X$ the wage has a density $f_0^1(w\mid X)$ bounded by a constant $\bar f$. (3) The quantile levels satisfy $\min\{\alpha_1,\alpha_2-\alpha_1\}>1-{\mathrm{P}}(D=1)$, and there exist $\bar\delta>0$ and $\underline f>0$ such that $\mathcal{J}:=[\underline\theta_1-\bar\delta,\bar\theta_2+\bar\delta]\subset(\underline w,\bar w)$ and $\mathrm{E}[p_0(X)f_0^1(w\mid X)]\geq\underline f$ for almost every $w\in\mathcal{J}$.
\end{assumption}
\begin{assumption}[Convergence Rates]
\label{ass:iqrrates}
For every fold $k$ and every $x$, the estimate $\widehat F^{1(k)}(\cdot\mid x)$ is non-decreasing on $[\underline w,\bar w]$ with values in $[0,1]$. There exist sequences $g_N^{p}$ and $g_N^{F}$ that satisfy the rate conditions of \Cref{ass:ex1rates} and such that, with probability approaching one, $\|\widehat p^{(k)}-p_0\|_{P,2}\leq g_N^{p}$ and $\sup_{w\in[\underline w,\bar w]}\|\widehat F^{1(k)}(w\mid\cdot)-F_0^1(w\mid\cdot)\|_{P,2}\leq g_N^{F}$ for every fold $k$.
\end{assumption}
\begin{corollary}[Interquantile Range under Selection]
\label{cor:iqr}
Suppose \Cref{ass:ex1ovl,ass:iqrid,ass:iqrrates} hold, and let $g$ collect the five moments in \eqref{eq:iqrmoments}. Then conclusions (1), (2) and (3) of \Cref{cor:ex1} hold.
\end{corollary}
\Cref{cor:iqr} shows that the conclusions of \Cref{cor:ex1} continue to hold for the non-affine moment system \eqref{eq:iqrmoments}. Because the lower-bound moments evaluate \(\widehat F^1\) at the unknown quantile locations, the first-stage rate must hold uniformly over thresholds. The lower density bound in \Cref{ass:iqrid}(3), together with the upper density bound in \Cref{ass:iqrid}(2), provides the local linear error bound that is automatic for the affine system of \Cref{cor:ex1}. In particular, the transformation
\[
\Phi(\theta)
=
\big(F_0^{\mathrm{up}}(\theta_1),
F_0^{\mathrm{up}}(\theta_2)\big)
\]is locally bi-Lipschitz on the relevant region, and the transformed identified set is a convex polygon, so Hoffman's bound yields a population moment violation proportional to \(d(\theta,\Theta_I)\). Let $r(\theta) := \theta_2 - \theta_1$ and $\mathcal{R}_I := r(\Theta_I)$. Although $\Theta_I$ need not be convex, it is nevertheless connected. Specifically, under the transformation $\Phi$ used in the proof of \cref{cor:iqr}, it is the inverse image of a convex polygon, and as a result, $\mathcal{R}_I$ is an interval. Since $r$ is Lipschitz, $d_H (r(\widehat{\Theta}_I), \mathcal{R}_I) \leq \sqrt{2} d_H (\widehat{\Theta}_I, \Theta_I)$, so that the estimated interquantile-range set converges at no slower rate than the rate in \cref{cor:iqr}(1). A confidence interval for the identified interquartile range is obtained from the interval hull
\begin{equation*}
\text{CI}_\tau := \left(\inf _{\theta \in C_N(\hat{c}_\tau; \hat{\xi})} r(\theta), \sup_{\theta \in C_N(\hat{c}_\tau; \hat{\xi})} r(\theta) \right).
\end{equation*}
Whenever $C_N(\hat{c}_\tau; \hat{\xi})$ contains $\Theta_I$, $\text{CI}_\tau$ contains $\mathcal{R}_I$. Therefore, $\lim \inf _{N \rightarrow \infty} {\mathrm{P}} (\mathcal{R}_I \subseteq \text{CI}_\tau) \geq 1 - \tau$. This projection is generally conservative for the scalar target because $\hat{c}_\tau$ is calibrated to cover the entire set $\Theta_I$. By contrast, \citet{KaidoMolinariStoye} calibrate inference directly for a component or smooth function of $\theta$.
\subsection{Wage gap with an interval-censored outcome}\label{sec:formal_vira}
\Cref{ass:ex2id} imposes a standard identification condition and ensures that the finite system of directional inequalities defines a bounded polyhedron. \Cref{ass:ex2width} provides bounded envelopes for the moments. The only condition specific to the non-smooth positive-part term is \Cref{ass:ex2kink}, which requires its population remainder to be quadratic in the propensity error. Finally, \Cref{ass:ex2rates} requires the resulting first-stage bias to be \(o(N^{-1/2})\) and the \(L^2\) nuisance errors to satisfy the continuity rate.
\begin{assumption}[Identification]
\label{ass:ex2id}
(1) The matrix $\Sigma_0 := \mathrm{E}[(D - \eta_0(X))(D - \eta_0(X))']$ is non-singular. (2) The finite set $q_1,\ldots,q_L \in \mathbb{S}^{d - 1}$ is symmetric about the origin, and spans $\mathbb{R}^d$. (3) $\Theta$ is compact and contains the polyhedron cut out by the $L$ half-spaces \eqref{eq:halfspace}.
\end{assumption}
\begin{assumption}[Boundedness]
\label{ass:ex2width}
For finite constants $\bar{D}$, $\bar{Y}$, $\bar{\Delta}$ it holds that $\|D\| \leq \bar{D}$, $|Y_L| \leq \bar{Y}$ and $\Delta \leq \bar{\Delta}$ almost surely.
\end{assumption}
\begin{assumption}[Quadratic kink remainder]
\label{ass:ex2kink}
For every $l = 1, \ldots, L$ there exists $C_K < \infty$ such that for every $\eta$ with $\|\eta(X)\|\leq\bar{D}$ almost surely,
\begin{equation}
|\mathcal{K}_l (\eta)| \leq C_K \|\eta - \eta_0\|_{P, 2} ^2,
\end{equation}
where
\begin{equation*}
\mathcal{K}_l (\eta) := \mathrm{E}\left[\Delta ((q_l' V(\eta))_+ - (q_l' V(\eta_0))_+ + q_l' (\eta - \eta_0)(X) \mathbf{1} \{q_l' V(\eta_0) > 0\})\right].
\end{equation*}
\end{assumption}
\begin{assumption}[Convergence rates]
\label{ass:ex2rates}
Every fold estimate satisfies $\|\hat{\eta}^{(k)}(X)\|\leq\bar{D}$, $|\hat{\ell}^{(k)}(X)|\leq\bar{Y}$ and $0\leq\hat{\pi}_{l}^{(k)}(X)\leq\bar{\Delta}$ almost surely, $l=1,\ldots,L$. There exist sequences $g_{\eta, N}$, $g_{\ell, N}$, $g_{\pi, N}$ such that with probability approaching one for every fold $k$, $\|\hat{\eta}^{(k)} - \eta_0\|_{P, 2} \leq g_{\eta, N}$, $\|\hat{\ell}^{(k)} - \ell_0\|_{P, 2} \leq g_{\ell, N}$, $\max_{l \leq L} \|\hat{\pi}_{l}^{(k)} - \pi_{0, l}\|_{P, 2} \leq g_{\pi, N}$, and
\begin{equation*}
g_{\eta, N} (g_{\eta, N} \vee g_{\ell, N} \vee g_{\pi, N}) = o(N^{-1/2}), \quad g_{\eta, N} \vee g_{\ell, N} \vee g_{\pi, N} = o((\log N)^{-1/2}).
\end{equation*}
\end{assumption}
\begin{corollary}[Wage Gap with an Interval-Censored Wage]
\label{cor:ex2}
Suppose \Cref{ass:ex2id,ass:ex2width,ass:ex2kink,ass:ex2rates} hold, and let $g=(g_1,\ldots,g_L)$ be built from \eqref{eq:emp-moment}. Then conclusions (1), (2) and (3) of \Cref{cor:ex1} hold.
\end{corollary}
The role of \cref{ass:ex2kink} is to control observations near the kink $q' V(\eta_0) = 0$. It can be verified in different ways depending on the design. For continuous designs, a conventional margin or anti-concentration condition around zero implies the quadratic bound. In discrete designs, the same condition may instead hold through sign stability: if perturbing $\eta$ does not change the sign of $q' V(\eta)$, then the kink remainder is exactly zero. The empirical application of \cref{sec:empirical} is a particularly simple special case. Conditional on $X$, $D$ has two support points, and the estimated propensity remains in \((0,1)\). Perturbing the propensity therefore moves the two residual support points without changing their signs, so the kink remainder vanishes exactly. Moreover, with constant bracket width $\Delta$, each $\pi_l$ is available as a closed-form function of the scalar propensity. Consequently, the general rate condition stated in \cref{ass:ex2rates} reduces to $g_{\eta, N} (g_{\eta, N} + g_{\ell, N}) = o(N^{-1/2})$.
\section{Empirical application}
\label{sec:empirical}
We use the interval-augmented March Supplement of the 2015 Current Population Survey data where we intentionally bracket an observed outcome (log wage) to demonstrate our procedure. The sample is restricted to white non-Hispanic individuals aged 25--64 who worked more than 35 hours per week for at least 50 weeks, resulting in a total of $N = 32{,}523$ observations. The outcome $Y$ is the log hourly wage, and the treatment vector is $D = (\text{female}, \text{female}\times \text{college})'$. Therefore, $\theta_1$ measures the gender wage gap among non-college workers, and $\theta_2$ measures the additional gender gap differential associated with college. The conditioning vector $X$ contains 259 demographic, region, and experience controls and their interactions. We construct the interval-censored wage while retaining the observed wage as an infeasible benchmark. For each bracket width $\Delta \in \{1, 2, 3\}$, we replace $Y$ by lower and upper endpoints $Y_L$ and $Y_U$ satisfying $\Delta = Y_U - Y_L$. Thus the information available to the estimator is only that $Y \in [Y_L, Y_U]$, while observed $Y$ is retained for the subsequent analysis.
For each $\Delta$, we apply the procedure of Sections \ref{sec:example2} and \ref{sec:overview}. The first-stage nuisance functions are the propensity $\eta_{0, 1} (X) = {\mathrm{P}}(\text{female} = 1 \mid X)$, and the conditional mean of the lower bracket, $\ell_0 (X) = \mathrm{E}[Y_L \mid X]$. Both nuisance components are estimated by $\ell_1$-penalized regressions with $K=2$ fold cross-fitting and the penalty selection rule of \cite{Program}. Note that the second component of the treatment propensity satisfies $\eta_{0,2} (X) = \eta_{0,1} (X) \cdot \text{college}$. We impose the orthogonal moment inequality \eqref{eq:emp-moment} at $L = 180$ equally spaced directions on the unit circle. For inference, we form the preliminary contour at $c_N = \log N$ and calibrate the final level by the balanced-block subsampling procedure of \cref{def:subsampling}. We set $b = 160$, giving $B_N = 203$ disjoint blocks, with each block drawing the same number of observations from each cross-fitting fold. The reported confidence region uses the 0.95 empirical quantile of the block statistics.
We contrast our method with two other procedures: the estimated set based on the non-orthogonal moment function and the infeasible regression using the observed wage. \cref{tab:cps} reports the results along the estimated identified set based on the orthogonal moment function and its 95\% confidence region. The comparison between the orthogonal and non-orthogonal set estimators reveals only a marginal difference. In this application, the numerical difference is small with at most 0.003$\Delta$ difference in each coordinate projection across all bracket widths. The negligible difference can be explained by the sufficiently accurate estimates of the propensity scores in the first step. Finally, a point-estimate resulting from the infeasible regression is well-covered by the estimated sets.
\cref{fig:cps} shows the estimated sets for the three bracket widths. First, it reflects how loss of information created by interval censoring influences the size of the estimated sets. Larger brackets necessarily result in larger estimated sets. Second, it describes how beneficial joint set estimation is compared to its coordinatewise projections imposed independently. Across all three values of $\Delta$, the area of the estimated set $\widehat{\Theta}_I$ is approximately one half of the area of the set resulting from the coordinatewise independent bounds. This discrepancy is a result of dependence between the two coefficient restrictions that is discarded when restrictions are imposed independently.
Two qualifications are important. First, the directional inequalities characterize the outer identified set associated with the restrictions developed in \cref{sec:example2}; as discussed in \cite{Semenova2023}, this set need not be sharp for the underlying structural coefficient $\theta_0$. Second, implementation replaces the continuum of directions by $L = 180$ directions, so the reported estimated set is itself a finite-direction outer approximation. Increasing the number of directions might therefore produce better approximations of the outer set.
\begin{figure}[htbp]
\centering
\includegraphics[width=\textwidth]{cps_identified_sets.pdf}
\caption{\label{fig:cps}Estimated identified set and $95\%$ confidence region for the bracket widths $\Delta\in\{1, 2, 3\}$.}
\end{figure}
\begin{table}[htbp]
\centering
\small
\caption{\label{tab:cps}Bounds on the gender wage gap and its college interaction with a bracketed log wage}
\begin{tabular}{clcccc}
\toprule
$\Delta$ & & $\theta_1$ (gender gap) & $\theta_2$ (female $\times$ college) & area\,/\,box & $\widehat{c}_{\tau}$ \\
\midrule
\multirow{3}{*}{$1$}
& $\widehat\Theta_I$ & $[-1.227,\ 0.805]$ & $[-1.960,\ 2.040]$ & $0.496$ & \multirow{3}{*}{$16.61$} \\
& $95\%$ region & $(-1.363,\ 0.941)$ & $(-2.103,\ 2.184)$ & & \\
& uncorrected moment & $[-1.229,\ 0.807]$ & $[-1.957,\ 2.037]$ & & \\
\midrule
\multirow{3}{*}{$2$}
& $\widehat\Theta_I$ & $[-2.327,\ 1.736]$ & $[-3.903,\ 4.097]$ & $0.496$ & \multirow{3}{*}{$23.41$} \\
& $95\%$ region & $(-2.511,\ 1.920)$ & $(-4.073,\ 4.268)$ & & \\
& uncorrected moment & $[-2.332,\ 1.741]$ & $[-3.896,\ 4.091]$ & & \\
\midrule
\multirow{3}{*}{$3$}
& $\widehat\Theta_I$ & $[-3.187,\ 2.909]$ & $[-6.050,\ 5.950]$ & $0.496$ & \multirow{3}{*}{$24.31$} \\
& $95\%$ region & $(-3.390,\ 3.112)$ & $(-6.224,\ 6.124)$ & & \\
& uncorrected moment & $[-3.194,\ 2.916]$ & $[-6.041,\ 5.941]$ & & \\
\midrule
\multicolumn{2}{l}{observed-wage estimate} & $-0.210\ (0.008)$ & $0.033\ (0.015)$ & & \\
\bottomrule
\end{tabular}
\caption*{\footnotesize Notes. CPS 2015, $N=32{,}523$ workers, $p=259$ controls. Rows labelled $\widehat\Theta_I$ and $95\%$ region report the coordinate projections of the estimated set and of the confidence region $C_N(\widehat{c}_{\tau};\widehat{\xi})$ with $\tau=0.05$. The column area\,/\,box is the ratio of the area of $\widehat\Theta_I$ to the area of that box. The critical value $\widehat{c}_{\tau}$ is the $(1-\tau)$-quantile of $B_N=203$ disjoint block statistics of size $b=160$, balanced across the two folds. The row ``uncorrected moment'' reports the projections obtained from the non-orthogonal specification. The observed-wage estimate is infeasible and is reported for reference, with heteroskedasticity-robust standard errors in parentheses.}
\end{table}
\section{Conclusion}
\label{sec:conclusion}
This paper develops estimation and inference for partially identified models in which the moment inequalities defining the identified set depend on point-identified nuisance functions estimated in a first stage. The approach combines Neyman-orthogonal moment inequalities, cross-fitting, and the criterion-function representation of partially identified sets. Under the stated conditions, estimation of the nuisance functions is second order and the feasible contour set inherits the first-order Hausdorff rate of its infeasible counterpart. For inference, we develop a disjoint-block subsampling procedure adapted to the cross-fitted first stage. It yields asymptotically valid pointwise coverage without strengthening the first-stage small-bias requirement used for set estimation. The examples illustrate different aspects of the framework. In the wage-selection application, integrating the conditional restrictions underlying \citet{BGIM} yields unconditional moment inequalities that can accommodate a rich, machine-learned conditioning set. The interquantile-range problem shows that the same framework extends beyond affine systems. The resulting identified set need not be convex, yet it can be estimated directly as the zero set of the criterion and projected onto the interquantile range without separately inverting estimated bound functions. Finally, the interval-outcome example illustrates a different source of nonregularity. The general theory proposed herein allows continuous, discrete, or mixed treatment designs provided the population remainder associated with the positive-part kink is quadratic in the first-stage error.
The results also indicate several limits of the present analysis. The subsampling guarantees are pointwise in the data-generating process, and the number of moment inequalities is fixed in the asymptotic theory. Developing uniformly valid calibration with regularized first stages, allowing the number of thresholds or directions to increase with the sample size, and extending the theory to moment systems with more general forms of parameter-dependent non-smoothness are natural directions for further work.
\newpage