EconBase
← Back to paper

Adaptive Estimation of Aggregated Values of Conditional Linear Programs

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

101,399 characters · 15 sections · 85 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Adaptive Estimation of Aggregated Values of Conditional Linear Programs

abstractWe develop a covariate-assisted approach to partially identified parameters that are solutions to an under-identified system of linear equations with known coefficients. Examples include bounds on treatment effects, models of unemployment with state dependence, choice-theoretic models of IV, and random utility models. The boundary (i.e., support function) of the proposed identified set is represented as an average of intersections of regression functions, aggregated over the covariate distribution. We show that the boundary is a regular parameter, propose asymptotic theory, and demonstrate using an empirical application to Jobs First.

\noindentKeywords: Duality, linear programming, support function, partial identification, intersection bounds, cross-fitting, asymptotic linear representation, stochastic programming \noindentJEL Numbers: C14, C31, C54

Introduction and Motivation

Linear programming problems are ubiquitous in economics. They arise in heterogeneous treatment analysis, multivalued treatments and instruments HeckmanPinto, SalanieLee, sample selection HorowitzManski, discrete choice TebaldiECMA2023, state dependence Torgovitsky, and random utility models KitamuraStoye, among many others. In many such settings, the number of identifying restrictions is smaller than the number of parameters, so point identification fails and worst-case bounds are the natural target of inference.

Inference on bounds is challenging for two reasons. First, closed-form expressions for the bounds are rarely available, motivating case-by-case derivations in individual applications, as in KT's analysis of Jobs First and kamat2021identifying's analysis of Head Start. Second, the linear program may admit multiple optimal solutions --- “flat faces” of the identified set Shapiro1991, Dumbgen2003, HsiehShiShum which create non-differentiabilities HiranoPorter2012 that invalidate standard asymptotic arguments. Both hurdles often discourage the use of baseline covariates, even when covariates are available and would plausibly tighten the bounds.

This paper develops a general framework for estimation and inference on aggregate values of conditional linear programs. The right-hand side of the linear system is allowed to vary with observed characteristics and is treated as an unknown nuisance function to be learned from data. The target parameter is the aggregated support function, obtained by averaging the covariate-specific bounds over the covariate distribution. We establish that, when at least one continuously distributed covariate enters the right-hand side, this aggregated parameter is regular and pathwise differentiable, with an influence function that takes a closed form. The cross-fitted plug-in estimator is root-N consistent and asymptotically normal, with confidence intervals available through a Gaussian multiplier-bootstrap procedure. As a leading special case, when the signal is taken to be the doubly robust signal of Robins, the influence function we derive coincides with the efficient influence function of LuedtkeLaan for the sharp upper bound on the always-takers' share. The framework accommodates first-stage regularized regression methods or other machine-learning estimators through cross-fitting.

The paper makes three theoretical contributions. The first is a closed-form identification result. We establish, via strong LP duality applied pointwise and a Jensen-type aggregation, that the aggregated support function equals the expectation, over the covariate distribution, of the minimum inner product between the conditional right-hand side and a finite collection of dual vertices. This aggregated bound is weakly tighter than the bound based on aggregate data alone. The representation connects the support-function approach to partial identification with an aggregated intersection-bound representation, a combination that is new to the literature. The second contribution is a regular asymptotic theory. We characterize the influence function of the aggregated support function under a margin condition that is generically satisfied when the conditional moments are continuously distributed, and develop a cross-fitted plug-in estimator together with a Gaussian multiplier-bootstrap inference procedure whose uniform coverage is established over an explicit class of distributions. The third contribution is a self-contained asymptotic theory for cross-fitted envelope-regression estimators, developed in the online supplement at the level of an abstract finite index set and a vector-valued nuisance function. This theory --- comprising an oracle expansion, a Gaussian approximation, and multiplier-bootstrap validity --- strictly generalizes Theorem 1 of LuedtkeLaan, which addresses the two-vertex case, and yields the asymptotic results of the present paper as a direct specialization to the conditional linear program.

We demonstrate the proposed method by revisiting the Jobs First study of KT, who report bounds on five response-probability parameters that summarize how Connecticut's 1996 welfare reform affected women's labor-supply and welfare participation decisions. Using twenty-eight baseline covariates and an $\ell_1$-penalized first stage with woman-id cross-fitting, we recover their bounds without relying on the closed-form derivations of their Online Appendix B, providing an alternative methodological route to their main results. We then use the LP framework to address a question that KT themselves posed: whether the above-FPL opt-in response reflects substantive labor-supply adjustment or trivial earnings reductions of a few dollars from just above the poverty line. Refining the earnings grid into nine bins and applying the same identification and inference machinery verbatim, we find that the largest opt-in lower bound occurs not at the smallest reduction but for women who would have had to reduce earnings by twenty to forty percent of the Federal Poverty Line --- a substantively large adjustment. The opt-in response is therefore inconsistent with trivial rounding. Beyond this specific finding, the empirical exercise illustrates how the LP framework makes outcome-grid refinements inferentially and computationally tractable.

A central conceptual point distinguishes this paper from earlier work on linear-program bounds. The existing literature on the support function approach to partial identification BM, BMM, CCMS studies the support function at a fixed, aggregate data vector. We instead study the aggregated support function, obtained by averaging the conditional support function over the covariate distribution. None of the papers in this strand considers this object. The distinction matters because aggregation, viewed as an integration operation, smooths the non-regularity that motivates the literature's specialized inference procedures FangSantosShaikh, HsiehShiShum: covariate values at which the binding dual vertex is non-unique form a measure zero set under a mild smoothness condition on the covariate distribution and therefore do not contribute to the first-order asymptotics. As a consequence, the aggregated support function is a regular, pathwise differentiable parameter, and we do not need further regularization to restore asymptotic normality.

The notion of “flat faces” that complicates inference in the existing literature merits clarification in the present setting. In its general use, the term refers to multiplicity of solutions to the primal linear program. By LP duality, primal-solution multiplicity is equivalent to the multiplicity of Lagrange multipliers --- the dual solutions --- which corresponds to a violation of the Linear Independence Constraint Qualification. Our setting is a special case of linear systems in which the left-hand-side coefficient matrix is deterministic and does not depend on the covariates. In this special case, the derivative of the value function with respect to the conditional right-hand side depends only on the dual solution and not on the primal solution. Multiplicity of primal solutions is therefore immaterial for inference; we need only exclude the possibility of multiple dual solutions, which is the content of our key assumption.

Literature review

This paper contributes to the growing literature on bounds arising from linear programming problems and affine moment inequalities HonoreTamer, AndrewsRothPakes, FangSantosShaikh, DongHsiehShum, HsiehShiShum, JLS. KitamuraStoye employ a dual approach to test whether observed demand is consistent with a random utility model. FangSantosShaikh and AndrewsRothPakes also rely on duality for inference procedures, in both cases at the level of the population LP rather than its conditional counterpart. The closest paper in this line is HsiehShiShum, who invoke duality arguments for a broad class of linear and quadratic programming problems and develop an inference method that accommodates ties. By contrast, we study a conditional version of the linear program in which the right-hand-side covariates are continuously supported, which generically rules out ties on a set of positive measure and renders the aggregated support function regular.

The paper also contributes to the literature on debiased inference with machine learning for trimmed and bounded functionals: the covariate-assisted Lee-type procedures of SemSupp2, the nonparametric truncated-mean estimator of olma2021nonparametric, the least-squares approach to heterogeneous effects of heiler2024, the intensive--extensive margin treatment evaluation of heiler2024intensive, and the continuous-treatment extensions of lee2025. These papers typically rely on closed-form representations of the target parameter. Our framework instead focuses on settings where covariates enter through a conditional linear program with no closed-form solution, providing a general LP-based route to regular and efficient inference in partially identified problems with rich covariates.

A second strand of related work is the support-function approach to partial identification, including BM, BMM, Gafarov, SemJoE. Following BM, we adopt the random-set approach to characterize the covariate-assisted identified set, with randomness induced by the covariate distribution. Gafarov proposes a different strategy, introducing regularization to handle the non-differentiability associated with flat faces of the identified set. Our approach instead leverages continuously distributed covariates, which generically eliminate flat faces and restore regularity without the need for regularization. The two approaches are complementary: regularization offers a way forward in environments with limited covariates, while covariate assistance provides a natural route to regular inference when richer covariates are available.

The most closely related concurrent work is JLS. Eight months after the first version of the present paper was posted, JLS developed a model-agnostic covariate-assisted inference framework for partially identified causal effects using optimal transport duality, with a central role for weak duality. Both papers exploit the same fundamental observation: conditioning on covariates and averaging covariate-specific dual bounds yields a bound that is weakly tighter than the bound obtained from population-average data. The two frameworks are structurally distinct and complementary, and they make different trade-offs between sharpness and robustness. The present paper delivers semiparametrically efficient, square-root-of-N inference for the sharp aggregated bound under a margin condition and consistent first-stage estimation, with an explicit closed-form influence function. JLS relax the consistency requirement on the first stage by exploiting weak duality directly: any dual-feasible selector delivers a valid bound in expectation, regardless of whether the first-stage estimator converges to the true conditional moments. The cost of this robustness is that the resulting bound is one-sided and need not be sharp. The two approaches are thus complementary: the optimal transport framework offers broader scope (continuous outcomes, model-agnostic validity under misspecification), while the finite-LP framework of the present paper provides a direct route to semiparametrically efficient inference with explicit influence functions in the large class of discrete-outcome economic models in which a linear-system structure is available.

Setup

Consider a system of linear equations

align[align omitted — 87 chars of source]

Here, “\(\bm\beta_0 \geq 0\)” indicates that all coordinates of the parameter vector \(\bm\beta_0 \) are nonnegative. The parameter \(\bm\beta_0 = \mathbb{E} [ \bm{\Pi} ]\) is a $d$-vector summarizing the unobserved heterogeneity by taking an expectation of an unobserved random vector \(\bm{\Pi} \in \mathbf{R}^d\). The $k$-vector \(\bm{b}_0 = \mathbb{E} [ \mathbf{B} ]\) is the expectation of a vector \(\mathbf{B}\) that is an unknown yet estimable parameter. The matrix \(A\) is a known, deterministic \(k\times d\) matrix. We focus on an empirically relevant case when \(k<d\), which implies that the vector \(\bm\beta_0\) may not be point-identified.

The paper studies projections of the partially identified vector $\bm\beta_0$ onto various directions of economic interest. For example, an upper bound on the \(j\)th coordinate (\(j\in\{1,\dots,d\}\)) is obtained by solving

align[align omitted — 156 chars of source]

where $\bm e_j$ is the $j$th standard basis vector in $\mathbf{R}^d$. A corresponding lower bound can be obtained by replacing \(\max_{\bm\beta_0} \bm e_j^\top \bm\beta_0\) with \(-\max_{\bm\beta_0} -\bm e_j^\top \bm\beta_0\) under the same constraints. More generally, for a given direction $q$ on a unit sphere $$\mathcal{S}^{d-1} = \{ q \in \mathbf{R}^d, \ \|q \| =1 \},$$ the projection of the identified set onto $q$ is given by

align[align omitted — 143 chars of source]

In order to fix the ideas, we discuss several empirically relevant examples of the linear system (ref).

Motivating Examples

example[Principal stratification] Let $X$ denote a vector of baseline covariates taking values in the set $\mathcal{X}$. Let $Z$ denote a randomly assigned instrument and $Y \in \mathcal{Y} \subset \mathbf{R}^{\dim Y}$ be the vector of endogenous variables, where $Z$ and $Y$ are assumed to have finite support. Let $\mathbf{U} \in \mathcal{U}$ denote a complete characterization of unobserved heterogeneity (e.g., a collection of counterfactual outcomes) with finite support $\text{supp} \mathbf{U} = \{ \mathbf{u}_1, \dots, \mathbf{u}_{N_U} \}$. Given $\mathbf{U}$ and $Z$, the endogenous variables $Y$ must be non-random. The instrument $Z$ is assumed to be completely independent of $\mathbf{U}$ and $X$, that is \begin{align} (\mathbf{U}, X) \perp\!\!\!\perp Z. \end{align} The target parameter \(\bm\beta_0\) is the $N_U$-dimensional vector of latent response-type probabilities, \(\Pr (\mathbf{U} = \mathbf{u})_{\mathbf{u} \in \text{supp} \mathbf{U}}\). For any set $R \subseteq \text{supp} Y$, \begin{align} \Pr (Y \in R \mid Z=z) &= \sum_{\mathbf{u} \in supp \mathbf{U}} \Pr( Y \in R \mid \mathbf{U} = \mathbf{u}, Z=z )\, \Pr (\mathbf{U} = \mathbf{u} \mid Z=z) \\ &= \sum_{\mathbf{u} \in supp \mathbf{U}} \Pr( Y \in R \mid \mathbf{U} = \mathbf{u}, Z=z )\, \Pr (\mathbf{U} = \mathbf{u}). \nonumber \end{align} Special cases of this Example include IV bounds in BalkePearl1994,BalkePearl1997, \citep*{LeeBound} bounds with discrete-valued outcomes.

The linear system in Example (ref) --- and, more generally, in most examples in this paper --- encodes a conservation-of-mass condition: each equation matches the probability of an observable outcome value to a sum of latent response-type probabilities consistent with that value. The number of equations therefore equals the cardinality of the outcome support, and a finite system requires $Y$ to have finite support. When $Y$ is continuously distributed, the RHS function becomes an infinite-dimensional object and the conditional LP is replaced by an infinite-dimensional moment problem requiring optimal-transport techniques JLS. Such extensions are outside the scope of the present paper. In applied work, researchers commonly discretize continuously distributed outcomes for tractability, as in KT and our Section (ref).

When both the treatment and the outcome are binary, Example (ref) reduces to a model studied by HSC,Manski. We spell it out as a separate example.

exampleLet $Z=D=1$ be an indicator of exogenous binary treatment, and let $S$ be a binary outcome (here, $Y=S$). The propensity score equation (ref) reduces to \begin{align} \Pr (\mathbf{U} =(1, 1)) + \Pr (\mathbf{U} =(1, 0)) &= \Pr (S=1 \mid D=1) \\ \Pr (\mathbf{U} =(1, 1)) + \Pr (\mathbf{U} =(0, 1)) &= \Pr (S=1 \mid D=0) . \end{align} This system is a special case of (ref) with \begin{align} A = \begin{pmatrix} 1 & 1 & 0 & 0 \\ 1 & 0 & 1 & 0 \\ 1 & 1 & 1 & 1 \\ \end{pmatrix}, \quad \bm{b}_0 = \begin{pmatrix} \Pr (S=1 \mid D=1) \\ \Pr (S=1 \mid D=0) \\ 1 \end{pmatrix} = \begin{pmatrix} s(1) \\ s(0) \\ 1 \end{pmatrix}. \end{align}

Example (ref) serves as the running example throughout the rest of the paper. Beyond this binary-outcome benchmark, the framework also accommodates a range of applied settings.

example[Choice-theoretic model of IV] Consider a model generated by a system of treatment and outcome equations as studied in HeckmanPinto: \begin{align} D &= f_D(X, Z, V), \nonumber \\ Y &= f_Y (X, D, V, \epsilon_Y), \nonumber \\ X, V , \epsilon_Y, Z & are mutually independent. \end{align} Here $Z$ is an exogenous instrument, $D$ is an endogenous treatment, and \[ Y = \sum_{d \in \mathcal{D}} Y(d)\, \bm{1}\{D=d\} \] is the observed outcome. The response vector $\mathbf{U} = (D(1), \dots, D(n_Z))$ encodes potential treatment assignments across instrument values $Z \in \mathcal{Z}$. The target parameter is the average potential outcome for a given response type $\mathbf{u}$: \begin{align} \mathbb{E}[ Y(d) \mid \mathbf{U} = \mathbf{u}] = \frac{\sum_{y \in supp Y} y \, \Pr(Y(d)=y, \mathbf{U}=\mathbf{u})}{\Pr(\mathbf{U}=\mathbf{u})}. \end{align} The independence assumption (ref) implies, for any $d \in \mathcal{D}$ and $z \in \mathcal{Z}$, \begin{align} \Pr(D=d \mid Z=z) = \sum_{\mathbf{u} \in supp(\mathbf{U})} \Pr(D=d \mid \mathbf{U}=\mathbf{u}, Z=z)\, \Pr(\mathbf{U}=\mathbf{u}), \end{align} which is parallel to (ref). It also implies \begin{align} &\Pr(\bm{1}\{Y \in R\}\cdot \bm{1}\{D=d\} \mid Z=z) \nonumber \\ &\quad = \sum_{\mathbf{u} \in supp(\mathbf{U})} \Pr(D=d \mid \mathbf{U}=\mathbf{u}, Z=z)\, \Pr(Y(d)\in R, \mathbf{U}=\mathbf{u}), \quad z \in \mathcal{Z}. \end{align} As discussed in HeckmanPinto, these equations can be written as linear programs. If the denominator of (ref) is identified, then bounds on $\mathbb{E}[Y(d) \mid \mathbf{U}=\mathbf{u}]$ are equivalent to bounds on the numerator $\sum_{y \in \text{supp} Y} y \Pr(Y(d)=y, \mathbf{U}=\mathbf{u})$, which is a special case of (ref) with $q=(y_1, \dots, y_{\#\text{supp} Y})$ and $\bm{\beta}_0=(\Pr(Y(d)=y, \mathbf{U}=\mathbf{u}))_{y \in \text{supp} Y}$.
example[Survey Response] Consider the setup of Example (ref). In addition to the instrument and treatment, suppose we observe a survey response variable \[ S \;=\; \sum_{d \in \mathcal{D}} \sum_{z \in \mathcal{Z}} \bm{1}\{D=d, Z=z\}\, S(d, z), \] where $S(d, z)$ is a binary indicator of whether the potential outcome $Y(d)$ is observed. In contrast to job training, where the employment indicator $S(d, z)=S(d)=\bm{1}\{Y(d)>0\}$ satisfies an exclusion restriction (e.g., ChenFlores) by design, the indicator $S(d, z)$ is unrestricted. The observed data consist of $(Z, X, D, S, S\cdot Y)$, where $Y$ is assumed to have finite support. Define the unobserved heterogeneity vector $\mathbf{U}$ as the $N_Z (N_D+1)$-dimensional random vector collecting the counterfactual treatment assignments and response indicators: \[ \mathbf{U} = \bigl(D(z_1), \dots, D(z_{N_Z}), (S(d, z))_{d \in \mathcal{D},\,z \in \mathcal{Z}}\bigr)'. \] Then, analogously to (ref), we obtain \begin{align} \Pr(S=1, D=d \mid Z=z) &=\sum_{\mathbf{u} \in \mathcal{U}} \Pr(S(d, z)=1, D=d \mid \mathbf{U}=\mathbf{u}, Z=z)\, \Pr(\mathbf{U}=\mathbf{u} \mid Z=z) \\ &=\sum_{\mathbf{u} \in \mathcal{U}} \Pr(S(d, z)=1, D=d \mid \mathbf{U}=\mathbf{u}, Z=z)\, \Pr(\mathbf{U}=\mathbf{u}), \nonumber \end{align} where the last equality uses the independence $\mathbf{U} \perp\!\!\!\perp Z$. In this setup, it is most informative to focus on treatment effect parameters for the always-observed principal strata. For example, one may consider the average potential outcome for always-observed compliers: \[ \mathbb{E}\!\big[ Y(d) \,\big|\, (S(d, z)=1)_{d \in \mathcal{D},\,z \in \mathcal{Z}}, \, D(1)=1, \, D(0)=0\big]. \]

We include below a final core example drawn from the random utility literature.

example[Random utility models, \citep*{KitamuraStoye,KitamuraStoye2019}] A random utility model (RUM, \citep*{KitamuraStoye,KitamuraStoye2019}) partially identifies counterfactual demand from a repeated cross-section of prices, budgets, and observed product choices. Suppose there are $K$ goods and $J$ budgets. Given a vector $p \in \mathbf{R}^K$, the budget set is $B(p) = \{y \in \mathbf{R}^K : p'y = 1\}$ and the consumption bundle satisfies $Y \in B(p)$. Let $A$ be a binary matrix whose columns represent possible non-stochastic demand systems, $\bm{\beta}_0 \ge 0$ a probability distribution over demand systems, and $\bm{b}_0$ a vector of observed factual shares $(\Pr(Y \in V \mid P = p_j))_{j \in J,\, V \in \mathcal{V}}$. KitamuraStoye2019 propose bounds on functions of counterfactual demand; our framework delivers debiased inference for such bounds when baseline covariates $X$ are available.

Covariates

This section extends the linear program (ref) to incorporate covariates. Specifically, suppose the system can be written as the conditional linear program

align[align omitted — 89 chars of source]

where $\bm\beta_0(x) = \mathbb{E}[\bm{\Pi} \mid X=x]$ is the $d$-dimensional vector function summarizing unobserved heterogeneity, and $\bm{b}_0(x) = \mathbb{E}[\bm{B} \mid X=x]$ is the $k$-dimensional vector function that is estimable. This assumption is high level and must be verified on a case-by-case basis. Below, we revisit several motivating examples and show how the conditional linear restriction (ref) arises from first principles.

continuance{(ref)} The independence assumption (ref) implies \begin{align} &\Pr (Y \in R \mid Z=z, X=x) \\ &= \sum_{\mathbf{u} \in supp \mathbf{U}} \Pr( Y \in R \mid \mathbf{U}=\mathbf{u}, Z=z, X=x )\, \Pr(\mathbf{U}=\mathbf{u} \mid Z=z, X=x) \nonumber \\ &= \sum_{\mathbf{u} \in supp \mathbf{U}} \Pr( Y \in R \mid \mathbf{U}=\mathbf{u}, Z=z )\, \Pr(\mathbf{U}=\mathbf{u} \mid X=x). \nonumber \end{align} The independence assumption (ref) implies that (ref) holds with $\bm{b}_0(x) = (s(1, x), s(0, x), 1)'$ where $s(d, x) = \Pr (S=1 \mid D=d, X=x)$ for $d \in \{1, 0\}$ and $A$ as in (ref).

For any $q \in \mathcal{S}^{d-1}$, consider the conditional linear program

align[align omitted — 155 chars of source]

Averaging over the distribution of covariates gives

align[align omitted — 69 chars of source]

which characterizes the boundary of the covariate-assisted identified set. This set is

align[align omitted — 124 chars of source]

where $\sigma(q)$ is given by (ref)--(ref). Invoking the argument of \citet*{Artstein} (cf.\ \citet*{BM}, Definition 5, p.\ 771) yields the following Proposition:

proposition[Covariate-assisted identified set] The set $\mathcal{B}$ is convex and compact and has support function $\sigma(q)$ in (ref). Equivalently, $\mathcal{B}$ consists of points of the form $\bm{\beta}_0 = \mathbb{E}[\bm{\beta}_0(X)]$, where $\bm{\beta}_0(x)$ satisfies (ref).

Proposition (ref) shows that $\mathcal{B}$ can be characterized in two equivalent ways: (i) as the average of partially identified solutions to (ref) (random set approach, \citep*{BM,BMM2}), or (ii) as the intersection of supporting hyperplanes defined by conditional projections of (ref) (support function approach, \citep*{BM,BMM,CCMS}).

Theoretical Results

Section (ref) gives an overview of strong duality and presents a dual representation of the boundary, which facilitates estimation and inference. Section (ref) derives an influence function for the boundary.

Dual Identification

Given a direction $q \in \mathbf{R}^d$, the linear system (ref) can be written as a standard-form linear program (LP):

align*[align* omitted — 130 chars of source]

The data enter the problem only through the expectation function $\bm{b}_0(x) = \mathbb{E}[ \bm{B} \mid X=x]$. The primal optimal value $\sigma(q, x)$ is assumed finite for every covariate value $x \in \mathcal{X}$. The dual LP is

align*[align* omitted — 132 chars of source]

where $\nu \in \mathbf{R}^k$ and $\lambda \in \mathbf{R}^d$ are the dual variables associated with the equality and inequality constraints. Eliminating the slack vector $\lambda$ via $\lambda = A'\nu - q \geq 0$ gives the inequality form of the dual feasible set

align[align omitted — 46 chars of source]

which is a data-free, covariate-free convex polytope with a finite vertex set $\mathcal{T}=\mathcal{T}(q)$, that is, the set of extreme points of $\{\nu \in \mathbb{R}^k : A'\nu \geq q\}$. Because most empirical questions fix a direction $q$, we suppress the argument and write $\mathcal{T}$ for $\mathcal{T}(q)$ whenever no ambiguity arises. For the dual variable as any minimizer

align[align omitted — 103 chars of source]

$\sigma_D(q, x) = \nu_0(x)' \bm{b}_0(x)$ is the dual optimal value. For linear programs, the primal and dual optimal values coincide:

align[align omitted — 90 chars of source]

that is, strong duality holds. Aggregating $\sigma(q, x)$ over the covariate distribution gives a dual representation for $\sigma(q)$:

align[align omitted — 109 chars of source]
continuance{(ref)} The upper Fréchet–Hoeffding bound on the always-takers’ share is \[ \mathbb{E}_X \big[\min (s(0, X), s(1, X))\big] = \mathbb{E} \big[\nu_0(X)' \bm{b}_0(X)\big], \] where the expectation function is $\bm{b}_0(x) = (s(1, x), \; s(0, x), \; 1)'$. The dual value $\nu_0(x)$ reduces to \begin{align} \nu_0(x) = \begin{cases} (0, 1, 0)' & s(1, x) > s(0, x), \\[6pt] (1, 0, 0)' & s(1, x) < s(0, x), \\[6pt] w (1, 0, 0)' + (1-w)(0, 1, 0)', & s(1, x) = s(0, x), \;\; w \in [0, 1]. \end{cases} \end{align} In the first two cases, where $s(1, x) \neq s(0, x)$, the binding vertex is unique. In the third case, when $s(1, x)=s(0, x)$, the binding vertex is not unique: the dual value $\nu_0(x)$ can take any value on the line segment $\{w (1, 0, 0)' + (1-w)(0, 1, 0)': w \in [0, 1]\}$.
propositionLet $\mathcal{T}$ be the set of vertices of the dual feasible set $\{\nu \in \mathbf{R}^k: A' \nu \geq q\}$ defined in (ref). The following statements hold: (1) The basic (i.e., no-covariate) boundary is an intersection bound \[ \bar{\sigma}(q) = \inf_{\nu \in \mathcal{T}} \nu' \mathbb{E}[\bm{b}_0(X)] = \inf_{\nu \in \mathcal{T}} \nu' \bm{b}_0, \] (2) The covariate-assisted boundary is an aggregated intersection bound \begin{align} \sigma(q) &= \mathbb{E} \big[ \inf_{\nu \in \mathcal{T}} \nu' \bm{b}_0(X)\big] = \mathbb{E} \big[\nu_0(X)' \bm{b}_0(X)\big] = \mathbb{E} \big[\nu_0(X)' \mathbf{B}\big]. \end{align} (3) The covariate-assisted boundary is weakly tighter than the basic one: \begin{align} \sigma(q) = \mathbb{E} \big[\inf_{\nu \in \mathcal{T}} \nu' \bm{b}_0(X)\big] \;\leq\; \inf_{\nu \in \mathcal{T}} \nu' \mathbb{E}[\bm{b}_0(X)] = \bar{\sigma}(q), \end{align} and the covariate-assisted set $\mathcal{B}$ is weakly contained in the basic set $\bar{\mathcal{B}}$.

Proposition (ref) characterizes the boundary of the identified set for $\bm{\beta}_0$ with and without covariates. It represents the boundary (i.e., support function) as (aggregated) intersection bounds for a general class of linear programs. As we demonstrate later on, this representation facilitates estimation and inference.

Influence Function

In this section, we derive the influence function for the support function. Our first assumption requires the dual variable to be unique a.s. in covariate space. This assumption rules out “flat faces” and establishes regularity of support function.

assumption[Unique Dual Vertex] The dual minimization problem (ref) has a unique minimizer $\nu_0(x)$ for each $x \in \mathcal{X}$. In other words, for any distinct $\nu_1, \nu_2 \in \mathbf{R}^k$, the probability that both achieve the same dual value as $\nu_0(x)$ is zero: \begin{align} \Pr\!\Big( X : \exists \nu_1 \neq \nu_2 \in \mathcal{T} \ s.t.\ \nu_1' \bm{b}_0(X) = \nu_2' \bm{b}_0(X) = \nu_0(X)' \bm{b}_0(X) \Big) = 0. \end{align}

Assumption (ref) requires that the binding vertex in the set $\mathcal{T}$ is unique almost surely under $P_X$. If this condition holds, we can define the dual function mapping $\mathcal{X}$ into $\mathcal{T}$ as

align[align omitted — 178 chars of source]

This assumption is plausible if the vector $\bm{b}_0(X)$ is continuously distributed, for example, if it has an a.s. bounded density.

The following Assumption (ref) is a mild technical condition on the random variable $\mathbf{B}$. For example, if the random variable $\mathbf{B}$ is bounded a.s., this assumption is trivially satisfied.

assumption[Bounded Second Moment] The variance of the signal vector $\mathbf{B}$ is bounded in operator norm: \begin{align} \sup_{x \in \mathcal{X}} \lambda_{\max}\!\left( \mathbb{E}_P[\mathbf{B}\mathbf{B}' \mid X=x] \right) \;\leq\; \bar B. \end{align}

A statistical functional $\theta(P)$ is pathwise differentiable at $P_0$ if, for every regular parametric submodel $\{P_\varepsilon : \varepsilon \in (-\delta, \delta)\}$ passing through $P_0$ at $\varepsilon = 0$ with score $S(W)$, the map $\varepsilon \mapsto \theta(P_\varepsilon)$ is differentiable at $\varepsilon = 0$ and the derivative can be represented as

align[align omitted — 159 chars of source]

for some mean-zero, finite-variance function $\phi(W)$ that does not depend on the choice of submodel.

Pathwise differentiability underwrites regular $\sqrt{N}$-inference: a regular, asymptotically linear estimator $\widehat\theta$ with influence function $\phi$ satisfies $\sqrt{N}(\widehat\theta - \theta) = N^{-1/2}\sum_{i=1}^N \phi(W_i) + o_P(1)$ and is asymptotically normal with variance $\mathbb{V}ar(\phi)$. When pathwise differentiability fails --- as it does for the pointwise support function at a fixed $b$ when flat faces are present HiranoPorter2012 --- regular $\sqrt{N}$-inference is generally impossible without additional smoothing or regularization.

propositionSuppose Assumptions (ref) and (ref) hold. Then $\sigma(q)$ is pathwise differentiable, and \begin{align} \phi_q(W) \;=\; \nu_0(X)'\mathbf{B} - \sigma(q) \end{align} is an influence function for $\sigma(q)$. Consequently, any regular, asymptotically linear estimator $\widehat\sigma(q)$ with influence function $\phi_q$ --- the existence of which is established for the cross-fitted plug-in of Definition (ref) in Proposition (ref) --- admits the representation \begin{align} \sqrt{N}\big(\widehat\sigma(q) - \sigma(q)\big) \;=\; \frac{1}{\sqrt{N}}\sum_{i=1}^N \phi_q(W_i) + o_P(1), \end{align} and is asymptotically normal with variance $\mathbb{V}ar(\phi_q(W)) = \mathbb{V}ar(\nu_0(X)'\mathbf{B})$, where the second equality uses the fact that $\sigma(q)$ is a deterministic constant for each fixed $q$.

Proposition (ref) shows that the influence function for the boundary of the aggregated identified set depends only on the dual variable $\nu_0(x)$ and on the chosen signal $\mathbf{B}$. The key finding is that the identity of the binding dual vertex $\nu_0(x)$ can be treated as known. No correction term is needed for estimating the $\arg\min$ in (ref), paralleling the “oracle property” in LuedtkeLaan.

remark[LuedtkeLaan as a special case] In Example (ref), Assumption (ref) reduces to \[ \Pr\!\big(s(1, X) = s(0, X)\big) = 0. \] If this condition holds, the semiparametric efficiency bound for $\mathbb{E}[\min(s(1, X), s(0, X))]$ is well defined LuedtkeLaan; otherwise, regular estimators may not exist HiranoPorter2012. Let $\pi(X) = \Pr(D=1 \mid X)$ be the propensity score. Taking $\mathbf{B}$ as the doubly robust signal of Robins, the dual signal $\nu_0(X)'\mathbf{B}$ entering the influence function $\phi_q(W) = \nu_0(X)'\mathbf{B} - \sigma(q)$ of Proposition (ref) reduces to \begin{align} \nu_0(X)'\mathbf{B} \;&=\; \bm{1}\{s(1, X) > s(0, X)\}\,\cdot\,\left[\, s(0, X) \;+\; \frac{1-D}{1-\pi(X)}\,\bigl(S - s(0, X)\bigr)\,\right] \nonumber\\ &\quad +\; \bm{1}\{s(1, X) < s(0, X)\}\,\cdot\,\left[\, s(1, X) \;+\; \frac{D}{\pi(X)}\,\bigl(S - s(1, X)\bigr)\,\right]. \end{align} The right-hand side coincides, up to centering, with the efficient influence function for $\mathbb{E}[\min(s(0, X), s(1, X))]$ established in LuedtkeLaan.
remark[Robins as a special case] In Example (ref), if covariates fail to detect the sign change, i.e. \[ s(1, x) > s(0, x) \quad \text{for all } x \in \mathcal{X}, \] then $\nu_0(x) = (0, 1, 0)'$ for all $x$, and (ref) reduces to the efficient influence function of Robins for $s(0) = \mathbb{E}[s(0, X)] = \Pr(S=1 \mid D=0)$.
remark[levis2023covariateassisted as a special case] levis2023covariateassisted studies the classical Balke--Pearl bounds BalkePearl1997, which can be written as a special case of Example (ref) with binary outcomes. In their formulation, the vector of latent shares is \(\bm{\beta}_0 = (\beta_{10}, \beta_{01}, \beta_{11}, \beta_{00})\), denoting the joint distribution of potential outcomes \((Y(1), Y(0))\). The revealed-preference consistency conditions with the observed marginals can be written compactly as \[ A \bm{\beta}_0 = \bm{b}_0 \] with \[ A = \begin{bmatrix} 1 & 0 & 1 & 0 \\ 0 & 1 & 0 & 1 \\ 1 & 1 & 0 & 0 \\ 0 & 0 & 1 & 1 \end{bmatrix}, \quad \bm{b}_0 = \begin{bmatrix} P(Y=1 \mid Z=0) \\ 1-P(Y=1 \mid Z=0) \\ P(Y=1 \mid Z=1) \\ 1-P(Y=1 \mid Z=1) \end{bmatrix}. \]

Riesz Representation is often used to construct moment functions obeying orthogonality conditions, see, e.g., Newey1994. The novelty of this work is to give an example of Riesz representer that is not available in a closed form but instead is given as a solution to linear program.

remark[Riesz Representation of Support Function] Let $P_{\theta}$ be a parametric submodel of $P_W$, and let $b_{\theta}(x)$ be the expectation function. Then \begin{align} \partial_{\theta} \sigma_{\theta}(q) = \mathbb{E} \big[ \nu_0(X)' \partial_{\theta} b_{\theta}(X) \big], \end{align} so $\nu_0(x)$ is the Riesz representer of $\bm{b}_0(x)$. Adding the correction term and invoking strong duality gives \[ \sigma(q, X) + \nu_0(X)' (\mathbf{B} - \bm{b}_0(X)) = \nu_0(X)' \mathbf{B}. \]

Remark (ref) shows that $\nu_0(x)$ is the Riesz representer of the nuisance vector-function $\bm{b}_0(X)$. The last element of representation (ref) provides the adjustment term needed for the construction of the influence function. Adding it to the original (primal) moment $\sigma(q) := \mathbb{E}[\sigma(q, X)]$ yields the influence function

align[align omitted — 153 chars of source]

Illustration: The Jobs First Welfare-to-Work Experiment

In this section, we develop the Jobs First linear program end-to-end, from the MDRC randomized trial through the coefficient matrix $A$ to the finer earnings grid that the empirical analysis of Section (ref) exploits. The treatment follows KT's data construction and discretization verbatim, and adds the mechanical refinement of $A$ that is new here. Numerical results appear in Section (ref); the estimation and inference theory used below is developed in Section (ref).

Jobs First was a welfare-to-work assistance program introduced in Connecticut in the late 1990s as an alternative to the federal Aid to Families with Dependent Children (AFDC) program. In 1996, the Manpower Demonstration Research Corporation (MDRC) conducted a randomized trial in which eligible female applicants were randomly assigned either to Jobs First (treatment group; we abbreviate as “JF” throughout) or left eligible for AFDC (control group). JF replaced AFDC's unlimited eligibility rules with a $21$-month time limit and a more generous earnings disregard: recipients could keep their welfare check while working until their earnings exceeded the Federal Poverty Line (FPL). The program thus provided stronger work incentives and enforced stricter participation limits.

\paragraph{Data description.} The data structure is a special case of Example (ref). The exogenous treatment indicator is $D_i = 1$ for women assigned to JF and $D_i = 0$ for women assigned to AFDC; since compliance with random assignment is perfect in the MDRC trial, we set $Z_i = D_i$ throughout. For each woman $i$, let $E_i$ denote quarterly earnings and $W_i$ an indicator of welfare receipt. Following KT, we form a discretized earnings outcome $Y^{\text{earn}}_i$ and a binary welfare-participation outcome $Y^{\text{wel}}_i$ by

equation[equation omitted — 297 chars of source]

The combined outcome $Y_i = (Y^{\text{earn}}_i, Y^{\text{wel}}_i)$ takes six values in $\{0p,\,1p,\,2p,\,0n,\,1n,\,2n\}$. Three reporting conventions, inherited from KT, govern how these values map to latent response types:

enumerate$2p \equiv 2u$: anyone reporting above-FPL earnings while on welfare is, by definition, under-reporting. • Under AFDC, $1p$ pools two latent types, which are observationally indistinguishable, $1r$ (truthful) and $1u$ (under-reporting below FPL). • Under JF, the earnings disregard removes the incentive to under-report below FPL, so $1p = 1r$.

The baseline covariate vector $X_i$ contains $28$ variables: age, education level, number of children, family and marital status, and a quarterly history of employment, earnings, AFDC participation, and food-stamp receipt over the eight quarters prior to random assignment. The continuously distributed earnings and AFDC-receipt histories are central to the validity of the margin condition (Assumption (ref)) in this application: the conditional shares $\bm b_0(x)$ inherit non-degeneracy from these continuous components, so near-ties between competing dual vertices occur with probability zero (cf.\ Remark (ref)). Following KT, we focus on the sample of $N = 4{,}641$ women whose child-count variable is not missing.

\paragraph{Linear-programming problem.} We represent this problem as a special case of the linear system (ref). Absent restrictions there are $7\times 6 = 42$ latent response margins, corresponding to every pairing of an AFDC state with a JF state. Revealed-preference arguments spelled out in detail in KT imply that only $10$ of these margins are feasible, and one of them, $\beta_{1u,1r}$, is degenerate at one. The number of free response probabilities therefore reduces to $9$, which we collect into the vector \[ \bm{\beta}_0 \equiv \big( \beta_{0n,1r},\, \beta_{0r,0n},\, \beta_{2n,1r},\, \beta_{0r,2n},\, \beta_{0r,1r},\, \beta_{0r,1n},\, \beta_{1n,1r},\, \beta_{0r,2u},\, \beta_{2u,1r} \big)'. \] These latent shares represent the fractions of women whose counterfactual AFDC and JF states are linked by revealed preference: $\beta_{s^a, s^j} = P(\text{AFDC state} = s^a,\ \text{JF state} = s^j)$ is the joint probability that a randomly selected woman would occupy AFDC state $s^a$ and JF state $s^j$. The empirical moments $\bm{b}_0$ are differences in observed shares under JF and AFDC, \[ \bm{b}_0 \equiv \big( p^{\,\text{JF}}_{0n} - p^{\,\text{AFDC}}_{0n},\, p^{\,\text{JF}}_{1n} - p^{\,\text{AFDC}}_{1n},\, p^{\,\text{JF}}_{2n} - p^{\,\text{AFDC}}_{2n},\, p^{\,\text{JF}}_{0p} - p^{\,\text{AFDC}}_{0p},\, p^{\,\text{JF}}_{2p} - p^{\,\text{AFDC}}_{2p} \big)', \] where $p^{\,\text{AFDC}}_{s}$ and $p^{\,\text{JF}}_{s}$ denote observed shares in state $s$ under AFDC and JF, respectively. Rather than working with joint probabilities $\beta_{s^a, s^j}$ directly, we reparametrize them as conditional probabilities \[ \pi_{s^a,s^j} \;=\; P(\text{JF state} = s^j \mid \text{AFDC state} = s^a) \;=\; \beta_{s^a,s^j} / p^{\,\text{AFDC}}_{s^a}, \] where $p^{\,\text{AFDC}}_{s^a}$ is the IPW-adjusted probability of state $s^a$ in the AFDC control group. This reparametrization ensures that the coefficient matrix $A$ is non-stochastic, and restrictions can be written as \[ A \bm{\beta}_0 = \bm{b}_0, \] with the $5 \times 9$ coefficient matrix

equation[equation omitted — 264 chars of source]

Each of the five rows enforces a conservation-of-mass condition: the observed difference in the share of women in a given observable state across the two regimes must be accounted for by the latent “flows” permitted by revealed preference. Unlike KT, who solve this system by deriving closed-form expressions of the conditional probabilities (their Online Appendix B), we work directly with the linear-programming representation.

The function to be maximized reduces to interpretable expressions of structural parameters. For example, to bound the transition probability $\pi_{2n, 1r} = P(\text{JF state} = 1r \mid \text{AFDC state} = 2n)$, the fraction of women who reduce their labor supply and opt into welfare in response to Jobs First, we choose $q = e_3 = (0, 0, 1, 0, 0, 0, 0, 0, 0)'$ in (ref), the third standard basis vector, and obtain an upper bound on the joint probability $\beta_{2n, 1r}$; the lower bound corresponds to $q = -e_3$. The combinatorial bound $|\mathcal{T}(q)| \le \binom{d}{k}$ of Remark (ref) is $\binom{9}{5} = 126$, but the dual vertex set is small in practice: for $q = \pm e_3$, $\mathcal{T}(q)$ contains only a handful of vertices, well within the regime where Section (ref)'s fixed-$|\mathcal{T}|$ asymptotic theory applies.

\paragraph{Discretizing outcomes.} KT discretize earnings into three bins - zero earnings, positive earnings at or below FPL, and earnings above FPL as shown in (ref). They derive bounds for conditional probabilities analytically using closed-form expressions (see their Online Appendix B).

In discussing their empirical results, KT themselves pose the following question:

quote“[The] finding of a significant opt-in response could hypothetically reflect trivial earnings reductions from \$1 above the poverty line to exactly the poverty line.”

Answering this question requires looking inside the coarse above-FPL bin: if the measured opt-in response is concentrated among women who reduced earnings by only a few dollars from just above FPL to the FPL threshold, the response is consistent with a trivial rounding; if instead the response is spread across larger reductions, it is consistent with substantive labor-supply adjustments. This question is beyond the scope of existing discretization and requires finer partition of the outcome bins.

Dividing the above-FPL category into two sub-bins proved to be analytically more involved than the baseline with three income bins as KT show in Online Appendix, Section 6. The linear programming framework makes this and further refinements easy to implement. Each new sub-bin introduces additional transition parameters (columns) and potentially additional conservation-of-mass equations (rows), so that the coefficient matrix $A$ is replaced by a block-expanded matrix $\widetilde{A}$. One might argue that under this finer granularity regime the construction of the non-stochastic matrix $\widetilde A$ is challenging. However, the refined partition is induced from the main (coarse) one and directly inherits the revealed preference arguments. These arguments underlying each row depend only on the welfare participation and reporting status of the woman, and do not depend on the specific level of earnings within a bin. We will elaborate with a setup and give an example construction below.

\paragraph{Block expansion of $A$.}

Let $\mathcal{S} = \{s_1, \ldots, s_k\}$ denote the original partition of the outcome support into coarse bins resulting in matrix $A$. A refinement $\widetilde{\mathcal{S}}$ is any partition that subdivides each $s \in \mathcal{S}$ into one or more disjoint sub-bins; we write $\tilde s \prec s$ to mean that $\tilde s \in \widetilde{\mathcal{S}}$ is contained in $s$. The refinement $\widetilde{\mathcal{S}}$ induces a refined linear system

align[align omitted — 129 chars of source]

where $\tilde{\bm b}_0$ collects the expectations of indicators for each refined observable state, $\widetilde A$ summarizes the flows into and out of each state at a finer partition level $\widetilde{\mathcal{S}}$ and $\tilde{\bm\beta}_0$ is the vector collecting corresponding finer transition probability parameters.

Revealed preference restricts transitions at the level of latent states, which do not depend on how the outcome is partitioned. Consequently, if the coarse model rules out the transition $s^a \to s^j$, then the refined model rules out every transition $\tilde s^a \to \tilde s^j$ with $\tilde s^a \prec s^a$ and $\tilde s^j \prec s^j$. The matrix $\widetilde A$ therefore inherits a block structure from $A$: each entry $A_{ij}$ expands into a block whose rows now should reflect the sub-bins of $s^a_i$ (if any) and columns should account for new transition parameters. In Section (ref), we report main results on a finer partition case with 9 income bins. For a coarse state $s \in \mathcal{S}$, let $m(s) = |\{\tilde s \in \widetilde{\mathcal{S}} : \tilde s \prec s\}|$ denote the number of sub-bins that refine it. Our main granular specification sets $m(s) = 1$ for zero earnings (which admits no further subdivision), $m(s) = 5$ for earnings below FPL, and $m(s) = 3$ for earnings above $\mathrm{FPL}$.

Index the rows of $A$ by coarse observable states $s_r$ and its columns by the pairs $(s^a, s^j)$ that label coarse $\beta$ parameters. The refined matrix $\widetilde A$ then decomposes into blocks $\widetilde A_{r, (s^a, s^j)}$ of size $m(s_r) \times m(s^a)\, m(s^j)$, one for each entry of $A$. Two cases determine the content of each block:

enumerate• If $A_{r, (s^a, s^j)} = 0$, the block is zero: no fine-level transition contributes to the flow equation for any sub-bin of $s_r$. • If $A_{r, (s^a, s^j)} \neq 0$, the block is sparse, with entries equal to $A_{r, (s^a, s^j)}$ in positions dictated by fine-level flow conservation. When $m(s_r) = 1$, every fine transition $(\tilde s^a, \tilde s^j) \prec (s^a, s^j)$ enters the single equation, so the block is the row vector $A_{r, (s^a, s^j)} \cdot \mathbf{1}'_{m(s^a)\, m(s^j)}$. When $m(s_r) > 1$, each sub-bin $\tilde s_r \prec s_r$ has its own flow equation, and the coarse coefficient appears only in positions $(\tilde s_r, (\tilde s^a, \tilde s^j))$ for which $\tilde s_r \in \{\tilde s^a, \tilde s^j\}$, i.e. on sub-diagonals selected by the fine-level conservation accounting.

To illustrate the block expansion of $A$ concretely, consider the transition parameter $\beta_{2n,\,1r}$, where state $2n$ (above-FPL, not on welfare) is the source and state $1r$ (below-FPL, on welfare) is the destination. In the main granular specification of Section (ref), $2n$ is refined into three sub-bins $\widetilde{\mathcal{S}}_{2n} = \{b_{6}n,\, b_{7}n,\, b_{8}n\}$ and $1r$ is refined into five sub-bins $\widetilde{\mathcal{S}}_{1r} = \{b_{1}r,\, b_{2}r,\, b_{3}r,\, b_{4}r,\, b_{5}r\}$. Here and below, subscripts $1$--$5$ index the five below-FPL sub-bins and subscripts $6$--$8$ index the three above-FPL sub-bins. We choose this transition to illustrate the general case in which both source and destination states are divided into sub-bins.

Under the coarse partition, the conservation-of-mass constraint for state $2n$ reads

equation[equation omitted — 108 chars of source]

The revealed-preference restriction encoded in this row permits exactly one outflow from $2n$, to $1r$: a woman earning above the poverty line while not on welfare under AFDC may reduce her earnings below the line and take up assistance under JF. The only permitted inflow to $2n$ is from state $0r$: a woman on welfare who exits and earns above the poverty line. Following KT, we drop the conservation constraint for the reference state $1r$; the $1r$ partition therefore affects only the column structure of $\widetilde{A}$, not its row structure\footnote{Women earning in range 1 on welfare face an incentive to underreport under AFDC, and hence the observable state $1r$ pools truthful reporters and underreporters ($1r$ and $1u$). A conservation-of-mass row for this group would require separating these latent types, which are empirically indistinguishable}.

Refining the partition introduces $3 \times 5 = 15$ new transition parameters $\{\beta_{b_{k}n,\,b_{i}r}\}$ for ${k\,\in\,\{6, 7, 8\},\;\, i\,\in\,\{1, \dots, 5\}}$, one for each pair of a source sub-bin of $2n$ and a destination sub-bin of $1r$. Because the revealed preference argument underlying equation (ref) --- that a woman earning above the poverty line may reduce earnings to qualify for assistance --- operates at the level of the coarse states $2n$ and $1r$ and is independent of how finely earnings are classified within each bin, every sub-bin transition $b_{k}n \to b_{i}r$ is admissible for all $k \in \{6, 7, 8\}$ and $i \in \{1, \dots, 5\}$. Consequently, the refined conservation constraint for each source sub-bin $b_{k}n$ with $k \in \{6,7,8\}$ takes the form

equation[equation omitted — 146 chars of source]

Equation (ref) is the direct granular analogue of equation (ref): the single outflow term $-\beta_{2n,\,1r}$ expands into the sum $-\sum_{i=1}^{5} \beta_{b_{k}n,\,b_{i}r}$ over all five destination sub-bins of $1r$, while the inflow term and the right-hand side update to reflect the finer observable states.

figure[figure omitted — 2,867 chars of source]

Figure (ref) illustrates how to apply this block-expanding procedure going from the original $5 \times 9$ matrix. The researcher is free to adopt whichever granular specification they want: however, the granularity of the columns cannot be coarser than those of the rows, since the conservation-of-mass equation of the row would not be possible to pin down using coarser flow parameters over the columns. As long as this criterion is satisfied, the number of columns (i.e. the granularity of the included transition parameters) can be freely modified. Granular refinement of the destination state need not be applied uniformly across parameters. For $\pi_{2n,1r}$ one may resolve the destination at the fine grid, while for $\pi_{1n,1r}$ the coarse grouping suffices; the choice is dictated by which transition probability is the object of interest. Figure (ref) shows sequentially more granular design specifications: Panel (b) shows an example of $13 \times 25$ matrix, where only 2n, 1n and 2p latent types are granularized, with a choice of $\beta_{2n, 1r}$, $\beta_{0r, 2n}$, $\beta_{0r, 1n}$, $\beta_{1n, 1r}$, $\beta_{0r, 2u}$ and $\beta_{2u, 1r}$ being split into $\beta_{b_kn, 1r}$, $\beta_{0r, b_kn}$, $\beta_{0r, b_in}$, $\beta_{b_in, 1r}$, $\beta_{0r, b_ku}$ and $\beta_{b_ku, 1r}$ with $k=6,7,8$ and $i=1,2,3,4, 5$. Panel (c) shows further splitting into $13 \times 33$ matrix, and Panel (d) corresponds to (ref) when both source and destination bins are split (group $G_3$).

\paragraph{Welfare Bounds.} The conditional linear program framework bounds any linear functional of $\bm\beta_0(x)$, not only individual transition probabilities. Bounding the aggregate welfare gain induced by the Jobs First reform is a leading example, corresponding to a particular choice of the objective vector $q$ in (ref), with the constraint set $A\bm\beta_0(x) = \bm b_0(x)$, $\bm\beta_0(x) \geq 0$ unchanged.

For a woman with covariates $x$, her monthly disposable income in state $s$ under regime $t \in \{a, j\}$ is $ I^{t}(s,\, x) \;\coloneqq\; y_s + G^{t}(y_s,\, x)$, where $y_s$ denotes the representative earnings in bin $s$ and $G^{t}(\cdot,\, x)$ is the transfer schedule as in KT under regime $t$. Let $\mathcal{M}$ denote the set of all transitions admissible by revealed preference restrictions. Abstracting from other forms of income and assistance, such as food stamps and the EITC, for each admissible transition $(s^a, s^j) \in \mathcal{M}$, let the welfare gain be the induced change in monthly disposable income $ \Delta_{(s^a,\, s^j)}(x) \;\coloneqq\; I^{j}(s^j,\, x) - I^{a}(s^a,\, x). $ The aggregate expected welfare gain attributable to the transitions in $\mathcal{M^*} \subseteq \mathcal{M}$ is

equation[equation omitted — 223 chars of source]

Because $q$ is fixed in our framework, we evaluate the welfare gain at a representative household, abstracting from heterogeneity in covariates in (ref). This approximation discards within-bin variation in disposable income, and its cost diminishes under designs with finer partitions: as bins become more homogeneous, the representative household evaluation more closely approximates the covariate-specific gains it replaces. The welfare weights are then fixed, and the functional is recovered by an objective vector $q$ that does not depend on $x$, indexed conformably with $\bm\beta_0$ and with coordinate \[ q_{(s^a,\,s^j)} \;=\; \Delta_{(s^a,\,s^j)}\; \mathbf{1}\!\left\{(s^a, s^j) \in \mathcal{M^*}\right\}. \] corresponding to welfare gain value of transition $(s^a, s^j)$.

\paragraph{From unconditional to conditional probabilities.} The LP (ref) is solved over the vector of joint probabilities $\bm{\beta}_0$, but the parameters of economic interest are the conditional transition probabilities $\pi_{s^a, s^j}$ introduced above, which is what we report in Section (ref). The two are linked by $\pi_{s^a, s^j} = \beta_{s^a, s^j} / p^{\,\text{AFDC}}_{s^a}$, where $p^{\,\text{AFDC}}_{s^a} = P(\text{AFDC state} = s^a)$ is a separately and consistently estimable scalar that is bounded away from zero in the data. A bound on $q' \bm{\beta}_0$ is therefore converted to a bound on the corresponding conditional probability by dividing through by $\widehat p^{\,\text{AFDC}}_{s^a}$.

This division has two consequences for the asymptotic theory of Section (ref). First, the influence function for the conditional probability $\pi_{s^a, s^j}$ is obtained from the influence function $\varphi_q(W)$ of $q' \bm{\beta}_0$ by the delta method: \[ \varphi_{\pi_{s^a, s^j}}(W) \;=\; \frac{1}{p^{\,\text{AFDC}}_{s^a}} \!\left[\, \varphi_q(W) \;-\; \pi_{s^a, s^j} \big(\mathbf{1}\{\text{AFDC state} = s^a\} - p^{\,\text{AFDC}}_{s^a}\big) \right]. \] The first term is the contribution of estimating $q' \bm{\beta}_0$; the second term accounts for the estimated denominator $\widehat p^{\,\text{AFDC}}_{s^a}$. Both contributions are $\sqrt{N}$-consistent and asymptotically Gaussian under the assumptions of Section (ref), since $p^{\,\text{AFDC}}_{s^a}$ is a sample mean of an indicator and is bounded away from zero in the data. Second, the cross-fitted plug-in estimator $\widehat{\sigma}(q) = N^{-1}\sum_{i=1}^N \widehat{\nu}(X_i)' \mathbf{B}_i$ of Definition (ref) is converted to an estimator of the bound on $\pi_{s^a, s^j}$ by dividing by $\widehat p^{\,\text{AFDC}}_{s^a}$. The reported bounds and confidence intervals in Section (ref) reflect this rescaling.

Estimation and Inference

This section develops the estimation and inference theory for the covariate-assisted boundary $\sigma(q)$. Section (ref) introduces the cross-fitted dual plug-in estimator and a multiplier-bootstrap procedure for constructing confidence intervals. Section (ref) states the assumptions and the main asymptotic result (Proposition (ref)), establishing $\sqrt{N}$-consistency, asymptotic normality, and uniform bootstrap coverage. Section (ref) discusses the assumptions, the role of the margin condition, and robust alternatives that remain valid when the margin condition fails.

The Estimator

In this Section, we introduce the dual estimator of the boundary assuming the conditions of Proposition (ref) hold. Let $\mathbf{B}$ be an observable function of data obeying $\bm{b}_0 = \mathbb{E}[\mathbf{B}]$, and let $(W_i)_{i=1}^N$ be an i.i.d.\ sample of size $N$. The first step is to construct the fitted values for the expectation function $\bm{b}_0(x)$. The second step is to construct an estimate of the boundary $\sigma(q)$.

definition[Primal and Dual Cross-Fitted Values] \begin{enumerate} • For a random sample of size $N$, denote a $K$-fold random partition of the sample indices $[N]=\{1, 2, \dots, N\}$ by $(J_k)_{k=1}^K$, where $K$ is the number of partitions and the sample size of each fold is $n=N/K$. For each $k \in [K] = \{1, 2, \dots, K\}$ define $J_k^c = \{1, 2, \dots, N\} \setminus J_k$. • For each $k \in [K]$, construct an estimator $\widehat{\bm{b}}_k = \widehat{\bm{b}}(W_{i \in J_k^c})$ of the nuisance parameter $\bm{b}_0$ using only the data $\{W_j: j \in J_k^c\}$. For any observation $i \in J_k$, define the primal fitted value as $\widehat{\bm{b}}_i = \widehat{\bm{b}}_k(X_i)$ and the dual plug-in fitted value \begin{align} \widehat{\nu}(X_i) = \sum_{\nu \in \mathcal{T}} \nu \, \bm{1}\!\left\{ \nu \in \arg \min_{\tilde\nu \in \mathcal{T}} \tilde\nu' \widehat{\bm{b}}_i \right\}. \end{align} \end{enumerate}
definition[Dual Plug-in Estimator] Let $(\widehat{\nu}'(X_i))_{i=1}^N$ be the dual cross-fitted values. Define \begin{align} \widehat{\sigma}(q) := N^{-1} \sum_{i=1}^N \widehat{\nu}'(X_i)\, \mathbf{B}_i. \end{align}

The plug-in $\widehat\nu(X_i)$ in (ref) is one of several reasonable estimators of $\nu_0(x)$; alternatives based on direct classification of the binding vertex or on smoothing of the inner minimization are natural extensions.

definition[Multiplier Bootstrap] Let $(e_i)_{i=1}^N$ be i.i.d.\ exponential random variables $e_i \sim \text{Exp}(1)$ independent of the data. Define the bootstrap analog of $\widehat{\sigma}(q)$ as \begin{align} \widetilde{\sigma}(q) := N^{-1} \sum_{i=1}^N \widehat{\nu}'(X_i)\, \mathbf{B}_i e_i. \end{align}

The multiplier bootstrap is computationally efficient: the $K$-fold cross-fitted nuisance $\widehat\nu(X_i)$ and the signal $\mathbf{B}_i$ are held fixed across bootstrap replications, and each replication reduces to a reweighted average.

Under sparsity or smoothness conditions on $\bm{b}_0(x)$, the dual estimator of the support function enjoys the following properties pointwise in $q \in \mathcal{S}^{d-1}$:

enumerate• For each fixed $q \in \mathcal{S}^{d-1}$, the estimator is consistent at the parametric rate: \begin{align} \big| \widehat{\sigma}(q) - \sigma(q) \big| = O_P(N^{-1/2}) = o_P(1). \end{align} • For each fixed $q \in \mathcal{S}^{d-1}$, the estimator $\widehat{\sigma}(q)$ is asymptotically Gaussian: \begin{align} S_N(q) := \sqrt{N}\,\big(\widehat{\sigma}(q) - \sigma(q)\big) = \mathbb{G}_N(q) + o_P(1). \end{align} • The estimator $\widehat{\sigma}(q)$ can be used to construct pointwise confidence intervals. Taking a $(1-\tau)$-pointwise confidence region (CR) as \begin{align*} [i(q), \bar{i}(q)] := \big[-\widehat{\sigma}(-q) - N^{-1/2} \widehat{c}_{1-\tau/2}(-q), \;\; \widehat{\sigma}(q) + N^{-1/2} \widehat{c}_{1-\tau/2}(q)\big], \end{align*} where the critical value $\widehat{c}_{1-\tau/2}(q)$ is the $(1-\tau/2)$ quantile of the absolute bootstrap statistic $|\widetilde{S}_N(q)|$, with \begin{align*} \widetilde{S}_N(q) := \sqrt{N}\,\big(\widetilde{\sigma}(q) - \widehat{\sigma}(q)\big). \end{align*} The lower endpoint subtracts the critical value from the estimated lower-bound boundary $-\widehat{\sigma}(-q)$, while the upper endpoint adds the critical value to the estimated upper-bound boundary $\widehat{\sigma}(q)$, giving a symmetric Wald-type interval for $q'\bm{\beta}_0 \in [-\sigma(-q), \sigma(q)]$. That is, for each fixed $q \in \mathcal{S}^{d-1}$ and each $P \in \mathcal{P}$, \begin{align} \liminf_{N \to \infty} \Pr_{P}\!\big(q' \bm{\beta}_0 \in [i(q), \bar{i}(q)]\big) \;\geq\; 1 - \tau. \end{align}

Asymptotic Theory

In this Section, we describe the assumptions and state the asymptotic results for the proposed estimation and inferential procedures.

assumption[First-Stage Rate] There exists a sequence $\phi_N = o(1)$ and a sequence of sets $\{B_N, N \geq 1\}$ such that the first-stage estimates $\widehat{\bm{b}}(x)$ of the true function $\bm{b}_0(x)$ belong to $B_N$ with probability at least $1 - \phi_N$. The sets $B_N$ shrink at the following rate: \begin{align} \sup_{b \in B_N}\;\sup_{x \in \mathcal{X}} \| \bm{b}(x) - \bm{b}_0(x) \| = o(N^{-1/4}). \end{align}

Assumption (ref) requires that the estimates $\widehat{\bm{b}}(x)$ of the vector function $\bm{b}_0(x)$ converge in $\ell_{\infty}$-norm at a sufficiently fast rate. A mean-square version of this assumption frequently occurs in the semiparametric literature (see, e.g., Newey1994, chernozhukov2016double). Examples of estimators achieving $\ell_{\infty}$ rates include $\ell_1$-regularized estimators in \citet*{Program} under sparsity conditions on the linear or logistic approximations of the coordinates of $\bm{b}_0(x)$.

remark[Verification of Assumption (ref) for Example (ref)] Each coordinate of $\bm{b}_0(x) = (b_0^{(1)}(x), \ldots, b_0^{(k)}(x))'$ in Example (ref) is a conditional probability, \[ b_0^{(j)}(x) \;=\; \Pr\!\big(Y \in R_j \,\big|\, Z = z_j,\, X = x\big), \qquad j = 1, \ldots, k, \] and is therefore bounded in $[0,1]$. Suppose each coordinate admits a sparse logistic-regression approximation on a $p$-dimensional dictionary $B(x) = (B_1(x),\ldots,B_p(x))'$ (possibly with $p \gg N$): \[ b_0^{(j)}(x) \;=\; \Lambda\!\big(\gamma_j'\,B(x) + r_j(x)\big), \qquad \Lambda(t) = \frac{1}{1 + e^{-t}}, \] where $\gamma_j \in \mathbb{R}^p$ is at most $s_\gamma$-sparse and the approximation error obeys, uniformly in $j$, \[ \sup_{j \le k}\;\Big(\frac{1}{N}\sum_{i=1}^N r_j^2(X_i)\Big)^{1/2} \;\lesssim_P\; \sqrt{\frac{s_\gamma^2\,\log p}{N}}. \] Estimate each coordinate by $\ell_1$-penalized logistic regression with a single, uniform-in-$j$ penalty level chosen as in \citet*{Program}; equivalently, run $k$ parallel logistic-Lasso regressions sharing the same regularizer, scaled so that the union bound over the $k$ score processes is absorbed into a $\log(pk)$ factor. As shown in \citet*{Program}, uniformly in $j$, \[ \sup_{x \in \mathcal{X}}\;\big|\widehat{b}_0^{(j)}(x) - b_0^{(j)}(x)\big| \;=\; O_P\!\left(\sqrt{\frac{s_\gamma^2\,\log(pk)}{N}}\right). \] For a vector with fixed dimension $k$, the Euclidean norm and the worst-coordinate sup-norm are equivalent up to a $\sqrt{k}$ factor: \[ \sup_{x \in \mathcal{X}} \|\widehat{\bm{b}}(x) - \bm{b}_0(x)\|_2 \;\le\; \sqrt{k}\,\sup_{x \in \mathcal{X}}\;\max_{j \le k}\,\big|\widehat{b}_0^{(j)}(x) - b_0^{(j)}(x)\big|, \] so the same rate carries over. Define $B_N$ as the set of vector-valued logistic models whose coefficients lie within an $\ell_1$-ball of radius $r_N = c\sqrt{s_\gamma^2 \log(pk)/N}$ around $(\gamma_1,\ldots,\gamma_k)$; with probability at least $1 - \phi_N$, $\widehat{\bm{b}} \in B_N$. Hence, the sparsity-and-design condition \begin{align} s_\gamma^2\,\log(pk) \;=\; o \big(N^{1/2}\big) \end{align} is sufficient for $\sup_{x}\|\widehat{\bm{b}}(x) - \bm{b}_0(x)\|_2 = o(N^{-1/4})$, which verifies Assumption (ref). The condition (ref) is the natural multivariate generalization of the standard high-dimensional logistic-Lasso requirement; for fixed $k$ it reduces to the familiar $s_\gamma^2 \log p = o(N^{1/2})$.
remark[Bounded image, unbounded $\mathcal{X}$] No support condition on $X$ is required for the verification above. Each coordinate satisfies $b_0^{(j)}(x) \in [0,1]$ regardless of the support of $X$, and logistic-Lasso fitted values lie in $(0,1)$, so the sup-norm rate is a property of the function class rather than of the geometry of $\mathcal{X}$. For nonparametric first stages on unbounded $\mathcal{X}$, the same conclusion follows after a $\sqrt{\log N}$-truncation onto a growing compact set $\mathcal{X}_N = \{x : \|x\| \le C\sqrt{\log N}\}$: under sub-Gaussian tails on $X$, $\Pr(X \notin \mathcal{X}_N) = o(N^{-1/2})$, so observations outside $\mathcal{X}_N$ contribute negligibly to the first-order asymptotics, and it suffices to verify the $\ell_\infty$ rate uniformly over $\mathcal{X}_N$.
assumption[Smooth Distribution of Covariates] The covariate distribution of $\bm{b}_0(X)$ satisfies the smoothness condition \begin{align} \Pr \left( 0 < \inf_{\nu_1 \neq \nu_2, \quad \nu_1, \nu_2 \in \mathcal{T}} \dfrac{| (\nu_1 - \nu_2)' \bm{b}_0(X) |}{\|\nu_1 - \nu_2\|} < t \right) \leq B_f t, \quad for some t \in [0, \eta). \end{align}

Assumption (ref) ensures that the distribution of $\bm{b}_0(X)$ is sufficiently smooth. In particular, it rules out degeneracy and requires the coordinates of $\bm{b}_0(X)$ to be linearly independent with probability one. Assumption (ref) is a version of margin assumption that is commonly imposed in debiased inference, e.g., QianMurphy,heiler2024intensive,SemSupp2.

remark[Verification of Assumption (ref) for Example (ref)] The dual polytope has a finite vertex set $\mathcal{T} = \{(0, 0, 0)',\,(1, 0, 0)',\,(0, 1, 0)'\}$, so the infimum in (ref) is a minimum over the three pairwise differences. For $q=(1, 0, 0)'$, the set $\mathcal{T}$ could be reduced to two elements $\{\,(1, 0, 0)',\,(0, 1, 0)'\}$ giving $|(\nu_1 - \nu_2)'\bm{b}_0(X)| = |s(1, X) - s(0, X)|$. Assumption (ref) therefore reduces to the margin condition \begin{align} \Pr\!\big(0 < |s(1, X) - s(0, X)| \le t\big) \;\le\; \bar B_f\,t, \qquad 0 \le t \le \eta, \end{align} which holds as long as the random variable $s(1, X) - s(0, X)$ has a Lebesgue density that is bounded by some $\bar B_f < \infty$ on $[-\eta, \eta]$.

We now state the main asymptotic theory result.

propositionSuppose Assumptions (ref), (ref), and (ref) hold. Then, for each fixed $q \in \mathcal{S}^{d-1}$, the dual cross-fitted plug-in estimator of the support function $\widehat{\sigma}(q)$ satisfies the consistency rate (ref), the Gaussian approximation (ref), and the coverage property (ref).

Proposition (ref) is our second main result. It guarantees that the dual plug-in estimator is consistent and asymptotically Gaussian, and that valid confidence regions can be constructed via the bootstrap procedure in Definition (ref). The proof is given in Online Supplement (ref) (Lemmas (ref)--(ref)).

Discussion of the Assumptions and the Results

An applied practitioner may have two distinct reasons to work with a conditional linear program rather than its unconditional counterpart. The first is validity: when identification of $\bm{b}_0$ requires conditioning on baseline covariates --- for example, when an instrument or treatment is exogenous only conditional on $X$ --- covariates are unavoidable for correct identification. The second is power: even when validity does not require conditioning, covariate aggregation tightens the identified set through Jensen's inequality. The remarks below focus on the second reason. Several of the points generalize a discussion of ponomarev2024 that was originally phrased for the optimal-welfare problem and applies, with minor modification, to the present partial-identification setting.

The next two remarks describe two distinct classes of procedures that remain valid without a margin condition, at the cost of power.

remark[Robust procedures via moment inequalities] Let $G$ be any measurable partition of $\mathcal{X}$, and let $\nu_G : \mathcal{X} \to \mathcal{T}$ be any selector that maps each covariate value to a candidate dual vertex. By dual feasibility, \[ \sigma(q) \;=\; \mathbb{E}\!\left[\min_{\nu \in \mathcal{T}} \nu' \bm{b}_0(X)\right] \;\le\; \mathbb{E}\!\left[\nu_G(X)' \bm{b}_0(X)\right] \;=\; \mathbb{E}\!\left[\nu_G(X)' \mathbf{B}\right], \] so each selector $\nu_G$ delivers a valid (typically non-sharp) bound on $\sigma(q)$. Stacking the inequalities induced by a finite collection of selectors $\{\nu_{G_1}, \ldots, \nu_{G_J}\}$ yields a moment-inequality system that can be tested using existing procedures CCK,canay2017practical,romano2014practical. The resulting confidence band is uniformly valid without any margin condition; the price is that the implied bound is no longer sharp. As discussed in ponomarev2024, in finite samples a non-sharp bound based on a well-chosen selector may even produce a tighter confidence band than the sharp bound based on $\nu_0(X)$, because the variance reduction from a more stable selector can outweigh the bias from non-sharpness.
remark[Robust procedures via smoothing] A complementary class of robust procedures replaces the $\min_{\nu \in \mathcal{T}}$ operator with a soft-min approximation, e.g.\ the log-sum-exp \[ \mathrm{softmin}_\beta(\nu' \bm{b}_0(x)) \;=\; -\beta^{-1} \log \!\left(\sum_{\nu \in \mathcal{T}} \exp\!\big(-\beta\, \nu' \bm{b}_0(x)\big)\right), \] with a temperature parameter $\beta = \beta_N \to \infty$. This route is taken by whitehouse2025 in policy learning and by levis2023covariateassisted for IV bounds. Smoothing restores pathwise differentiability and yields regular $\sqrt{N}$-inference without a margin condition, at the cost of a regularization bias of order $\beta_N^{-1} \log |\mathcal{T}|$. The bias scales logarithmically with the size of the dual vertex set, which makes this approach attractive when $|\mathcal{T}|$ is moderate but increasingly costly as $|\mathcal{T}|$ grows.
remark[Margin condition and the choice of inferential target] When the margin condition (Assumption (ref)) fails, the sharp bound $\sigma(q)$ may not be the power-optimal inferential target. ponomarev2024 formalize this point in the welfare context: they exhibit a class of DGPs at which the sharp bound is first-order dominated --- in expected confidence-band length --- by a bound based on a strictly suboptimal selector, and they show that such first-order dominance is possible if and only if the margin condition fails. The same logic applies to the conditional LP setting: when $\bm{b}_0(X)$ concentrates near a tie between competing dual vertices, the variance of the plug-in $\widehat\nu(X)' \mathbf{B}$ is inflated by the instability of the argmin, and a smoother target may dominate. We therefore view the margin condition not as a mild technicality but as a substantive precondition for the sharp bound to be the right thing to aim at.
remark[Cross-fitting and the margin condition] Cross-fitting and the margin condition serve distinct purposes and are not substitutes. The margin condition (Assumption (ref)) ensures pathwise differentiability of $\sigma(q)$ and the existence of a regular, $\sqrt{N}$-asymptotically normal estimator. Cross-fitting controls overfitting bias from the first-stage estimate $\widehat{\bm b}(\cdot)$, ensuring that the $o(N^{-1/4})$ rate threshold of Assumption (ref) is met when nonparametric or high-dimensional methods are used. For parametric or $\ell_1$-regularized first stages with standard sparsity and Donsker-type complexity conditions, in-sample estimation is sufficient and cross-fitting is optional; for general machine-learning first stages, cross-fitting is essential. The key point is that cross-fitting cannot rescue Wald-type inference when the margin condition fails. As ponomarev2024 establish, if $\bm{b}_0(X)$ admits a tie among competing dual vertices on a set of positive measure, then no regular estimator of $\sigma(q)$ exists HiranoPorter2012, and the cross-fitted plug-in has a non-standard, heavy-tailed limit. Wald-type confidence intervals centered at $\widehat{\sigma}(q)$ are then generally invalid regardless of the first-stage quality, and the alternatives in Remarks (ref) and (ref) should be used instead. Conversely, the margin condition without cross-fitting is also insufficient when the first stage is high-dimensional: overfitting bias enters at first order and breaks the influence-function representation. Both ingredients are needed.
remark[Large-scale linear programs] The asymptotic theory of Section (ref) treats the dual vertex set $\mathcal{T}$ as fixed in $N$. The combinatorial bound $|\mathcal{T}| \le \binom{d}{k}$, which holds for an LP in standard form with $k$ equality constraints in $d$ non-negative variables under the genericity condition that every $k$ active columns of $A$ are linearly independent, is sharp; in the Jobs First application of Section (ref), $d = 9$ and $k = 5$, so $|\mathcal{T}|$ can be as large as $126$. As the LP grows --- through finer outcome discretization --- two things happen. First, Assumption (ref) becomes more demanding, because the union bound across the $|\mathcal{T}|$ candidate vertices enters the first-stage rate through a $\log |\mathcal{T}|$ factor (cf.\ Remark (ref)). Second, the margin condition (Assumption (ref)) imposes one tail bound for every pair of distinct vertices, so the practical chance of a near-tie grows with $|\mathcal{T}|$. Both effects argue for keeping the LP no larger than the science of the problem requires; when finer discretization is desired but $|\mathcal{T}|$ becomes large, the soft-min smoothing of Remark (ref) or a moment-inequality relaxation of Remark (ref) are natural retreats. A formal analysis allowing $|\mathcal{T}|$ to grow with $N$ is left for future work.
remark[Strong and weak duality] Section (ref) established strong LP duality: at every $x$, the primal value $\sigma(q,x)$ equals the dual value $\nu_0(x)' \bm{b}_0(x)$, and aggregation gives the representation $\sigma(q) = \mathbb{E}[\nu_0(X)' \mathbf{B}]$ of Proposition (ref). This identity holds without further restriction; what requires unique identification of the binding dual vertex $\nu_0(X)$ --- i.e., the margin condition of Assumption (ref) --- is regular $\sqrt{N}$-inference on $\sigma(q)$, not sharpness of the bound itself. Weak duality is a strictly weaker but more robust property: any dual-feasible selector $\nu_G : \mathcal{X} \to \mathcal{T}$ --- not necessarily the true argmin --- satisfies \[ \sigma(q) \;=\; \mathbb{E}\!\left[\min_{\nu \in \mathcal{T}} \nu' \bm{b}_0(X)\right] \;\le\; \mathbb{E}\!\left[\nu_G(X)' \bm{b}_0(X)\right], \] which is the source of validity for the moment-inequality procedure of Remark (ref). The two regimes encode the sharpness--robustness trade-off discussed in Remark (ref): strong duality delivers the sharp bound together with regular inference under the margin condition, while weak duality delivers a non-sharp but uniformly valid bound without it. A practical consequence of weak duality is robustness to first-stage misspecification. JLS exploit this in their optimal-transport framework to obtain inference that remains valid even when the first-stage estimator of $\bm{b}_0(\cdot)$ is inconsistent. The same property holds here: because $\widehat\nu(X_i) \in \mathcal{T}$ by construction, the cross-fitted plug-in $N^{-1} \sum_{i=1}^N \widehat\nu(X_i)' \mathbf{B}_i$ remains a valid upper bound on $\sigma(q)$ in expectation regardless of whether $\widehat{\bm b}(\cdot)$ converges to $\bm{b}_0(\cdot)$. What inconsistent first stages give up is sharpness, not validity --- a useful guarantee in finite samples and under model uncertainty. For the discrete-outcome examples of Section (ref), the LP dual and the optimal-transport dual of JLS coincide, and the two procedures are numerically identical.

Empirical Application

Section (ref) provides bound results in the empirical setup of Jobs First dataset studied by KT.

Jobs First Application

\paragraph{Coarse partition.} We revisit the empirical setup of KT, illustrated earlier in Section (ref). We apply the proposed framework to the same sample of $4{,}641$ women who were randomized into Jobs First program. All parameters follow the convention introduced in Section (ref), with this section reporting on $5 \times 9$ matrix design. Table (ref) reproduces Table 5 bounds under the three-bin original partition with an OLS first-stage estimator on the full sample with no cross-fitting (see Remark (ref)). Our results are consistent with a substantial intensive-margin opt-in response. The parameter of primary economic interest is the intensive-margin labor-supply response $\pi_{2n,\,1r}$: we report that at least $28.8\%$ of women earning above-FPL under AFDC decrease their labor supply as a response to the reform to qualify for transfers under JF. This lower bound is marginally tighter than the corresponding KT estimate of $28\%$; accounting for sampling uncertainty yields a conservative 95 percent confidence interval lower limit of $14.6\%$. The upper bound tightens from the uninformative value of $100\%$ to $89\%$, though it remains economically wide.

A second opt-in response is identified among women who would not have worked under AFDC. The estimated bounds for $\pi_{0n,\,1r}$ are $\{0.149,\, 0.581\}$, with a conservative 95 percent confidence interval of $[0.052,\, 0.891]$. The lower bound of $0.149$ exceeds the corresponding KT estimate of $0.055$, a difference attributable to the conditioning on covariates.

JF also generated a substantial participation response among the below-FPL earners. The estimated bounds for $\pi_{1n,\,1r}$ are $\{0.372,\, 0.834\}$, implying that at least $37.2\%$ of women who would have worked off assistance at below-FPL earnings under AFDC were induced to participate at eligible earnings levels under JF.

The remaining response probabilities---$\pi_{0r,\,0n}$, $\pi_{0r,\,1n}$, $\pi_{0r,\,2n}$, $\pi_{0r,\,1r}$, $\pi_{0r,\,2u}$ and $\pi_{2u,\,1r}$---are each tightened on at least one bound relative to KT. Interestingly, $\pi_{2u,\,1r}$ now admits a strictly positive lower bound of $5.4\%$ providing new evidence that a strictly positive fraction of women who underreported above-FPL earnings under AFDC responded by reducing their earnings to below-FPL and reporting truthfully under JF.

For composite margins, the identified set for $\pi_{n,\,p}$, the fraction of women induced by JF to take up welfare, narrows about by $3\%$ on both sides to $\{0.265,\, 0.428\}$. The extensive-margin probability $\pi_{0,\,1+}$ is point-identified at $0.161$, consistent with the KT estimate of $0.167$ and within its reported confidence interval.

Table (ref) in Appendix reports results with LASSO first-stage estimator, across two covariate set regimes and cross-fitting schemes.

table[table omitted — 3,509 chars of source]

\paragraph{Granular partition.} Table (ref) presents identified sets under the nine-bin partition of Section (ref), estimated with a LASSO first stage, GroupKFold cross-fitting and two covariate regimes, baseline set of 28 and an extended set of 255 variables constructed from pairwise interactions and polynomial transformations.

The granular exercise is motivated by the concern, articulated by KT themselves, that the opt-in response identified under the coarse partition “could hypothetically reflect trivial earnings reductions from \$1 above the poverty line to exactly the poverty line” KT. This can be tackled by decomposing the identified set for the above-FPL nonparticipation state ($2n$) across three sub-bins indexed $b_6n$, $b_7n$, and $b_8n$, corresponding to monthly earnings in the ranges $[1.0,\, 1.2)\times\mathrm{FPL}$, $[1.2,\, 1.4)\times\mathrm{FPL}$, and ${\geq}\,1.4\times\mathrm{FPL}$, respectively. Section (ref) shows how to adjust the linear system to account for finer partitions. We find that at least $30.0\%$, $35.5\%$, and $20.6\%$ of women who, under AFDC, worked off assistance in each respective sub-bin reduced their earnings below the poverty line in response to the JF reform. Under the baseline covariate regime, the corresponding lower bounds are $30.0\%$, $32.8\%$, and $20.4\%$; the qualitative ordering is preserved, with the attenuation consistent with the extended covariate set delivering higher first-stage predictive power and thereby tightening the identified set, though the tightening is not drastic, which can point to covariates being not very informative.

The ordering of lower bounds across sub-bins is consistent with trivial earnings rounding as the primary behavioral mechanism. The opt-in behavior, however, reflects almost similar scale in $[1,\, 1.2)\times\mathrm{FPL}$ and $[1.2,\, 1.4)\times\mathrm{FPL}$ range. These findings jointly imply that the identified opt-in response is concentrated among women who undertook both marginal and somewhat substantive ($20-40\%$) labor-supply adjustments.

For the five sub-bins of the below-FPL nonparticipation state ($1n$), indexed $b_1n$ through $b_5n$ and corresponding to quintiles of the interval $(0,\, \mathrm{FPL}]$, the lower bounds exhibit substantial dispersion across the earnings distribution with strongest participation response at intermediate below-FPL earnings levels of $b_3n$ sub-bin with $55\%$ of women taking up welfare. Table (ref) in Appendix reports results with different granularity specifications.

table[table omitted — 3,400 chars of source]

\paragraph{Welfare bounds.}

Two competing forces govern the choice of partition for welfare analysis. Finer bins reduce the discretization error in $\Delta_{(s^{a},\,s^{j})}$ defined in Section (ref). When earnings are evaluated at bin midpoints or IPW-weighted means, a coarser partition averages over a wider interval of true earnings, potentially conflating transitions with substantially different welfare consequences. At the same time, a finer partition expands the matrix $\widetilde{A}$ further and introduces additional transition parameters, which widens the identified set. Importantly, even under very granular regimes the lower bound on an individual granular cell can be zero even though the lower bound on the composite flow it belongs to, such as $2n \to 1r$, can be strictly positive. This reflects a general property of the LP that the lower bound on a sum weakly exceeds the sum of the component lower bounds, since the joint minimum must hold at a single feasible point. Any cell can be zeroed by shifting mass to other destination sub-bins, but the conservation-of-mass constraints forbid zeroing the aggregate since the observed contraction of the $2n$ population can only be rationalized by opting into welfare. For confirmation, we report the composite bounds for granular specifications in Table (ref) in Appendix, which includes examples of trivial individual bounds but non-trivial composite ones. Because of this we opt for using $11 \times 53$ design to bound the composite welfare gains. Among the specifications we examine, its cells deviate the least from the actual monetary welfare value of each transition. For the opt-in margin $2n \to 1r$ the resulting bound on the welfare gain is of indeterminate sign. The identified interval is [−-$178.6\$$, $14.6\$$] per month and includes zero, so the data are uninformative as to whether opting into the program raises or lowers the welfare as we define it in Section~\ref{sec:illustration}. Table~\ref{tab:welfare_bounds} in Appendix shows welfare bounds for both granular bins and composites under different granularity specifications. Under less granular specifications, welfare losses for $2n \to 1r$ range from at least --$21.05\$$ to --$61.47\$$ per month.

center[center omitted — 316 chars of source]