EconBase
← Back to paper

On the Lower Confidence Band for the Optimal Welfare in Policy Learning

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

60,966 characters · 15 sections · 67 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

On the Lower Confidence Band for the Optimal Welfare in Policy Learning

abstractWe study inference on the optimal welfare in a policy learning problem and propose reporting a lower confidence band (LCB). A natural approach to constructing an LCB is to invert a one-sided $t$-test based on an efficient estimator for the optimal welfare. However, we show that for an empirically relevant class of DGPs, such an LCB can be first-order dominated by an LCB based on a welfare estimate for a suitable suboptimal treatment policy. We show that such first-order dominance is possible if and only if the optimal treatment policy is not “well-separated” from the rest, in the sense of the commonly imposed margin condition. When this condition fails, standard debiased inference methods are not applicable. We show that uniformly valid and easy-to-compute LCBs can be constructed analytically by inverting moment-inequality tests with the maximum and quasi-likelihood-ratio test statistics. As an empirical illustration, we revisit the National JTPA study and find that the proposed LCBs achieve reliable coverage and competitive length.

\noindentKeywords: Statistical decision theory, policy learning, optimal welfare, lower confidence band, partial identification, sensitivity analysis, cross-fitting, uniformity \\ \noindentJEL Numbers: C14, C31, C54

Introduction

Treatment assignment problems are ubiquitous in economics, including governments providing subsidies to disadvantaged households, firms offering job training opportunities to their employees, colleges allocating scholarships to students, and online retailers offering discounts to customers. In such settings, a decision-maker (DM) aims to design a treatment rule that determines who should --- and who should not --- be treated, based on observable individual characteristics, to maximize welfare Manski2004. Since developing good treatment rules may be costly and time-consuming, the DM might want to quantify the potential welfare gains. To this end, the DM may conduct a preliminary experiment and test a hypothesis that the optimal welfare (or welfare gain) exceeds a certain threshold.

Conducting inference for the optimal welfare (and welfare gain) is a challenging task. From a practical perspective, it may require solving complicated non-convex optimization problems, estimating functions of high-dimensional inputs non-parametrically, and dealing with noisy welfare estimates due to suboptimal experiment design. Theoretically, a major complication is the potential non-uniqueness of the optimal policy, which makes standard debiased inference methods inapplicable HiranoPorter2012, LuedtkeLaan.

In this paper, we show that good estimators and tight lower confidence bands (LCBs) for the optimal welfare (and welfare gain) can be obtained by leveraging suboptimal policies. Our first contribution is to demonstrate a possible trade-off between the welfare level and the precision with which it can be estimated in finite samples. For empirically relevant data-generating processes (DGPs), we provide an example of a slightly suboptimal policy, whose welfare can be estimated substantially more precisely than the optimal one. As a result, an LCB targeting such suboptimal welfare can be first-order tighter --- at the $N^{-1/2}$ scale for sample size $N$ --- than the LCB targeting the optimal welfare directly. Additionally, such suboptimal policy yields a better estimator of the optimal welfare in terms of mean-squared error, for all $N$ large enough. In particular, this example shows that incorporating asymptotically redundant information can yield first-order improvements for estimators and inference procedures in finite samples.

Our second contribution is to characterize the class of DGPs for which the first-order trade-off between welfare and precision is possible. Intuitively, if the optimal policy is “well-separated” from the rest, the precision gain of any suboptimal policy cannot compensate for the welfare loss. We formalize this intuition using a local asymptotic approximation around a DGP at which “separation” fails, and derive minimax rates for the gap between the two LCBs. As a result, we show that the first-order trade-off is possible if and only if the margin condition of MammenTsybakov and Tsybakov fails to hold uniformly over the relevant DGPs. In such settings, standard debiased inference procedures may be invalid, so alternative inference methods are needed.

To this end, we propose LCBs that address the aforementioned welfare-precision tradeoff and remain valid regardless of the margin condition. The idea is to construct a (possibly large but) finite subclass of test policies, based on economic intuition, within which a “good” suboptimal policy may be found. Each of these policies provides a lower bound on the optimal welfare, yielding a collection of moment inequalities that can be tested using existing methods andrews2010inference, CLR, romano2014practical, canay2017practical. The existing tests combine self-normalization CLR with moment selection, leading to tight LCBs that remain valid under relatively weak conditions. For the problem at hand, the tests can often be inverted analytically, so the LCBs are easy to compute in practice.

To illustrate our theoretical results, we revisit the U.S. National Job Training Partnership Act (JTPA) experiment Bloom1997. The experiment randomly assigned individuals with distinct education levels and baseline earnings to a job training program and recorded their post-treatment salary. For most education years --- apart from graduation thresholds --- the respective conditional average treatment effect is statistically insignificant, indicating a violation of the margin condition. Standard procedures that either ignore education or use a holdout sample to estimate the first-best policy suffer from substantial power loss. We consider several classes of test policies based only on education and construct the corresponding LCBs by inverting moment-inequality tests as described above. In line with the theoretical predictions, the LCBs are substantially shorter than the available alternatives.

Related Literature \, This paper contributes to a large cross-disciplinary literature on optimal treatment choice, following Manski2004. In econometrics, contributions range from early program-evaluation and partial-identification approaches to modern policy learning DEHEJIA2005141,HiranoPorter2009,Stoye,Chamberlain2011,BhatacharyaDupas2012,Tetenov2012,Rai2019,KitagawaTetenov,MbakopTabord, AtheyWager2, Sun,SasakiUra,KitagawaLeeChen2022,Yata2021,ArmstrongShen2023,chernozhukov2025polece,Moon:25. In statistics, optimal treatment regimes are commonly learned via Q-learning and A-learning QianMurphy,Murphy2003,Robins2004,ShiEtAl2018. This literature focuses primarily on obtaining treatment rules that perform well in terms of expected regret.

In this paper, we consider a complementary problem of inference on the optimal welfare, also studied in LuedtkeLaan.\footnote{This problem is distinct from the “inference on winners” considered in andrews2024inference, andrewschen2025, and chernozhukov2025polece, and the proposed LCBs are generally not valid in those settings. } In the absence of ties among the best policies, the authors showed that the optimal welfare is a regular parameter and derived the semiparametric efficiency bound for it. The bound turns out to be the same as if the best policy was known ex ante. When ties are present, the optimal welfare is no longer regular HiranoPorter2012, but in view of the above, an oracle efficient estimator based on one of the optimal policies still provides a natural benchmark for our analysis. We complement the results of LuedtkeLaan by studying MSEs of the estimators and expected length of the associated LCBs in finite samples, formalizing the necessity of the margin condition for one-sided inference, and proposing simple robust inference procedures. The proposed procedures provide alternatives to the approaches based on smoothing, as in chen2023inference, levis2023covariateassisted and whitehouse2025, or entropic regularization, as in benmichael2025. They also relate to a broader literature on robust policy learning, including decisions under ambiguity benmichael2021safepolicy,cui2024policy and concerns about external validity adjaho2023externally. Although we focus on the utilitarian (linear) formulation of welfare throughout, the proposed approach also applies in non-linear settings, such as inequality-sensitive welfare studied in kasy2016partial, kitagawa2021equality, terschuur2025locally, among others.

This paper also contributes to the literature on inference for partially identified parameters. We show that in finite samples, inference based on sharp bounds may be less precise than inference based on loose bounds, giving rise to a first-order trade-off between sharpness and precision. We argue that existing inference methods are able to address this trade-off by combining self-normalization (precision-correction) and moment selection, while retaining uniform validity andrews2010inference,CLR,romano2014practical,canay2017practical,bai2022two.\footnote{A related question of inference with over-identifying inequality constraints is studied, e.g., in cox2024simple and ketz2025short. Our setting is different in that the target parameter may not be asymptotically Gaussian even when the constraints are not binding.}

The rest of the paper is organized as follows. Section (ref) introduces the policy learning problem and motivates our target parameters. Section (ref) gives a sequence of DGPs exhibiting the first-order dominance. Section (ref) discusses the role of the margin assumption. Section (ref) proposes robust inference procedures. Section (ref) contains an empirical application. Section (ref) concludes. Appendix (ref) contains proofs. Appendix (ref) contains auxiliary theoretical results. Appendix (ref) contains auxiliary empirical details.

Setup

Policy Learning Problem

Consider a population of individuals characterized by their potential outcomes in treated and untreated states, $Y(1), Y(0) \in \mathcal{Y} \subseteq \mathbf{R}$, and characteristics $X \in \mathcal{X} \subseteq \mathbf{R}^{d_X}$. A decision-maker (DM) aims to maximize the average welfare in the population by subjecting some individuals to treatment, depending on their observable characteristics $X$. That is, the DM chooses a treatment rule $G \in \mathcal{G} \subseteq 2^\mathcal{X}$ to maximize

equation[equation omitted — 108 chars of source]

where $G^c = \mathcal{X} \backslash G$ denotes the complement of $G$. The class of feasible treatment rules $\mathcal{G}$ may be ex ante restricted for institutional reasons, such as transparency or non-discrimination in treatment, or practical reasons, such as computation and implementation.

We assume that the DM has access to experimental data that identifies $W_G$. The observable data vector $Z = (D, Y, X)$ contains the assigned treatment $D \in \{0, 1\}$, realized outcome $Y \in \mathcal{Y}$, and covariates $X \in \mathcal{X}$, so that $Y = DY(1) + (1-D) Y(0)$ and $D \perp (Y(1), Y(0)) \,|\, X.$ The propensity score will be denoted by $\pi(x) = P(D = 1 \,|\,X = x)$. The conditional mean and variance functions of the potential outcomes are non-parametrically identified as $ m(d, x) = \mathbb{E}[Y(d) \,|\,X = x] = \mathbb{E}[Y \,|\, D = d, X = x]$ and $\sigma^2(d, x) = \mathbb{V}ar(Y(d)\,|\,X = x) = \mathbb{V}ar(Y \,|\,D = d, X = x)$, for $d \in \{0, 1\}$, and the conditional average treatment effect (CATE) function as $\tau(x) = m(1, x) - m(0, x)$. As a result, the average welfare function is identified as $W_G = \mathbb{E}[m(0, X) + \bm{1}\{ X \in G \} \tau(X)]$ and can be non-parametrically estimated. To this end, the DM observes a random sample $(Z_i)_{i = 1}^N$ distributed i.i.d. $Z_i \sim P \in \mathbf{P}$, for a class of distributions $\mathbf{P}$ specified below.

The objects of interest throughout the paper are the maximum (or first-best, or optimal) welfare, denoted by

equation[equation omitted — 92 chars of source]

where $G^*$ denotes any policy attaining the maximum,\footnote{For simplicity, we assume that the maximum is well-defined.} and the corresponding welfare gain,

equation[equation omitted — 89 chars of source]

which is non-negative as long as the policy class $\mathcal{G}$ includes the status quo policy $\varnothing$ of not treating anyone.

Lower Confidence Bands

In many settings, the DM would naturally be interested in lower confidence bands (LCBs) for the maximum welfare or the corresponding welfare gain. For example, consider a firm deciding whether to build a job-training center. Suppose the firm maximizes the net welfare subject to a “safety” constraint that the risk of false adoption (i.e., incurring negative welfare) must be below level $\alpha$, for some $\alpha \in (0,1)$. This leads to testing \[ H_0:\; W_{G^*} \leq 0 \qquad\text{vs}\qquad H_1:\; W_{G^*} > 0. \] In such settings, LCBs are natural inputs to threshold decision rules LehmannRomano.

As another example, consider an online retailer deciding whether to offer a discount for a certain type of good to its customers. The retailer may first run a small-scale randomized experiment to explore whether any discount rule can lead to increase in profits. This corresponds to testing \[ H_0:\; W_{G^*}^{\text{gain}} = 0 \qquad\text{vs}\qquad H_1:\; W_{G^*}^{\text{gain}} > 0, \] which is equivalent to comparing a \(100(1-\alpha)\%\) LCB for \(W_{G^*}^{\text{gain}}\) with zero.

The main input in the construction of LCBs is an estimator for the welfare function $W_G$. For each policy $G$, we can express $W_G = \mathbb{E}[\psi_G(Z)]$, where

align[align omitted — 214 chars of source]

is the efficient, doubly robust, moment function Robins, Hahn98. For suitable first-stage estimators $\widehat{m}(d, x)$ and $\widehat{\pi}(x)$, a regular semiparametrically efficient estimator $\widehat{W}_G$ can be constructed using cross-fitting, so that

equation[equation omitted — 112 chars of source]

where $\sigma_G^2 = \mathbb{V}ar(\psi_G(Z))$. Given a significance level $\alpha \in (0,1)$, a $100(1-\alpha) \%$ LCB for $W_G$ can be formed as

align[align omitted — 112 chars of source]

where $z_{1-\alpha}$ is the $(1-\alpha)$ quantile of $\mathcal{N}(0,1)$ and $ \widehat{\sigma}_{G}$ is a consistent estimator of the asymptotic standard deviation $\sigma_G$.

Since $W_{G} \leq W_{G^*}$, for any $G \in \mathcal{G}$, an LCB based on any suboptimal policy $G \in \mathcal{G}$ provides valid one-sided coverage for the optimal welfare,

align[align omitted — 149 chars of source]

As a result, $\widehat{LCB}_{G}$ can be meaningfully compared across distinct policies. As an ideal benchmark, we consider an LCB based on an infeasible efficient estimator of the welfare under a first-best policy,

align[align omitted — 123 chars of source]

As discussed in the introduction, such LCB is a valid reference point even when the optimal policy is not unique. Since $\widehat{LCB}_{G^*}$ is based on an efficient estimator for $W_{G^*}$ and the standard deviation is rescaled by $N^{-1/2}$, one might expect that $\widehat{LCB}_{G^{*}}$ always be preferred to $\widehat{LCB}_{G}$ in large samples, for any suboptimal policy $G$. We show, however, that this is not the case. Given the direction of the intended comparison, considering an oracle LCB as a benchmark only strengthens our point. Of course, our recommended inference procedures in Section (ref) account for the first-best policy being unknown.

Asymptotic Criterion for LCB ranking

To compare the candidate LCBs, we consider the LCB gap, defined as

align[align omitted — 111 chars of source]

A positive sign of $\Delta_G$ indicates that the policy $G$ is nearly optimal yet the corresponding welfare is substantially more precisely estimated. Consequently, the corresponding ${LCB}_{G}$ may be preferred to ${LCB}_{G^{*}}$ in large samples.

The motivation for studying LCB gap comes from a local asymptotic approximation along smooth parametric sub-models, standard in the semi-parametric efficiency theory. To elaborate, let $\mathbf{P}$ denote the class of all admissible distributions of the data. Consider a distribution $P_0 \in \mathbf{P}$ such that $W_{G^*(P_0)} = W_{G}$ for $G \ne G^*(P_0)$. Let $T(P_0)$ denote the tangent space at $P_0$,\footnote{See Hahn98 for the derivation of $T(P)$ in the present setting.} and $P_{N, h} = P_{1/\sqrt{N}, h}$, for $h \in T(P_0)$, be a sequence of distributions following a smooth parametric submodel $\{t \mapsto P_{t, h}\} \subseteq \mathbf{P}$. Denote \[

array[array omitted — 115 chars of source]

\] where the dependence of $\mu(h)$ and $s(h)$ on $N$ is suppressed for notational convenience, and note that \[ \Delta_{G}(P_{N, h}) = N^{-1/2}(z_{1-\alpha} s(h) - \mu(h)). \] The assumed regularity of $\widehat{W}_{G}$, consistency of $\widehat{\sigma}_{G}$, and contiguity of $P_{N, h}$ with respect to $P_0$ imply that, under $P_{N, h}$, \[ \sqrt{N}(\widehat{LCB}_{G} - \widehat{LCB}_{G^*(P_{N, h})}) \;\;\; \Rightarrow_d \;\;\; \mathcal{N}(z_{1-\alpha} s(h) - \mu(h), \sigma_{\Delta}^2(P_0)), \] for some $\sigma_{\Delta}^2(P_0) \geqslant 0$. That is, the distribution of $\sqrt{N}(\widehat{LCB}_{G} - \widehat{LCB}_{G^*(P_{N, h})})$ under any sequence of “perturbations” $P_{N, h}$ of $P_0$, is determined by $z_{1-\alpha}s(h) - \mu(h) = \sqrt{N} \Delta_{G}(P_{N, h})$. Moreover, under further regularity conditions, \[ \mathbb{E}[\widehat{LCB}_{G} - \widehat{LCB}_{G^*(P_{N, h})}] = \Delta_{G}(P_{N, h}) + o(1), \] so the LCB gap can be interpreted as a large-sample analog to the difference of expected LCBs.\footnote{An ideal way to rank LCBs is in terms of the first-order dominance; See Lehmann1959. Unfortunately, since distinct policies typically result in LCBs with distinct large-sample variances, this criterion does not apply in a Gaussian limit. A natural alternative is to compare LCBs in terms of their expected values, as suggested, e.g., in LeonHarter. While the exact expectations may not exist without further restrictions or be distorted by the biases in first-stage estimators, their large sample analogs remain tractable.} For these reasons, we consider the LCB gap in the formal results below.

First-Order Dominance

In this section, we give an example of a model in which the welfare-precision trade-off is of the first order, and discuss the implications of this phenomenon. We focus on welfare throughout, but similar considerations apply to welfare gain. See Remark (ref) for the details.

The Data Generating Process

First, we specify a suitable DGP for $(Y(1), Y(0), D, X)$. It suffices to specify the marginal distribution of $X$, the propensity score, and the conditional distributions of $Y(1) \mid X$ and $Y(0) \mid X$. Let $X$ be a binary covariate distributed as

equation*[equation* omitted — 94 chars of source]

Denote the propensity score by $$ P (D=1 \mid X=1) = \pi(1); \quad P (D=1 \mid X=0) = \pi(0), \quad \text{ for some } \pi(1), \pi(0) \in (1/4, 3/4). $$ Let $F(\mu, \sigma^2)$ be any distribution with mean $\mu$ and variance $\sigma^2$. Suppose the potential outcomes are distributed as

align[align omitted — 277 chars of source]

where $\epsilon \in (0, 1/2)$ is a vanishing sequence to be specified. Since we focus on the average welfare, the joint distribution of $(Y(1), Y(0)) \,|\, X$ is immaterial, so we leave it unspecified\footnote{ With variance parameters $\sigma^2(1,1) = 1/4, \quad \sigma^2 (1,0)=1, \quad \sigma^2(0,1) = 200, \quad \sigma^2(0,0) = 1$, the statement holds for all sample sizes exceeding $1745$. For the variances in the main text, the minimal cutoff sample size $N$ is approximately $6 000$.}

Simple algebra shows that the CATE function takes the form $$ \tau(1) = -\epsilon < 0; \;\;\;\;\;\; \tau(0) =\epsilon>0, $$ the unique first-best policy is

align[align omitted — 42 chars of source]

and the corresponding welfare is

align[align omitted — 75 chars of source]

In addition, consider the “treat everyone” policy, $G=\mathcal{X}$ whose welfare is

align[align omitted — 109 chars of source]

Note that the welfare gap between the two policies scales linearly with $\epsilon$

align[align omitted — 84 chars of source]

while the standard deviation gap does not depend on $\epsilon$,

align[align omitted — 74 chars of source]

Estimators and Lower Confidence Bands

Since $X$ is binary, the average welfare under any fixed policy $G$ can be efficiently estimated using the regression-adjusted estimator. For each \((d,x) \in \{0, 1\}^2\), denote

align[align omitted — 89 chars of source]

and define the esitmators

equation[equation omitted — 220 chars of source]

where one is added to the denominator throughout to prevent division by zero.\footnote{This step introduces bias of order $O(N^{-1})$ which is negligible for a sufficiently large sample. An alternative is to work with unadjusted denominators on the event where both of them are strictly positive.} Recalling from (ref) that $G^* = \{0\}$, the first-best welfare is estimated as

align[align omitted — 133 chars of source]

where $\widehat p = \sum_{i=1}^N X_i/N$. Similarly,

align[align omitted — 134 chars of source]

The mean squared errors of the two estimators, with respect to $W_G^*$, are given by

equation[equation omitted — 279 chars of source]

The asymptotic variances of $\widehat{W}_G$, for $G \in \{G^*, \mathcal{X}\}$, can be estimated as $$ \widehat \sigma^2_G = \dfrac{1}{N} \sum_{i=1}^N (\widehat \psi_G(Z_i) - \widehat {W}_G)^2, $$ where $\widehat{\psi}_G(Z_i)$ is obtained by plugging the estimated propensity score and regression functions from (ref) in (ref). The corresponding LCBs are obtained as

align[align omitted — 263 chars of source]

Following the discussion of Section (ref), we compare ${LCB}_{\mathcal{X}}$ and ${LCB}_{G^{*}}$ in terms of LCB gap

align[align omitted — 175 chars of source]

First-Order Dominance

Our first main result shows that $\widehat{W}_{\mathcal{X}}$ dominates $\widehat{W}_{G^{*}}$ in terms of MSE, and the respective LCB gap is positive.

proposition[First-Order Dominance] For all $N$ large enough, for the DGP (ref) and estimators (ref) and (ref), the following statements hold: \begin{enumerate} • Both MSEs in (ref) are finite and \begin{align} MSE (\widehat{W}_{\mathcal{X}}) < MSE (\widehat{W}_{G^{*}}); \end{align} • For any significance level $\alpha \in (0,1)$, there is a constant $C_{\alpha}>0$ such that \begin{align} \Delta_{\mathcal{X}} > C_{\alpha} N^{-1/2}. \end{align} \end{enumerate}

Proposition (ref) makes three points. First, the trade-off between welfare and precision may be first-order. As a result, suboptimal policies may yield better point estimates and tighter, on average, lower confidence bands for the optimal welfare. That is, the first-best policy --- the policy that is best to implement --- may differ from the policy whose estimated welfare is best to report.\footnote{DGPs with treatment effects vanishing at the $N^{-1/2}$ rate have been employed to obtain a meaningful limiting experiment HiranoPorter2009 or establish minimax rates for expected regret KitagawaTetenov, AtheyWager2. In this paper, we use DGPs with similar conditional means and carefully chosen variances to establish a lower bound on the LCB gap.} Similar observations apply to inference on partially-identified parameters, as we further discuss in Remark (ref).

Second, there is a distinction between the two-sided and one-sided inferential objectives. In the two-sided case, the bias typically must vanish faster than the standard deviation to ensure valid coverage of the confidence intervals. In the one-sided case, coverage remains valid as long as the direction of the bias matches the direction of the confidence band, which allows bias and variance to be potentially of the same order. Proposition (ref) gives a concrete, empirically relevant example of this distinction\footnote{The one-sided dominance result echoes findings in one-sided nonparametric inference: in adaptive tests and multiscale procedures, directed smoothing bias can be exploited to lower variance while preserving size DumbgenSpokoiny2001, Armstrong2015. Our setting differs in the target parameter (optimal welfare rather than a function at a point) and mechanism (policy-induced bias $W_{G^{*}} - W_G$ rather than smoothing bias). }.

Third, efficiency arguments in near non-regular settings may be problematic. For each $\epsilon>0$, the oracle efficient estimator $\widehat{W}_{G^*}$ attains the semiparametric efficiency bound LuedtkeLaan, but in the limit, $\epsilon=0$, the optimal welfare is a non-regular parameter, and semiparametric efficiency bounds do not apply HiranoPorter2012. Proposition (ref) demonstrates that, for distributions within a $N^{-1/2}$-neighborhood of $\epsilon=0$ (excluding zero), $\widehat{W}_{\mathcal{X}}$ dominates $\widehat{W}_{G^{*}}$ in terms of MSE, for all $N$ large enough. Thus, the familiar notion of efficiency fails not only at the point of non-regularity, but already in a $N^{-1/2}$-neighborhood around it.

remark[Implications for welfare gain] The above example could be modified to obtain a first-order dominance statement for the welfare gain in (ref). Consider the DGPs \begin{align} Y(1)\mid X=1 &\sim F\bigl(\tfrac12 - \epsilon,\;1\bigr), & Y(1)\mid X=0 &\sim F\bigl(\tfrac12,\;10\bigr), \\ Y(0)\mid X=1 &\sim F\bigl(\tfrac12,\;1\bigr), & Y(0)\mid X=0 &\sim F\bigl(\tfrac12 - \epsilon,\;10\bigr). \end{align} where asymptotic variance is small for $X=1$ and large for $X=0$. Let $G^{*} = \{0\}$ be the optimal policy and $G = \{1\}$ be the suboptimal policy. Simple algebra shows that the welfare gap and variance gap satisfy \begin{align} W^{gain}_{G^{*}}-W^{gain}_{G} &\leq \epsilon, \\ (\sigma^{gain}_{G^{*}})^2-(\sigma^{gain}_{G})^2_ &> 7. \end{align} As a result, an analog of (ref) holds for the LCB gap for welfare gain. $\blacksquare$
remark[Redundant moment inequalities] The above discussion applies to inference for partially identified parameters. For example, consider the setting of Section (ref) with binary potential outcomes and unconditional treatment exogeneity, i.e. $(Y(1), Y(0), X) \perp D$. The share of “always-takers”, $\theta=P(Y(1) = Y(0) = 1),$ can be bounded from above by either $\delta_1 = P(Y=1 \mid D=0)$ or $\delta_2 = \mathbb{E} [\min (P(Y=1 \mid D=1,X), P(Y=1 \mid D=0,X))]$. By Jensen's inequality, $\delta_2$ gives a tighter bound than $\delta_1$. A $100(1-\alpha) \%$ Upper Confidence Band (UCB) for $\theta$ can be formed using either of the two bounds $$ \widehat{UCB}_j = \widehat{\delta}_j + N^{-1/2} z_{1-\alpha} \widehat{\sigma}_j, $$ where $\widehat \sigma_j$ are consistent estimators of the asymptotic standard deviations $\sigma_j$ of $\hat{\delta}_j$, for $j = 1, 2$. Similar to Proposition (ref), there exist DGPs such that \begin{align} \delta_2+N^{-1/2} z_{1-\alpha} \sigma_2>\delta_1+N^{-1/2} z_{1-\alpha} \sigma_1, \end{align} for all $N$ large enough. As a result, a UCB based on a non-sharp bound first-order dominates its sharp counterpart in terms of the average length. In other words, inference based on a sharp bound may be less informative in finite samples. $\blacksquare$

Margin Condition and Higher-Order Dominance

Next, we investigate whether the conclusions of Proposition (ref) carry over when the model is restricted by the following additional assumptions.

assumption[Regularity] (i) The propensity score $\pi(x)$ satisfies $\kappa < \pi(x) < 1-\kappa$, for almost all $x \in \mathcal{X}$, for some $\kappa \in (0, 1/2)$; (ii) The outcome is bounded so that $P (|Y| \leq M/2) = 1$, for some $M < \infty$.
assumption[Margin Condition] For some $\eta \in (0,M)$ and $\delta \in (0,\infty)$, \begin{align} P(|\tau(X)| < t) \leq (t/\eta)^{\delta}, \quad \forall t \in [0, \eta). \end{align}

Assumption (ref) imposes regularity conditions common in the policy learning literature KitagawaTetenov,MbakopTabord. Assumption (ref) is the margin condition of Tsybakov. In addition to requiring uniqueness of the first-best policy, it controls the intensity with which $\tau(X)$ concentrates in a neighborhood of zero. When the optimal policy is unique, the existence of suitable values of $\delta$ and $\eta$ is a matter of mild regularity conditions. For example, if $|\tau(X)|$ is continuous and has a density bounded at zero, then (ref) holds for any $\delta<1$ with $\eta$ small enough. If $\tau(X)$ has finite support and $P(\tau(X) = 0) = 0$, then (ref) holds for any $\delta > 0$ and a sufficiently small $\eta$.

The sequence of DGPs in Proposition (ref) can be chosen to satisfy Assumption (ref), but it fails to satisfy Assumption (ref) with uniform lower bounds on $\eta$ and $\delta$. As we show below, this is precisely what drives the first-order dominance phenomenon. To state the formal result, we assume that any $G \subseteq \mathcal{X}$ is feasible.\footnote{The upper bound in Proposition 2 holds for all $G \subseteq \mathcal{X}$, so it applies to any restricted class $\mathcal{G}$ as well. The lower bound holds within restricted classes $\mathcal{G}$ as long as they include threshold policies based on each covariate.} Proposition (ref) below characterizes the order of magnitude of the worst-case LCB gap $\Delta_G$ over all policies $G \subseteq \mathcal{X}$.

proposition[Higher-Order Dominance] Let $\mathbf{P}$ denote the class of DGPs obeying Assumptions (ref)--(ref) for some $ 0 < \underline{\delta} \leq \delta \leq \overline{\delta} < \infty$, $\eta = \eta(\delta) > 0$, and $\inf_{x \in \mathcal{X}, d \in \{1,0\}} \sigma^2(d, x) \geq \underline{\sigma}^2>0$. There exist constants $0<\underline C<\overline C<\infty$, depending on $(M,\kappa, \underline{\delta}, \overline{\delta}, \underline{\sigma})$, such that \begin{equation} C N^{-(1 + \delta)/2} \leq \sup_{P \in \mathbf{P}} \sup_{G \subseteq \mathcal{X} } \Delta_G \leq \overline{C} N^{-(1 + \delta)/2}. \end{equation}

Proposition (ref) demonstrates that once uniform lower bounds on $\delta$ and $\eta$ are imposed, no suboptimal policy $G$ can lead to first-order dominance in the sense of Proposition (ref). The smaller the value of $\delta$, the more $\tau(X)$ concentrates near zero, the looser the upper bound in (ref). In the limit, $\delta=0$, which corresponds to failure of the margin condition, the lower bound in (ref) recovers the first-order dominance result (ref). In the absence of uniform bounds on the margin parameters, Propositions (ref) and (ref) imply that the first-best welfare may not be the optimal, or relevant, inferential target. The following remarks discuss testable implications of the margin condition and possible testing procedures, as well as further connections with the literature.

remark[Testing uniqueness of the optimal policy] Let $X$ be a discrete covariate taking $J$ distinct values with positive probabilities. Then, the conditional average treatment effect reduces to a vector $(\tau(j))_{j=1}^J$. The first-best policy is non-unique if (and only if) $\tau(j) = 0$ for some $j \in \{1, 2, \dots, J\}$. The null hypothesis \begin{align} H_0: \exists j:  \tau(j) = 0 \end{align} is a union of $J$ simple hypotheses $H_{0j}: \tau(j)=0$. Then, letting $R_j$ denote the rejection region for testing $H_{0j}$, the test with a rejection region $$ R=\cap_{j=1}^J R_j, $$ is valid for $H_0$, although typically conservative Berger1997. $\blacksquare$
remark[Testing the margin assumption] In the general case where both discrete and continuous covariates are present, Assumption (ref) is no longer equivalent to uniqueness of the optimal policy. We describe a testable implication that we find empirically relevant in Section (ref). Let $P(G^* \triangle G)$ denote the share of people treated differently under the optimal policy $G^*$ and an alternative $G$. This share links welfare and standard deviation gaps. Specifically, the welfare gap is lower bounded as \begin{align} W_{G^*} - W_{G} \geq C_1P (G^{*} \triangle G)^{1+\frac{1}{\delta}} \end{align} for $C_1 = C_1(\delta)=\eta\delta(\frac{1}{1+\delta})^{1 + \frac{1}{\delta}} > 0$ Tsybakov. Given a lower bound $\underline{\delta}>0$ and fixing $\eta>0$, consider a null hypothesis $H_0: \delta \geqslant \underline{\delta}$. Since both functions $\delta \rightarrow C_1(\delta)$ and $\delta \rightarrow c^{1+\frac{1}{\delta}}$ are increasing in $\delta$, the lower bound (ref) on welfare gap implies that, for any policy $G$, \begin{align} C_1(\delta) P(G^* \triangle G)^{1 + \frac{1}{\delta}} - (W_{G^*} - W_G) \leq 0. \end{align} In particular, if the welfare gap $W_{G^*} - W_G$ of some policy $G$ vanishes with sample size, the share of people treated differently under $G$ and $G^*$, must vanish, too. Existing methods from the moment inequality literature, such as CLR and CherNeweySantos, can then be applied to construct a test. Pursuing this formally is left for future work.\footnote{The lower bound in (ref) plays a role analogous in spirit to the polynomial minorant condition used in partial identification literature, e.g., Condition C.2 in CHT, Condition V in CLR, and Assumption 4.2 in Armstrong2014. In its general form, this condition relates the difference in the criterion function to the distance metric on the parameter of interest. In policy learning settings, the decision set $\mathcal{G}$ is a collection of partitions of covariate space. In both CLR and KitagawaTetenov, this condition is imposed to tighten convergence guarantees for the proposed estimators. In contrast to prior work, this paper uses the (failure of) margin assumption to motivate the use of suboptimal policies for constructing lower confidence bands for welfare. } $\blacksquare$
remark[Implications for debiased inference] Propositions (ref) and (ref) imply that sharp bounds may not be optimal, or relevant, inferential targets in the absence of uniform margins, highlighting the tightness of this condition in the context of covariate-assisted bounds; see kallus2020assessing, kallus2022whats,kallus2022treatment, levis2023covariateassisted, SemSupp2, Semenova2024. We expect this insight to imply the tightness of the margin condition in other settings, such as support function analysis CCMS and algorithmic fairness liu2025, and other policy-relevant metrics. $\blacksquare$

Robust Testing Procedures

In this section we discuss testing procedures that address the welfare-precision trade-off and remain valid regardless of the margin assumption. Let $\mathcal{G}_{\text{test}} \subseteq \mathcal{G}$ be a class of policies, which, based on economic intuition, may contain a good lower bound for the optimal welfare. We look for a LCB of the form

equation[equation omitted — 187 chars of source]

where $\hat{c}_\alpha$ is as small as possible to guarantee the desired coverage. We show that such LCB naturally arise from testing moment inequalities, which allows to use a host of existing testing procedures. Our results take the form of finite-sample algebraic identities, so the coverage properties of the resulting LCBs are inherited from validity of the underlying tests. The latter relies only on the uniform CLT-type assumptions and holds regardless of the margin condition. We refer the reader to CLR and canay2017practical for the details.

Lower Confidence Bands via Testing Moment Inequalities

Suppose $\mathcal{G}_{\text{test}}$ is finite (potentially growing with sample size). Let $\theta = W_{G^*}$ denote the parameter of interest, and consider testing

equation[equation omitted — 116 chars of source]

Suppose the estimator \((\widehat W_G)_{G\in\mathcal{G}_{\text{test}}} \) for \( (W_G)_{G\in\mathcal{G}_{\text{test}}}\) satisfies

equation[equation omitted — 140 chars of source]

for a positive definite covariance matrix $\Sigma= (\Sigma_{G_1G_2})_{G_1, G_2 \in \mathcal{G}}$, and a consistent estimator $\widehat{\Sigma}$ is available. A test for (ref) can then be constructed as

equation[equation omitted — 123 chars of source]

with, e.g., the maximum test statistic

equation[equation omitted — 145 chars of source]

where $\hat{\sigma}_{G} = (\widehat{\Sigma}_{GG})^{1/2}$ and $\hat{c}_{\alpha}(\theta)$ is suitable a critical value. A common computationally simple choice is the least-favorable critical value, corresponding to

equation[equation omitted — 204 chars of source]

where the quantile can be estimated using bootstrap or Gaussian approximation.

Given the direction of the inequalities in (ref), the set of all values of $\theta$ for which the test in (ref) does not reject, $\{\theta \in \mathbf{R}: \hat{\phi}_N(\theta) = 0\}$, provides a LCB for $W_{G^*}$. For the least-favorable critical value, the test compares the value of a partially linear decreasing function of $\theta$ with a constant, which allows to obtain a simple closed form for the LCB.

proposition[LCB by test inversion] The LCB obtained by inverting a test in (ref) with the least-favorable critical value in (ref) is given by \begin{equation} \widehat{LCB}_{\max}^{LF} = \max_{G\in\mathcal G_{\normalfont test}} \Bigl\{ \widehat W_G - \hat{c}_{\alpha, \max}^{LF}\frac{\widehat \sigma_G}{\sqrt{N}}\, \Bigr\}. \end{equation}

Intuitively, the above procedure corresponds to constructing a candidate LCB for $W_{G^*}$ using each suboptimal policy $G \in \mathcal{G}_{\text{test}}$ separately and taking the shortest one, thus explicitly resolving the welfare-precision trade-off. The least-favorable critical value $\hat{c}_{\alpha, \max}^{LF}$ ensures that the resulting LCB has the desired coverage, but it essentially assumes that all of the moment inequalities in (ref) are binding, which may be too conservative. The critival value can be reduced using moment selection procedures, such as the Generalized Moment Selection (GMS) of andrews2010inference, or pre-testing, as in romano2014practical. Although both procedures perform well in practice, we focus on GMS because it allows for closed-form test inversion.

The critical value for the GMS procedure is computed as follows. Define the set of inequalities that are “close to binding,” \[ I_{N}(\theta) = \left\{G \in \mathcal{G}_{\text{test}}: \frac{\sqrt{N}(\hat{W}_G - \theta)}{\hat{\sigma}_G} > -\kappa_{N} \right\}, \] where $\kappa_N > 0$ is a sequence of tuning parameters such that $\kappa_N \to \infty$ and $\kappa_N / \sqrt{N} \to 0$, for example, $\kappa_N = \sqrt{\log N}$. Then, the GMS critical value is

equation[equation omitted — 200 chars of source]

where the quantile can be estimated using bootstrap or Gaussian approximation. As $\theta$ increases, the set $I_N(\theta)$ shrinks, so $\hat{c}_{\alpha, \max}^{GMS}(\theta)$ is a decreasing step-function of $\theta$. Thus, the test in (ref) with the GMS critical value compares a partially linear decreasing function of $\theta$ with a step-function. Since there may be multiple intersections, the confidence region obtained by test inversion may not be convex, although it can be shown that the probability of such an event approaches zero as $N$ increases. In the statement below, we conservatively define the LCB starting from the lowest intersection point.

proposition[LCB by test inversion with GMS] The LCB obtained by inverting the test in (ref) with the critical value (ref) can be computed as follows. For $j \in \{1, \dots, |\mathcal{G}_{\text{test}}|\}$, let $t^{(j)}$ denote the $j$-th largest value among $\widehat{W}_G + \kappa_N \widehat{\sigma}_G / \sqrt{N}$, and set $t^{(|\mathcal{G}_{\text{test}}| + 1)} = -\infty$. Let $I^{(j)} \subseteq \mathcal{G}_{\text{test}}$ collect the policies $G$ corresponding to $t^{(1)}, \dots, t^{(j)}$, and $\widehat{c}_{\alpha}^{(j)}$ be computed as in (ref) with $I^{(j)}$ instead of $I_N(\theta)$. Denote $ \hat{\theta}^{(j)} = \max_{G \in \mathcal{G}}(\widehat{W}_G - \widehat{c}^{(j)}_{\alpha} \widehat{\sigma}_G /\sqrt{N}).$ Then, \begin{align} \widehat{LCB}_{\max}^{GMS} = \min\{\theta^{(j)}: t^{(j)} \geqslant \hat{\theta}^{(j)} > t^{(j+1)}\}. \end{align}

The LCB in (ref) uses a weakly smaller critical value than the LCB in (ref), so it is always shorter. Yet, the two LCBs are uniformly valid over the same set of distributions. Taken together, self-normalization of the test statistic and a moment selection procedure allow to resolve the welfare-precision trade-off while ensuring that the resulting LCB is robust to violations of the margin condition.

Lower Confidence Bands via Intersection Bounds

Inference methods for intersection-bounds-type parameters, such as $\max_{G \in \mathcal{G}}W_G$, have been introduced by CLR (CLR for short). The authors pointed out that inference based on the plug-in estimator $\max_{G \in \mathcal{G}}\widehat{W}_G$ may be distorted for two reasons: upward bias and large differences in precision of estimates $\widehat{W}_G$ across $G \in \mathcal{G}$. To address these issues, they introduced a “precision corrected” LCB of the form (ref) and proposed a different moment selection device, tailoring the analysis to an infinite number of intersection parameters (i.e., infinite $\mathcal{G}$). In what follows, we derive a new duality result between the procedure of CLR and test inversion in the spirit of Section (ref) and use it to obtain a computationally simpler LCB.

In the preceding section, to find a good lower bound on $W_{G^*}$, we restricted attention to policies in the test class $\mathcal{G}_{\text{test}} \subseteq \mathcal{G}$. A better lower bound may potentially be obtained by taking convex combinations of $(W_G)_{G \in \mathcal{G}_{\text{test}}}$, which is equivalent to randomizing over $G \in \mathcal{G}_{\text{test}}$. Specifically, let $\Lambda = \bigl\{\lambda\in \mathbf{R}^{|\mathcal{G}_{\text{test}}|}_+:\,\mathbf{1}'\lambda=1\bigr\}$, where $\mathbf{1}=(1,\ldots,1)'\in \mathbf{R}^{|\mathcal{G}_{\text{test}}|}$, denote the probability simplex, $W_{\text{test}} = (W_G)_{G \in \mathcal{G}_{\text{test}}}$ collect the test policies into a finite vector, and $\widehat{W}_{\text{test}} = (\widehat{W}_G)_{G \in \mathcal{G}_{\text{test}}}$ denote the corresponding estimator vector. Each $\lambda \in \Lambda$ yields a lower bound $\lambda'W_{\text{test}}\leq W_{G^*}$ for the optimal welfare. Therefore, following CLR, we look for a LCB of the form

equation[equation omitted — 229 chars of source]

where the critical value $\hat{c}_{\alpha}$ is chosen to ensure correct coverage.

A version of CLR's procedure calibrates $\hat{c}_{\alpha}$ by approximating the supremum of the self-normalized Gaussian process $(\sqrt{N}(\lambda'\widehat{W}_{\text{test}} - \lambda'W_{\text{test}}) / (\lambda'\widehat{\Sigma}\lambda)^{1/2} )_{\lambda \in \Lambda}$ in simulations, which can be computationally heavy. We replace that step with a finite-dimensional convex program using convex duality. We show that for any vector $T$ and positive definite matrix $\Sigma$,

equation[equation omitted — 222 chars of source]

Consequently, an LCB of the form (ref) actually arises from inverting a test for (ref) using the so-called Quasi-Likelihood-Ratio (QLR) test statistic,

equation[equation omitted — 242 chars of source]

also considered in andrews2010inference. The least-favorable critical value,

equation[equation omitted — 296 chars of source]

can be estimated using bootstrap or Gaussian approximation and requires solving one convex program per simulation. Our final Proposition summarizes this discussion.

proposition[$LCB$ by test inversion with QLR] The LCB obtained by inverting a test in (ref) with the QLR test statistic (ref) and least-favorable critical value (ref) takes the form \begin{equation} \widehat{LCB}_{\normalfont mix} = \max_{\lambda\in\Lambda} \left\{ \lambda^{\prime} \widehat{W}_{\normalfont test} - (\hat{c}_{\alpha, QLR}^{LF})^{1/2} \frac{\sqrt{\lambda'\,\widehat\Sigma\,\lambda}}{\sqrt{N}} \right\}. \end{equation}

In practice, \(\widehat{LCB}_{\max}^{LF}\) in (ref) (and its GMS version (ref)) are computationally simpler and employ a less conservative critical value than \(\widehat{LCB}_{\text{mix}}\). However, \(\widehat{LCB}_{\text{mix}}\) involves searching over all convex mixtures of test policies which creates more scope to trade off mean welfare against precision. As discussed in Example 4.1 in canay2017practical, both tests are admissible, so the corresponding LCBs cannot generally be ranked. Depending on the underlying DGP, either of the LCBs may be tighter.

Empirical Application

To illustrate the welfare-precision trade-off in practice and showcase the proposed procedures, we revisit the National Job Training Partnership Act (JTPA) study, considered in HeckmanIchimuraTodd and AbadieAngrist and recently revisited in the context of policy learning by KitagawaTetenov, MbakopTabord, and AtheyWager2, among others. A detailed description of the study is available in Bloom1997.

The study randomized whether applicants would be eligible to receive job training and related services for a period of eighteen months. The treatment $D$ is the indicator of program eligibility. The outcome $Y$ is the applicant's cumulative earnings thirty months after assignment. Two baseline covariates $X=(PreEarn, Educ)$ include pre-program earnings (in USD) and years of education. By design, unconditional independence holds, $$ (Y(1), Y(0), X) \perp D, $$ so the first-best welfare and the corresponding welfare gain are identified in each of the models $(D,Y)$, $(PreEarn, D,Y)$, $(Educ, D,Y)$ and $(X,D,Y)$. This fact allows us to compare the estimated optimal welfare gains and corresponding LCBs across the models and highlight connections with our theoretical results.

Table (ref) presents the estimates and LCBs for the welfare gain based on first-best policy rules in three different policy classes: no covariates (Row 1), only $PreEarn$ (Rows 2--3), and both covariates $X$ (Row 4). The welfare gain from treating everyone (Row 1) corresponds to the Average Treatment Effect. It provides a robust lower bound for the optimal welfare gain based on more complex policy classes, so we use it as a reference point. In Rows 2--3, given that $PreEarn$ is continuously distributed and the margin assumption is plausible, we adopt the cross-fitted efficient-score estimator. We consider estimating the CATE function of $PreEarn$ via series regression (Row 2) and random forest (Row 3). To estimate the propensity score, we bin $PreEarn$ into five cells of similar size and use cell-specific averages as an input into the regression adjustment estimator of the form (ref).

First, in the full model $(X, D, Y)$, we find that the margin assumption likely fails. For eleven out of twelve education groups, the CATE is not significant at the $5\%$ level, so ties among the first-best treatment rules based on $Educ$ are very likely. Although the continuous covariate $PreEarn$ may alleviate the concern, violation of margin assumption can still be detected based on the heuristic in Remark (ref). The estimated welfare gap (Row 4, Column 4) is negative yet 23% of individuals would be treated differently than under the optimal policy (Row 4, Column 1), so the inequality (ref) is violated in-sample. To ensure validity of the reported LCB, we do not cross-fit. Using only two-thirds of the sample to compute the point estimate and its $95\%$ LCB incurs substantial efficiency loss, resulting in the lowest LCB in the Table.

Second, in the model $(PreEarn, D,Y)$, we do not detect sufficient heterogeneity to warrant personalized treatment assignment. Comparing Rows 2--3 with Row 1 yields welfare gaps of $-246$ and $-220$, relative to treating everyone. Since the estimated sign is negative,\footnote{ The estimated sign is negative due to the use of the efficient/doubly robust estimators (ref), which are not necessarily ordered in-sample.} the true gap is likely of the order sampling error. Moreover, the LCB in Row 1 exceeds those in Rows 2--3 by $40\%$ and $28\%$, respectively. We attribute these findings to potential biases in the first-stage estimators of regression functions and/or lack of heterogeneity in CATE function of $PreEarn$.

Next, we implement the LCBs proposed in Section (ref), choosing the test policies based on education level. We expect the treatment effects to be non-increasing in education level, with possible jumps at graduation years, $Educ=12$ and $Educ=16$. Thus, we limit the focus on cutoff policies of the form $\{Educ \leq C\}$. In particular, the policy $\{Educ \leq 11\}$ corresponds to treating only those who did not graduate from high-school (37.3% of the sample); $\{Educ \leq 12\}$ adds those who graduated from high-school but did not attend college (80.0% of the sample); $\{Educ \leq 15\}$ adds those who attended but did not graduate from college (95.9% of the sample); $\{Educ \leq 16\}$ adds college graduates (98.7% of the sample); and $\{Educ \leq 18\}$ corresponds to treating everyone.

Table (ref) presents LCBs obtained with the maximum test statistic and different test sets $\mathcal{G}_{\text{test}}$ determined by the cutoffs. In Row 1, the cutoff set corresponds to those who attended but did not graduate from high school and college, as well as everyone in the sample; and Row 2 includes all possible cutoffs. The first-best policy in both classes is $\{Educ \leq 15\}$ with the estimated welfare gain of 1440.25 USD, which exceeds all point estimates in Table (ref). For the first test class, the least-favorable and GMS confidence bands coincide and exceed all of the LCBs in Table (ref). The second test class contains policies that are far from optimal and thus provide loose lower bounds. As a result, the least-favorable test is conservative, while GMS leads to a tighter LCB.

table[table omitted — 1,843 chars of source]

\newcolumntype{Y}{>{\arraybackslash}X}

table[table omitted — 960 chars of source]

Conclusion

In this paper, we addressed the question of reporting a Lower Confidence Band on the optimal welfare in a policy learning problem. First, we documented the trade-off between welfare and precision and showed that it can be first-order. Second, we connected the first-order trade-off to the lack of uniformity in the margin condition of MammenTsybakov, Tsybakov. Finally, we proposed procedures for reporting Lower Confidence Bands that address the trade-off and remain valid regardless of the margin condition.