EconBase
← Back to paper

Nonparametric Uniform Inference in Binary Classification and Policy Values

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

101,004 characters · 22 sections · 81 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Nonparametric Uniform Inference in Binary Classification and Policy Values

abstract{5.9mm} We develop methods for nonparametric uniform inference in cost-sensitive binary classification, a framework that encompasses maximum score estimation, predicting utility maximizing actions, and policy learning. These problems are well known for slow convergence rates and non-standard limiting behavior, even under point identified parametric frameworks. In nonparametric settings, they may further suffer from failures of identification. To address these challenges, we introduce a strictly convex surrogate loss that point-identifies a representative nonparametric policy function. We then estimate this representative policy function to conduct inference on both the optimal classification policy and the optimal policy value. This approach enables Gaussian inference, substantially simplifying empirical implementation relative to working directly with the original classification problem. In particular, we establish root-$n$ asymptotic normality for the optimal policy value and derive a Gaussian approximation for the optimal classification policy at the standard nonparametric rate. Extensive simulation studies corroborate the theoretical findings. We apply our method to the National JTPA Study to conduct inference on the optimal treatment assignment policy and its associated welfare. \vskip0.4cm \noindentKeywords: cost-sensitive binary classification, Gaussian approximation and limit, nonparametric inference, statistical decision, surrogate loss. \vskip1cm \baselineskip=15pt

Introduction

Cost-sensitive binary classification problems encompass important classes of econometric models: e.g., the maximum score estimator manski1975maximum, manski1985semiparametric; utility maximization with a binary action elliott2013predicting; and a recent strand of research on treatment choice problems manski2000identification, manski2004statistical, stoye2009minimax, stoye2012minimax, hirano2009asymptotics, bhattacharya2012inferring, tetenov2012statistical, kitagawa2018should, athey2021policy, mbakop2021model, manski2021econometrics, viviano2025policy. See hirano2020asymptotic for a review of statistical decision analysis in econometrics.

A large majority of the existing literature focuses on analyzing learning performance, such as consistency, convergence rates, limit distribution, regret bounds, rather than on statistical inference. Inference, however, is well known to be a challenging problem. For $M$-estimation with a discontinuous objective, kim1990cube show that the estimated policy parameters exhibit a nonstandard (non-Gaussian) limit distribution with a cube-root-$n$ rate of convergence. Moreover, a naive bootstrap is not valid in this setting abrevaya2005bootstrap, leger2006bootstrap. Consequently, statistical inference for cost-sensitive binary classification problems remains challenging in practice, even when one considers a parametrized policy class.\footnote{With this said, it is worth emphasizing that there have been some developments toward practical inference in the parametric case, such as the smoothing approach horowitz1992smoothed, the subsampling approach delgado2001subsampling, seo2018local, the $m$-out-of-$n$ bootstrap approach lee2006m, bickel2011resampling, and a modified bootstrap approach cattaneo2020bootstrap, cheng2024inference.}

If a researcher further wishes to allow for a nonparametric policy class to accommodate more flexible classifiation policies, the problem becomes even more severe. Parametric problems are still point-identified after a suitable normalization manski1975maximum, manski1985semiparametric. By contrast, there is no obvious normalization that guarantees point identification when the policy function is nonparametric. Statistical inference under partially identified parameters poses significant challenges, and the development of valid and practical inference methods in this setting remains an open question.

This paper proposes a novel method of Gaussian inference for cost-sensitive binary classification problems with a nonparametric class of classification policy functions. The framework encompasses the aforementioned maximum score estimation, utility maximization under binary action, and treatment choice problems, among others. While this may seem daunting, if not infeasible, to those familiar with the literature, we attempt to develop a formal framework that establishes the validity of the proposed Gaussian inference procedure and demonstrate its practical ease.

Our proposed method of inference is easy to implement and hence practical, despite the inherently difficult nature of the problem. Unlike the method of horowitz1992smoothed, we do not require the choice of a tuning parameter to smooth the indicator. Unlike the method of delgado2001subsampling, we do not require the choice of a subsampling size. Furthermore, our inference method is based on the Gaussian distributions, which also facilitates an easy-to-implement bootstrap procedure.

Given the difficulty of the problem with partially identified nonparametric optimal policy functions, readers may wonder how we could enable Gaussian inference. The main novelty of this paper lies in the “identification” of a representative optimal classification policy function from among an infinite set of optimal classification policy functions. This is accomplished through the use of a strictly convex surrogate loss in place of the 0--1 loss. The use of surrogate losses has been explored in the statistical learning literature lugosi2004bayes, zhang2004statistical, steinwart2005consistency, bartlett2006convexity, zhao2012estimating, and kitagawa2023constrained adopt and extend this approach in the welfare maximization context. The motivations of these existing papers differ from ours: the existing work primarily aims to facilitate computation by convexifying the objective function while guaranteeing risk equivalence, whereas our objective is to establish the aforementioned “identification” of a representative optimal policy. While some of the existing papers advocate the use of the hinge loss for their respective purposes, we explicitly rule out the hinge loss as a candidate surrogate loss for our purpose. In this sense, our use and objective of the surrogate loss is orthogonal to those of the seemingly related papers in the literature.

Once the representing nonparametric optimizer is “identified,” we can proceed with econometric analyses for point-identified functional parameters, even though the underlying model admits infinitely many parameters in the identified set. We estimate this nonparametric representative policy function using a sieve approach. Our procedure goes beyond standard sieve M-estimation by accommodating infinite-dimensional nuisance parameters, such as the propensity score and conditional mean functions arising in the augmented inverse propensity score weighting estimator. In this setting, we achieve the standard nonparametric rate of convergence. Moreover, we establish a strong Gaussian approximation for the representative policy by applying Yurinskii's coupling inequality pollard2002, which in turn enables Gaussian inference for the optimal classification policy.

With our estimator for the optimal representative binary classification policy, we can estimate the optimal policy value by evaluating its (cross-fitted) empirical analog at our optimal policy estimator. We establish its root-$n$ consistency and limiting Gaussian distribution. This result is achieved through a combination of several techniques and notable properties of our problem to be outlined ahead, and it shares a similar spirit to chen2003estimation, who derive root-$n$ semiparametric results in related settings with non-smooth objective functions. In our setting, we show that a margin condition kitagawa2018should,ponomarev2024lower plays a crucial role in ensuring that the population policy value function is sufficiently smooth at the optimal policy, thereby laying the foundation for Gaussian inference for the optimal policy values. As we accommodate infinite-dimensional nuisance parameters in the estimation of a finite-dimensional parameter, our approach bears resemblance to double machine learning chernozhukov2018double. However, the asymptotic analysis differs from DML because we additionally employ a nonparametrically estimated functional optimizer, i.e., the classification policy.

matzkin1992nonparametric and mbakop2021model introduce nonparametric classification policy functions for estimation and learning in the contexts of (A) maximum score estimation and (B) empirical welfare maximization, respectively. We contribute to this literature by further developing inference methods and establishing their theoretical properties. Table (ref) summarizes the gaps in the existing literature that are filled by the present paper.

table[table omitted — 1,117 chars of source]

As already noted, in the maximum score literature (A), kim1990cube and horowitz1992smoothed develop limit distributions of the orignal maximum score estimator and the smoothed maximum estimator, respectively. chen2018best construct a best subset maximum score prediction rule that is linear in covariates by maximizing a penalized objective function. Recently, chen2025relu reformulate the maximum score estimator by rectified linear unit (ReLU) functions to achieve computation advantage and convergence rate that is faster than cube-root-n. All these work consider a linear index of finite number of covariates. Under the empirical welfare maximization framework (B), rai2018statistical conduct inference on optimal policy functions and maximized welfare under finite VC dimension. chernozhukov2025policy introduce a framework for selecting policies that maximize a penalized empirical welfare function with the penalty being proportional to a consistent estimator of the risk of each candidate policy. They primarily focus on cases in which the set of policies has finite carnality, with brief discussion of extensions to infinity. We contribute to the both literature, (A) and (B), by allowing the VC dimension of the classes of classification policy functions to diverge to infinity, so as to accommodate nonparametric policy functions.

In the context of treatment assignment, several plug-in approaches to inference exist, leveraging nonparametric nuisance parameters, such as the conditional average treatment effects, that are directly estimable bhattacharya2012inferring,armstrong2023inference,park2025debiased,chen2025inference. In contrast, our approach is aligned with empirical welfare maximization as a special case of cost-sensitive binary classification, and is therefore complementary to these plug-in methods.\footnote{See hirano2020asymptotic for details about the conceptual distinction between the plug-in approach and the empirical welfare maximization approach in statistical decisions. Our approach is also conceptually different from xu2025asymptotic, who analyzes the optimal point decision from a continuum of alternatives in limit experiments.} Furthermore, our general binary classification framework extends beyond treatment assignment settings to encompass maximum score and utility maximization frameworks.

Finally, it is worth noting several related results from the statistics and biostatistics literature. luedtke2016statistical develop root-$n$ rate confidence intervals for the optimal welfare even when the parameter is not pathwise differentiable. wu2021resampling propose a smoothing-based procedure for model-free inference on the optimal treatment. liang2022estimation employ surrogate convex losses to conduct inference on low-dimensional parameters. wu2023model develop a uniform inference method by quantifying the unknown link function, thereby enabling flexible modeling of interactions between treatment and covariates. The present paper, which adopts a surrogate-loss approach to uniform inference for nonparametric policies, complements specific aspects of each of these contributions.

The paper is organized as follows. Section (ref) introduces the setup and presents motivating examples. Section (ref) discusses the key issues arising from the discontinuous nature of the loss function. Section (ref) develops the “identification” of the representative classification policy function. Section (ref) presents estimation and inference procedures for the representative classification policy function, and Section (ref) establishes estimation and inference results for the optimal policy value. Section (ref) examines the finite-sample performance of the proposed methods through simulation studies. Section (ref) provides an empirical analysis that demonstrate the application of our method. Section (ref) concludes.

Setup and Examples

Let \(Z_i\) and \(X_i\), $i=1,2,\cdots, n$, denote random vectors of unit $i$ observed by researchers or policymakers. Some variables may be common to both \(Z_i\) and \(X_i\). Let \(\mathcal{G}\) denote a class of functions. With these notations, a broad class of econometric frameworks can be represented as a cost-sensitive binary classification problem:

align[align omitted — 227 chars of source]

Here, \(\psi^{+}\) and \(\psi^{-}\) denote non-negative functions that may depend on a possibly infinite-dimensional nuisance parameter \(\eta\), whose true value is denoted by \(\eta_0\). This general risk-minimization framework is motivated by three examples. Sections (ref), (ref), and (ref) below introduce the maximum score estimator, utility maximization problem, and welfare maximization framework, respectively.

Example I: Maximum Score

Let \( Y_i \) be a binary random variable and \( X_i \) a random vector such that \[ Y_i=\mathbf{1}\{g(X_i)+v_i\geq 0\}, \] where the error term $v_i$ satisfies $\text{Median}(v_i|X_i)=0$ almost surely. The maximum score estimator manski1975maximum, manski1985semiparametric for $g$ is defined as the sample analogue of

align*[align* omitted — 144 chars of source]

for some function class \( \mathcal{G} \). This can be equivalently reformulated as

align*[align* omitted — 168 chars of source]

where $Z_i := Y_i$, \( \psi^{+}(z,\eta) := 1-z\), and \( \psi^{-}(z,\eta) := z \). Therefore, the maximum score estimation can be viewed as a special case of the general risk-minimization framework with the risk given in (ref). $\blacktriangle$

Example II: Expected Utility Maximization with Binary Actions

elliott2013predicting consider the problem of maximizing the expected utility: $$ \max_{a(\cdot)}\mathbb{E}\left[U(a(X_i),Y_i,X_i)\right] $$ under a $\{-1,1\}$-valued binary outcome $Y_i$ with respect to a binary action $a(x) \in \{-1,1\}$ for each $x$. Imposing the condition that $ U(1,1,x) > U(-1,1,x) $ and $ U(-1,-1,x) > U(1,-1,x) $ for all $x$ in the support of $X$ to ensure correct classification yields higher utility, elliott2013predicting show that the problem can be reformulated as $\max_{g \in \mathcal{G}}V(g)$, where $$ V(g) = \mathbb{E}[ b(X_i) [Y_i+1-2c(X_i)] \cdot \mathbf{1}\{g(X_i) \ge 0\} -b(X_i) [Y_i+1-2c(X_i)] \cdot \mathbf{1}\{g(X_i)<0\}], $$ $b(x) = U(1,1,x) - U(-1,1,x) + U(-1,-1,x) - U(1,-1,x)$, and $c(x) = [U(-1,-1,x)-U(1,-1,x)] / b(x)$. The optimal action $a^*(x)$ can then be recovered from the optimal classifier $g^\ast$ as $a^\ast(x) = \mathrm{sign}(g^*(x))$ for each $x$. Like Example I (Section (ref)), the above maximization problem can be further reformulated as a special case of the general risk-minimization framework with the risk given in (ref) by taking $Z_i:=(Y_i,X_i')'$, $\psi^{+}(z,\eta) := |b(x)(y+1-2c(x))|-b(x)(y+1-2c(x))$, and $\psi^{-}(z,\eta) := |b(x)(y+1-2c(x))|+b(x)(y+1-2c(x))$. $\blacktriangle$

Example III: Welfare Maximization

Let \( Z_i = (Y_i, A_i, X_i')' \), where \( Y_i \) denotes a scalar outcome, \( A_i \) a binary treatment indicator, and \( X_i \) a vector of covariates supported on $\mathcal{X}$. Let \( Y_i(1) \) and \( Y_i(0) \) denote the potential outcomes under treatment and control, respectively. The observed outcome is generated according to \[ Y_i = Y_i(1) \cdot A_i + Y_i(0) \cdot (1 - A_i). \] Define the propensity score function as \( \pi(x) = \mathbb{P}(A_i = 1 \mid X_i = x) \).

A policymaker is interested in selecting a policy function \( g \) from \( \mathcal{G} \) that maximizes welfare, measured by the mean outcome:

align[align omitted — 125 chars of source]

To analyze this objective, the literature commonly imposes the following set of assumptions.

enumerate[(WM.i)] • {\bf Unconfoundedness:} $(Y_i(1), Y_i(0)) {\mathrel{\perp\!\!\!\perp}} A_i~\vert~X_i$; • {\bf Moments:} $\mathbb{E}[\vert Y_{i}(a)\vert]<\infty$ for each $a=0, 1$ and; • {\bf Overlap:} There exists $\underline{C} \in (0,1/2)$ such that $\mathbb{P}(\underline{C}\leq \pi_{0}(x)\leq 1-\underline{C})=1$ for all $x\in\mathcal{X}$.

Under these assumptions (WM.(ref))--(WM.(ref)), the mean outcome (ref) can be identified by the inverse propensity score weighting (IPW) as in kitagawa2018should, which in turn may be rewritten in the augmented-IPW (AIPW) as in athey2021policy:

align*[align* omitted — 277 chars of source]

where $\mu(l, x)=\mathbb{E}[Y_{i} \vert A_{i}=l, X_{i}=x]$ for $l \in \{0,1\}$. Collecting the nuisance parameters \( \eta := (\mu(1,\cdot), \mu(0,\cdot), \pi(\cdot)) \), we express the \( g \)-relevant component of the identified objective, \( V(g) \), as

align[align omitted — 395 chars of source]

Denote the positive components of $\psi_1(Z_i,\eta)$ and $\psi_0(Z_i,\eta)$ by $\psi_1^{+}(Z_i,\eta)$ and $\psi_0^{+}(Z_i,\eta)$, respectively, and similarly denote the negative parts of $\psi_1(Z_i,\eta)$ and $\psi_0(Z_i,\eta)$ by $\psi_1^{-}(Z_i,\eta)$ and $\psi_0^{-}(Z_i,\eta)$, respectively. Then, the welfare (ref) can be rewritten as

align[align omitted — 844 chars of source]

where $\psi^{+}(Z_i,{\eta}) := \psi_1^{-}(Z_i,\eta)+ \psi_0^{+}(Z_i,\eta)$ and $\psi^{-}(z,{\eta}) := \psi_1^{+}(Z_i,\eta)+ \psi_0^{-}(Z_i,\eta)$. Note that the second expectation in the last line takes the form of (ref). Thus, the welfare maximization problem can be viewed as a special case of the general risk-minimization framework with a risk of the form (ref). $\blacktriangle$

Issues

As Sections (ref)--(ref) demonstrate, the cost-sensitive binary classification problem $\min_{g \in \mathcal{G}} L(g)$ with the risk function given by (ref) encompasses several important econometric frameworks. However, this problem is more challenging to address than standard econometric problems. The difficulty becomes even more pronounced when allowing for a nonparametric class $\mathcal{G}$ of policy functions.

The primary challenge arises from the discontinuous nature in the sample analog of the objective function. In practice, this discontinuity induces computational difficulties: the sample objective function is non-convex, making numerical optimization difficult, if not infeasible. Beyond these practical concerns, the discontinuity also has adverse theoretical implications, giving rise to non-standard limiting behavior and non-identification, as detailed in the following two subsections.

Non-Standard Limit under Identifying Normalizations

To simplify the discussion, suppose for the moment that \(g\) belongs to a parametric class specified as \(g(x) = x^\prime \gamma\), with appropriate normalizations. In this case, the problem \(\min_{g \in \mathcal{G}} L(g)\) reduces to an \(M\)-estimation of \(\gamma\) with a discontinuous sample objective function. In this setting, kim1990cube show that the resulting \(M\)-estimator converges at the cube-root-\(n\) rate and admits a non-standard asymptotic distribution. Consequently, naive bootstrap procedures are invalid in this context abrevaya2005bootstrap,leger2006bootstrap.

Non-Identification in General

Although Section (ref) focuses on a simple parametric class, it already highlights the inherent difficulty of the problem. When \(\mathcal{G}\) is a nonparametric class, the challenge becomes even more severe, as the optimal policy function \(g\) is no longer point-identified even after normalization. This is because only the sign of \(g\) affects the risk in (ref). In particular, if \(g^{\ast} \in \mathcal{G}\) is a solution, then any \(g^{\ast\ast} \in \mathcal{G}\) such that \(\text{sign}(g^{\ast\ast}(x)) = \text{sign}(g^{\ast}(x))\) for all \(x\) is also a solution. In a nonparametric class \(\mathcal{G}\), one can readily construct infinitely many such functions \(g^{\ast\ast}\) that are sign-equivalent to \(g^{\ast}\), thereby rendering the optimal policy function unidentifiable. Theoretical analysis of these partially identified nonparametric models poses even greater challenges than those encountered in the parametric case of Section (ref).

Representative Identification

The difficulties associated with (ref), as discussed in the previous section, stem from the discontinuity of the 0–1 loss function. We remedy them by replacing the 0–1 loss with a strictly convex surrogate loss. This substitution not only yields a continuous and convex objective to ease computation, but also ensures point identification of the optimal policy function even when $\mathcal{G}$ is nonparametric.

Let $\phi: \mathbb{R} \rightarrow \mathbb{R}$ be a surrogate loss function, and consider the following modification of (ref):

align[align omitted — 149 chars of source]

The following theorem presents the utility of using this function $\phi(\cdot)$ in place of the original 0-1 loss.

thm[Identification] Suppose that $\mathbb{E}[\psi^+(Z_i,\eta_0)|X_i] -\mathbb{E}[\psi^-(Z_i,\eta_0)|X_i] \neq 0$ almost surely and $\mathcal{G}$ includes all measurable functions. Let $\phi(\cdot)$ be differentiable at zero and $\partial \phi(0)<0$. \begin{enumerate}[(i)] • If $\phi(\cdot)$ is convex, then any $g_\ast \in\arg\min_{g\in\mathcal{G}} L^\ast(g)$ satisfies $g_\ast\in \arg\min_{g \in \mathcal{G}} L(g)$. • If $\phi(\cdot)$ is strictly convex and $\mathcal{G}$ is convex, then there is unique $g_\ast$ such that $g_\ast = \arg\min_{g\in\mathcal{G}} L^\ast(g)$.\footnote{The uniqueness is up to the equivalence class in $\mathcal{G}$ identified by the underlying probability measure.} \end{enumerate}

Theorem (ref)(i) guarantees that minimizing the surrogate risk \( L^\ast \) yields an optimal policy function with respect to the original risk \( L \), as established in the literature bartlett2006convexity. Furthermore, Theorem (ref)(ii) ensures that the surrogate risk admits a unique minimizer (up to \( L_1 \)-equivalence). In other words, by employing a strictly convex surrogate function \( \phi(\cdot) \), we achieve point identification of a nonparametric solution \( g_\ast \) to the original problem \( \min_{g \in \mathcal{G}} L(g) \). This, in turn, facilitates the application of econometric methods based on point estimation and their associated asymptotic properties, which are typically more tractable than those relying on set identification.

The surrogate loss approach has been explored in the literature bartlett2006convexity,zhao2012estimating,kitagawa2023constrained; however, the motivation behind our use of it is fundamentally different. While existing studies primarily leverage surrogate losses to enable numerical optimization by convexifying the originally non-convex objective, our primary aim is to establish identification of the optimal policy function. Moreover, while zhao2012estimating and kitagawa2023constrained advocate the hinge loss for their purposes, we explicitly rule out the hinge loss in part (ii) of our Theorem (ref) and instead promote alternative loss functions that better align with our identification objective -- see below for a concrete list. In this respect, both our goals and the methodological approaches are largely orthogonal to those in the literature. To the best of our knowledge, the use of surrogate losses for the purpose of identification (or representation), rather than computational convenience, is novel.

Although we identify only one representative solution \( g_\ast \), this suffices for conducting inference on the optimal classification policy. Indeed, \( g_\ast \) is sign-equivalent to any solution of $\min_{g \in \mathcal{G}}L(g)$ up to a null set. Accordingly, the optimal classification rule defined by \(x \mapsto \mathbf{1}\{x \in \mathcal{X} \mid g_\ast(x) \geq 0\} \) delivers the same policy decisions as the rule \( x \mapsto \mathbf{1}\{x \in \mathcal{X} \mid g(x) \geq 0\} \) associated with any solution \( g \in \arg\min_{g \in \mathcal{G}} L(g) \), possibly except on a set of measure zero. Thus, we may perform inference on the optimal classification policy and the corresponding optimal policy value (e.g., welfare) using the representative, point-identified representative policy function \( g_\ast \), effectively treating it as the unique solution.

Various choices for the function \( \phi \) have been proposed in the literature: the hinge loss, the exponential loss, the logistic loss, the squared loss, and the sigmoid loss, among others. While part (i) of Theorem (ref) accommodates all of these examples, part (ii) rules out the hinge loss.

The use of different loss functions leads to different representers $g_*$. However, as argued above, these representers yield the same optimal classification rule, since they agree in sign except on a null set. As such, the choice of loss function, and hence the resulting representer, may be viewed as analogous to selecting different parameter normalizations proposed by manski1975maximum, manski1985semiparametric in the context of the maximum score framework (Example I) introduced in Section (ref).

The assumption that

align*[align* omitted — 116 chars of source]

plays a crucial role. If there exists an event with positive probability on which

align*[align* omitted — 96 chars of source]

then the function $g$ could take either positive or negative values on that event, thereby failing to point-identify the solution. This condition is essential not only for point identification but also for the inference on the optimal policy value; see, for example, luedtke2016statistical for related discussions.

Estimation and Inference for the Representative Policy Function

Given the `identification' result from Section (ref), one can obtain a point estimate of the representative classification policy function $g_\ast$. Although $g_\ast$ is only one among infinitely many solutions, it suffices for constructing the optimal classification rule $x \mapsto \mathbf{1}\{g_\ast(x) \geq 0\}$, since only the sign of $g_\ast$ matters and any other solution has an equivalent sign almost everywhere. Consequently, the optimal classification policy and the optimal policy value (e.g., welfare) can be estimated.

Estimation of the Representative Policy Function

Let $\mathcal{X}$ denote the support of $X_i$, and $\{p_1(\cdot),p_2(\cdot),...\}$ be a sieve basis consisting of measurable functions $p_k: \mathcal{X} \rightarrow \mathbb{R}$, $k \in \mathbb{N}$. For each $k \in \mathbb{N}$, let $\mathcal{G}_k = \{x \mapsto p_{1:k}(x)'\beta \ : \ \beta \in \mathbb{R}^k\}$ denote the sieve space, where $p_{1:k}(\cdot) := (p_1(\cdot),...,p_k(\cdot))'$. Let $g_k$ denote the $L_2(\mathbb{P})$-projection of $g_\ast \in \mathcal{G}$ onto $\mathcal{G}_k$, and let $r_g = g_\ast-g_k$ be its approximation error. Denote the complexity measure by $\xi_k := \sup_{x \in \mathcal{X}}\|p_{1:k}(x)\|$.

We now introduce the cross-fitting method for estimating \(g_\ast\). Partition the index set \([n] = \{1,\ldots,n\}\) into \(m\) folds \(I_{1}, \ldots, I_{m}\), and let \(I^{c}_{(\ell)} := [n] \setminus I_{(\ell)}\) denote the complement of the \(\ell\)-th fold. Assume that the folds are approximately balanced, so that $ n_{(1)} \asymp \cdots \asymp n_{(m)}, $ where \(n_{(\ell)} := |I_{(\ell)}|\), and the number of folds \(m\) is fixed.

For a given fold \(\ell \in [m] = \{1,\ldots,m\}\), obtain an estimator \(\widehat{\eta}_{(\ell)}\) of the nuisance parameter \(\eta_0\) using the complementary subsample \(I_{(\ell)}^c\). Then, estimate the representative policy function \(g_\ast\) on the held-out fold \(I_{(\ell)}\). Specifically, the fold-wise estimator of \(g_\ast\) is given by

align*[align* omitted — 369 chars of source]

\(p_{1:k}(\cdot) := (p_1(\cdot), \ldots, p_k(\cdot))^\prime\), \(\widehat{\psi}^{\pm}_{(\ell)}(\cdot) := \psi^{\pm}(\cdot, \widehat{\eta}_{(\ell)})\) for shorthand, and \(\mathbb{E}_{n,(\ell)}[f] := \frac{1}{|I_{(\ell)}|} \sum_{i \in I_{(\ell)}} f(Z_i)\) denotes the sample mean over the \(\ell\)-th fold.

Finally, aggregate the \(m\) fold-wise estimators to obtain the cross-fitted estimator

align[align omitted — 98 chars of source]

For later use, define \(\widehat{\eta}\) as the full-sample estimator of the nuisance parameter \(\eta_0\), and let \(\widehat{\psi}^\pm(\cdot) := \psi^\pm(\cdot, \widehat{\eta})\). This estimator will be employed for variance estimation and bootstrap inference. The conditions required for both the cross-fitted and full-sample estimators are presented in Section (ref).

Large Sample Theory for the Representative Policy Function Estimation

We next establish the uniform convergence rate of \(\widehat{g}\) toward \(g_\ast\) and prove a Gaussian coupling principle that characterizes the Gaussian approximation of its estimation error.

We now introduce a set of regularity conditions.

as[Sieves] Let the following conditions be satisfied: \begin{enumerate}[(i)] • (Bounded Moments) $\mathbb{E}[\vert\psi^+(Z_i, \eta)\vert^3\vert X_i=x] < \infty$ and $\mathbb{E}[\vert\psi^-(Z_i, \eta)\vert^3\vert X_i=x] < \infty$ uniformly for all $x \in \mathcal{X}$. • (Eigenvalues) For $Q := \mathbb{E}[p_{1:k}(X_i)(\psi^{+}(Z_i)\cdot \partial^2\phi(-g_\ast(X_i))+\psi^{-}(Z_i)\cdot \partial^2\phi(g_\ast(X_i))) p_{1:k}(X_i)^{\prime}]$ and $\Sigma := \mathbb{E}[p_{1:k}(X_i)(-\psi^{+}(Z_i)\cdot \partial\phi(-g_\ast(X_i))+\psi^{-}(Z_i)\cdot \partial\phi(g_\ast(X_i)))^2 p_{1:k}(X_i)^{\prime}]$, there exist constant values $C_{\min}$, $C_{\max} \in (0,\infty)$ such that $C_{\min} < \lambda_{\min}(Q) < \lambda_{\max}(Q) < C_{\max}$ and $C_{\min} < \lambda_{\min}(\Sigma) < \lambda_{\max}(\Sigma) < C_{\max}$ with $ \lambda_{\min}(\cdot)$ and $\lambda_{\max}(\cdot)$ indicating the minimum and maximum eigenvalues. • (Growth Condition) The complexity measure $\xi_k$ grows sufficiently slowly such that it satisfies $\sqrt{\xi_k^2 \log(k) / n} \to 0$ and $ \log(\xi_k) \lesssim \log(k)$ • (Approximation Error) There is a sequence $\{B_k\}_k$ of bounds on approximation errors associated with the sequence $\{\mathcal{G}_k\}_k$, such that $\sqrt{n}B_k\log(k)\rightarrow0$ and \begin{align*} \Vert r_g\Vert_{\mathbb{P},\infty}:=\sup _{x \in \mathcal{X}}\vert r_g(x)\vert \leq B_k, \end{align*} \end{enumerate}

Assumption (ref) collects standard regularity conditions commonly used in the series-based nonparametric estimation literature and adapts them to our cost-sensitive binary classification problem. Assumption (ref)(ref) requires the conditional moments of the weights to be bounded and serves multiple purposes. In particular, the existence of the third moment is needed to establish a coupling inequality. Assumption (ref)(ref), which rules out collinearity in the population signal matrix, is a conventional requirement in nonparametric analysis newey1997convergence. Assumptions (ref)(ref) and (ref) control the complexity of the function class and the nonparametric approximation error, respectively.

To facilitate our subsequent writings, we define some short-hand notations:

align*[align* omitted — 353 chars of source]

where $\partial\phi(g)$ and $\partial^2\phi(g)$ denote the first- and second-order derivatives of $\phi$, respectively. With these notations, we state the following conditions on the estimators of the nuisance parameter that enter the weights.

as[Nuisance Parameter Estimators] There exists a sequence $\mathcal{S}_n$ of shrinking subsets such that $g_\ast, g_k, \widehat g \in \mathcal{S}_n$ with probability at least $1-\Delta_n$ such that $\Delta_n \to 0$ and the following conditions are satisfied. \begin{enumerate}[(i)] • (Cross-Fitting)For each $\ell \in [m]$ and any $g_1, g_2 \in \mathcal{S}_n$, the cross-fitted nuisance parameter estimator satisfies \begin{align} &\left\Vert\mathbb{E}_{n,(\ell)} \left[\left(\Psi_1(Z_i, \widehat{\eta}_{(\ell)},g_1)- \Psi_1(Z_i,\eta_0,g_\ast)\right)p_{1:k}(X_i) \right]\right\Vert= o_p(\sqrt{1/(n\log(k)^2)}),\\ &\left\Vert\mathbb{E}_{n,(\ell)} \left[\left(\psi^{\pm}(Z_i, \widehat{\eta}_{(\ell)})- \psi^{\pm}(Z_i,\eta_0)\right)p_{1:k}(X_i)p_{1:k}(X_i)^\prime\right]\right\Vert= o_p(\xi_k\sqrt{\log(k)/n}),\\ &\left\Vert\mathbb{E}_{n,(\ell)} \left[\left(\Psi_2(Z_i, \widehat{\eta}_{(\ell)},g_1,g_2)- \Psi_2(Z_i,\eta_0,g_\ast,g_\ast)\right)p_{1:k}(X_i)p_{1:k}(X_i)^\prime\right]\right\Vert= o_p(\xi_k\sqrt{\log(k)/n}),\\ &\left\Vert\mathbb{E}_{n,(\ell)} \left[\left(\Psi_1(Z_i,\widehat{\eta}_{(\ell)},g_1)^2- \Psi_1(Z_i,\eta_0,g_\ast)^2\right)p_{1:k}(X_i)p_{1:k}(X_i)^\prime\right]\right\Vert= o_p(\xi_k\sqrt{\log(k)/n}). \end{align} • (Full-Sample) For any $g_1, g_2 \in \mathcal{S}_n$ of $g_\ast$, the full-sample nuisance parameter estimator satisfies \begin{align} &\left\Vert\mathbb{E}_{n} \left[\left(\Psi_1(Z_i, \widehat{\eta},g_1)- \Psi_1(Z_i,\eta_0,g_\ast)\right)p_{1:k}(X_i) \right]\right\Vert= o_p(\sqrt{1/(n\log(k)^2)}),\\ &\left\Vert\mathbb{E}_{n} \left[\left(\psi^{\pm}(Z_i, \widehat{\eta})- \psi^{\pm}(Z_i,\eta_0)\right)p_{1:k}(X_i)p_{1:k}(X_i)^\prime \right]\right\Vert= o_p(\xi_k\sqrt{\log(k)/n}),\\ &\left\Vert\mathbb{E}_{n} \left[\left(\Psi_2(Z_i, \widehat{\eta},g_1,g_2)- \Psi_2(Z_i,\eta_0,g_\ast,g_\ast)\right)p_{1:k}(X_i)p_{1:k}(X_i)^\prime \right]\right\Vert= o_p(\xi_k\sqrt{\log(k)/n}),\\ &\left\Vert\mathbb{E}_{n} \left[\left(\Psi_1(Z_i,\widehat{\eta},g_1)^2- \Psi_1(Z_i,\eta_0,g_\ast)^2\right)p_{1:k}(X_i)p_{1:k}(X_i)^\prime\right]\right\Vert= o_p(\xi_k\sqrt{\log(k)/n}). \end{align} \end{enumerate}

Assumption (ref) imposes rate restrictions on the estimation error of the subsample and full-sample estimators for $\eta_0$. Conditions (ref) and (ref) are used for the Gaussian coupling of large-dimensional score vectors, Conditions (ref)--(ref) and (ref)--(ref) ensure the convergence rate of the sample Gram matrix, and Conditions (ref) and (ref) guarantee the convergence rate of the variance estimator.\footnote{The sample Gram matrix and the variance estimator will be introduced later in (ref) and (ref) respectively.}

In the context of Example I (Maximum Score; Section (ref)) and Example II (Expected Utility Maximization; Section (ref)), there is no nuisance parameter \(\eta\), so Assumption (ref) is trivially satisfied. In the context of Example III (Welfare Maximization; Section (ref)), one can adopt alternative estimators subject to suitable low-level conditions to ensure that Assumption (ref) holds. Such examples include the $\ell_1$-penalized and related methods in sparse models belloni2013inference and other machine learning methods under appropriate conditions. For the current work, we focus on $\ell_1$-penalized methods for estimating our nuisance parameters using full-sample and sample-splitting techniques.

Finally, we impose a smoothness condition on the surrogate loss and the representative policy function to facilitate the derivation of their asymptotic properties.

as[Surrogate Loss] The surrogate loss function $\phi$ is twice continuously differentiable.
as[Representative Policy] The representative classification policy function $g_\ast$ is uniformly bounded on compact $\mathcal{X}$ and belongs to a Hölder ball $\Lambda^h(L)$ of smoothness order $h> 1/2$ and radius $L \in (0,\infty)$.\footnote{See (ref) in the appendix for the definition of the Hölder ball $\Lambda^h(L)$ of smoothness order $h$ and radius $L$.}

For standardization, we introduce the variance process $\sigma^2(x)=p_{1:k}(x)^{\prime}Q^{-1}\Sigma Q^{-1}p_{1:k}(x)$, where

align*[align* omitted — 334 chars of source]

We may estimate it by the sample counterpart $\widehat{\sigma}^2(x)=p_{1:k}(x)^{\prime}\widehat{Q}^{-1}\widehat{\Sigma}\widehat{Q}^{-1}p_{1:k}(x)$ where

align[align omitted — 466 chars of source]

With the above definitions and notation in place, the following theorem establishes the uniform convergence rate and Gaussian coupling for the representative policy function estimator \(\widehat{g}\).

thmIn addition to the conditions for Theorem (ref) (ii), suppose that Assumptions (ref), (ref), (ref) and (ref) hold. Then, we have \begin{align*} \sup_{x\in\mathcal{X}}\vert\widehat{g}(x)-g_\ast(x)\vert\lesssim_p \xi_{k} \sqrt{\log (k)/n}+B_{k}. \end{align*} In addition, letting $\mathcal{N}_{k}\sim\mathcal{N}(\mathbf{0}_{k},I_{k\times k})$, we have \begin{align} \sqrt{n}\frac{(\widehat{g}(x)-g_\ast(x))}{\Vert \widehat{\sigma}(x)\Vert}=_d\frac{\sigma(x)^{\prime}}{\Vert \sigma(x)\Vert}\mathcal{N}_k+o(1/\log(k)) uniformly for all $x\in\mathcal{X}$, \end{align} provided that $\sup _{x \in \mathcal{X}} \sqrt{n}\vert r_g(x)\vert /\Vert \sigma(x)\Vert=o(1/\log(k))$ and $\log(k)^{6}k^4\xi^{2}_{k}\log(n)^2/n\rightarrow0$.

A few remarks are in order. First, the point identification result from Section (ref) is the key enabler of our asymptotic analysis and plays a crucial role in constructing valid inference procedures. Second, as a consequence, our estimator achieves the standard nonparametric rate of convergence established in the sieve literature, despite the originally non-standard problem involving partial identification. Third, in addition to the requirements stated in Theorem (ref)(ii), Assumption (ref) requires the surrogate loss to be twice continuously differentiable. This condition still accommodates all of the aforementioned options, again except the hinge loss.

Since our nonparametric policy function estimator is strongly approximated by a Gaussian process whose distribution depends on the data only through the covariance, it is feasible to conduct uniform nonparametric inference using bootstrapped critical values. The details of this inference procedure for \(g(\cdot)\) are presented in Section (ref).

Bootstrap Inference for the Representative Policy Function

Given Theorem (ref), it is natural to consider the following \(t\)-statistic process:

align[align omitted — 114 chars of source]

The sieve score bootstrap chen2015sieve can be employed to compute the relevant critical values. Specifically, let \(\{\omega_{i}\}_{i}\) be i.i.d. standard Gaussian random variables, independent of the data. Consider the score-bootstrap \(t\)-statistic process

align[align omitted — 301 chars of source]

for \(x \in \mathcal{X}\), where \(\widehat{\psi}^\pm(\cdot)\) is defined in Section (ref), \(\widehat{Q}\) and \(\widehat{\sigma}\) are defined in Section (ref), and $\mathbb{G}_n[f]=\sqrt{n}(\mathbb{E}_n[f]-\mathbb{E}[f])$ denotes the empirical process under the i.i.d. sampling scheme.

To compute the critical value \(cv_{n}(1-\alpha)\) of the supremum statistic \(\sup_{x \in \mathcal{X}} |t(x)|\) in (ref), one can evaluate \(\sup_{x \in \mathcal{X}} |t^{b}_{g}(x)|\) based on a large number of independent draws of \(\{\omega_{i}\}_{i}\). The critical value \(cv_{n}(1-\alpha)\) is then approximated as

align[align omitted — 165 chars of source]

Construct the uniform confidence band (UCB) for the representative policy function \(g_\ast\) as

align[align omitted — 251 chars of source]

where \(cv^{b}_{n}(1-\alpha)\) guarantees that \(g_\ast(x) \in [\widehat{g}^{b}_{l}(x), \widehat{g}^{b}_{u}(x)]\) uniformly for all \(x \in \mathcal{X}\) with asymptotic confidence level \(100(1-\alpha)\%\), as formally stated in the following theorem.

We introduce the notation $\Psi^\ast_1(Z_i,\omega_i,\eta,g):=\omega_i(-\psi^{+}(Z_i,\eta)\cdot\partial\phi(-g(X_i))+\psi^{-}(Z_i,\eta)\cdot\partial\phi(g(X_i)))$, and the following assumption on the nuisance parameter estimation analogously to Assumption (ref):

asFor any $g$ belonging to a shrinking neighborhood $\mathcal{S}_n$ of $g_\ast$, the full-sample nuisance parameter estimator satisfies \begin{align*} &\left\Vert\mathbb{E}_{n} \left[\left(\Psi_1^\ast(Z_i,\omega_i, \widehat{\eta},g)- \Psi_1^\ast(Z_i,\omega_i,\eta_0,g_\ast)\right)p_{1:k}(X_i) \right]\right\Vert= o_p(\sqrt{1/(n\log(k)^2)}). \end{align*}

The following theorem establishes the uniform probability coverage.

thm[Uniform Inference via Two-Sided UCB] In addition to the conditions of Theorem (ref), suppose that Assumption (ref) holds. With the critical value $cv^{b}_{n}\left(1-\alpha\right)$ constructed in (ref), the uniform confidence band (UCB) defined in (ref) satisfies \begin{align} &\mathbb{P}\left\{g_\ast(x)\in[\widehat{g}^{b}_{l}(x),\widehat{g}^{b}_{u}(x)] for all x\in\mathcal{X}\right\}=1-\alpha+o\left(1\right). \end{align}

There has been progress in developing inference methods for parametric policy functions wu2021resampling,liang2022estimation,cheng2024inference. However, to the best of our knowledge, inference for nonparametric policy functions is first addressed in this paper, thereby filling an important gap in the literature.

To compute the one-sided critical values, one can evaluate \(\sup_{x \in \mathcal{X}} t^{b}_{g}(x)\) and \(\inf_{x \in \mathcal{X}} t^{b}_{g}(x)\) based on a large number of independent draws of \(\{\omega_{i}\}_{i}\), as before. The upper and lower critical values, \(cv_{n}^{1,b}(1-\alpha)\) and \(cv_{n}^{1,b}(\alpha)\), are then computed as

align[align omitted — 357 chars of source]

We construct the one-sided uniform confidence bands (UCBs) for testing the uniform null hypotheses \(\mathcal{H}_0: g_\ast(x) \leq 0 \;\;\forall x \in \mathcal{X}\) or \(\mathcal{H}_0: g_\ast(x) \geq 0 \;\;\forall x \in \mathcal{X}\):

align[align omitted — 293 chars of source]

Here, (ref) and (ref) ensure that \(\sup_{x \in \mathcal{X}} \widehat{g}^{b}_{1,l}(x) \leq 0\) with asymptotic confidence level \(100(1-\alpha)\%\) under \(\mathcal{H}_0: g_\ast(x) \leq 0 \;\;\forall x \in \mathcal{X}\). Similarly, (ref) and (ref) ensure that \(\sup_{x \in \mathcal{X}} \widehat{g}^{b}_{1,u}(x) \geq 0\) with asymptotic confidence level \(100(1-\alpha)\%\) under \(\mathcal{H}_0: g_\ast(x) \geq 0 \;\;\forall x \in \mathcal{X}\). This result is formalized in the theorem below.

thm[Uniform Inference via One-Sided UCB] In addition to the conditions of Theorem (ref), suppose that Assumption (ref) holds. With the critical values $cv^{1,b}_{n}(1-\alpha)$ constructed in (ref), the one-sided lower UCB defined in (ref) satisfies \begin{align} &\mathbb{P}\left\{\widehat{g}^{b}_{1,l}(x)\leq 0 for all x\in\mathcal{X}\right\}\geq 1-\alpha+o\left(1\right) under \mathcal{H}_0:g_\ast(x)\leq 0 for all x \in \mathcal{X}. \end{align} Similarly, with the critical values $cv^{1,b}_{n}\left(\alpha\right)$ constructed in (ref), the one-sided upper UCB defined in (ref) satisfies \begin{align} &\mathbb{P}\left\{\widehat{g}^{b}_{1,u}(x)\geq 0 for all x\in\mathcal{X}\right\}\geq 1-\alpha+o\left(1\right) under \mathcal{H}_0:g_\ast(x)\geq 0 for all x \in \mathcal{X}. \end{align}

Estimation and Inference for the Optimal Policy Value

Suppose that a policymaker is interested in making inference on the optimal policy value of the form

align*[align* omitted — 205 chars of source]

with $ \psi_1(z,\eta) := \overline{\psi}(z,\eta) - \psi^{+}(z,\eta) $ and $ \psi_0(z,\eta) := \overline{\psi}(z,\eta) - \psi^{-}(z,\eta) $ for some function \(\overline{\psi}\).

Note that under this setup, the welfare function \(V\) can be rewritten as

align*[align* omitted — 68 chars of source]

where $L$ is defined in (ref). Therefore, the representative classification policy function \(g_\ast\) identified via the surrogate loss approach attains the maximum level of welfare within \(\mathcal{G}\); that is, \(V_0 = V(g_\ast)\). Although the representative policy function \(g_\ast\) is obtained by minimizing the surrogate risk \(L^\ast\) in (ref) to ensure point identification, this surrogate risk \(L^\ast\) does not coincide with the value of the original risk \(L\) in (ref). To estimate the optimal value \(V_0\), we need to evaluate the original risk \(L\), rather than the surrogate risk \(L^\ast\), at an estimate of \(g_\ast\).

For instance, in the context of Example III discussed in Section (ref), \(V_0\) corresponds to the optimal welfare value in (ref), with the functions \(\psi_1\) and \(\psi_0\) defined in (ref) and \(\overline{\psi}(z,\eta) = \psi_1^{+}(z,\eta) + \psi_0^{-}(z,\eta)\).

Estimation of the the Optimal Policy Value

Let \(\widehat{g}\) denote the estimator of the representative classification policy function \(g_\ast\), as obtained in Section (ref). To account for the estimation of \(\eta\), we employ a cross-fitting procedure to estimate the optimal policy value. Partition \([n] = \{1,\ldots,n\}\) into \(m\) folds \(I_{1}, \ldots, I_{m}\). Assume that the folds are approximately balanced, i.e., \(n_{(1)} \asymp \cdots \asymp n_{(m)}\), where \(n_{(\ell)} := |I_{(\ell)}|\), and that \(m\) is fixed. For a given fold \(\ell \in [m] = \{1,\ldots,m\}\), obtain an estimator \(\widehat{\eta}_{(\ell)}\) of \(\eta_0\) using the complementary subsample \(I_{(\ell)}^c\). We then estimate the optimal policy value \(V_0\) via the plug-in cross-fitted estimator

align[align omitted — 264 chars of source]

where \(\widehat{\psi}_{1,(\ell)}(Z_i) := \psi_1(Z_i,\widehat{\eta}_{(\ell)})\) and \(\widehat{\psi}_{0,(\ell)}(Z_i) := \psi_0(Z_i,\widehat{\eta}_{(\ell)})\) for all \(i \in I_{(\ell)}\) and each \(\ell \in [m]\).

In the context of Example III in Section (ref), the estimated weights take the form

align*[align* omitted — 345 chars of source]

where \(\widehat{\eta}_{(\ell)} = \big(\widehat{\mu}_{(\ell)}(1,\cdot),\, \widehat{\mu}_{(\ell)}(0,\cdot),\, \widehat{\pi}_{(\ell)}(\cdot)\big)\) denotes the estimator of the conditional mean functions \(\mu(1,\cdot)\) and \(\mu(0,\cdot)\), as well as the propensity score function \(\pi(\cdot)\), constructed using the complementary subsample \(I_{(\ell)}^c\) for all \(i \in I_{(\ell)}\).

Asymptotic Properties for the Optimal Policy Value

The optimal policy value estimator (ref) is constructed via a plug-in rule using our estimator $\widehat{g}$ of the representative policy function $g_\ast$, as given in Section (ref). For valid inference, $\widehat{g}$ must satisfy conditions ensuring that both its estimation error and its surrogate misspecification error are dominated and hence do not affect the limiting distribution of the policy value estimator. Theorem (ref) shows that our estimator $\widehat{g}$ converges at the rate $\xi_{k}\sqrt{\log(k)/n}$. We verify below that, under a set of mild conditions, this nonparametric rate suffices to ensure its impact on the distribution of the policy value estimator is asymptotically negligible.

as[Margin Condition] The random variable $U_i := \mathbb{E}[\psi_1(Z_i,\eta_0)|X_i] -\mathbb{E}[\psi_0(Z_i,\eta_0)|X_i]$ admits a continuous probability density function $f_U$ over a small neighborhood $[-\delta,\delta]$ of zero.

Assumption (ref) corresponds to the standard margin condition widely used in the empirical welfare maximization literature (see, for example, kitagawa2018should and ponomarev2024lower). This assumption requires the data to exhibit a “low-noise” property. Specifically, if the random variable $U_i$ has a continuous density around zero, then $\mathbb{P}(|U_i|<t) = O(t)$ for sufficiently small $t>0$, which further implies $\mathbb{E}[|U_i|1\{|U_i|<t\}] = O(t^2)$. Intuitively, when $|U_i|$ is small---meaning that the welfare difference between the two actions is nearly tied---the estimated optimal rule is more likely to misclassify the optimal decision, but each such misclassification incurs only a small welfare loss. If these near-tie events occur infrequently, as ensured by the margin condition, the expected welfare loss regarding the optimal rule is even smaller. So, the margin condition ensures that the population welfare functional $V(g)$ is not too sensitive to small deviations of $g$ from the optimal policy $g_\ast$.\footnote{Assumption (ref) corresponds to the margin condition in kitagawa2018should, where $\mathbb{P}(|U_i|<t)=O(t^\alpha)$ with $\alpha=1$. Our results remain valid under any distribution of $U_i$ satisfying $\mathbb{P}(|U_i|<t)=O(t^\alpha)$ for $\alpha\ge 1$, but generally fail when $\alpha<1$.}

For convenience in stating the next assumption on nuisance parameter estimation, define the \(L_2(\mathbb{P})\) norm by $\| f \|_{\mathbb{P},2} := \big(\mathbb{E}[\| f(X) \|^2]\big)^{1/2}$ where the norm \(\|\cdot\|\) inside the expectation denotes the Euclidean norm.

as\begin{enumerate}[(i)] • (Nuisance Parameter Estimation) $\Vert \widehat{\eta}_{(\ell)}-\eta_0\Vert_{\mathbb{P},2}=o_p(n^{-1/4})$ for each $\ell\in[m]$. Furthermore, with probability at least $1-\epsilon_n$ for $\epsilon_n=o(1)$, $\widehat{\eta}_{(\ell)}$ belongs to a shrinking neighborhood $\mathcal{T}_n$ of $\eta_0$ for each $\ell \in[m]$, such that the following convergence results hold: \begin{align*} &\sqrt{n} \sup _{\eta \in \mathcal{T}_n}\left\vert \mathbb{E} \left[\left(\psi_1(Z_i,\eta)-\psi_1(Z_i,\eta_0)\right)\mathbf{1}(-g_k(X_i)<0)\right]\right\vert=o(1),\\ &\sqrt{n} \sup _{\eta \in \mathcal{T}_n}\left\vert \mathbb{E} \left[\left(\psi_0(Z_i,\eta)-\psi_0(Z_i,\eta_0)\right)\mathbf{1}(g_k(X_i)<0)\right]\right\vert=o(1),\\ &\sqrt{n}\sup _{\eta \in \mathcal{T}_n}\left(\mathbb{E}\left\vert \left[\left(\psi_1(Z_i,\eta)-\psi_1(Z_i,\eta_0)\right)\mathbf{1}(-g_k(X_i)<0)\right]\right\vert^2\right)^{1 / 2}=o(1), and \\ &\sqrt{n}\sup _{\eta \in \mathcal{T}_n}\left(\mathbb{E}\left\vert \left[\left(\psi_0(Z_i,\eta)-\psi_0(Z_i,\eta_0)\right)\mathbf{1}(g_k(X_i)<0)\right]\right\vert^2\right)^{1 / 2}=o(1). \end{align*} • (Boundedness) For $j=0,1$, and uniformly over the shrinking neighborhood $\mathcal{T}_n$ of $\eta_0$$, \mathbb{E}[\vert\psi_j(Z_i, \eta)\vert^q]$ is bounded for some $q > 4$, and $\mathbb E[\psi_j(Z_i, \eta)|X_i=x]$ is bounded over the support of $X$. \end{enumerate}

Assumption (ref).(ref) requires that the cross-fitted estimator \(\widehat{\eta}\) of the nuisance parameter induces asymptotically negligible estimation errors. This condition has been also employed by athey2021policy in the context of Example III, and can be achieved by employing a Neyman-orthogonal score, which ensures that first-step estimation errors are asymptotically negligible provided that \(\widehat{\eta}_{(\ell)}\) converges at the \(n^{-1/4}\) rate, rather than the parametric \(n^{-1/2}\) rate. Such convergence rates are attainable with a variety of machine learning methods. For example, the Lasso achieves the \(n^{-1/4}\) rate under approximate sparsity and restricted eigenvalue conditions bickelritovtsybakov2009,belloni2015some. Random forests and boosted trees are valid provided the regression function is sufficiently smooth and sample splitting is employed buehlmannYu2002,wagerAthey2018. Deep neural networks with controlled depth and bounded weight norms can approximate smooth functions at nearly minimax rates chenwhite1999,yarotsky2017,schmidtHieber2020, among others. Assumption (ref).(ref) further imposes uniform bounds on the (conditional) expectations of weight functions, $\psi_1$ and $\psi_0$. Note that this condition allows for unbounded outcome variable $Y$ in the welfare maximization (Example III) and is also satisfied in the expected utility maximization with binary action (Example II) under the Condition 1 of elliott2013predicting. It holds trivially in the maximum score estimation (Example I).

Under these rate conditions and the boundedness conditions, we obtain the following asymptotic normality result for our estimator of the optimal policy value.

thmIn addition to the conditions for Theorem (ref), suppose that Assumptions (ref) and (ref) hold, and the sieve order $k$ satisfies $k^{4+\varepsilon}/n\lesssim1$ for some $\varepsilon>0$. Then, for any sequence of data generating processes $\mathbb{P}_n$, we have \begin{align*} \sqrt{n}\{\widehat{V}_n(\widehat{g})-V_0\}\rightarrow_d\mathcal{N}(0,\sigma^2_v), \end{align*} where $\sigma^2_v=\mathbb{E}[s_L(Z_i, \eta_0,g_\ast)^2]$ is the variance of the score $$s_L(Z_{i},\eta_0,g_\ast)=\psi_{1}(Z_i,\eta_0)\cdot\mathbf{1}(g_\ast(X_i)\geq0)+\psi_{0}(Z_i,\eta_0)\cdot\mathbf{1}(g_\ast(X_i)<0)-V_0.$$

The asymptotic Gaussianity established in Theorem (ref) enables inference on the optimal policy value \(V_0\) in the standard manner, i.e., using normal critical values such as 1.96. To approximate the asymptotic variance \(\sigma_v^2\) in practice, we introduce a bootstrap procedure in Section (ref).

While the proof in the appendix formally establishes the results of Theorem (ref), we provide the intuition here, particularly illustrating why the estimation effect of our classification policy function estimator \(\widehat{g}\) as given in Section (ref) is asymptotically negligible despite the fact that it appears in the indicator function and converges at a slower rate than the root-$n$ rate. Consider the decomposition:

align[align omitted — 388 chars of source]

where the term in (ref) captures the estimation error of \(\widehat{g}\), the term in (ref) collects the sampling error from \(\mathbb{E}_n\) and the estimation error of \(\widehat{\eta}\), the term (ref) indicates the sieve approximation error of \(g_k\) for \(g_\ast\), and the term in (ref) measures the identification error. Theorem (ref) implies that the identification error in (ref) is zero. The component in (ref) corresponds to the standard semiparametric part and is asymptotically Gaussian under standard assumptions, in particular under ours. The remaining task is to show that the first component in (ref) is asymptotically negligible. Since this argument is not entirely obvious, we provide the intuition here in the main text, while leaving the formal proof to the appendix.

The component in (ref) in question can be further decomposed as

align[align omitted — 288 chars of source]

The two bracket groups on the right-hand side of (ref) can be shown to be asymptotically negligible by applying the maximal inequality for VC classes massart2006risk, Sauer's Lemma lugosi2002pattern, and the dominated sieve approximation errors. The term in (ref) is more delicate, since \(\widehat{g}\) converges only at the nonparametric rate as obtained in Theorem (ref). Nevertheless, this term is also negligible because its first-order effect vanishes due to the first-order condition for the optimality of \(g_\ast\) with respect to \(V\). As a result, the second-order term becomes the leading component, and its existence is guaranteed by the margin condition (Assumption (ref)). We further illustrate this point through numerical simulations in Section (ref). In particular, we show that the plug-in estimator $\widehat{V}_n(\widehat{g})$ and the oracle estimator $\widehat{V}_n(g_\ast)$ yield nearly identical confidence intervals.

In summary, the component in (ref) drives the asymptotic normality in the same way as in standard semiparametric analysis, while all other components are asymptotically negligible. See Appendix (ref)--(ref) for further details. A couple of remarks are in order.

remAlthough the plug-in value estimator in (ref) is constructed using the estimator $\widehat{g}$ described in Section (ref), our inferential results do not rely on the specific form of this estimator. What is required for valid second-stage inference is that the first-stage policy function estimator converges sufficiently fast so that its estimation error becomes asymptotically negligible. Indeed, with slight modifications on the proof, we can show that any estimator $\widetilde{g}$ would be admissible if it converges uniformly over $\mathcal X$ to a limit $g_{\ast\ast}$ satisfying $V(g_{\ast\ast}) = V_0$ at a rate faster than $o_p(n^{-1/4})$. That is, our framework accommodates a broad class of first-stage methods, including $\ell_1$-penalized estimators athey2021policy, parametric approaches based on the 0--1 loss kitagawa2018should, and nonparametric CATE-based methods bhattacharya2012inferring,armstrong2023inference,park2025debiased, provided they achieve the uniform convergence rate $o_p(n^{-1/4})$ to $\mathbb{E}[\psi_1(Z,\eta_0)\mid X=x] - \mathbb{E}[\psi_0(Z,\eta_0)\mid X=x]$. Thus, while our construction employs the estimator $\widehat g$ for the representative classification policy $g_\ast$, the resulting limit theory supports valid inference for a wide range of first-stage estimators used in the policy learning literature.
remOur optimal policy value estimation is related to semiparametric estimation with a non-smooth criterion chen2003estimation. Our framework extends this by allowing not only non-smooth but also discontinuous criteria. Even when the loss, and hence the sample criterion, is discontinuous, the asymptotically Gaussian term dominates provided that the population criterion is smooth, as illustrated in the intuition above. In particular, the first-order condition for the optimality of \(g_\ast\) plays a crucial role in ensuring that the potentially problematic term in (ref) is in fact negligible. $\blacktriangle$
remhirano2012impossibility highlight that non-differentiability in a population criterion function prevents regular estimation. In particular, such an impossibility phenomenon arises in estimating \(\phi(\theta) = \max\{\theta_1, \theta_2\}\), which is non-differentiable at the kink point \(\theta_1 = \theta_2\). This kink causes the plug-in estimator \(\max\{\widehat{\theta}_1, \widehat{\theta}_2\}\) to have a non-Gaussian distribution when \(\theta_1 = \theta_2\), even if \((\widehat{\theta}_1,\widehat{\theta}_2)\) were Gaussian and centered at \((\theta_1,\theta_2)\). In contrast, our welfare criterion can be rewritten as \begin{align*} \sup_{g \in \mathcal{G}} V(g) &= \sup_{g \in \mathcal{G}} \mathbb{E}\!\left[ \mathbb{E}[\psi_1(Z_i,\eta_0)\mid X_i] \cdot \mathbf{1}\{g(X_i)\geq0\} + \mathbb{E}[\psi_0(Z_i,\eta_0)\mid X_i] \cdot \mathbf{1}\{g(X_i)<0\} \right] \\ &= \mathbb{E}\!\left[ \max \left\{ \vartheta_1, \vartheta_2 \right\} \right], \end{align*} rather than \(\max\{\theta_1,\theta_2\}\), where \(\vartheta_1 := \mathbb{E}[\psi_1(Z_i,\eta_0)\mid X_i]\) and \(\vartheta_2 := \mathbb{E}[\psi_0(Z_i,\eta_0)\mid X_i]\) are random variables. Hence, the two problems differ mathematically. As emphasized in Remark (ref), given the smoothness of the population objective, asymptotic normality is attainable even with a non-smooth sample criterion, with the first-order condition for the optimality of \(g_\ast\) playing the crucial role. $\blacktriangle$

Bootstrap Inference for the Optimal Policy Value

Theorem (ref) facilitates standard inference based on Gaussian critical values, e.g., \(\approx 1.96\). We test the null hypothesis \(\mathcal{H}_0: V_0 = v_0\) using the statistic

align*[align* omitted — 84 chars of source]

To approximate the variance \(\sigma_v^2\) of its Gaussian limit, we employ the bootstrap counterpart

align*[align* omitted — 136 chars of source]

where the estimated score takes the form of

align*[align* omitted — 257 chars of source]

\(\widehat{\eta}_{(\ell(i))}\) denotes the nuisance parameter estimator based on the complementary subsample \(I^c_{(\ell(i))}\) of the fold \(I_{(\ell(i))}\) that contains observation \(i\), and the perturbation process \(\{\delta_i\}_{i \in [n]}\) is a sequence of i.i.d. standard Gaussian bootstrap weights independent of the data.

The following theorem establishes the consistency of the score bootstrap method for approximating the limiting Gaussian distribution of the estimated optimal policy value.

thmIn addition to the conditions of Theorem (ref), suppose that the perturbation process $\{\delta_i\}_{i\in[n]}$ is a sequence of i.i.d. standard Gaussian bootstrap weights independent of the data. Given the data generating process $\mathbb{P}_n$, the bootstrapped score converges in distribution to a normal distribution: \begin{align*} \widetilde{Z}_n\rightarrow_d\mathcal{N}(0,\sigma^2_v), \end{align*} where $\sigma^2_v$ was defined in Theorem (ref).

For a fixed \(\alpha \in (0,1)\), let \(\widehat{c}_{1-\alpha}\) denote the \((1-\alpha)\)-quantile of \(\widetilde Z_n\). A consequence of Theorem (ref) is the asymptotic coverage property

align*[align* omitted — 150 chars of source]

under the null hypothesis \(\mathcal{H}_0: V_0 \leq v_0\) against alternative hypothesis \(\mathcal{H}_1: V_0>v_0\).

Simulation Studies

In this section, we use numerical simulations to evaluate the finite-sample performance of our proposed inference methods for cost-sensitive binary classification. As an illustrative example, we focus on the welfare maximization problem (Example III) introduced in Section (ref). Within this context, we examine both the performance of our proposed uniform inference procedure for the optimal treatment assignment (classification) policy and the inference method for the optimal welfare (policy value).

The outcome is modeled as $ Y_i = A_i \cdot \Delta(X_i) + S(X_i) + u_i, $ where the idiosyncratic error term is generated as \(u_i \stackrel{\text{i.i.d.}}{\sim} \mathcal{N}(0,1)\). The binary treatment assignment is given by $ A_i = 2 \cdot \text{Bernoulli}(\pi(X_i)) - 1, $ so that \(A_i \in \{-1,1\}\). The treatment effect is specified as $ \Delta(X_i) = \tanh\!\big((1, X_i^\prime)\gamma\big), $ the baseline function as $ S(X_i) = \sin(X_i^\prime \beta_S), $ and the propensity score as $ \pi(X_i) = \frac{\exp(X_i^\prime \beta_\pi)}{1 + \exp(X_i^\prime \beta_\pi)}. $

We consider a nonlinear design featuring two covariates. The first covariate, \(X_{i1} \sim \text{Uniform}[0,1]\), represents normalized pre-treatment income. The second covariate, \(X_{i2}\), is constructed by rescaling a discrete variable that mimics years of education. Specifically, for \(X_{i2}\), we draw a categorical variable taking values in \(\{7, \ldots, 18\}\) with probabilities calibrated to the empirical distribution of education levels in the JTPA data, and divide by 18 to ensure support on the unit interval.

The parameter values of \(\gamma\) are set to be consistent with the null hypotheses under consideration. For the one-sided null hypothesis \(\mathcal{H}_0: g_\ast(x) \leq 0 \;\; \forall x \in \mathcal{X}\), we consider the setting \(\gamma = (0, -1/\sqrt{n}, -1/\sqrt{n})^\prime\), with \(\beta_S = (-1, 1)^\prime\) and \(\beta_\pi = (1, -1)^\prime\). For the null \(\mathcal{H}_0: g_\ast(x) \geq 0 \;\; \forall x \in \mathcal{X}\), on the other hand, we instead use the setting \(\gamma = (0, 1/\sqrt{n}, 1/\sqrt{n})^\prime\), with \(\beta_S = (-1, 1)^\prime\) and \(\beta_\pi = (1, -1)^\prime\). Extensive simulations are conducted with sample sizes \(n \in \{250, 500, 1000\}\), and each configuration is replicated \(S = 1,000\) times through Monte Carlo iterations. The number of the cross-fitting folds is set to $m=2$ both for estimating the policy and welfare.

For nuisance parameters, the simulation design incorporates the conditional mean outcome functions \(\mu(a; x)\) for \(a \in \{-1, 1\}\), estimated via Lasso, and the propensity score function \(\pi(x)\), estimated via \(\ell_1\)-penalized logistic regression. The tuning parameters for both procedures are chosen by five-fold cross-validation.

The representative classification policy function is estimated using a sieve approximation based on a tensor-product Legendre polynomial basis. Specifically, for \(k \in \{2,3,4\}\), we employ the tensor-product basis constructed from the orthonormal Legendre polynomials \(\{p_j(x)\}_{j=0}^\infty\) of $k$-dimensions defined on the support \(\mathcal{X}\). The sieve basis is given by $ p_{1:k}(x_1, x_2) = \big\{\, p_j(x_1)\, p_l(x_2) : j, l \in [k] \,\big\}, $ and the policy function is approximated as $ g_k(x_1, x_2) = \sum_{j=1}^{k} \sum_{l=1}^{k} \beta_{jl}\, p_j(x_1) p_l(x_2), $ or, equivalently, $ g_k(x_1, x_2) = p_{1:k}(x_1, x_2)^\prime \beta. $

To remain consistent with the theoretical conditions required for the asymptotic analysis, the simulation study focuses exclusively on twice continuously differentiable surrogate loss functions. We consider the following three examples:

enumerate[(a)] • Logistic loss: $\phi(x) = \log(1 + e^{-x})$; • Exponential loss: $\phi(x) = e^{-x}$; and • Squared loss: $\phi(x) = (1 - x)^2$.

Uniform Inference for Optimal Policy

This subsection presents the finite-sample performance of our proposed uniform inference method for the optimal policy rule. We consider the two null hypotheses: \(\mathcal{H}_0: g_{\ast}(x) \leq 0 \;\; \forall x \in \mathcal{C}\) and \(\mathcal{H}_0: g_{\ast}(x) > 0 \;\; \forall x \in \mathcal{C}\), where the set \(\mathcal{C}\) is defined as \(\mathcal{C} := [0.05, 0.95] \times \{10\}\).

table[table omitted — 1,619 chars of source]

Table (ref) reports the finite-sample uniform coverage frequencies based on the one-sided tests across sample sizes \(n \in \{100, 250, 500\}\) and the three surrogate loss functions. The top panel (I) corresponds to the null hypothesis \(\mathcal{H}_0: g_{\ast}(x) \leq 0 \;\; \forall x \in \mathcal{C}\), while the bottom panel (II) corresponds to the null hypothesis \(\mathcal{H}_0: g_{\ast}(x) > 0 \;\; \forall x \in \mathcal{C}\).

The results confirm that our bootstrap uniform confidence bands successfully control the test size. The size performance is slightly conservative, which is attributable to the fact that the true data-generating functions lie in the interior of the null, far from the least favorable boundary points of the null.

Gaussian Limit and Inference for the Optimal Welfare

This subsection demonstrates the Gaussian limit of our estimated welfare estimator via numerical studies. In addition to (a) the plug-in estimated optimal policy rule, we also study the Gaussian limit theory of welfare under three comparative benchmarks: (b) an oracle treatment assignment rule, (c) a treat-everyone policy, and (d) a random-design policy. The random-design policy (d) assigns treatment according to an independent Bernoulli\((0.5)\) distribution.

figure[figure omitted — 570 chars of source]

For each of the four scenarios (a)--(d), we normalize the estimated optimal welfare using the Monte Carlo variance estimate computed across 1,000 simulation replications. We then compare the simulated densities of these normalized estimates with that of the standard normal distribution for each scenario. The sample size is fixed at \(n=500\) throughout. Figure (ref) displays the simulation density of the estimated optimal welfare in solid black, alongside the standard normal density in dashed red.

Each plotted density from (a) to (d) closely resembles the standard normal distribution. In particular, the result in panel (a) for the plug-in estimated optimal welfare corroborates the theoretical property of Gaussian limit behavior established in Theorem (ref). Panel (a) is comparable to panel (b) for the oracle counterpart, and this observation is consistent with the asymptotic negligibility of the preliminary representative classification policy estimation emphasized in Section (ref). Although we report one set of results based on the logistic loss, these findings hold robustly across different surrogate loss functions.

Figure (ref) reports the median, interquartile, and interdecile ranges of empirical welfare estimates under four treatment assignment rules: (a) the plug-in optimal policy estimated using the logistic loss (Est. Optimal); (b) the oracle assignment rule (Oracle Optimal); (c) the treat-everyone rule (Everyone); and (d) the random assignment rule with estimated variance (Random Assign). The results are based on $n = 500$ observations and $S = 1{,}000$ Monte Carlo replications. The distribution under the estimated optimal policy closely resembles that under the oracle rule, corroborating the asymptotic negligibility of $\widehat g$ for $\sqrt{n}\bigl(\widehat V_n(\widehat g) - V_0\bigr)$ as argued in Section (ref). Moreover, the welfare estimates under both the estimated optimal policy and the oracle policy substantially exceed those attained under the treat-everyone and random assignment rules.

figure[figure omitted — 285 chars of source]

Finally, we conduct hypothesis tests to compare the welfare achieved under the estimated optimal policy to that under the benchmark policies $g_\dagger$.\footnote{See Appendix (ref) for details on this testing procedure.} Specifically, we perform a two-sided test of $\mathcal{H}_0: V_0 = V(g_\dagger)$ against $\mathcal{H}_1: V_0 \neq V(g_\dagger)$, as well as one-sided tests of $\mathcal{H}_0: V_0 \leq V(g_\dagger)$ versus $\mathcal{H}_1: V_0 > V(g_\dagger)$ and $\mathcal{H}_0: V_0 \geq V(g_\dagger)$ versus $\mathcal{H}_1: V_0 < V(g_\dagger)$. Table (ref) reports the Monte Carlo rejection frequencies. Both the two-sided and right-sided tests are rejected in most Monte Carlo draws, indicating that the estimated optimal policy delivers a significantly higher level of welfare.

table[table omitted — 509 chars of source]

Empirical Analysis

With the “identification”, estimation, and inference methods developed in this paper, we now revisit the problem of policy learning using the experimental data from the National JTPA Study. While prior studies such as kitagawa2018should and mbakop2021model have analyzed this problem, they did not conduct formal inference for the optimal treatment assignment (classification) policy or its associated welfare (policy value). Our goal is to build on this literature and provide inference for these important features.

The JTPA study randomly assigned applicants to be eligible for a job training program for a period of 18 months. It collected background information on the applicants prior to random assignment, as well as administrative and survey data on their earnings over the 30 months following the assignment. The dataset consists of 11,008 observations from the adult sample (ages 22 and older).

Following benchmark studies in the literature, we adopt the following variable definitions. The outcome variable, \(Y_i\), is defined in two ways: (i) total individual earnings during the 30 months following program assignment, or (ii) total individual earnings during the 30 months following program assignment minus the average cost of services per treatment assignment (\$774). The covariates, \(X_i\), used to define the policy class are years of education and pre-program earnings.\footnote{For both the outcome variable (total individual earnings over the 30-month period) and the covariate (pre-program earnings), we divide their values by 100.}

We consider two types of policy functions: (a) a linear rule \(g(x) = \gamma^{\prime}x\); and (b) a class of nonlinear functions defined as follows: Let \(\mathcal{X}_1\) denote the covariate space of years of education, and let \(\mathcal{X}_2\) denote the covariate space of pre-program earnings. The set of nonlinear assignment rules is given by \[ \mathcal{G} = \bigl\{\, g(x_1, x_2) = f(x_1) - x_2 \;\big|\; f: \mathcal{X}_1 \to \mathcal{X}_2 \,\bigr\}. \]

Linear and nonlinear rules have been studied in the literature kitagawa2018should, mbakop2021model. In particular, while the nonlinear class follows mbakop2021model, we allow for a more general specification without imposing their sign restriction ($f(x_1) \geq x_2$). We can accommodate this generality by leveraging our method’s ability to achieve nonparametric point “identification,” along with the associated estimation and inference procedures. That said, the main contribution of this paper lies in its inference component, which has not been addressed by these preceding papers.

The optimal policy function is approximated using a sieve expansion based on tensor-product Legendre polynomial basis functions with sieve order \(k = 3\). The number of folds is set to \(m = 2\). The nuisance functions are estimated via machine learning methods: the propensity score is obtained using \(\ell_1\)-penalized logistic regression, and the conditional mean outcome is estimated using Lasso. In both cases, tuning parameters are selected via cross-validation. For inference, we implement a score-based bootstrap procedure with 1,000 repetitions to construct confidence bands for the optimal policy and a confidence interval for its corresponding welfare.

To ensure that the approximation error of Legendre polynomials diminishes in the \(L_\infty\) space, we normalize both covariates to lie within the unit interval \([0,1]\). Specifically, we define \(x_1\) as the years of education divided by its maximum value, and \(x_2\) as normalized pre-treatment earnings, computed as \[ x_2 = \frac{\text{pre-treatment earnings} - \min(\text{pre-treatment earnings})} {\max(\text{pre-treatment earnings}) - \min(\text{pre-treatment earnings})}. \] Similarly, we normalize the outcome variable \(Y_i\) using the same min--max transformation: \[ Y_i = \frac{\text{earnings}_i - \min(\text{earnings})} {\max(\text{earnings}) - \min(\text{earnings})}. \]

The surrogate loss function used in our analysis is the logistic loss. Our objective is to construct a left-sided uniform confidence band for (ref) to test the null hypothesis \[ \mathcal{H}_0:\; g_\ast(x_1,x_2) \leq 0, \] over \(x_1 \in [7/18, 1]\), with \(x_2\) fixed at the 25th, 50th, and 75th quantiles of its empirical distribution.

figure[figure omitted — 864 chars of source]

First, we conduct inference results for the optimal policy rule. We compute the one-sided uniform confidence bands of (ref) using the critical values obtained from the score bootstrap method described in Section (ref). The sieve-based nonparametric estimators and their corresponding uniform confidence bands are computed for the four cases described above: (i) no training cost with a linear policy rule, (ii) no training cost with a nonlinear policy rule, (iii) training cost with a linear policy rule, and (iv) training cost with a nonlinear policy rule. For each case, we construct one-sided uniform confidence bands over years of education, holding pre-program earnings fixed at the 25th, 50th, and 75th quantiles.

Figure (ref) presents the estimation and inference results for the optimal policy. When training costs are not considered and years of education are below 12, Figures (ref) and (ref) show that the 95% uniform confidence bands allow us to reject the uniform null hypothesis, indicating that assigning individuals in this range to the control group is statistically significant at the 5% level. In contrast, when training costs are taken into account, this statistical significance disappears, as shown in Figures (ref) and (ref). The economic intuition behind this finding is straightforward: as the cost of the National JTPA program increases, the net benefit of participation declines, reducing its appeal and thereby lowering the expected welfare for the representative individual. Overall, our inference method for the optimal policy rule complements the existing findings kitagawa2018should,mbakop2021model by providing formal statistical inference as well as more general nonlinearity.

table[table omitted — 1,085 chars of source]

We now report the estimation and inference results for (optimal) welfare. Table (ref) presents the 95% confidence intervals (CIs) for the estimated welfare under three different treatment assignment rules: (i) the estimated optimal policy rule based on the logistic surrogate loss, (ii) the policy that assigns all individuals to the treatment group, and (iii) the random assignment rule following a Bernoulli(0.5) distribution. Since Theorem (ref) establishes that the empirical welfare of a suitable plug-in treatment assignment rule is asymptotically normal, constructing CIs for welfare functions only requires appropriate variance estimates. Specifically, for the estimated optimal policy rule, we obtain the variance estimate using our score bootstrap procedure (see Section (ref)), while for the other two cases, we use the variance formula provided in Theorem (ref).

When linear rules are considered, the estimated welfare of the estimated optimal rule is \$1,285 in the case of no training cost and \$533 in the case of training costs. These findings are very similar to those reported by kitagawa2018should, who obtain estimated welfare values of \$1,180 and \$404, respectively. While they do not develop a formal theory for inference, kitagawa2018should also report confidence intervals based on bootstrap for this application. Our confidence intervals for the two aforementioned cases, (\$901, \$1{,}668) and (\$120, \$946), are comparable to theirs but notably narrower, thereby offering greater precision and enhancing the reliability of statistical decisions.

Additionally, Table (ref) compares the estimated optimal rule, the “treat-everyone” rule, and the random assignment rule across the cases. The estimated welfare of the random assignment rule lies outside the 95% confidence intervals for the estimated optimal rule, clearly indicating that our estimated rule achieves higher welfare than the random assignment rule, which has been widely used in empirical economic research. Moreover, the estimated welfare of the optimal rule is consistently higher than that of the “treat-everyone” rule, further demonstrating the practical usefulness of our policy learning method.

table[table omitted — 854 chars of source]

Finally, we conduct hypothesis tests to compare the welfare achieved under the estimated optimal policy with that under the benchmark policies $g_\dagger \in \{\text{``Treat Everyone''}, \text{``Random Assign''}\}$.\footnote{See Appendix (ref) for details on this testing procedure.} Specifically, we perform a two-sided test of $\mathcal{H}_0: V_0 = V(g_\dagger)$ against $\mathcal{H}_1: V_0 \neq V(g_\dagger)$, as well as one-sided tests of $\mathcal{H}_0: V_0 \leq V(g_\dagger)$ versus $\mathcal{H}_1: V_0 > V(g_\dagger)$ and $\mathcal{H}_0: V_0 \geq V(g_\dagger)$ versus $\mathcal{H}_1: V_0 < V(g_\dagger)$. We consider two settings, depending on whether treatment costs are incorporated. Figure (ref) summarizes the resulting $p$-values.

Several findings emerge. First, when comparing against the random assignment rule, both the two-sided and right-sided tests are rejected at the $5\%$ significance level, providing strong empirical evidence that the optimal policy rule achieves higher welfare than the random assignment rule. Second, when compared with the “treat-everyone” rule, the optimal policy rule does not provide sufficiently strong evidence to demonstrate higher welfare. This result holds robustly across the cases with and without treatment costs. This finding is unsurprising, as assigning all individuals to treatment can raise the overall level of welfare, especially among less educated individuals. However, when highly educated individuals represent only a small share of the population, the difference between the optimal treatment rule and the “treat-everyone” rule can be small, making it difficult to reject the null hypothesis that treats the two rules as equivalent.

Conclusion

In this paper, we have developed a toolkit for nonparametric uniform inference in cost-sensitive binary classification, a broad framework that encompasses maximum score estimation, utility-maximizing choice prediction, policy learning, and related problems. These settings are well known for slow convergence and nonstandard limiting behavior under parametric classification rules. Furthermore, nonparametric formulations of classification policy functions introduce the additional challenge of failures of identification.

To address these difficulties, we proposed a strictly convex surrogate loss that point-identifies a representative nonparametric classification policy function. Estimating this representative classification policy function provides a unified path to inference for both the optimal classification rule and its associated policy value. In particular, this strategy delivers a Gaussian approximation for the optimal classification policy and a Gaussian limit for the optimal policy value, enabling standard inferential procedures and substantially simplifying empirical implementation despite the intrinsic difficulty of the underlying problem.

Simulation evidence corroborates the theoretical results, showing that the proposed methods perform well in finite samples. Finally, our application to the National JTPA Study demonstrates how the procedure can be used to recover an optimal treatment assignment rule and to evaluate its welfare implications relative to alternative policies at hand.

While the primary contributions of this paper lie in developing inference methods for nonparametric classification policies, our approach also provides, as a by-product, practical benefits for estimating these policies. This improvement arises from two sources. First, the use of a strictly convex surrogate loss facilitates routine numerical optimization for sieve estimation. Second, and more importantly, our point identification of the representative policy function enables a straightforward estimation of the resulting classification rules.