EconBase
← Back to paper

Identification and Inference for Algorithmic Frontiers with Selective Labels

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

200,802 characters · 14 sections · 87 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Identification and Inference for Algorithmic Frontiers with Selective Labels

frontmatter\begin{aug} \address[id=add1]{ \orgdiv{Department of Economics}, \orgname{Cornell University}} \address[id=add2]{ \orgdiv{Department of Economics}, \orgname{Cornell University}} \address[id=add3]{ \orgdiv{Department of Economics}, \orgname{Cornell University}} \end{aug} \support{Preliminary version. This draft: June 12, 2026. We thank Jos{\'e} Montiel-Olea for helpful comments.} \begin{abstract} This paper provides identification results to characterize a fairness-accuracy (FA) frontier, and statistical inference tools to test hypotheses and build a confidence set for the FA-frontier, when outcomes are observed only for selected individuals. When the selection process is unrestricted but loss is measured in specific ways, we provide a characterization of the sharp identification region of the FA-frontier. Under an assumption of unconfoundedness conditional on observables (and unrestricted loss functions), we obtain point identification and propose a debiased machine learning estimator, derive its asymptotic distribution, and show how this can be used to carry out inference for the FA-frontier. In work in progress, we extend the partial identification results to a broader class of loss functions. \end{abstract} \begin{keyword} \kwd{Algorithmic fairness} \kwd{selective labels} \kwd{statistical inference} \kwd{support function} \end{keyword}

Introduction

Over the last decade, algorithms have come to play an increasingly prevalent role in everyday life, often by supporting high-stake decisions through predictions of quantities such as re-offense risk, repayment likelihood, and health care needs. These predictions enter institutional decision processes that determine who is detained or released, who is granted credit, and which patients receive outreach and care-management interventions. The increasing reliance on algorithms has prompted the emergence of a set of regulatory and oversight processes, known as AI governance, that rule the design, deployment, documentation, and monitoring of AI systems, including, in some settings, restrictions on the use of protected attributes as model inputs. A distinctive feature of these regimes is their evidentiary orientation: for high-risk AI systems, the EU AI Act requires technical documentation, logging capabilities, and appropriate accuracy and robustness over the lifecycle eu_ai_act, while related governance proposals emphasize dataset documentation, model reporting, and internal auditing gebru_etal_2021,mitchell_etal_2019,raji_etal_2020.

This paper extends the research agenda set out in liu:mol25v2 to provide a toolkit for nonparametric estimation of and statistical inference on the fairness-accuracy trade-off that characterizes algorithms. Fairness and accuracy evaluations play a central role in governance, as they determine both the validity of model-supported decisions and whether those decisions impose disproportionate harms on legally protected groups.\footnote{Algorithms may exhibit systematic differences in performance across legally protected groups, both in predictive accuracy and in the decisions they induce ang:lar:mat:kir16, arn:dob:hul21, obe:pow:vog:mul19, cow:tuc20,ber:jei:jab:kea:rot21. Yet, improving parity in expected losses often requires sacrificing accuracy for at least one group, and conversely increasing accuracy for one group may worsen disparities.} At the same time, governance questions are rarely about a single fixed model. In public oversight and non-discrimination doctrine, the relevant question is often whether a challenged procedure can be replaced by an alternative that achieves comparable accuracy with less adverse impact, shifting attention from a single algorithm to a potentially large class of admissible decision rules, e.g., those obtained under constraints on inputs, model classes, or the use of sensitive attributes. Regulators and oversight bodies are called to adjudicate (i) whether a deployed rule is Pareto efficient, (ii) whether design restrictions enable less discriminatory alternatives (LDAs) with comparable accuracy, and (iii) whether such LDAs exist within an admissible class of decision rules.\footnote{For example, US employment discrimination law prohibits selection procedures that cause disparate impact unless the employer demonstrates they are “job related for the position in question and consistent with business necessity.” Even then the practice is unlawful if an alternative practice would serve the employer's legitimate interests with less discriminatory impact civil_rights_act_1991.} As in liu:mol25v2, we build on the theoretical framework put forward by \citet*[LLMO henceforth]{lia:lu:mu:oku24}, which in our view is well-suited to formalize these questions because it characterizes a fairness-accuracy (FA) frontier under broad preferences over group-specific expected losses as well as the implications of banning attributes as inputs to the algorithms.

Our first innovation relative to liu:mol25v2's groundwork is to allow for accuracy and fairness of an algorithm to be assessed through two different loss functions, thereby aligning our treatment with the literature in computer science corbettdavies2017cost,menon2018cost,kim2020fact,gillis2024lda. We show that the support function of a convex feasible set $\mathcal{E}\subset{\mathbb R}^3$ of expected accuracy losses for the two groups and the difference in expected fairness losses between the groups plays a crucial role in characterizing the FA-frontier and carrying out inference.\footnote{In liu:mol25v2's analysis and \citetalias{lia:lu:mu:oku24}'s main text, the same loss function is used to measure fairness and accuracy, so that the feasible set $\mathcal{E}$ is in ${\mathbb R}^2$. liu:mol25v2 propose to use the support function of $\mathcal{E}\subset{\mathbb R}^2$ as the unifying tool for statistical inference.} We extend the results in \citetalias{lia:lu:mu:oku24}'s Appendix O by providing a simple and testable necessary and sufficient condition under which the FA-frontier coincides with the Pareto frontier. When this condition is not satisfied, a policymaker who values fairness may choose to deploy an algorithm that is Pareto dominated in accuracy loss, to ameliorate its performance in fairness loss.

Our second innovation addresses the reality that in many high-stakes settings, outcomes are observed only for selected individuals: pretrial misconduct (or lack thereof) is observed only if bail is granted; repayment behavior is observed only when credit is extended; certain health outcomes are observed only for patients who receive treatment or remain under follow-up within a system; etc. Hence, instead of observing the outcome vector of interest $Y^*$, one observes $\ZY^*$, where $Z\in\{0,1\}$ indicates observability. This selective labels problem creates challenges for governance compliance similar to the ones long recognized in the vast econometrics literature on selectively observed data heckman_1979,manski_1989. Accuracy metrics computed using only observed outcomes may misrepresent true performance, potentially overstating accuracy where it matters most (e.g., if those denied bail would not have re-offended). Hence, assessments of whether a system achieves “appropriate levels of accuracy” eu_ai_act or whether LDAs exist civil_rights_act_1991 may be incorrect and required documentation may systematically exclude the most policy-relevant counterfactual outcomes.

When classification loss is used to measure accuracy and statistical parity is used to measure fairness menon2018cost, and no restrictions are imposed on the selection process, we characterize the sharp identification region for the FA-frontier in the spirit of manski_1989 mol20,can:ill:velez26. Doing so is challenging because the FA-frontier is ultimately about algorithms: it equals the set of expected losses $\varepsilon(a)\in\mathcal{E}$ associated with algorithms $a\in\mathcal{A}(\mathcal{X})$ (measurable functions mapping covariates $X\in\mathcal{X}$ to $[0,1]$) such that no other algorithm $\tilde{a}\in\mathcal{A}(\mathcal{X})$ yields expected losses $\varepsilon(\tilde{a})$ that FA-dominate $\varepsilon(a)$ in a sense made precise in Definition (ref). When labels are selectively observed, both $\varepsilon(a)$ and $\mathcal{E}$ are partially but not point identified. Hence, to pass judgment on whether algorithm $a$ yields expected losses $\varepsilon(a)$ on the frontier, care needs to be taken to couple the same candidate distribution for the missing labels in constructing both $\varepsilon(a)$ and $\mathcal{E}$, and then assess membership across all admissible distributions for the missing labels. We reduce this daunting task to solving a finite dimensional optimization problem.

When the selection process is restricted by a missing-at-random (MAR) assumption, where conditional on attributes, outcome observability does not convey additional information on $Y^*$, we obtain point identification of the FA-frontier. For any loss function, we provide a debiased machine learning (DML) estimator for the FA-frontier based on inverse probability weighting. Our DML approach is particularly appropriate for governance contexts, where high-dimensional administrative data enable flexible nonparametric modeling of selection and formal inference tools are needed for legally defensible determinations. While ideas transfer from the program evaluation and high-dimensional inference literature CHKSS, we extend the methods to carry out inference on a new set-valued estimand, the FA-frontier. We derive the asymptotic distribution of our estimator and test statistics building on liu:mol25v2,sem23,cha:che:mol:sch18,ChernozhukovDML,ber:mol08, and fan:san19, provide inverse-probability weighting correction in the DML estimator, and asymptotically valid confidence sets for the FA frontier and tests for LDA existence.

Accompanying software implementing both the MAR-based and partial identification approaches is under development, along with empirical applications of our methods.

Related Literature. A large literature in computer science, statistics, and economics studies algorithmic fairness; see cho:rot18, bar:har:nar23, and Corbett24 for overviews. Much of this work treats fairness as a constraint or regularizer in an objective that prioritizes predictive performance dwork12, ber:hei:jab:jos:kea:mor:jam:nee:roth17, or studies impossibility results and trade-offs among fairness criteria kleinberg2016inherent. A related literature computes fairness-accuracy Pareto frontiers or ranges of disparities for specific model classes, typically as an optimization or auditing exercise rather than as a problem of statistical inference wei:nie21, lit:wey:all22, cos:ram:cho21. When labels are perfectly observed, aue:lia:tab:oku24 provide a sample-splitting based test for LDA existence, and fal:jor:uli26 study minimax-optimality of the sample analog of a Pareto-optimal linear rule and propose a uniform high-probability bound on its empirical error.

Within the algorithmic fairness literature, other work has focused on selective labels lak:kle:les:lud:mul17, kle:lak:les:lud:mul17, ram:cos:ken25, kha:tam:yao. ram:cos:ken25 develop a general framework for robust evaluation and design of predictive algorithms under selective labels and unobserved confounding, delivering bounds for performance measures under alternative restrictions on the selection process. We view their contribution as addressing a distinct question from ours: how to evaluate predictive performance in the presence of counterfactual outcomes. We focus instead on identification and inference for the fairness-accuracy trade-off of group losses across all feasible algorithms using a given information set, and on hypothesis tests and confidence statements about frontier properties such as LDA existence that directly map to governance questions and design restrictions.

Outline. Section (ref) lays out notation and derives the FA-frontier extending \citetalias{lia:lu:mu:oku24} and liu:mol25v2. Section (ref) formalizes the selective labels problem, derives the sharp identification region for the FA-frontier for specific loss functions and unrestricted selection process, and obtains point identification under a MAR assumption for any loss function. Section (ref) puts forward our DML estimator and derives its asymptotic distribution. Section (ref) proposes a test for equality between the Pareto and FA-frontier. It then puts forward our test for existence of an LDA in the presence of selectively observed labels and derives its asymptotic distribution, building on liu:mol25v2. Section (ref) concludes. Our main proofs are in Appendix (ref); Appendix (ref) reports auxiliary results.

Setup

Let a population of individuals be described by an outcome vector $Y^*\in\mathcal{Y}\subset{\mathbb R}^{d_Y}$, a binary group identity $G\in\{r,b\}$ (red or blue), and a vector of covariates $X\in\mathcal{X}\subset\mathbb{R}^{d_X}$. For example, one component of $Y^*$ may denote an individual's number of active chronic illnesses in the subsequent year and the other component the cost of care; $G$ may denote the individual's race, and $X$ may include age, gender, biomarkers, comorbidity, and medication variables. As in liu:mol25v2, we leave the relation between $G$ and $X$ unspecified, but require $G$ not to be a deterministic function of $X$. Each individual receives a binary decision $D\in\{0,1\}$, e.g., whether they are automatically enrolled in a high-risk care management program; or granted bail; or granted or denied a loan (while we assume $G$ and $D$ to be binary, the results can be extend to multiple groups and multiple decisions). An algorithm $a:\mathcal{X} \mapsto [0,1]$ assigns a probability distribution to $D$ conditional on $X$; e.g., the algorithm assigns each patient a health risk score in $[0,1]$; or a recidivism probability; or a repayment probability. For simplicity, we take $a(X)$ to be the only input to the decision, and hence $D\sim a(X)$. We let $\mathcal A(\mathcal{X})$ denote the set of all (measurable) algorithms that map from the input space $\mathcal{X}$ to a probability distribution over $D$. Extending the framework in liu:mol25v2, we let $\ell^A:\{0,1\}\times\mathcal{Y}\mapsto\mathbb R$ be a loss function used to measure the accuracy of an algorithm for an individual with outcome $y\in\mathcal{Y}$, and $\ell^F:\{0,1\}\times\mathcal{Y}\mapsto\mathbb R$ a loss function used to measure the algorithm's fairness. For example, one could use classification error to measure accuracy ($\ell^A(d,y)=\mathds{1}(d\neq y_1)$), which in a health care example returns the value $1$ if the algorithm mistakenly enrolls a healthy person in the high-risk care program or fails to enroll someone who is very sick. And one could use statistical parity to measure fairness ($\ell^F(d,y)=d$), hence judging an algorithm more fair if the proportion of either group receiving the two decisions is closer.

This paper is concerned with the case where instead of observing $Y^*$, we observe $Y=\ZY^*$, where $Z \in \{0,1\}$ is an indicator variable that records when the outcome of interest $Y^*$ is observed in the testing data. In other words, we allow for selectively observed labels. Before analyzing this problem, however, we provide a tractable characterization of the FA-frontier, under the assumption that $(Y^*,X,G,Z)\sim\mathbb{P}^*$ is observed, when the loss function used to measure accuracy differs from that used to measure fairness. In doing so, we extend the analysis in Appendix O.1 of \citetalias{lia:lu:mu:oku24}. Armed with our novel characterization, we then study what can be learned about the FA-frontier and how can valid statistical inference be carried out, when in fact only the selectively observed labels $Y$ are available.

The Fairness-Accuracy Frontier

We denote $\mu_g \equiv \mathbb{P}^*(G=g)$ the population proportion of group $g \in \{r,b\}$ and, for $d \in \{0,1\}$, the (unobserved) ideal labels associated with loss function $\iota\in\{A,F\}$ by

align[align omitted — 103 chars of source]

We denote the conditional expectation of $L_d^{g,\iota}$ taken with respect to $\mathbb{P}^*(Y^*,G|X)$ by

align*[align* omitted — 142 chars of source]
defn[Ideal labels and conditional expectations] For $d\in\{0,1\}$, define \begin{align} \mathbf{L}_d &\equiv \begin{bmatrix} L_d^{r,A}\\ L_d^{b,A}\\ L_d^{r,F}-L_d^{b,F} \end{bmatrix}\in {\mathbb R}^3, \end{align} and let $\boldsymbol{\theta}_d(X)\equiv \mathbb{E}^*[\mathbf{L}_d|X]\in{\mathbb R}^3$ and $\Delta\boldsymbol{\theta}(X)\equiv \boldsymbol{\theta}_1(X)-\boldsymbol{\theta}_0(X)\in{\mathbb R}^3$.

To make sure that the ideal labels are well-defined, we assume:

asm[Moment restrictions for both losses] For all $d\in\{0,1\}, g\in\{r,b\}$, and $\iota\in\{A,F\}$, for some constants $0<c_1<1$, $0<c_2<\infty$: (i) $\mu_g\in(c_1,1-c_1)$ and $\operatorname{ess}\sup_{X\in\mathcal{X}}\mathbb{E}^*\left[\left(L_d^{g,\iota}\right)^2\big|X\right]<c_2$; (ii) $\operatorname{ess}\inf_{X\in\mathcal{X}}\mathrm{eig}_{\min}\left[Var^*(\mathbf{L}_d|X)\right]>c_1$, with $\mathrm{eig}_{\min}$ denoting the smallest eigenvalue.
defn[Algorithm space] An algorithm is a Borel-measurable map $a: \mathcal{X} \to [0,1]$, where $a(x)$ represents the probability of choosing $D=1$ given $X=x$. Denote $\mathcal{A}(\mathcal{X})$ the set of all such measurable maps. Write $\mathcal{A}$ when the covariate space is clear from context.

For each $a \in \mathcal{A}$ and $g \in \{r,b\}$, define the group-wise expected loss associated with loss function $\ell^\iota, \iota \in \{A,F\}$, when $D\sim a(X)$ as:

align[align omitted — 158 chars of source]

\citetalias{lia:lu:mu:oku24} (Appendix O.1) derive the FA-frontier and study its properties when the loss function used to measure accuracy differs from that used to measure fairness, which is the framework used in this paper. Here we adapt their definitions to our notation.\footnote{They measure accuracy with $\ell: \{0,1\} \times \mathcal{Y} \to {\mathbb R}$ (our $\ell^A$) and fairness with $\tilde{\ell}: \{0,1\} \times \mathcal{Y} \to {\mathbb R}$ (our $\ell^F$).} They work in ${\mathbb R}^2$ with a feasible set defined with respect to the accuracy loss:

align[align omitted — 116 chars of source]

\citetalias{lia:lu:mu:oku24} measure the unfairness of an algorithm $a$ through $|f(a)|$, with $f:\mathcal A\to{\mathbb R}$ a linear functional. Their framework includes as a special case $f(a) = e_r^F(a)-e_b^F(a)$, the disparity in expected fairness loss between groups, and for simplicity we specialize our discussion to that case. For each pair of expected accuracy losses $c = (c_r, c_b) \in \mathcal{E}^A$, define the minimal achievable unfairness as

equation[equation omitted — 105 chars of source]
defn[\citetalias{lia:lu:mu:oku24}'s FA-dominance and FA-frontier ] For $c,c' \in \mathcal{E}^A$, we say $c' \succ_{FA^{\text{\tiny LLMO}}} c$ if $c_r' \leq c_r,~c_b' \leq c_b,~d(c') \leq d(c)$, with at least one inequality strict. Then the FA-frontier relative to the accuracy loss is: \begin{align} \mathcal{F}^A\equiv\{c \in \mathcal{E}^A: \nexists c' \in \mathcal{E}^A such that c' \succ_{FA^{\tiny LLMO}} c\}\subset{\mathbb R}^2. \end{align}

Given the group-wise expected losses induced by algorithm $a\in\mathcal A$ for $\iota\in\{A,F\}$ in (ref), we instead propose to work with the feasible set for $\left\{\left(e^A_r(a),e^A_b(a),e^F_r(a)-e^F_b(a)\right):a\in\mathcal A\right\}$, where we explicitly include a coordinate measuring the difference in expected fairness loss for the two groups. Doing so allows us to characterize the properties of the FA frontier, extending \citetalias{lia:lu:mu:oku24}'s results, and to provide a representation of it that is amenable to estimation and inference, even in the presence of selectively observed data.

defn[Feasible set] Given $\mathbb{P}^*$ and $\mathcal A$, as $a$ ranges in $\mathcal A$, the feasible set for $\left(e^A_r(a),e^A_b(a),e^F_r(a)-e^F_b(a)\right)$ is \begin{align*} \mathcal{E} \equiv \left\{(\varepsilon_1,\varepsilon_2,\varepsilon_3) \in {\mathbb R}^3: \exists a \in \mathcal{A} such that \varepsilon_1 = e_r^A(a), \varepsilon_2 = e_b^A(a), \varepsilon_3 = e_r^F(a)-e_b^F(a)\right\}. \end{align*}

To simplify notation, we omit the dependence of $\mathcal{E}$ on $(\mathbb{P}^*,\mathcal A)$. In Lemma (ref), we show that under Assumptions (ref)-(ref), $\mathcal{E}$ is non-empty, compact, and convex. Definition (ref) leads to an alternative formulation of FA-dominance and to an FA-frontier in ${\mathbb R}^3$, as follows:\footnote{aue:lia:tab:oku24 use a similar definition of dominance.}

defn[FA-dominance and FA-frontier] Given $\varepsilon,\varepsilon' \in \mathcal{E}$, $\varepsilon' \succ_{FA} \varepsilon$ if $\varepsilon'_1 \leq \varepsilon_1,~\varepsilon'_2 \leq \varepsilon_2,~|\varepsilon'_3| \leq |\varepsilon_3|$, with at least one inequality strict. The FA-frontier is: \begin{align} \mathcal{F}\equiv\{\varepsilon \in \mathcal{E}: \nexists \varepsilon' \in \mathcal{E} such that \varepsilon' \succ_{FA} \varepsilon\}\subset{\mathbb R}^3. \end{align}

We also define the Pareto frontier, which ignores considerations of fairness:

align[align omitted — 249 chars of source]

and the points in ${\mathbb R}^3$ that minimize, respectively, $e_r^A$, $e_b^A$, and $|e_r^F-e_b^F|$:

align[align omitted — 218 chars of source]
figure[figure omitted — 1,535 chars of source]

In Appendix (ref) we show that the shape of our FA-frontier $\mathcal{F}$ depends on the position of $\mathcal{E}$, $\mathcal{PF}$, $R$ and $B$ relative to the hyperplane $\{\varepsilon_3=0\}=\{e_r^F=e_b^F\}$. Figure (ref) illustrates several examples. In the next section, we establish that our FA-frontier and \citetalias{lia:lu:mu:oku24}'s FA-frontier coincide in what we argue is the only way that matters, and we provide a characterization of $\mathcal{F}$ that does not depend on the position of $\mathcal{E}$, $\mathcal{PF}$, $R$ and $B$.

Support Function-Based Characterization of the FA-Frontier

To begin, we characterize liu:mol25v2 the support function of $\mathcal{E}$, denoted $h_\mathcal{E}(q)$ for $q$ a direction vector in $\mathbb{S}^2 \equiv \{q\in{\mathbb R}^3:\|q\|=1\}$. This functional turns out to be crucial for all our results. We assume:

asm[Margin condition] There exists a constant $m\in(0,1]$ such that for every $\delta>0$, $\sup_{q\in\mathbb{S}^2}\mathbb{P}^*\big(|q^\intercal \Delta\boldsymbol{\theta}(X)|\le \delta\big)\lesssim\delta^m$.
theorem[Support function of $\mathcal{E}$] Under Assumptions (ref)(i)-(ref), the support function of $\mathcal{E}$, denoted $h_\mathcal{E}(q)$, and its gradient with respect to $q \in \mathbb{S}^2$ are, respectively, \begin{align} h_\mathcal{E}(q)&\equiv\sup_{\varepsilon\in\mathcal{E}}q^\intercal\varepsilon=\mathbb{E}^*[q^\intercal \mathbf{L}_0 + q^\intercal(\mathbf{L}_1-\mathbf{L}_0)\mathds{1}\{q^\intercal \Delta\boldsymbol{\theta}(X)>0\}\Big],\\ \nabla_qh_{\mathcal{E}}(q)&= \mathbb{E}^*\bigl[\mathbf{L}_0+(\mathbf{L}_1-\mathbf{L}_0)\mathds{1}\{q^\intercal \Delta\boldsymbol{\theta}(X)>0\}\bigr], \end{align} and the set $\mathcal{E}$ is strictly convex.

Eq. (ref) implies that $\mathcal{S}_\mathcal{E}(q)\equiv\{\varepsilon\in\mathcal{E}:q^\intercal\varepsilon=h_\mathcal{E}(q)\}$, the support set of $\mathcal{E}$ in direction $q\in\mathbb{S}^2$, equals the singleton $\nabla_qh_{\mathcal{E}}(q)$ sch93. It yields that

align*[align* omitted — 285 chars of source]

The next corollary, proven in liu:mol25v2, characterizes precisely which algorithms yield expected losses on the boundary of $\mathcal{E}$:

corrLet Assumptions (ref)(i)-(ref) hold. Let $\partial\mathcal{E}=\{\mathcal{S}_\mathcal{E}(q):q\in\mathbb{S}^2\}$ denote the boundary of $\mathcal{E}$. Then for any algorithm $a\in\mathcal{A}(\mathcal{X})$, $\varepsilon(a)\equiv\left(e^A_r(a),e^A_b(a),e^F_r(a)-e^F_b(a)\right)\in\partial\mathcal{E}$ if and only if for some $q\in\mathbb{S}^2$, $a(X)=\mathds{1}\{q^\intercal\Delta\boldsymbol{\theta}(X)>0\},~\mathbb{P}^*_X\text{-a.s.}$

Corollary (ref) echos the generalized Neyman-Pearson Lemma leh:rom08.\footnote{ To apply that theorem, in its notation set $m=1,\,c_1=0,\,\mu=\mathbb{P}$, and let $f_1(x)=0$ for all $x\in\mathcal{X}$ and $f_2(x)=q^\intercal\Delta\boldsymbol{\theta}(x)$. The critical functions $\phi$ are our algorithms $a\in\mathcal{A}(\mathcal{X})$.} One of its important consequences is that no algorithm such that $a(X)\in(0,1)$ for a set of $X$ of positive probability can yield group risks on $\partial\mathcal{E}$.

As the goal of this section is to characterize $\mathcal{F}$, given $\varepsilon^*=(\varepsilon_1^*,\varepsilon_2^*,\varepsilon_3^*)\in \mathcal{E}$, we define

align[align omitted — 251 chars of source]

the set of group-wise expected accuracy losses and fairness disparity $\varepsilon\in{\mathbb R}^3$---whether or not they are feasible---that are both weakly more accurate and weakly fairer than $\varepsilon^*$. In Lemma (ref), we show that the support function of $\mathcal{C}(\varepsilon^*)$ equals:

align[align omitted — 302 chars of source]

Given an algorithm $a^*\in\mathcal A$ and $\varepsilon^*\equiv(e^A_r(a^*),e^A_b(a^*),e^F_r(a^*)-e^F_b(a^*))$, we establish in Theorem (ref) that $\varepsilon^*\in\mathcal{F}$ if and only if the set $\mathcal{C}(\varepsilon^*)$, depicted in the two panels of Figure (ref) as the shaded green hyper-rectangle corresponding to two different values of $\varepsilon^*$, can be properly separated from $\mathcal{E}$ sch93. Panel (a) depicts a case where $\varepsilon^*\in\mathcal{F}$ and proper separation occurs; panel (b) depicts a case where $\varepsilon^*\notin\mathcal{F}$ and proper separation fails. We also show that $\varepsilon^*\in\mathcal{F}$ if and only if $(\varepsilon_1^*,\varepsilon_2^*)\in\mathcal{F}^A$ with $\varepsilon_1^*=e^A_r(a^*)$ and $\varepsilon_2^*=e^A_b(a^*)$, and $|\varepsilon_3^*|$ equals the minimum absolute disparity in (ref) associated with $(\varepsilon_1^*,\varepsilon_2^*)$. To do so, we define the projection $\pi: {\mathbb R}^3 \to {\mathbb R}^2$ by $\pi(x_1, x_2, x_3) = (x_1, x_2)$, and let $\pi(\mathcal{E})=\{\pi(\varepsilon):\varepsilon\in\mathcal{E}\} = \mathcal{E}^A$.

figure[figure omitted — 833 chars of source]
theoremLet $\varepsilon^*=(\varepsilon_1^*,\varepsilon_2^*,\varepsilon_3^*)\in \mathcal{E}$ and define $\widetilde{\mathbb{S}}^2\equiv\{q\in \mathbb{S}^2:q_1\ge 0,\ q_2\ge 0\}$. Under Assumptions (ref)(i)-(ref), the following are equivalent: \begin{enumerate}[label=(\roman*)] • $\varepsilon^*\in \mathcal{F}$; • $\mathcal{E}\cap\mathcal{C}(\varepsilon^*)=\{\varepsilon^*\}$; • $\min_{q\in \widetilde{\mathbb{S}}^2} \left( h_{\mathcal{C}(\varepsilon^*)}(q)+h_{\mathcal{E}}(-q) \right)=0$; • $\pi(\varepsilon^*)=(\varepsilon_1^*,\varepsilon_2^*)\in \mathcal{F}^A$ and $|\varepsilon_3^*|=d\big((\varepsilon_1^*,\varepsilon_2^*)\big)$. \end{enumerate}

An immediate consequence of Theorem (ref) is that $\mathcal{F}$ captures as many features of the trade-off between accuracy and fairness as $\mathcal{F}^A$ does in \citetalias{lia:lu:mu:oku24}'s original contribution, while preserving convexity of the feasible set and yielding a simple characterization of the FA-frontier.\footnote{In fact, $\mathcal{F}$ yields a finer characterization than $\mathcal{F}^A$. Consider algorithms $a,a'\in\mathcal A$ such that $\varepsilon:=\varepsilon(a)=(0.5,0.5,1)$ and $\varepsilon':=\varepsilon(a')=(0.5,0.5,0)$. Then $\varepsilon'\succ_{FA}\varepsilon$ but $c'\not\succ_{FA^{\text{\tiny LLMO}}} c$ for $c=(\varepsilon_1,\varepsilon_2)$ and $c'=(\varepsilon_1',\varepsilon_2')$.} Indeed, using the equivalence $\ref{thm:FequivLLMO_e_in_sF}\Leftrightarrow\ref{thm:FequivLLMO_support}$ in Theorem (ref), similarly to liu:mol25v2, we have

equation[equation omitted — 204 chars of source]

When the same loss function is used to measure accuracy and fairness, as in \citetalias{lia:lu:mu:oku24}'s main text and in liu:mol25v2, it is simple to characterize when the FA-frontier coincides with the Pareto frontier in (ref).\footnote{In that case, $\ell^A=\ell^F$ and $e_g^A=e_g^F,\,g\in\{r,b\}$, so $\mathcal{E}$ can be defined in ${\mathbb R}^2$ and \citetalias{lia:lu:mu:oku24} show that the Pareto and the FA frontiers coincide if and only if $R$ and $B$, defined in (ref), lie on opposite sides of the $45^o$ line.} When different loss functions are used, such characterization is harder to obtain, and \citetalias{lia:lu:mu:oku24} provide only a sufficient condition for it. Working with the feasible set $\mathcal{E}\subset{\mathbb R}^3$ and its support function, we are able to provide a simple, necessary and sufficient, testable condition under which $\mathcal{PF}=\mathcal{F}$. The characterization is valuable because when $\mathcal{F}$ does not coincide with $\mathcal{PF}$, the trade-off between fairness and accuracy becomes substantial: a policymaker who values fairness may choose to deploy an algorithm that is Pareto dominated in accuracy loss, to ameliorate its performance relative to fairness loss. To derive the result, it is useful to distinguish between cases in which $\mathcal{E}$ has kinks, and cases where it does not, through the following condition.

asm[No Kinks] For any $q,v\in\mathbb{S}^2,q\neqv$, \begin{align} \mathbb{P}\left(q^\intercal\Delta\boldsymbol{\theta}(X)>0,v^\intercal\Delta\boldsymbol{\theta}(X)<0\right)>0. \end{align}

Assumption (ref) is a necessary and sufficient condition under which $\mathcal{E}$ has no kinks (see Supplemental Appendix B.2.3 in bon:mag:mau12, bon:mag:mau12 and Remark 3.3 in liu:mol25v2, liu:mol25v2). It yields injectivity of the support map of $\mathcal{E}$, $\mathcal{S}_\mathcal{E}(q) = \arg \max_{\varepsilon \in \mathcal{E}} q^\intercal \varepsilon$ (see Theorem (ref) and the discussion following it).

The next result characterizes when $\mathcal{PF}$ and $\mathcal{F}$ are exactly the same.

theoremUnder Assumptions (ref)(i)-(ref), $\mathcal{PF}=\mathcal{F}$ if and only if for all $\varepsilon^*\in\mathcal{F}$ there exists $\alpha\in[0,1]$ such that for $q(\alpha)=\tfrac{[\alpha,1-\alpha,0]^\intercal}{\Vert[\alpha,1-\alpha,0]^\intercal\Vert}$, $\varepsilon^*=\arg\min_{\eta\in \mathcal{E}} q(\alpha)^\intercal \eta$. If Assumption (ref) also holds, then $\mathcal{PF}=\mathcal{F}$ if and only if $\mathcal{PF}\subset\{\varepsilon_3=0\}$.

Intuitively, $\mathcal{PF}=\mathcal{F}$ when every point on the FA-frontier is Pareto-optimal. When each point on the boundary of $\mathcal{E}$ has a unique supporting direction (which is guaranteed by Assumption (ref)), this happens if and only if $\mathcal{PF}\subset\{\varepsilon_3=0\}=\{\varepsilon:e_r^F=e_b^F\}$; in this case, all points along the Pareto frontier are fair and there is no conflict between improving fairness and improving accuracy. On the other hand, when $\mathcal{E}$ has kinks, one may have $\mathcal{PF}=\mathcal{F}$ even though $\mathcal{PF}\nsubseteq\{\varepsilon_3=0\}$. The complexity of the geometry of $\mathcal{F}$ in relation to the geometry of $\mathcal{E}$ and the usefulness of our Theorem (ref) are illustrated in Figure (ref). If $\mathcal{E}$ has no kinks and lies completely to one side of the hyperplane $\{\varepsilon_3=0\}$ (Panel (a), where $\mathcal{E}\subset\{\varepsilon_3>0\}=\{e_r^F>e_b^F\}$) or $R$ and $B$ lie on opposite sides of that hyperplane (Panel (c)), then $\mathcal{PF}\subsetneq\mathcal{F}$ since $\mathcal{PF} \not\subset \{\varepsilon_3=0\}$. On the other hand, if $\mathcal{E}$ has kinks properly facing the origin, then it can be that $\mathcal{PF}=\mathcal{F}$ even when $\mathcal{E}$ lies completely to one side of the hyperplane $\{\varepsilon_3=0\}$ (Panel (b)). If $\mathcal{E}$ crosses the hyperplane $\{\varepsilon_3=0\}=\{e_r^F=e_b^F\}$ with $\mathcal{PF}$ fully contained in one of $\{\varepsilon_3>0\}$ or $\{\varepsilon_3<0\}$, then $\mathcal{PF}\subsetneq\mathcal{F}$ regardless of the presence or absence of kinks (Panel (d)), see Corollary (ref). If $\mathcal{E}$ has no kinks and $\mathcal{PF}\subset\{e_3=0\}$, then $\mathcal{PF}=\mathcal{F}$ (Panel (e)), yet having only $\{R,B\} \in\{e_3=0\}$ but not the entire Pareto frontier may yield $\mathcal{PF}\subsetneq\mathcal{F}$ (Panel (f)). Hence, the characterization in Theorem (ref) is particularly useful to avoid case-by-case distinctions and pre-tests.

Selectively observed Labels

Having fully characterized the FA-frontier in the idealized case when $(Y^*,X,G,Z)\sim\mathbb{P}^*$ is observed, we turn to our goal of providing identification results and inference procedures for $\mathcal{F}$ when only $(Y,X,G,Z)\sim\mathbb{P}$ is observed, with $Y=\ZY^*$ the selected labels and $Z$ a selection indicator. This problem has been studied, e.g., by lak:kle:les:lud:mul17,kle:lak:les:lud:mul17,ram:cos:ken25,kha:tam:yao. When evaluating whether algorithmic policies can improve upon human decision-making, selectively observed labels are compounded by the possibility that $Z$ is determined by a decision maker (e.g., a judge) who observes information beyond the recorded covariates made available to the algorithm. In this case, for a specific choice of loss functions, we report partial identification results. We then provide point identification results and statistical inference procedures for the case where label observability is induced by deployment of a decision rule in $\mathcal A(\mathcal{X})$, so that $Z$ is determined by an algorithm that uses only the inputs in $\mathcal{X}$.

Partial Identification of the FA-Frontier

In this section, we allow $Z$ to depend on variables unavailable to the algorithms in $\mathcal A(\mathcal{X})$. This creates substantial challenges for identification and inference, because not only the feasible set $\mathcal{E}$ is partially identified, but even for a fixed algorithm $a^*\in\mathcal A(\mathcal{X})$, the vector of expected losses $\varepsilon^*\equiv\varepsilon(a^*):=(e_r^A(a^*),e_b^A(a^*),e_r^F(a^*)-e_b^F(a^*))$ is also only partially identified. Consequently, extra care needs to be taken, when characterizing what can be learned about $\mathcal{F}$ in (ref), to correctly couple the possible values of $\mathcal{E}$ and $\varepsilon^*$ associated with each candidate distribution for the unobserved labels.

To make progress on this task, we specialize our analysis to the case where $Y^*$ is a binary scalar variable, the accuracy loss measures classification error, $\ell^A(d,y) = \mathds{1}\{d \neq y\}$, and the fairness loss measures statistical parity, $\ell^F(d,y) = \mathds{1}\{d=1\}$. Write $Y=\ZY^*$ for the observed binary outcome. In this case the fairness coordinate depends only on the decision rule and the observable distribution of $(X,G)$, so lack of point identification comes entirely from the selectively observed outcomes entering the two accuracy coordinates.

Denote the unobserved conditional expectation of $Y^*$ when $Z=0$ by

align*[align* omitted — 89 chars of source]

and $\lambda^*(x)\equiv[\lambda_r^*(x),\lambda_b^*(x)]^\intercal$. Without restrictions on the selection process, $\lambda^*$ is unknown. As $Y^*\in\{0,1\}$, every measurable map $\lambda\equiv(\lambda_r,\lambda_b):\mathcal{X}\to[0,1]^2$ is compatible with $\mathbb{P}$, the observed law of $(Y,X,G,Z)$.\footnote{By “compatible with the observed data” we mean that it yields a law of $(Y^*,X,G,Z)$ that reproduces the observed distribution of $(Y,X,G,Z)$ and satisfies $Y=ZY^*$ almost surely.} Let $\Lambda\equiv \left \{ \lambda: \mathcal{X} \to [ 0,1]^2 \right \}$ denote this class of functions.

We next derive the sharp identification region for $h_\mathcal{E}(q)$. To do so, we need to introduce some notation. Define the observable functions

align[align omitted — 685 chars of source]

Lemma (ref) shows that the functions $\boldsymbol{\theta}_d(\cdot)$ in Definition (ref) associated with distribution $\lambda\in\Lambda$ satisfy $\boldsymbol{\theta}_0(x;\lambda)=A_0(x)+B(x)\lambda(x)$ and $\boldsymbol{\theta}_1(x;\lambda)=A_1(x)-B(x)\lambda(x)$. Hence, the feasible set associated with distribution $\lambda$ and its support function are, respectively,

align[align omitted — 396 chars of source]

Accordingly, the sharp identified set for $h_\mathcal{E}(q)$ in direction $q$ is $\{h_\mathcal{E}(q;\lambda):\lambda\in\Lambda\}$.

For a fixed algorithm $a^*$, let $A_{d,j}(x)$ denote the $j$th coordinate of $A_d(x)$, $d\in\{0,1\}$, in (ref)-(ref) and define

equation[equation omitted — 191 chars of source]

Then we can express $h_{\mathcal{C}(\varepsilon(a^*;\lambda))}$ in (ref) associated with $\lambda\in\Lambda$ as

equation[equation omitted — 173 chars of source]

We now present an assumption on the observable distribution $\mathbb{P}$ under which the margin condition (Assumption (ref)) holds for all distributions consistent with the data.

asmThere exists a constant $m\in(0,1]$ such that, for every $\delta>0$, \begin{align*} \sup_{q\in\mathbb{S}^2} \mathbb{P}_X\left( \inf_{s\in[0,1]^2} \left|q^\intercal\{A_1(X)-A_0(X)-2B(X)s\} \right| \le\delta \right) \lesssim\delta^m. \end{align*}

Under Assumption (ref), Theorem (ref) characterizes frontier membership for a given $\lambda$, where crucially we use the same candidate unobserved distribution $\lambda$ both in obtaining $h_\mathcal{E}(\cdot)$ and $h_{\mathcal{C}(\varepsilon^*)}(\cdot)$. Recall $\varepsilon^*(a^*;\lambda)$ denotes the vector of expected losses induced by $a^*$ under unobserved distribution $\lambda$; then $a^*$ lies on the FA-frontier under $\lambda$ if and only if

align*[align* omitted — 135 chars of source]

The next result shows how to check whether $a^*$ can generate expected losses on the frontier for some distribution of the unobserved labels, through a finite dimensional optimization problem that involves functionals of the observed distribution $\mathbb{P}$.

theoremLet Assumptions (ref)(i) and (ref) hold. Suppose $Y^*\in\{0,1\}$, $\ell^A(d,y)=\mathds{1}\{d\neq y\}$, and $\ell^F(d,y)=\mathds{1}\{d=1\}$. Then a given algorithm $a^*$ generates expected losses on the FA-frontier for some distribution consistent with the observed data if and only if \begin{equation} \min_{q \in \tilde\mathbb{S}^2} \mathbb{E}\left[ \max\{J_0(\lambda_1;q,a^*),J_1(\lambda_0;q,a^*)\} \right] = 0. \end{equation} where $\lambda_0=(0,0)^\intercal$, $\lambda_1=(1,1)^\intercal$, and \begin{align} J_0(\lambda;q,a^*) &\equiv C(q;a^*)-q^\intercal A_0(X)-2a^*(X)q^\intercal B(X)\lambda,\nonumber\\ J_1(\lambda;q,a^*) &\equiv C(q;a^*)-q^\intercal A_1(X)+2(1-a^*(X))q^\intercal B(X)\lambda. \end{align}

Theorem (ref) shows that once the source of partial identification is reduced to the functional $\lambda$, the search over the infinite-dimensional class $\Lambda$ collapses to a finite-dimensional criterion involving only the corner distributions $\lambda_0$ and $\lambda_1$.

Our next goal is to characterize the sharp identification region of $\mathcal{F}$ using a finite dimensional optimization problem that does not involve $a\in\mathcal{A}(\mathcal{X})$ and $\lambda\in\Lambda$. Recall $\mathcal{E}(\lambda)$ in (ref) and define the FA frontier associated with distribution $\lambda$ and the FA-frontier envelope as

align[align omitted — 307 chars of source]

For a given $q\in\widetilde{\mathbb{S}}^2$, collect all expected loss vectors achievable by an algorithm that takes the threshold form specified in Corollary (ref) for some distribution $\lambda$ and thus are the boundary points of $\mathcal{E}(\lambda)$, in the following set\footnote{The minimization in the definition of $\mathcal{M}_q$ is a sign convention that we use because the first two coordinates are losses (we could alternatively take $\arg\max_d -q^\intercal\boldsymbol{\theta}_d(x;\lambda)$ for $q\in\widetilde{\mathbb{S}}^2$).}

multline[multline omitted — 295 chars of source]

In Lemma (ref) we show that for every $q\in\widetilde{\mathbb{S}}^2$, the set $\mathcal{M}_q$ is nonempty, compact, and convex. In Lemma (ref) we show that its support function, $h_{\mathcal{M}_q}(v),v\in\mathbb{S}^2$, is given by

align*[align* omitted — 111 chars of source]

where $\kappa_0(v,q;x)$ and $\kappa_1(v,q;x)$, formally defined in (ref)-(ref), are known functions of $A_0(x),A_1(x),B(x)$. We then have the following characterization of $\mathcal{F}^\exists$:

theoremLet Assumptions (ref) and (ref) hold. Then \begin{align} \mathcal{F}^\exists=\bigcup_{q\in\widetilde{\mathbb{S}}^2}\left(\mathcal{M}_q\cap\{\varepsilon\in{\mathbb R}^3:q_3\varepsilon_3\ge0\}\right). \end{align} For $\varepsilon\in{\mathbb R}^3$, define \begin{align} T_{\mathrm{env}}(\varepsilon)&\equiv\inf_{q\in\widetilde{\mathbb{S}}^2:\,q_3\varepsilon_3\ge0}[\sup_{v\in\mathbb{S}^2}\{v^\intercal\varepsilon-h_{\mathcal{M}_q}(v)\}]_+. \end{align} Then, for every $\varepsilon\in{\mathbb R}^3$, $\varepsilon\in\mathcal{F}^\exists\iff T_{\mathrm{env}}(\varepsilon)=0$.
remarkIn work in progress, we obtain a debiased machine learning estimator for $J_d(\lambda;q,a^*)$, $d\in\{0,1\}$, and we develop an inference procedure to test the hypothesis that a given algorithm generates expected losses on the FA-frontier. We also explore extending Theorem (ref) to other loss functions.

Point Identification Under Missing at Random Assumptions

While the distribution $\mathbb{P}$ of $(Y,G,X,Z)$ is point identified by the observed data, in the absence of additional assumptions the distribution $\mathbb{P}^*$ of $(Y^*,G,X,Z)$ is not. The previous section shows that, as a consequence, the support function $h_\mathcal{E}(q)$ and $\varepsilon^*(a),\,a\in\mathcal{A}(\mathcal{X})$, are not point identified either, as the ideal labels and their conditional expectations ($\mathbf{L}_d$ and $\boldsymbol{\theta}_d(X)$, respectively) are a function of $Y^*$ and $\mathbb{P}^*$.

We leverage the fact that all algorithms in $\mathcal A(\mathcal{X})$ can only use the covariates $X$ for training (and no other unobserved variables) to argue plausibility of a selection on observables assumption that we use to point identify the distribution $\mathbb{P}^*$.

asm[Ideal labels missing at random (MAR)] $(Y^*,G) \perp Z | X$ and $\pi(X)\equiv \mathbb{E}[ Z |X] \in (0,1)$ a.s..

Assumption (ref) may look stronger than the typical missing at random (unconfoundedness or selection on observables) assumption, which might instead require $Y^* \perp Z | X,G$. Our approach remains valid under this weaker condition, but here we work with Assumption (ref) because it yields lighter notation and simpler expressions. We argue that, in many settings of interest, the additional requirement $G \perp Z | X$ is a natural consequence of how $Z$ is generated. Specifically, suppose that $Z=D$, where $D\in\{0,1\}$ is the binary decision an individual receives (e.g., granted or denied a loan; granted or denied bail), and that this decision is produced by a deployed decision rule in $\mathcal A(\mathcal{X})$ that takes $X$ as input and, following regulations, does not use $G$ or apply any group-dependent post-processing. Then, for all $g\in\{r,b\}$, $\mathbb{P}(D=1| X,G=g)=\mathbb{P}(D=1| X)= a(X)$, so that $G \perp Z | X$ holds. Combined with the standard selection on observables condition, $Y^* \perp Z | X,G$, this yields Assumption (ref). The more restrictive part of Assumption (ref) is therefore the requirement $Y^* \perp Z | X,G$. This condition is plausible when the assignment mechanism that determines $Z$ depends on recorded variables capturing the main drivers of both selection and outcomes, and when discretion at the decision-making stage is limited or can be accounted for. In such settings, conditioning on $(X,G)$ may reasonably approximate conditioning on the information that drives selection, so that $Z$ does not convey additional information about $Y^*$ beyond $(X,G)$. When instead the assignment process exploits unrecorded “soft information” (e.g., interviews, free text, documents, or unlogged history) or when discretion is systematically related to expected outcomes, $Y^* \perp Z | X,G$ may fail, motivating our partial identification analysis in Section (ref).

The next lemma shows that, under Assumption (ref) the support function of $\mathcal{E}$ is point identified through a standard inverse propensity weighting identity:

lemmaLet Assumptions (ref)(i) and (ref) hold. Then for $\Delta\mathbf{L}\equiv\mathbf{L}_1-\mathbf{L}_0$, \begin{align} \boldsymbol{\theta}_d(X) &= \mathbb{E} \left[ \left. \tfrac{Z\mathbf{L}_d }{\pi(X)} \right| X \right], \\ h_\mathcal{E}(q) &= \mathbb{E} \left[ \tfrac{q^\intercal \mathbf{L}_0Z }{\pi(X)} + \left ( \tfrac{q^\intercal\Delta\mathbf{L} Z}{\pi(X)}\right)\mathds{1}\{q^\intercal \Delta\boldsymbol{\theta}(X)>0\} \right],\quadfor all q\in\mathbb{S}^2,\\ e_g^\iota(a) & = \mathbb{E}\left[ \tfrac{a(X)L_1^{g,\iota} Z}{\pi(X)} + \tfrac{(1-a(X))L_0^{g,\iota} Z}{\pi(X)} \right],\quad g=r,b and \iota=A,F. \end{align}

A Debiased Machine Learning Estimator and Its Asymptotic Distribution

The DML Estimator

Our goal is to estimate the support function $h_\mathcal{E}(q)$ using a random sample $\{(Y_i,G_i,X_i,Z_i): 1 \le i \le n \}$ drawn from the distribution $\mathbb{P}$. By Lemma (ref), $h_\mathcal{E}(q)$ is identified by (ref), which involves nuisance functions $\pi(X)$ and $\Delta \boldsymbol{\theta}(X)$ that need to be estimated in the first step. As the ideal labels $\mathbf{L}$ depend on $\boldsymbol{\mu} \equiv [\mu_r,\mu_b]^\intercal$ (see (ref)-(ref)), we also need to estimate $\boldsymbol{\mu}$. In order to build an estimator for $h_\mathcal{E}(q)$ that is locally robust to the first-step estimation error, we propose using the debiased machine learning (DML) approach. Consider the estimand

equation[equation omitted — 141 chars of source]

where $\boldsymbol{\eta}\equiv(\Delta \boldsymbol{\theta},\pi,\boldsymbol{\theta}_0)$ is a vector of nuisance functions and

align[align omitted — 507 chars of source]

The local effect of the first-step function $\pi$ on the original estimand $h_\mathcal{E}(q)$ in (ref) is accounted for by $ \alpha^h(q;X_i) \left(1 - \tfrac{Z_i}{\pi(X_i)}\right)$. As we show in the proof of Theorem (ref), under Assumptions (ref) and (ref) presented below, it is not necessary to account for the local effect of the first-step function $\Delta \boldsymbol{\theta}$ on $h_\mathcal{E}(q)$. Finally, we note that the estimand in (ref) is based on $\mathbf{L}$, which by (ref) relies on the population means $\mu_g, g\in\{r,b\}$.

Following the DML approach ChernozhukovDML,velez2024asymptotic, we first randomly split the indices $[n]\equiv\{1,\ldots,n\}$ into $K$ equal-sized folds $\mathcal{I}_k$, i.e., $\cup_{k=1}^K \mathcal{I}_k = [n]$. We denote by $n_k$ the size of $\mathcal{I}_k$.\footnote{When $n$ is not divisible by $K$, the number of observations in some folds will be $\lfloor n/K \rfloor$ while in others $\lfloor n/K \rfloor + 1$, where $\lfloor n/K \rfloor$ is the greatest integer less than or equal to $n/K$.} We then estimate $\widehat{\boldsymbol{\eta}}(X_i)$ as $\widehat{\boldsymbol{\eta}}_k(X_i)$ for $i \in \mathcal{I}_k$ and $k=1,\ldots,K$, where $\widehat{\boldsymbol{\eta}}_k(\cdot)$ is estimated using all data except the portion with indices in fold $\mathcal{I}_k$. Finally, the first-stage estimators for the nuisance functions are used to form a second-stage estimator for $h_\mathcal{E}(q)$ based on the expression of the estimand defined in (ref):

equation[equation omitted — 362 chars of source]

where $\widehat\mathbf{L}$ is defined as in (ref)-(ref) with $\mu_g$ replaced by $\widehat{\mu}_g \equiv n^{-1} \sum_{i=1}^n \mathds{1}\{G_i=g\}$.

We also propose a DML estimator for $\varepsilon^* = \varepsilon(a^*) \in \mathcal{E}$. By (ref) in Lemma (ref) and Definition 3, we can write $\varepsilon^*= \mathbb{E}[ \xi_i(a^*;\mathbf{L},\boldsymbol{\eta})]$, where for a given algorithm $a\in\mathcal A$,

align[align omitted — 368 chars of source]

and again the local effect of the first-step function $\pi$ on the original estimand $\varepsilon^*$ is accounted for by $\boldsymbol{\alpha}^e(X_i) \big(1 - \tfrac{Z_i}{{\pi}(X_i)}\big)$. Using this notation, we propose to estimate $\varepsilon^*$ by

equation[equation omitted — 310 chars of source]

where $\widehat{\mathbf{L}}$ stacks the feasible labels and $\{\widehat{\boldsymbol{\eta}}_k\}_{k=1}^K$ are estimators for the nuisance functions using cross-fitting, as we described above.

Asymptotic Distribution of the DML Estimator

For the nuisance functions' estimation error to be asymptotically negligible for the DML estimator, the nuisance estimators must satisfy suitable regularity conditions set out next.

asmThere is a known partition of $X$, $ X = (X_1, X_2),$ with $X_1 \in \mathcal{X}_1 \subset {\mathbb R}^{p}$ and $X_2 \in \mathcal{X}_2 \subset {\mathbb R}^{d_X-p}$ such that: \begin{itemize} • For every $\delta>0$, $ \sup_{q\in\mathbb{S}^2}\mathbb{P}\left( |q^\intercal \Delta \boldsymbol{\theta}(X)| \le \delta | X_2 \right) \lesssim \delta. $ • There is a constant $c>0$ such that $\inf_{ q \in \mathbb{S}^2} Var[ \,|q^\intercal \Delta \boldsymbol{\theta}(X)| \,| X_2] \ge c~.$ • There are unknown functions $\nu:\mathcal{X}_1\to{\mathbb R}^{d_\nu}$ and $\gamma:\mathcal{X}_2\to{\mathbb R}^{d_\gamma}$, and a known function $F: {\mathbb R}^{d_\nu} \times {\mathbb R}^{d_\gamma} \times {\mathbb R}^2 \to [\epsilon,1-\epsilon] \times {\mathbb R}^3$, such that: \begin{align*} (\pi(X), \Delta \boldsymbol{\theta}(X)) &= F(\nu(X_1), \gamma(X_2),\mu) \end{align*} and the partial derivatives of $F$ are uniformly bounded for any $\nu \in {\mathbb R}^{d_\nu}$, $\gamma \in {\mathbb R}^{d_\gamma}$, and $\mu \in \mathbb{B}_\epsilon(\mu_r,\mu_b) \equiv\{ \mu \in {\mathbb R}^2: \|\mu - (\mu_r,\mu_b)\|<\epsilon\}$, that is $$\sup_{\nu \in {\mathbb R}^{d_\nu}} \sup_{ \gamma\in {\mathbb R}^{d_\gamma}} \sup_{\mu \in \mathbb{B}_\epsilon(\mu_r,\mu_b) }\| D F(\nu, \gamma,\mu)\| < \infty~.$$ \end{itemize}

A sufficient condition for Assumption (ref) is given in liu:mol25v2, where a partially linear structure is assumed for $\Delta\boldsymbol{\theta}$.

Let $\mathfrak{F}_k \equiv \sigma( (Y_i,G_i,X_i,Z_i): i \notin \mathcal{I}_k)$ be the $\sigma$-algebra generated by the data used for constructing the estimator $\widehat{\boldsymbol{\eta}}_k$ for the nuisance ${\boldsymbol{\eta}} \equiv (\Delta \boldsymbol{\theta},\pi,\boldsymbol{\theta}_0)$. Let $\widehat{\boldsymbol{\mu}}_k$ be an estimator for $\boldsymbol{\mu} \equiv [\mu_r,\mu_b]^\intercal$ that uses only the data outside fold $\mathcal{I}_k$ (similar to how we define $\widehat{\boldsymbol{\eta}}_k$): $$\widehat{\boldsymbol{\mu}}_k = \left(\tfrac{1}{n-n_k} \sum_{i \notin \mathcal{I}_k} \mathds{1}\{G_i=r\},~\tfrac{1}{n-n_k} \sum_{i \notin \mathcal{I}_k} \mathds{1}\{G_i=b\} \right)^\intercal~.$$

We assume that the estimators for $\Delta \boldsymbol{\theta}(X)$ and $\pi(X)$ share the same structure as in Assumption (ref). Furthermore, we treat $\nu(X_1)$ as a low-dimensional nonparametric component that can be estimated uniformly well over $\mathcal{X}_1$, whereas $\gamma(X_2)$ captures a high-dimensional component that can be estimated sufficiently well in the sense that its mean squared error converges to zero sufficiently fast. We assume that $\boldsymbol{\theta}_0$ can be estimated sufficiently well.

asmThere exist estimators $\widehat{\nu}_k$, $\widehat{\gamma}_k$, and $\widehat{\boldsymbol{\theta}}_{0,k}$ for $k=1,\ldots,K$ such that \begin{enumerate} • $(\widehat{\pi}_k(X),~\Delta \widehat{\boldsymbol{\theta}}_k(X)) = F(\widehat{\nu}_k(X_1), \widehat{\gamma}_k(X_2),\widehat{\boldsymbol{\mu}}_k)$. • $\sup_{x_1 \in \mathcal{X}_1} \|\widehat{\nu}_k(x_1) - \nu(x_1)\| = o_p(n^{-1/4})$. • $\mathbb{E} \left[ \|\widehat{\gamma}_k(X_2)-{\gamma}(X_2)\|^2 | \mathfrak{F}_k \right] = o_p(n^{-1/2})$ . • $\mathbb{E} \left[ \|\widehat{\boldsymbol{\theta}}_{0,k}(X)-{\boldsymbol{\theta}}_0(X)\|^2 | \mathfrak{F}_k \right] = o_p(n^{-1/2})$. \end{enumerate}

Example machine learning methods that, under suitable conditions, are $n^{1/4}$-consistent in the $L^2$-norm include $\ell_1$-penalized methods, boosting, and neural nets ChernozhukovDML.

Similarly to liu:mol25v2, our main asymptotic result shows that $\widehat{h}_{\mathcal{E}}(q;\widehat{\mathbf{L}}, \widehat{\boldsymbol{\eta}})$ converges to a Gaussian process uniformly in $q\in\mathbb{S}^2$, where the score $\zeta_i(q;\mathbf{L},\boldsymbol{\eta})$ in (ref) is the influence function that governs the part of the limit distribution of $\widehat{h}_{\mathcal{E}}(q; \widehat{\mathbf{L}},\widehat{\boldsymbol{\eta}})$ due to the uncertainty in $\widehat{\boldsymbol{\eta}}$, and the remaining part is attributed to estimating $\boldsymbol{\mu}$:

theoremLet Assumptions (ref)-(ref) and (ref)-(ref) hold. Then, $$ \sqrt{n}\left( \widehat{h}_\mathcal{E}(q; \widehat{\mathbf{L}}, \widehat{\boldsymbol{\eta}}) - h_\mathcal{E}(q) \right) \Rightarrow \mathbb{G}[\zeta_i^{*}(q;\mathbf{L},\boldsymbol{\eta})] \quad \text{in} \quad \ell^{\infty}(\mathbb{S}^2)~,$$ where \begin{align} \zeta_i^*(q;\mathbf{L},{\boldsymbol{\eta}}) \equiv \zeta_i(q;\mathbf{L},{\boldsymbol{\eta}}) + \sum_{g \in \{r,b\}} \Gamma_g^h(q) \left(1-\tfrac{\mathds{1} \{ G_i = g\}}{\mu_g}\right), \end{align} and $\Gamma_g^h(q)\equiv \mathbb{E}\left[ \mathds{1}\{G_i=g\} \left(q^\intercal \mathbf{L}_{0,i}+ q^\intercal\Delta\mathbf{L}_i\mathds{1}\{q^\intercal \Delta\boldsymbol{\theta}(X)>0\} \right) \tfrac{Z_i}{\pi(X_i)} \right] $
remarkOne could alternatively construct a DML estimator of $h_\mathcal{E}(q)$ using $\zeta_i^{*}(q;\mathbf{L},\boldsymbol{\eta})$ in (ref) instead of $\zeta_i(q;\mathbf{L},\boldsymbol{\eta})$ in (ref). Doing so would be valid because $\zeta_i^{*}(q;\mathbf{L},\boldsymbol{\eta})$ satisfies the Neyman orthogonality condition with respect to $\boldsymbol{\eta}$ and $\boldsymbol{\mu}$. In practice, when $\widehat{\boldsymbol{\eta}}$ is obtained via cross-fitting and $\boldsymbol{\mu}$ is estimated by sample means, $\widehat{h}_\mathcal{E}(q; \widehat{\mathbf{L}}, \widehat{\boldsymbol{\eta}})$ is numerically equivalent to $\frac{1}{n}\sum_{i=1}^n \zeta_i^{*}(q; \widehat{\mathbf{L}},\widehat{\boldsymbol{\eta}})$. Hence, our proposed estimator implicitly accounts for first-step estimation error in $\widehat\boldsymbol{\mu}$.

Testing Procedures

We provide test statistics and their asymptotic distribution to test whether (i) the FA-frontier coincides with the Pareto frontier; and (ii) a less discriminatory alternative to a given algorithm exists.

Testing $\mathcal{PF}=\mathcal{F}$

Theorem (ref) provides a necessary and sufficient condition for $\mathcal{PF}=\mathcal{F}$. Here we put forward an equivalent characterization based on the support function process, which is amenable to testing the null hypothesis $\mathcal{PF}=\mathcal{F}$ against the alternative $\mathcal{PF}\neq\mathcal{F}$.

propLet Assumptions (ref)-(ref) hold. For $C$ a constant pinned down by Assumption (ref), $\mathcal{E}\subset \mathbb{B}_C$, with $\mathbb{B}_C$ the ball in ${\mathbb R}^3$ of radius $C$. Define \begin{align*} T^{LDA}(\varepsilon)&\equiv\left[\sup_{q\in\mathbb S^2}\{q^\intercal\varepsilon-h_\mathcal{E}(q)\}\right]_+ +\left[\inf_{q\in\widetilde{S}^2}\{h_{\mathcal C(\varepsilon)}(q)+h_\mathcal{E}(-q)\}\right]_+,\\ \psi_\mathcal{PF}(\varepsilon)&\equiv\left[\inf_{\alpha\in[0,1]}\{q(\alpha)^\intercal\varepsilon+h_\mathcal{E}(-q(\alpha))\}\right]_+, \qquad q(\alpha)\equiv\tfrac{[\alpha,1-\alpha,0]^\intercal}{\Vert[\alpha,1-\alpha,0]^\intercal\Vert}. \end{align*} Then $\mathcal{PF}=\mathcal{F} \iff \sup_{\varepsilon\in \mathbb{B}_C:T^{\texttt{LDA}}(\varepsilon)=0}\psi_\mathcal{PF}(\varepsilon)=0$. If Assumption (ref) also holds, fix any $\bar s>0$ and define \begin{align} \Delta_{\mathcal{PF}}^{h}(\bar s) \equiv \sup_{\alpha\in[0,1]} \left[ h_\mathcal{E}(-q(\alpha)) - \inf_{s\in[-\bar s,\bar s]} \sqrt{1+s^2}\, h_\mathcal{E}\left( \tfrac{-q(\alpha)+s u_3}{\sqrt{1+s^2}} \right) \right], \end{align} where $u_3=[0,0,1]^\intercal$. Then $\mathcal{PF}=\mathcal{F}\iff \Delta_{\mathcal{PF}}^{h}(\bar s)=0$.

Eq. (ref) translates the geometric condition $\mathcal{PF}\subset\{\varepsilon_3=0\}$ in Theorem (ref) into a criterion involving only the support function. For each Pareto-supporting direction $-q(\alpha)$, the expression inside the supremum checks whether $-q(\alpha)$ already minimizes the support value over fairness-direction tilts $\{-q(\alpha)+s u_3:s\in[-\bar s,\bar s]\}$. In the proof of Proposition (ref), we show that the derivative of this tilted support function at $s=0$ equals the fairness coordinate of the Pareto point; by convexity, $s=0$ minimizes the support value over $[-\bar s,\bar s]$ if and only if this derivative---and thus the fairness coordinate---is zero. Hence, the criterion in (ref) is zero exactly when all Pareto points have zero fairness disparity, i.e. $\mathcal{PF}\subset\{\varepsilon_3=0\}$, which is equivalent to $\mathcal{PF}=\mathcal{F}$ by Theorem (ref).

remarkIn work in progress, we derive the limit distribution of sample analogs of $\psi_{\mathcal{PF}}(\varepsilon)$ and $\Delta_{\mathcal{PF}}^{h}$, and valid bootstrap critical values.

Assessing the Existence of a Less Discriminatory Alternative

When evaluating whether there exists a less discriminatory alternative (LDA) to a given algorithm, regulators and policymakers need to confront the fact that they can only rely on finite data to do so. Concretely, given an algorithm $a^*\in\mathcal{A}(\mathcal{X})$, we call another algorithm $\tilde{a} \in\mathcal{A}$ an LDA if it yields expected losses that are at least as accurate as those associated with $a^*$ for both groups, and at least as fair, with one of these inequalities strict. In terms of the FA-dominance notion and FA-frontier reported in Definition (ref), the absence of an LDA is equivalent to the fact that $a^*$ yields expected losses with

align[align omitted — 140 chars of source]

By Theorem (ref), (ref) and (ref) below are equivalent, hence we can use the latter to test the hypothesis that no LDA exists against the alternative that it does, as follows:

align[align omitted — 304 chars of source]

where $h_{\mathcal{C}(e^*)}(q)$ is defined in (ref). We refer to (ref) as the LDA hypothesis.

We propose to reject the LDA hypothesis for large values of the following test statistic:

align[align omitted — 459 chars of source]

where $\widehat{h}_\mathcal{E}^\omega$ is a weighted estimator for $h_\mathcal{E}$ defined in (ref) in Appendix (ref) and $\widehat{\varepsilon}^{\omega}$ is a weighted estimator for $\varepsilon^*$ defined in (ref) in Appendix (ref). The former uses weights $\omega_i^h = 1 - \Xi_i/2$ and the latter uses weights $\omega_i^\varepsilon = 1 + \Xi_i/2$, $i=1,\dots,n$, with $\Xi_i$ a Rademacher random variable taking values in $\{-1,1\}$ with uniform probability and independent of the sample data and of the first stage training split. The test statistic in (ref) was proposed in liu:mol25v2. The reason why it resorts to using reweighted estimators for $h_\mathcal{E}(q)$ and $\varepsilon^*$ is that under the null hypothesis that $\varepsilon(a^*)\in\mathcal{F}$, Corollary (ref) yields that there exists a direction vector $q^*\in\tilde{\mathbb{S}}$ such that $a^*(X)=\mathds{1}\{q^{*\intercal}\Delta\boldsymbol{\theta}(X)>0\},\,\mathbb{P}^*_X$-a.s. As a consequence, one can show that under Assumption (ref), for $\varepsilon^*$ a point at which $h_{\mathcal{C}(\varepsilon^*)}(q^*)$ is locally differentiable, the limit distribution of a the test statistic like the one in (ref) but with $\omega_i^h=\omega_i^e=1$ for all $i=1,\dots,n$, has a degenerate limit distribution. liu:mol25v2 show that this test statistic with degenerate limit distribution asymptotically rejects a true null with probability zero, and a false null under local ($1/\sqrt n$) alternatives motivated by practical significance criteria with probability one bla:spi22. The use of Rademacher weights regularizes the limit distribution of the test statistic, trading asymptotic size $\alpha$ for the null and local asymptotic unbiasedness at local alternatives closer to the null than $1/\sqrt n$ times the practical significance threshold, for power above nominal size instead of one just above that threshold.

Our contribution is deriving the test's asymptotic distribution as well as a valid bootstrap procedure to estimate the associated critical values, in the presence of selectively observed labels. We consider a bootstrap-based critical value $\widehat{c}_{1-\alpha}$ defined in Procedure (ref) below. To simplify notation, let $\boldsymbol{\beta} \equiv ( {h}_\mathcal{E}(q),{\varepsilon}^*) \in \ell^{\infty}(\mathbb{S}^2) \times {\mathbb R}^3$ denote the vector of population values and let $\widehat{\boldsymbol{\beta}}^\omega \equiv ( \widehat{h}_\mathcal{E}^\omega(q),\widehat{\varepsilon}^\omega) \in \ell^{\infty}(\mathbb{S}^2) \times {\mathbb R}^3$ denote the vector of their estimators. We can write $T_n^{\texttt{LDA}} = \sqrt{n} \left( \phi(\widehat{\boldsymbol{\beta}}^\omega) - \phi(\boldsymbol{\beta}) \right)$, where $\phi$ is defined in (ref) in Appendix (ref).

proc[Bayesian Bootstrap for the Quantiles of $T_n^{\texttt{LDA}}$] \begin{enumerate} • Draw $\{W_i\}_{i=1}^n$ i.i.d. from the exponential distribution with mean $1$ independent of the sample $\{(Y_i,G_i,X_i,Z_i)\}_{i=1}^n$ and the weights $\{\Xi_i\}_{i=1}^n$. Let the bootstrap analogue of $\widehat{\boldsymbol{\beta}}^\omega$ be $\widetilde{\boldsymbol{\beta}}^\omega \equiv ( \widetilde{h}_\mathcal{E}^\omega,\widetilde{\varepsilon}^\omega)$, where $ \widetilde{h}^\omega_\mathcal{E}(q;\widetilde{\mathbf{L}}^h,\widehat{\boldsymbol{\eta}}) \equiv \tfrac{1}{n} \sum_{i =1}^n \left( \tfrac{W_i \omega_i^h}{\overline{W^h}} \right) \zeta_i(q;\widetilde{\mathbf{L}}^h,\widehat{\boldsymbol{\eta}}) $, $ \widetilde{\varepsilon}^\omega \equiv \tfrac{1}{n} \sum_{i =1}^n \left( \tfrac{W_i \omega_i^\varepsilon}{\overline{W^\varepsilon}} \right) \xi_i(a;\widetilde{\mathbf{L}}^\varepsilon,\widehat{\boldsymbol{\eta}})$, $\widetilde\mathbf{L}^h$ and $\widetilde\mathbf{L}^\varepsilon$ are defined as in (ref)-(ref) with $\mu_g$ replaced by $\widetilde{\mu}_g^h = n^{-1} \sum_{i=1}^n \left( \tfrac{W_i \omega_i^h}{\overline{W^h}} \right) \mathds{1}\{G_i=g\}$ and $\widetilde{\mu}_g^\varepsilon = n^{-1} \sum_{i=1}^n \left( \tfrac{W_i \omega^\varepsilon}{\overline{W^\varepsilon}} \right) \mathds{1}\{G_i=g\}$, respectively, and $\overline{W^h} = n^{-1} \sum_{i=1}^n W_i \omega_i^h$ and $\overline{W^\varepsilon} = n^{-1} \sum_{i=1}^n W_i \omega_i^\varepsilon$. • Numerically approximate $\phi'_{\boldsymbol{\beta}}(\cdot)$, the directional derivative of $\phi(\cdot)$ at $\boldsymbol{\beta}$, by \begin{align*} \widehat{\phi}'_{\boldsymbol{\beta}}( \Ddot{\boldsymbol{\beta}})=\tfrac{1}{s_n}\left(\phi\left(\widehat{\boldsymbol{\beta}}^\omega+s_n(\Ddot{\boldsymbol{\beta}})\right)-\phi\left(\widehat{\boldsymbol{\beta}}^\omega\right)\right), \end{align*} where $\Ddot{\boldsymbol{\beta}}\in\ell^\infty(\mathbb{S}^2)\times\mathbb{R}^3$ is a candidate direction at which we evaluate $\phi'_{\boldsymbol{\beta}}(\cdot)$ and $s_n$ is a vanishing sequence of step sizes such that $\sqrt{n}s_n\to\infty$. • Obtain $\widehat{\phi}'_{\boldsymbol{\beta}}\big(\sqrt{n}\{\widetilde{\boldsymbol{\beta}}^\omega-\widehat{\boldsymbol{\beta}}^\omega\}\big)$ and calculate \begin{align*} \widehat{c}_{1-\alpha}\equiv\inf\bigg\{c:\mathbb{P}\bigg(\left.\widehat{\phi}'_{\boldsymbol{\beta}}\big(\sqrt{n}\{\widetilde{\boldsymbol{\beta}}^\omega-\widehat{\boldsymbol{\beta}}^\omega\}\big)\leq c \,\,\right|\, \{(Y_i,G_i,X_i,Z_i,\Xi_i)\}_{i=1}^n\bigg)\geq 1-\alpha\bigg\} . \end{align*} \end{enumerate}

Let us define the LDA test by $\varphi_n^{\texttt{LDA}} = \mathds{1} \{ T_n^{\texttt{LDA}} > \widehat{c}_{1-\alpha + \kappa}+\kappa\}$ for a given significance level $\alpha \in (0,1)$, where $\kappa>0$ is an arbitrarily small positive constant. The constant $\kappa$ can be taken equal to $10^{-6}$ as in and:shi13. We can take $\kappa = 0$ when the asymptotic distribution of $T_n^{\texttt{LDA}}$ is continuous at its $(1-\alpha)$-quantile; $\kappa>0$ appears in the critical value used in the LDA test to handle the technicalities arising from possible discontinuity points of the asymptotic distribution of $T_n^{\texttt{LDA}}$.

The next result establishes the consistency of the Bayesian bootstrap outlined in Procedure (ref) and guarantees the asymptotic correct size of the proposed test $\varphi_n^{\texttt{LDA}}$.

theoremLet Assumptions (ref)--(ref) hold. Then, the Bayesian bootstrap outlined in Procedure (ref) is consistent. Furthermore, if the LDA hypothesis holds, the test $\varphi_n^{\texttt{LDA}}$ has asymptotically correct size, $\limsup_{n \to \infty}\mathbb{E}[\varphi_n^{\texttt{LDA}}] \le \alpha~.$

Given Theorem (ref), we can build an asymptotically valid confidence set for $\mathcal{F}$ by test inversion, similarly to liu:mol25v2.

Conclusions

In this paper, we consider the problem of identification and inference for the FA-frontier put forward in \citetalias{lia:lu:mu:oku24}, when labels (outcomes) are selectively observed. We build on the methodology developed by liu:mol25v2, who assume away the selective labels problem to focus on deriving moment-inequality representations for the frontier. Selective labels challenge one's ability to (point) identify features of the conditional distribution of the true labels, including the FA-frontier. Naively computed accuracy metrics using only observed outcomes may misrepresent true performance, potentially overstating accuracy where it matters most and distorting conclusions on whether LDAs exist.

Under the assumption that labels are missing at random conditional on a rich set of observed covariates, we obtain point identification of the FA-frontier using inverse propensity score weighting and derive the asymptotic distribution of the test statistic that we use to characterize whether an algorithm is on the frontier. The inference method that we propose is based on a debiased machine learning estimator, which is particularly appropriate in this context, where high-dimensional administrative data enable flexible nonparametric modeling of the selection process and hence make the missing at random assumption more credible, but formal inference with valid confidence statements remains necessary for legally defensible determinations, making the use of methods that control the bias of the first-step nonparametric estimation crucial.

When the selection process is left completely unrestricted (hence, we dispense with the missing at random assumption) we provide a tractable characterization of the sharp identification region for the FA-frontier, for a specific class of loss functions. In work in progress, we derive a DML estimator and its asymptotic theory (along with valid bootstrap methods) to build a confidence set for the partially identified FA-frontier and to test the hypothesis that no LDA exists to a given algorithm.

appendix\section{Proofs of the Main Results} Assume throughout that $(\mathcal{X},\mathcal B(\mathcal{X}))$ is a standard Borel space. \subsection{Proofs for Section (ref)} \begin{proof}[Proof of Theorem (ref)] Recall $\mathcal{E}=\left\{\big(e_r^A(a),\,e_b^A(a),\,e_r^F(a)-e_b^F(a)\big): a\in\mathcal A\right\}$, so \begin{align*} h_{\mathcal{E}}(q)=\sup_{a\in\mathcal A}q^\intercal\left(e_r^A(a),\,e_b^A(a),\,e_r^F(a)-e_b^F(a)\right). \end{align*} Using the representation $\left(e_r^A(a),\,e_b^A(a),\,e_r^F(a)-e_b^F(a)\right) =\mathbb{E}^*[\boldsymbol{\theta}_0(X)]+\mathbb{E}^*\left[a(X)\Delta\boldsymbol{\theta}(X)\right]$, we obtain $q^\intercal\left(e_r^A(a),\,e_b^A(a),\,e_r^F(a)-e_b^F(a)\right)=\mathbb{E}^*\left[q^\intercal \boldsymbol{\theta}_0(X)\right]+\mathbb{E}^*\left[a(X)q^\intercal \Delta\boldsymbol{\theta}(X)\right]$. Hence, \begin{align*} h_{\mathcal{E}}(q)= \mathbb{E}^*\left[q^\intercal \boldsymbol{\theta}_0(X)\right] + \sup_{a\in\mathcal A}\mathbb{E}^*\left[a(X)q^\intercal \Delta\boldsymbol{\theta}(X)\right]. \end{align*} Because $a(X)\in[0,1]$ and $q^\intercal \Delta\boldsymbol{\theta}(X)$ is $\sigma(X)$-measurable, the pointwise maximizer is \begin{align*} a^q(X)\in \arg\max_{t\in[0,1]} t\,q^\intercal \Delta\boldsymbol{\theta}(X) = \begin{cases} 1, & q^\intercal \Delta\boldsymbol{\theta}(X)>0,\\ 0, & q^\intercal \Delta\boldsymbol{\theta}(X)<0, \end{cases} \end{align*} with arbitrary tie-breaking on $\{q^\intercal \Delta\boldsymbol{\theta}(X)=0\}$. Therefore, \begin{align*} \sup_{a\in\mathcal A}\mathbb{E}^*\left[a(X)q^\intercal \Delta\boldsymbol{\theta}(X)\right]= \mathbb{E}^*\left[\left(q^\intercal \Delta\boldsymbol{\theta}(X)\right)_+\right]. \end{align*} Next, note that $\mathds{1}\{q^\intercal \Delta\boldsymbol{\theta}(X)>0\}$ is $\sigma(X)$-measurable. Therefore, by iterated expectations, \begin{multline*} \mathbb{E}^*\left[q^\intercal \mathbf{L}_0+q^\intercal(\mathbf{L}_1-\mathbf{L}_0)\mathds{1}\{q^\intercal \Delta\boldsymbol{\theta}(X)>0\}\right] =\mathbb{E}^*\left[ q^\intercal \boldsymbol{\theta}_0(X)+q^\intercal \Delta\boldsymbol{\theta}(X)\mathds{1}\{q^\intercal \Delta\boldsymbol{\theta}(X)>0\} \right]. \end{multline*} This proves (ref). To prove (ref), fix $q\in \mathbb{S}^2$ and let $v\in{\mathbb R}^3$ be arbitrary; denote $u_+ \equiv \max\{u,0\}$. For $t\in{\mathbb R}$ such that $q+t v\neq 0$, using the support-function formula, we have \begin{align*} \tfrac{h_{\mathcal{E}}(q+t v)-h_{\mathcal{E}}(q)}{t}=\mathbb{E}^*\left[v^\intercal \boldsymbol{\theta}_0(X)+\tfrac{ \left((q+t v)^\intercal \Delta\boldsymbol{\theta}(X)\right)_+- \left(q^\intercal \Delta\boldsymbol{\theta}(X)\right)_+}{t} \right]. \end{align*} Hence it is enough to study the difference quotient of the map $u\mapsto u_+$. For each realization $X=x$ such that $q^\intercal \Delta\boldsymbol{\theta}(x)\neq 0$, the function $t\mapsto \left((q+t v)^\intercal \Delta\boldsymbol{\theta}(x)\right)_+$ is differentiable at $t=0$, with derivative $v^\intercal \Delta\boldsymbol{\theta}(x)\mathds{1}\{q^\intercal \Delta\boldsymbol{\theta}(x)>0\}$. Therefore, pointwise on the event $\{q^\intercal \Delta\boldsymbol{\theta}(X)\neq 0\}$, as $t\rightarrow 0$, \begin{align} \tfrac{\left((q+t v)^\intercal \Delta\boldsymbol{\theta}(X)\right)_+-\left(q^\intercal \Delta\boldsymbol{\theta}(X)\right)_+}{t}\rightarrow v^\intercal \Delta\boldsymbol{\theta}(X)\mathds{1}\{q^\intercal \Delta\boldsymbol{\theta}(X)>0\}. \end{align} Assumption (ref) implies $\mathbb{P}^*\bigl(q^\intercal \Delta\boldsymbol{\theta}(X)=0\bigr)=0$ for every $q\in \mathbb{S}^2$, because for every $\delta>0$, \begin{align*} \mathbb{P}^*\left(q^\intercal \Delta\boldsymbol{\theta}(X)=0\right) \le \sup_{\tilde q\in \mathbb{S}^2}\mathbb{P}^*\bigl(|\tilde q^\intercal \Delta\boldsymbol{\theta}(X)|\le \delta\bigr) \lesssim\delta^m \end{align*} and letting $\delta\downarrow 0$ yields the claim. Thus, the convergence in (ref) holds $\mathbb{P}^*$-almost surely. Next we obtain an integrable dominating function. Since $u\mapsto u_+$ is $1$-Lipschitz, \begin{align*} \left|\tfrac{\left((q+t v)^\intercal \Delta\boldsymbol{\theta}(X)\right)_+ - \left(q^\intercal \Delta\boldsymbol{\theta}(X)\right)_+}{t}\right| \le |v^\intercal \Delta\boldsymbol{\theta}(X)| \end{align*} for all $t\neq 0$. Therefore, \begin{align*} \left|v^\intercal \boldsymbol{\theta}_0(X)+\tfrac{\left((q+t v)^\intercal \Delta\boldsymbol{\theta}(X)\right)_+ - \left(q^\intercal \Delta\boldsymbol{\theta}(X)\right)_+}{t}\right| \le |v^\intercal \boldsymbol{\theta}_0(X)| + |v^\intercal \Delta\boldsymbol{\theta}(X)| \end{align*} By Assumption (ref), the random vectors $\boldsymbol{\theta}_0(X)$ and $\Delta\boldsymbol{\theta}(X)$ are integrable, so the right-hand side is integrable. Hence, by dominated convergence, \begin{align*} \lim_{t\rightarrow 0}\tfrac{h_{\mathcal{E}}(q+t v)-h_{\mathcal{E}}(q)}{t} &= v^\intercal\mathbb{E}^*\left[\boldsymbol{\theta}_0(X)+\Delta\boldsymbol{\theta}(X)\mathds{1}\{q^\intercal \Delta\boldsymbol{\theta}(X)>0\}\right]. \end{align*} Thus the directional derivative of $h_{\mathcal{E}}$ at $q$ in direction $v$ exists and is linear in $v$. It follows that $h_{\mathcal{E}}$ is differentiable at $q$, with gradient \begin{align*} \nabla_{q}h_{\mathcal{E}}(q)=\mathbb{E}^*\left[\boldsymbol{\theta}_0(X)+\Delta\boldsymbol{\theta}(X)\mathds{1}\{q^\intercal \Delta\boldsymbol{\theta}(X)>0\}\right]. \end{align*} Using again the law of iterated expectations yields (ref). Finally, by sch93, $\mathcal{E}$ is strictly convex because by (ref) its support set is a singleton in each direction $q\in\mathbb{S}^2$. \end{proof} \begin{proof}[Proof of Theorem (ref)] Recall $\widetilde{\mathbb{S}}^2\equiv\{q\in \mathbb{S}^2:q_1\ge 0,\ q_2\ge 0\}$ and that under Assumptions (ref)-(ref), $\mathcal{E}$ has a nonempty interior, is compact, and is strictly convex. Step 1: (ref) $\iff$ (ref). By definition, \begin{align*} \mathcal{C}(\varepsilon^*)=\left\{\varepsilon=(\varepsilon_1,\varepsilon_2,\varepsilon_3)\in{\mathbb R}^3:\varepsilon_1\le \varepsilon_1^*,\ \varepsilon_2\le \varepsilon_2^*,\ |\varepsilon_3|\le |\varepsilon_3^*|\right\}. \end{align*} Hence $\mathcal{E}\cap \mathcal{C}(\varepsilon^*)=\left\{\varepsilon\in \mathcal{E}:\varepsilon_1\le \varepsilon_1^*,\ \varepsilon_2\le \varepsilon_2^*,\ |\varepsilon_3|\le |\varepsilon_3^*|\right\}.$ Next, recall that $\varepsilon^*\in \mathcal{F}$ if and only if there does not exist $\varepsilon\in \mathcal{E}$ with $\varepsilon_1\le \varepsilon_1^*$, $\varepsilon_2\le \varepsilon_2^*$, and $|\varepsilon_3|\le |\varepsilon_3^*|$, with at least one strict inequality. Since $\varepsilon^*\in \mathcal{E}\cap \mathcal{C}(\varepsilon^*)$, this is equivalent to $\mathcal{E}\cap \mathcal{C}(\varepsilon^*)=\{\varepsilon^*\}$. Step 2: (ref) $\iff$ (ref). We begin showing $\ref{thm:FequivLLMO_E_single_inters_C}\Rightarrow \ref{thm:FequivLLMO_proj}$. Assume $\mathcal{E}\cap \mathcal{C}(\varepsilon^*)=\{\varepsilon^*\}$. We prove first that $|\varepsilon_3^*|=d((\varepsilon_1^*,\varepsilon_2^*))$. If not, then $d((\varepsilon_1^*,\varepsilon_2^*))<|\varepsilon_3^*|$, and by definition of $d$, there exists $\tilde \varepsilon_3\in{\mathbb R}$ such that $(\varepsilon_1^*,\varepsilon_2^*,\tilde \varepsilon_3)\in \mathcal{E}$ and $|\tilde \varepsilon_3|<|\varepsilon_3^*|$. But then $(\varepsilon_1^*,\varepsilon_2^*,\tilde \varepsilon_3)\in \mathcal{E}\cap \mathcal{C}(\varepsilon^*)$ and this point is distinct from $\varepsilon^*$, contradicting $\mathcal{E}\cap \mathcal{C}(\varepsilon^*)=\{\varepsilon^*\}$. Hence, $|\varepsilon_3^*|=d((\varepsilon_1^*,\varepsilon_2^*))$. We next show that $\pi(\varepsilon^*)=(\varepsilon_1^*,\varepsilon_2^*)\in \mathcal{F}^A$. Suppose not. Then there exists $(c_r,c_b)\in \mathcal{E}^A$ such that $c_r\le \varepsilon_1^*$, $c_b\le \varepsilon_2^*$, and $d((c_r,c_b))\le d((\varepsilon_1^*,\varepsilon_2^*))$, with at least one strict inequality. As $(c_r,c_b)\in \mathcal{E}^A$, by definition of $d$ there exists $\tilde \varepsilon_3\in{\mathbb R}$ such that $(c_r,c_b,\tilde \varepsilon_3)\in \mathcal{E}$ and $|\tilde \varepsilon_3|=d((c_r,c_b))$. Using the equality already proved, $d((\varepsilon_1^*,\varepsilon_2^*))=|\varepsilon_3^*|$, we obtain $c_r\le \varepsilon_1^*$, $c_b\le \varepsilon_2^*$, and $|\tilde \varepsilon_3|\le |\varepsilon_3^*|$, with at least one strict inequality. Thus $(c_r,c_b,\tilde \varepsilon_3)\in \mathcal{E}\cap \mathcal{C}(\varepsilon^*)$ with $(c_r,c_b,\tilde \varepsilon_3)\neq \varepsilon^*$, contradicting $\mathcal{E}\cap \mathcal{C}(\varepsilon^*)=\{\varepsilon^*\}$. Hence, $\pi(\varepsilon^*)\in \mathcal{F}^A$, proving $\ref{thm:FequivLLMO_E_single_inters_C}\Rightarrow \ref{thm:FequivLLMO_proj}$. We now prove $\ref{thm:FequivLLMO_proj}\Rightarrow \ref{thm:FequivLLMO_E_single_inters_C}$. Assume $\pi(\varepsilon^*)=(\varepsilon_1^*,\varepsilon_2^*)\in \mathcal{F}^A$ and $|\varepsilon_3^*|=d((\varepsilon_1^*,\varepsilon_2^*))$. Let $\varepsilon=(\varepsilon_1,\varepsilon_2,\varepsilon_3)\in \mathcal{E}\cap \mathcal{C}(\varepsilon^*)$. Then $\varepsilon_1\le \varepsilon_1^*$, $\varepsilon_2\le \varepsilon_2^*$, and $|\varepsilon_3|\le |\varepsilon_3^*|$. Since $\varepsilon\in \mathcal{E}$, we have \begin{align} d((\varepsilon_1,\varepsilon_2))\le |\varepsilon_3|\le |\varepsilon_3^*|=d((\varepsilon_1^*,\varepsilon_2^*)). \end{align} Thus, $(\varepsilon_1,\varepsilon_2)\in \mathcal{E}^A$ weakly improves on $(\varepsilon_1^*,\varepsilon_2^*)$ in both group-accuracy losses and in the fairness index $d$. As $(\varepsilon_1^*,\varepsilon_2^*)\in \mathcal{F}^A$, it follows that $(\varepsilon_1,\varepsilon_2)=(\varepsilon_1^*,\varepsilon_2^*)$ (else we would contradict $(\varepsilon_1^*,\varepsilon_2^*)\in \mathcal{F}^A$). Substituting back $(\varepsilon_1^*,\varepsilon_2^*)$ for $(\varepsilon_1,\varepsilon_2)$ in (ref) yields $|\varepsilon_3|=|\varepsilon_3^*|$. If $\varepsilon_3\neq \varepsilon_3^*$, then necessarily $\varepsilon_3=-\varepsilon_3^*$. If $\varepsilon_3^*=0$, this is impossible, so $\varepsilon_3=\varepsilon_3^*$. If $\varepsilon_3^*\neq 0$, by convexity of $\mathcal{E}$ the midpoint $\bar \varepsilon\equiv \tfrac{\varepsilon+\varepsilon^*}{2}=(\varepsilon_1^*,\varepsilon_2^*,0)$ belongs to $\mathcal{E}$. Therefore, $d((\varepsilon_1^*,\varepsilon_2^*))\le 0$, which implies $d((\varepsilon_1^*,\varepsilon_2^*))=0$. But then $|\varepsilon_3^*|=d((\varepsilon_1^*,\varepsilon_2^*))=0$, contradicting $\varepsilon_3^*\neq 0$. Hence $\varepsilon_3=\varepsilon_3^*$, so $\varepsilon=\varepsilon^*$. We conclude that $\mathcal{E}\cap \mathcal{C}(\varepsilon^*)=\{\varepsilon^*\}$, proving $\ref{thm:FequivLLMO_proj}\Rightarrow \ref{thm:FequivLLMO_E_single_inters_C}$. Step 3: (ref) $\Rightarrow$ (ref). Assume $\mathcal{E}\cap \mathcal{C}(\varepsilon^*)=\{\varepsilon^*\}$. As $\mathcal{E}$ is compact and convex with nonempty interior, and $\mathcal{C}(\varepsilon^*)$ is closed and convex, the two sets admit a common supporting hyperplane at their unique common point. Hence, there exists $q\neq 0$ and $\alpha\in{\mathbb R}$ such that \begin{align*} q^\intercal \varepsilon\ge \alpha\ge q^\intercal c \qquad \forall \varepsilon\in \mathcal{E},\ \forall c\in \mathcal{C}(\varepsilon^*). \end{align*} As $\varepsilon^*\in \mathcal{E}\cap \mathcal{C}(\varepsilon^*)$, necessarily $\alpha=q^\intercal \varepsilon^*$, hence $\inf_{\varepsilon\in \mathcal{E}}q^\intercal \varepsilon = \sup_{c\in \mathcal{C}(\varepsilon^*)}q^\intercal c$. Equivalently, $-h_{\mathcal{E}}(-q)=h_{\mathcal{C}(\varepsilon^*)}(q)$. Since $h_{\mathcal{C}(\varepsilon^*)}(q)<+\infty$, we must have $q_1\ge 0$ and $q_2\ge 0$, because $\mathcal{C}(\varepsilon^*)$ is unbounded in the negative first- and second-coordinate directions. Since $q\neq 0$, positive homogeneity of support functions implies that, after replacing $q$ by $q/\|q\|$, we may assume $q\in \widetilde{\mathbb{S}}^2$ while preserving the equality $h_{\mathcal{C}(\varepsilon^*)}(q)+h_{\mathcal{E}}(-q)=0$. Since $\varepsilon^*\in \mathcal{E}\cap \mathcal{C}(\varepsilon^*)$, for every $q\in \widetilde{\mathbb{S}}^2$, $h_{\mathcal{C}(\varepsilon^*)}(q)\ge q^\intercal \varepsilon^*$ and $h_{\mathcal{E}}(-q)\ge -q^\intercal \varepsilon^*$, and hence $h_{\mathcal{C}(\varepsilon^*)}(q)+h_{\mathcal{E}}(-q)\ge 0$. Therefore, $\min_{q\in \widetilde{\mathbb{S}}^2} \left(h_{\mathcal{C}(\varepsilon^*)}(q)+h_{\mathcal{E}}(-q)\right)=0$, proving $\ref{thm:FequivLLMO_E_single_inters_C}\Rightarrow \ref{thm:FequivLLMO_support}$. Step 4: (ref) $\Rightarrow$ (ref). Assume $\min_{q\in \widetilde{\mathbb{S}}^2}\left(h_{\mathcal{C}(\varepsilon^*)}(q)+h_{\mathcal{E}}(-q)\right)=0$. As $\widetilde{\mathbb{S}}^2$ is compact, $\mathcal{E}$ is compact, and $h_{\mathcal{C}(\varepsilon^*)}$ is finite and continuous on $\widetilde{\mathbb{S}}^2$, the minimum is attained at some $\bar q\in \widetilde{\mathbb{S}}^2$. Thus, $\sup_{c\in \mathcal{C}(\varepsilon^*)}\bar q^\intercal c=\inf_{\varepsilon\in \mathcal{E}}\bar q^\intercal \varepsilon$. Let $\tilde \varepsilon\in \mathcal{E}\cap \mathcal{C}(\varepsilon^*)$. Then \begin{align*} \bar q^\intercal \tilde \varepsilon \le \sup_{c\in \mathcal{C}(\varepsilon^*)}\bar q^\intercal c= \inf_{\varepsilon\in \mathcal{E}}\bar q^\intercal \varepsilon \le \bar q^\intercal \tilde \varepsilon. \end{align*} Hence, every $\tilde \varepsilon\in \mathcal{E}\cap \mathcal{C}(\varepsilon^*)$ satisfies $\bar q^\intercal \tilde \varepsilon=\inf_{\varepsilon\in \mathcal{E}}\bar q^\intercal \varepsilon$. Therefore, \begin{align*} \mathcal{E}\cap \mathcal{C}(\varepsilon^*) \subset \left\{ \varepsilon\in \mathcal{E}:\bar q^\intercal \varepsilon=\inf_{\tilde \varepsilon\in \mathcal{E}}\bar q^\intercal \tilde \varepsilon \right\}, \end{align*} that is, the intersection is contained in the support set of $\mathcal{E}$ in direction $-\bar q$. Since $\mathcal{E}$ is strictly convex, every support set of $\mathcal{E}$ is a singleton. Because $\varepsilon^*\in \mathcal{E}\cap \mathcal{C}(\varepsilon^*)$, the above inclusion implies $\mathcal{E}\cap \mathcal{C}(\varepsilon^*)=\{\varepsilon^*\}$, proving $\ref{thm:FequivLLMO_support}\Rightarrow \ref{thm:FequivLLMO_E_single_inters_C}$. Combining Steps 1-4, we obtain $\ref{thm:FequivLLMO_e_in_sF}\iff\ref{thm:FequivLLMO_E_single_inters_C}\iff\ref{thm:FequivLLMO_support}\iff\ref{thm:FequivLLMO_proj}$. \end{proof} \begin{proof}[Proof of Theorem (ref)] We divide the proof of this theorem into four steps. The first two steps prove the first part of the theorem using Lemma (ref), while the last two steps additionally use the injectivity of the support map $\mathcal{S}_\mathcal{E}$ implied by the non-kink condition. Step 1: Suppose that for any $\varepsilon^* \in \mathcal{F}$ there exists $q^* = q(\alpha) \in \widetilde{\mathbb{S}}^2$ (defined in the statement of the theorem) such that $\varepsilon^*= \mathcal{S}_\mathcal{E}(-q^*)$ and $q^*_3 = 0$. By part 1 of Lemma (ref), this implies that $\varepsilon^* \in \mathcal{PF}$; therefore, $\mathcal{F} \subset \mathcal{PF}$. Since $\mathcal{PF} \subset \mathcal{F}$ by Lemma (ref), we conclude $\mathcal{PF}= \mathcal{F}$. Step 2: Suppose that $\mathcal{PF}= \mathcal{F}$. Since $\varepsilon^* \in \mathcal{F}$ implies $\varepsilon^* \in \mathcal{PF}$, we conclude that there exist $q^* \in \widetilde{\mathbb{S}}^2$ such that $\varepsilon^*= \mathcal{S}_\mathcal{E}(-q^*)$ and $q^*_3 = 0$ by part 1 of Lemma (ref). We conclude the proof by setting $q(\alpha) = q^*$. Recall that Assumption (ref) (no kinks) and the argument in bon:mag:mau12 imply the injectivity of the support map $\mathcal{S}_\mathcal{E}$. \textbf{Step 3:} Suppose that $\mathcal{PF} = \mathcal{F}$. Note that the injectivity of $\mathcal{S}_\mathcal{E}$ implies that any $\varepsilon^* \in \mathcal{F}$ has a unique support direction $q$ whose third coordinate is $0$ due to $\mathcal{PF} = \mathcal{F}$ and part 1 of Lemma (ref). Now, suppose that $\mathcal{PF} \subset \{ \varepsilon_3 = 0 \}$ is false, that is, there exist $\varepsilon^* \in \mathcal{PF}$ such that $\varepsilon_3^* \neq 0$. In what follows, we will find $\varepsilon^t \in \mathcal{F}$ such that $\varepsilon^t = \mathcal{S}_\mathcal{E}(-q^t)$ and $q^t_3 \neq 0$, which will be a contradiction to $\mathcal{PF} = \mathcal{F}$ and the injectivity of $\mathcal{S}_\mathcal{E}$. By part 1 of Lemma (ref), there exist $q = (q_1,q_2,0) \in \widetilde{\mathbb{S}}^2 $ such that $\varepsilon^* = \mathcal{S}_\mathcal{E}(-q)$. Let $s\equiv\mathrm{sgn}(\varepsilon_3^*)\in\{-1,1\}$ and for $t>0$ define \begin{align*} q^t=\tfrac{(q_1,q_2,st)} {\sqrt{q_1^2+ q_2^2+t^2}}. \end{align*} Then $q^t\to q$ as $t\downarrow 0$. By continuity of $\mathcal{S}_\mathcal{E}$, \begin{align*} \varepsilon^t=\mathcal{S}_\mathcal{E}(-q^t)\to \mathcal{S}_\mathcal{E}(-q) = \varepsilon^*. \end{align*} Hence, for sufficiently small $t>0$, $\mathrm{sgn}(\varepsilon^t_3)=s=\mathrm{sgn}(q^t_3)$, so $q^t_3\varepsilon^t_3>0$. By part 2 of Lemma (ref), this implies $\varepsilon^t\in\mathcal{F}$ such that $\varepsilon^t = \mathcal{S}_\mathcal{E}(-q^t)$ and $q^t_3 \neq 0$, which is a contradiction. \textbf{Step 4:} Suppose that $\mathcal{PF} \subset \{ \varepsilon_3 = 0 \}$. Note that $\mathcal{PF} \subset \{ \varepsilon_3 = 0 \}$ implies that the third coordinate of both $R=\mathcal{S}_\mathcal{E}(-u_1)$ and $B=\mathcal{S}_\mathcal{E}(-u_2)$ is zero, where $u_1=(1,0,0)^\intercal$ and $u_2 = (0,1,0)^\intercal$; therefore, $u_3^\intercal R = 0$ and $u_3^\intercal B = 0$ where $u_3 = (0,0,1)$. Since $R \neq B$, it follows that $\min_{\varepsilon \in \mathcal{E}} u_3^\intercal \varepsilon <0$ due to strict convexity of $\mathcal{E}$. This implies that $\mathcal{S}_\mathcal{E}(-u_3) \notin \mathcal{F}$ by part 2 of Lemma (ref). Similarly, we can conclude $\mathcal{S}_\mathcal{E}(u_3) \notin \mathcal{F}$. Since $q \notin \{u_3,-u_3\}$, it follows that $q_1 + q_2 >0$. Let $\varepsilon^* \in \mathcal{F}$. By part 2 of Lemma (ref), there exist $q \in \widetilde{\mathbb{S}}^2 $ such that $\varepsilon^* = \mathcal{S}_\mathcal{E}(-q)$ and $\varepsilon^*_3 q_3 \ge 0$. Let $\tilde{q} = (q_1, q_2,0)/\sqrt{q_1^2 + q_2^2}$ and $\rho = \mathcal{S}_\mathcal{E}(-\tilde{q} )$. Consider the following derivations: \begin{align*} q^\intercal \rho &\overset{(1)}{=} q_1\rho_1+q_2\rho_2 \overset{(2)}{\le} q_1\varepsilon_1^*+q_2\varepsilon_2^* \overset{(3)}{\le} q_1\varepsilon_1^*+q_2\varepsilon_2^*+q_3\varepsilon_3^* \overset{(4)}{\le} q^\intercal \rho, \end{align*} where (1) holds since $\rho \in \mathcal{PF}$ by part 1 of Lemma (ref) and $\rho_3 = 0$ due to $\mathcal{PF} \subset \{ \varepsilon_3 = 0 \}$, (2) holds since $\rho$ minimizes $q_1\varepsilon_1+q_2\varepsilon_2$ over $\mathcal{E}$, (3) holds since $\varepsilon^* \in \mathcal{F}$ and $\varepsilon^*_3 q_3 \ge 0$ by part 2 of Lemma (ref), and (4) holds since $\varepsilon^*$ minimizes $q_1\varepsilon_1+q_2\varepsilon_2 + q_3 \varepsilon_3$ over $\mathcal{E}$. Therefore, $\rho$ also minimize $q_1\varepsilon_1+q_2\varepsilon_2 + q_3 \varepsilon_3$ over $\mathcal{E}$. As a result, due to strict convexity of $\mathcal{E}$, we have that $\rho = \varepsilon^*$, which implies that $\varepsilon^* \in \mathcal{PF}$. Therefore, $\mathcal{F} \subset \mathcal{PF}$ which implies $\mathcal{F} = \mathcal{PF}$ since $\mathcal{PF} \subset \mathcal{F}$ due to Lemma (ref). \end{proof} \subsection{Proofs for Section (ref)} \subsubsection{Proofs for Section (ref)} For $\lambda\in\Lambda$, let $\boldsymbol{\theta}_d(x;\lambda)$ denote the completion-dependent analogue of $\boldsymbol{\theta}_d(x)$. Under the special losses considered here, \begin{align*} L_0^{g,A} &= \tfrac{Y_i^* \mathds{1}\{G=g\}}{\mu_g}, L_1^{g,A} = \tfrac{(1-Y_i^*) \mathds{1}\{G=g\}}{\mu_g} \\ L_0^{r,F} - L_0^{b,F} &= 0, L_1^{r,F} - L_1^{b,F} = \tfrac{\mathds{1}\{G=r\}}{\mu_r} - \tfrac{\mathds{1}\{G=b\}}{\mu_b} \end{align*} The next lemma gives the $\lambda$-dependent representations used in Section (ref). \begin{lemma} \textit{For every $\lambda\in\Lambda$ and every $x\in\mathcal{X}$, \begin{align} \boldsymbol{\theta}_0(x;\lambda)&=A_0(x)+B(x)\lambda(x), & \boldsymbol{\theta}_1(x;\lambda)&=A_1(x)-B(x)\lambda(x). \end{align} Consequently, \begin{align} h_\mathcal{E}(q;\lambda) &= \mathbb{E}\!\left[ \max\left\{ q^\intercal\!\left(A_0(X)+B(X)\lambda(X)\right), q^\intercal\!\left(A_1(X)-B(X)\lambda(X)\right) \right\} \right], \\ h_{\mathcal{C}(\varepsilon^*(a;\lambda))}(q) &= C(q;a)+\mathbb{E}\!\left[(1-2a(X))q^\intercal B(X)\lambda(X)\right], \qquad q\in\widetilde{\mathbb{S}}^2, \end{align} where \begin{equation} C(q;a) \equiv \sum_{j=1}^2 q_j\mathbb{E}\!\left[a(X)A_{1,j}(X)+(1-a(X))A_{0,j}(X)\right] +|q_3|\,\left|\mathbb{E}\!\left[a(X)A_{1,3}(X)\right]\right|. \end{equation}} \end{lemma} \begin{proof} Fix $g\in\{r,b\}$. Using $Y=\ZY^*$ and the definition of $\lambda_g(x)$, \begin{align*} \theta_0^{g,A}(x;\lambda) &= \mathbb{E}^*\!\left[\left.\tfrac{Y^*\mathds{1}\{G=g\}}{\mu_g}\right|X=x\right] \\ &= \mathbb{E}\!\left[\left.\tfrac{\ZY\mathds{1}\{G=g\}}{\mu_g}\right|X=x\right] + \mathbb{E}\!\left[\left.\tfrac{(1-Z)Y^*\mathds{1}\{G=g\}}{\mu_g}\right|X=x\right] \\ &= \mathbb{E}\!\left[\left.\tfrac{\ZY\mathds{1}\{G=g\}}{\mu_g}\right|X=x\right] + \lambda_g(x)\,\mathbb{E}\!\left[\left.\tfrac{(1-Z)\mathds{1}\{G=g\}}{\mu_g}\right|X=x\right], \end{align*} which is the $g$th accuracy coordinate of $A_0(x)+B(x)\lambda(x)$. Likewise, \begin{align*} \theta_1^{g,A}(x;\lambda) &= \mathbb{E}^*\!\left[\left.\tfrac{(1-Y^*)\mathds{1}\{G=g\}}{\mu_g}\right|X=x\right] \\ &= \mathbb{E}\!\left[\left.\tfrac{(Z(1-Y)+(1-Z))\mathds{1}\{G=g\}}{\mu_g}\right|X=x\right] - \lambda_g(x)\,\mathbb{E}\!\left[\left.\tfrac{(1-Z)\mathds{1}\{G=g\}}{\mu_g}\right|X=x\right], \end{align*} which is the $g$th accuracy coordinate of $A_1(x)-B(x)\lambda(x)$. For the fairness coordinate, \[ \theta_0^F(x;\lambda)=0, \qquad \theta_1^F(x;\lambda)=\mathbb{E}\!\left[\left.\tfrac{\mathds{1}\{G=r\}}{\mu_r}-\tfrac{\mathds{1}\{G=b\}}{\mu_b}\right|X=x\right], \] which does not depend on $\lambda$. This proves (ref). Equation (ref) now follows directly from Theorem (ref) applied to the $\lambda$-dependent pair $(\boldsymbol{\theta}_0(\cdot;\lambda),\boldsymbol{\theta}_1(\cdot;\lambda))$. Next fix an algorithm $a$. By construction, \[ \varepsilon^*(a;\lambda)=\mathbb{E}\!\left[a(X)\boldsymbol{\theta}_1(X;\lambda)+(1-a(X))\boldsymbol{\theta}_0(X;\lambda)\right]. \] For $j\in\{1,2\}$, using (ref), \[ \varepsilon_j^*(a;\lambda) = \mathbb{E}\!\left[a(X)A_{1,j}(X)+(1-a(X))A_{0,j}(X)\right] +\mathbb{E}\!\left[(1-2a(X))B_{jj}(X)\lambda_j(X)\right]. \] For the third coordinate, $\varepsilon_3^*(a) = \mathbb{E}\!\left[a(X)A_{1,3}(X)\right]$, so it is point identified. Equation (ref) then gives, for $q\in\widetilde{\mathbb{S}}^2$, \begin{align*} h_{\mathcal{C}(\varepsilon^*(a;\lambda))}(q) &= q_1\varepsilon_1^*(a;\lambda)+q_2\varepsilon_2^*(a;\lambda)+|q_3|\,|\varepsilon_3^*(a)| \\ &= C(q;a)+\mathbb{E}\!\left[(1-2a(X))q^\intercal B(X)\lambda(X)\right], \end{align*} which is (ref). \end{proof} \begin{proof}[Proof of Theorem (ref)] By Assumption (ref), all distributions consistent with the data also satisfy the conditions of Theorem (ref); hence, algorithm $a^*$ yields $\varepsilon^* \in \mathcal{F}$ if and only if \begin{equation*} \min_{q \in \tilde\mathbb{S}^2} \left[h_{\mathcal{C}(\varepsilon^*)}(q) + h_{\mathcal{E}}(-q)\right] = 0. \end{equation*} Given the definition of $J_d(\lambda;q,a)$, $d=0,1$, in Theorem (ref), this is equivalent to \begin{equation} \min_{q \in \tilde\mathbb{S}^2} \mathbb{E}^*\left[ \max \left \{ J_1(\lambda^*(X);q,a^*),J_0(\lambda^*(X);q,a^*) \right\} \right] =0 \end{equation} Since $\lambda^*(X)$ is not identified by the data, we say that $a^*$ is on the fairness-accuracy frontier for some distribution consistent with the data if we can find a $\lambda^{(a^*)}(\cdot) \in \Lambda$ such that (ref) holds with $\lambda^*(X) = \lambda^{(a^*)}(X)$. We can write the existence equivalently as follows: \begin{equation} \min_{\lambda(\cdot) \in \Lambda} \min_{q \in \tilde\mathbb{S}^2} \mathbb{E}\left[ \max \left \{ J_1(\lambda(X);q,a^*),J_0(\lambda(X);q,a^*) \right\} \right] =0 \end{equation} To simplify notation we use $J_d(\lambda)$ instead of $J_d(\lambda(X);q,a^*)$. Since $J_d(\lambda)$ is linear in $\lambda$, we write $J_d(\lambda) = \tilde{A}_d + \tilde{B}_d^\intercal \lambda$, where $\tilde{B}_d=(\tilde{B}_d^r,\tilde{B}_d^b)^\intercal$. We use the previous notation to write the left-hand side of (ref) as follows \begin{align} &\overset{(1)}{=} \min_{q \in \tilde\mathbb{S}^2} \min_{\lambda(\cdot) \in \Lambda} \mathbb{E}^*\left[ \max_{t \in [0,1]} \left \{ t J_1(\lambda) + (1-t) J_0(\lambda) \right\} \right] \notag \\ &\overset{(2)}{=} \min_{q \in \tilde\mathbb{S}^2} \mathbb{E}\left[ \min_{\lambda(\cdot) \in \Lambda} \max_{t \in [0,1]} \left \{ t J_1(\lambda) + (1-t) J_0(\lambda) \right\} \right] \notag \\ &\overset{(3)}{=} \min_{q \in \tilde\mathbb{S}^2} \mathbb{E}\left[ \max_{t \in [0,1]} \min_{\lambda(\cdot) \in \Lambda} \left \{ t \left( \tilde{A}_1 + \tilde{B}_1^\intercal \lambda \right) + (1-t) \left( \tilde{A}_0 + \tilde{B}_0^\intercal \lambda \right) \right\} \right] \notag \\ &\overset{(4)}{=} \min_{q \in \tilde\mathbb{S}^2} \mathbb{E}\left[ \max_{t \in [0,1]} \left \{ t \tilde{A}_1 + (1-t)\tilde{A}_0 + \min\{0, t \tilde{B}_1^r + (1-t) \tilde{B}_0^r\} + \min\{0, t \tilde{B}_1^b + (1-t) \tilde{B}_0^b\} \right\} \right] \notag\\ &\overset{(5)}{=} \min_{q \in \tilde\mathbb{S}^2} \mathbb{E}\left[ \max \left \{ \tilde{A}_0 + \min\{0, \tilde{B}_0^r\} + \min\{0, \tilde{B}_0^b \} , a^*(X) \tilde{A}_1 + (1-a^*(X))\tilde{A}_0, \tilde{A}_1 \right\} \right] \end{align} where (1) holds because $\max\{a,b\}=\max_{t \in [0,1]} \{t a + (1-t)b\}$, (2) holds by the measurability selection theorem, (3) holds by the minimax theorem since the objective function is linear in $t$ and $\lambda$ and both domains are convex, (4) by solving in $\lambda(X) \in [0,1]^2$ conditional on $X$, and (5) holds by claims 1--3 below. Finally, (ref) is sufficient to conclude the proof of the proposition since $\tilde{A}_0 + \min\{0, \tilde{B}_0^r\} + \min\{0, \tilde{B}_0^b \}= J_0(\lambda_1;q,a^*)$, $\tilde{A}_0 = J_0(\lambda_0;q,a^*)$, $\tilde{A}_1 = J_1(\lambda_0;q,a^*)$, and $a^*(X) \in \{0,1\}$ due to Corollary (ref). \textbf{Claim 1}: $t \tilde{B}_1^g + (1-t) \tilde{B}_0^g \le 0 $ if $ t \le t^* = a^*(X)$. Recall that $\tilde{B}_d = (\tilde{B}_d^r,\tilde{B}_d^b)^\intercal$ and \begin{align*} \tilde{B}_0 &= -2a^*(X) q^\intercal B(X) \\ \tilde{B}_1 &= 2(1-a^*(X)) q^\intercal B(X) . \end{align*} Therefore, $t \tilde{B}_1 + (1-t) \tilde{B}_0 = 2\left(t - a^*(X) \right) q^\intercal B(X) $, which implies the claim since $q_1 \ge 0$, $q_2 \ge 0$, and $B(X)$ defined in (ref) is nonnegative. \textbf{Claim 2}: $t \tilde{B}_1^g + (1-t) \tilde{B}_0^g \ge 0 $ if $ t \ge t^* = a^*(X)$. Similar to claim 1. \textbf{Claim 3:} $J(t) $ is maximized at $t=0$, $a^*(X)$, and $1$, where $$J(t) \equiv t \tilde{A}_1 + (1-t)\tilde{A}_0 + \min\{0, t \tilde{B}_1^r + (1-t) \tilde{B}_0^r\} + \min\{0, t \tilde{B}_1^b + (1-t) \tilde{B}_0^b\} ~.$$ Furthermore, $J(a^*(X)) = a^*(X) \tilde{A}_1 + (1-a^*(X))\tilde{A}_0$ and $J(1) = \tilde{A}_1$. The proof of this claim follows by linearity on $t$ and claims 1 and 2. \end{proof} \begin{proof}[Proof of Theorem (ref)] By definition, $\mathcal{F}^\exists=\bigcup_{\lambda\in\Lambda}\mathcal{F}(\lambda)$. First let $\varepsilon\in\mathcal{F}^\exists$. Then $\varepsilon\in\mathcal{F}(\lambda)$ for some $\lambda\in\Lambda$. By Lemma (ref), there exists $q\in\widetilde{\mathbb{S}}^2$ such that $\varepsilon=\mathcal{S}_{\mathcal{E}(\lambda)}(-q)$ and $q_3\varepsilon_3\ge0$. Since $\mathcal{S}_{\mathcal{E}(\lambda)}(-q)$ is generated by a rule minimizing $q^\intercal\boldsymbol{\theta}_d(x;\lambda)$ pointwise, the corresponding expected loss vector belongs to $\mathcal{M}_q$. Hence \begin{align*} \varepsilon\in\mathcal{M}_q\cap\{\eta\in{\mathbb R}^3:q_3\eta_3\ge0\}. \end{align*} Conversely, suppose $\varepsilon\in\mathcal{M}_q\cap\{\eta\in{\mathbb R}^3:q_3\eta_3\ge0\}$ for some $q\in\widetilde{\mathbb{S}}^2$. By definition of $\mathcal{M}_q$, there exist $\lambda\in\Lambda$ and a $q$-optimal rule $a$ such that \begin{align*} \varepsilon=\mathbb{E}\left[(1-a(X))\boldsymbol{\theta}_0(X;\lambda)+a(X)\boldsymbol{\theta}_1(X;\lambda)\right]. \end{align*} This vector minimizes $q^\intercal\eta$ over $\eta\in\mathcal{E}(\lambda)$, so $\varepsilon=\mathcal{S}_{\mathcal{E}(\lambda)}(-q)$. Since $q_3\varepsilon_3\ge0$, the representation of $\mathcal{F}(\lambda)$ due to Lemma (ref) implies $\varepsilon\in\mathcal{F}(\lambda)$. Therefore, $\varepsilon\in\mathcal{F}^\exists$, establishing (ref). Using the result in (ref), suppose $\varepsilon\in\mathcal{F}^\exists$. Then there exists $q\in\widetilde{\mathbb{S}}^2$ such that $q_3\varepsilon_3\ge0$ and $\varepsilon\in\mathcal{M}_q$. By Lemma (ref), $\mathcal{M}_q$ is compact and convex, hence representing set membership through support function dominance roc97, we have \begin{align*} \rho_q(\varepsilon)\equiv\sup_{v\in\mathbb{S}^2} \{v^\intercal\varepsilon-h_{\mathcal{M}_q}(v)\}\le0,\quad\text{hence}\quad[\rho_q(\varepsilon)]_+=0. \end{align*} Since $q$ is feasible in the infimum defining $T_{\mathrm{env}}(\varepsilon)$ in (ref), $0\le T_{\mathrm{env}}(\varepsilon)\le[\rho_q(\varepsilon)]_+=0$. Therefore $T_{\mathrm{env}}(\varepsilon)=0$. Conversely, suppose $T_{\mathrm{env}}(\varepsilon)=0$. By Lemma (ref), the infimum is attained. Hence there exists $q^*\in \{q\in\widetilde{\mathbb{S}}^2:q_3\varepsilon_3\ge0\}$ such that $[\rho_{q^*}(\varepsilon)]_+=0$. Thus, $\rho_{q^*}(\varepsilon)\le0$. Using again the support function representation of set membership for the compact convex set $\mathcal{M}_{q^*}$, $\varepsilon\in\mathcal{M}_{q^*}$. Since $q^*\in \{q\in\widetilde{\mathbb{S}}^2:q_3\varepsilon_3\ge0\}$, we also have $q_3^*\varepsilon_3\ge0$. Therefore, \begin{align*} \varepsilon\in\mathcal{M}_{q^*}\cap\{\eta\in{\mathbb R}^3:q_3^*\eta_3\ge0\}\subseteq\mathcal{F}^\exists, \end{align*} where the last inclusion follows from (ref). This proves $\varepsilon\in\mathcal{F}^\exists\iff T_{\mathrm{env}}(\varepsilon)=0$. \end{proof} \subsubsection{Proofs for Section (ref)} \begin{proof}[Proof of Lemma (ref)] Since $\pi(X)\in(0,1)$ a.s., the ratio $Z\mathbf{L}_d/\pi(X)$ is well-defined. Applying the law of iterated expectations, \begin{align*} \mathbb{E}\left[\left.\tfrac{Z\mathbf{L}_d}{\pi(X)}\right|X\right]&= \mathbb{E}\left[\left. \mathbb{E}\left[\left.\tfrac{Z\mathbf{L}_d}{\pi(X)}\right|X,Y^*,G\right]\right|X\right]= \mathbb{E}\left[\left.\tfrac{\mathbf{L}_d}{\pi(X)}\mathbb{E}[Z|X,Y^*,G]\right| X\right]. \end{align*} By Assumption (ref), $\mathbb{E}[Z| X,Y^*,G]=\pi(X)$ a.s., and hence $\tfrac{\mathbf{L}_d}{\pi(X)}\mathbb{E}[Z|X,Y^*,G]=\mathbf{L}_d$, a.s., and the characterization of $\boldsymbol{\theta}_d(X)$ in (ref) follows. Repeating the same argument and using (ref) yields the characterizations of $h_\mathcal{E}(q)$ and $e_g^\iota(a)$ in (ref)-(ref). \end{proof} \subsection{Proofs for Section (ref)} Let $\{\omega_i\}_{i=1}^n$ be random positive weights independent of the data. Define \begin{align} \widehat{h}_\mathcal{E}^\omega(q; \widehat{\mathbf{L}}^\omega, \widehat{\boldsymbol{\eta}}) \equiv \tfrac{1}{n} \sum_{i =1}^n \left( \tfrac{\omega_i}{\overline{\omega}} \right) \zeta_i(q; \widehat{\mathbf{L}}^\omega,\widehat{\boldsymbol{\eta}}) , \end{align} where $\overline{\omega} = n^{-1} \sum_{i=1}^n \omega_i$ and $\widehat{\mathbf{L}}^\omega$ is defined as in (ref)--(ref) with $\mu_g$ replaced by $\widehat{\mu}_g^\omega = n^{-1} \sum_{i=1}^n \left( \tfrac{\omega_i}{\overline{\omega}} \right) \mathds{1}\{G_i=g\}$. Note that $\widehat{h}_\mathcal{E}(q; \widehat{\mathbf{L}}, \widehat{\boldsymbol{\eta}}) = \widehat{h}_\mathcal{E}^\omega(q; \widehat{\mathbf{L}}^\omega, \widehat{\boldsymbol{\eta}})$ by taking $\omega_i = 1 $, $\forall ~i$. The next result presents a general version of Theorem (ref). It presents the limiting distribution for the class of weighted DML estimator $\widehat{h}_\mathcal{E}^\omega$ of the support function $h_\mathcal{E}$. \begin{theorem} \textit{Let Assumptions (ref)--(ref) hold and let $\{\omega_i\}_{i=1}^n$ be random positive weights independent of the data such that $\mathbb{E}[\omega_i] = 1$ and $\mathbb{E}[\omega_i^2] < C_\omega$. Then, $$ \sqrt{n}\left( \widehat{h}_\mathcal{E}^\omega(q; \widehat{\mathbf{L}}^\omega, \widehat{\boldsymbol{\eta}}) - h_\mathcal{E}(q) \right) \Rightarrow \mathbb{G}[\omega_i \{\zeta_i^{*}(q;\mathbf{L},\boldsymbol{\eta})-h_\mathcal{E}(q)\}] \quad \text{in} \quad \ell^{\infty}(\mathbb{S}^2)~,$$ where $\zeta_i^{*}(q;\mathbf{L},\boldsymbol{\eta})$ and $\widehat{h}_\mathcal{E}^\omega(q; \widehat{\mathbf{L}}^\omega, \widehat{\boldsymbol{\eta}}) $ are defined in (ref) and (ref), respectively. } \end{theorem} \begin{proof} The proof has two parts. We first show in part 1 that \begin{equation} \sup_{q \in \mathbb{S}^2 } \left| \sqrt{n}(\widehat{h}_\mathcal{E}^\omega(q,\widehat{\mathbf{L}}^\omega; \widehat{\boldsymbol{\eta}}) - h_\mathcal{E}(q) ) - \mathbb{G}_n[ \omega_i \{\zeta_i^{*}(q;\mathbf{L},\boldsymbol{\eta})-h_\mathcal{E}(q)\}] \right| = o_p(1) . \end{equation} We then use van:wel13 to conclude in part 2 that \begin{equation} \mathbb{G}_n[\omega_i \{\zeta_i^{*}(q;\mathbf{L},\boldsymbol{\eta})-h_\mathcal{E}(q)\}] \Rightarrow \mathbb{G}[\omega_i \{\zeta_i^{*}(q;\mathbf{L},\boldsymbol{\eta})-h_\mathcal{E}(q)\}] \quad \text{in} \quad \ell^\infty(\mathbb{S}^2) . \end{equation} \textbf{Part 1:} Our goal is to establish (ref). Define $$ \widehat{h}_\mathcal{E}^{\omega,*}(q; {\mathbf{L}}, {\boldsymbol{\eta}}) = \tfrac{1}{n} \sum_{i =1}^n \left( \tfrac{\omega_i}{\overline{\omega}} \right) \zeta_i^*(q;\mathbf{L},{\boldsymbol{\eta}})~.$$ By Lemma (ref), we have $$ \sup_{q \in \mathbb{S}^2 } | \sqrt{n}(\widehat{h}_\mathcal{E}^\omega(q,\widehat{\mathbf{L}}^\omega; \widehat{\boldsymbol{\eta}}) - \widehat{h}_\mathcal{E}^{\omega,*}(q; {\mathbf{L}}, {\boldsymbol{\eta}}) ) | = o_p(1)~. $$ This implies that (ref) holds if the equation below holds, \begin{equation} \sup_{q \in \mathbb{S}^2 } | \sqrt{n}(\widehat{h}_\mathcal{E}^{\omega,*}(q;{\mathbf{L}}, {\boldsymbol{\eta}}) - h_\mathcal{E}(q) ) - \mathbb{G}_n[ \omega_i \{\zeta_i^{*}(q;\mathbf{L},\boldsymbol{\eta})-h_\mathcal{E}(q)\}] | = o_p(1) . \end{equation} Now, we rewrite the left-hand size of (ref) using the identity $ \mathbb{G}_n[ \omega_i \{\zeta_i^{*}(q;\mathbf{L},\boldsymbol{\eta})-h_\mathcal{E}(q)\}] = \overline{\omega} \sqrt{n}(\widehat{h}_\mathcal{E}^{\omega,*}(q,{\mathbf{L}}; {\boldsymbol{\eta}}) - h_\mathcal{E}(q) ) $. That is $$ |1 - \tfrac{1}{\overline{\omega}}| \sup_{q \in \mathbb{S}^2 } | \mathbb{G}_n[ \omega_i \{\zeta_i^{*}(q;\mathbf{L},\boldsymbol{\eta})-h_\mathcal{E}(q)\}] |~. $$ Since $ |1 - \tfrac{1}{\overline{\omega}}| = o_p(1)$ by the definition of the weights $\omega_i$ and the law of large numbers. Therefore, (ref) holds if $ \sup_{q \in \mathbb{S}^2 } | \mathbb{G}_n[ \omega_i \{\zeta_i^{*}(q;\mathbf{L},\boldsymbol{\eta})-h_\mathcal{E}(q)\}] | = O_p(1)$, which holds by (ref) and the Continuous Mapping Theorem. Note that we can use (ref) since the proof of part 2 is independent of the proof of part 1. This completes the proof of part 1. \textbf{Part 2}: Our goal is to establish (ref). Consider the random vector \begin{align*} W_i = (\mathbf{L}_{0,i}, \Delta \mathbf{L}_{i}, Z_i, \pi(X_i),\Delta \boldsymbol{\theta}(X_i), \boldsymbol{\theta}_0(X_i),\omega_i) \in \mathcal{W} \end{align*} and the class of functions $\mathcal{Q} \equiv \left\{ f: \mathcal{W} \to {\mathbb R} : f(W_i) = \omega_i \{\zeta_i^{*}(q;\mathbf{L},\boldsymbol{\eta})-h_\mathcal{E}(q)\},q \in \mathbb{S}^2 \right\}$. Here, we are using the true value of the nuisance function $\boldsymbol{\eta}$---which is fixed---and treating $\boldsymbol{\eta}(X_i)$ as a random vector. In order to use van:wel13 to complete the proof of part 2, we need to verify that (i) an envelope function of the class $\mathcal{Q}$ has finite second-moment, and (ii) the uniform entropy bound in van:wel13 holds. To verify (i), we note that the minimum envelope $\sup_{q \in \mathbb{S}^2} | \omega_i \{\zeta_i^{*}(q;\mathbf{L},\boldsymbol{\eta})-h_\mathcal{E}(q)\}|$ has finite second-moment due to Assumptions (ref) and (ref). To verify (ii), we rely on and94 that states that the class of functions of type I and their mixing (by addition and product operation) verify the uniform entropy bound. We conclude by noting that $\mathcal{Q}$ is of type I, since any function in $\mathcal{Q}$ can be written as the sum of (i) linear functions in $q$, (ii) indicators of linear functions in $q$, and (iii) product of (i) and (ii). \end{proof} \begin{proof}[Proof of Theorem (ref)] It follows from Theorem (ref) by taking $\omega_i=1,~ \forall ~i$ . \end{proof} \subsection{Proofs for Section (ref)} \begin{proof}[Proof of Proposition (ref)] Theorem (ref) yields $T^{\texttt{LDA}}(\varepsilon)=0 \iff \varepsilon\in \mathcal{F}$. Next, we show that for $\varepsilon\in\mathcal{E}$, $\psi_\mathcal{PF}(\varepsilon)=0 \iff \varepsilon\in\mathcal{PF}$. Indeed, for every $\alpha\in[0,1]$, \begin{align*} q(\alpha)^\intercal\varepsilon+h_\mathcal{E}(-q(\alpha))=q(\alpha)^\intercal\varepsilon-\inf_{\eta\in\mathcal{E}}q(\alpha)^\intercal \eta \ge 0. \end{align*} So, $\psi_\mathcal{PF}(\varepsilon)=0$ if and only if there exists $\alpha\in[0,1]$ with $q(\alpha)^\intercal\varepsilon=\inf_{\eta\in\mathcal{E}}q(\alpha)^\intercal \eta$, i.e., $\varepsilon=\arg\min_{\eta\in\mathcal{E}}q(\alpha)^\intercal \eta$. By definition, this is equivalent to $\varepsilon\in\mathcal{PF}$. Since $\{\varepsilon\in \mathbb{B}_C:T^{\texttt{LDA}}(\varepsilon)=0\}=\mathcal{F}$, we have $\sup_{\varepsilon\in \mathbb{B}_C:T^{\texttt{LDA}}(\varepsilon)=0}\psi_\mathcal{PF}(\varepsilon)=\sup_{\varepsilon\in\mathcal{F}}\psi_\mathcal{PF}(\varepsilon)$. Because $\psi_\mathcal{PF}(\varepsilon)\ge 0$ for every $e\in\mathcal{E}$, the right hand side equals $0$ if and only if $\psi_\mathcal{PF}(\varepsilon)=0$ for all $\varepsilon\in\mathcal{F}$. For $v\neq 0$, by positive homogeneity of $h_\mathcal{E}(v)$, $h_\mathcal{E}(v)=\|v\|h_\mathcal{E}\left(\tfrac{v}{\|v\|}\right)$, so that, for every $\alpha\in[0,1]$ and $s\in[-\bar{s},\bar{s}]$, $\sqrt{1+s^2}\, h_\mathcal{E}\left(\tfrac{-q(\alpha)+s u_3}{\sqrt{1+s^2}} \right)=h_\mathcal{E}(-q(\alpha)+s u_3)$ for $u_3=[0,0,1]^\intercal$. Indeed, $\| -q(\alpha)+s u_3\|=\sqrt{1+s^2}$, since $q(\alpha)_3=0$ and $\|q(\alpha)\|=1$. Fix $\alpha\in[0,1]$, set $q=q(\alpha)$, and consider the scalar function \begin{align*} g_\alpha(s)\equiv h_\mathcal{E}(-q+s u_3) = \sup_{\eta\in\mathcal{E}}(-q+s u_3)^\intercal\eta . \end{align*} The function $g_\alpha$ is convex because it is the supremum of affine functions of $s$. By Assumptions (ref)-(ref), $\mathcal{E}$ is compact and strictly convex, so each support set is a singleton. Hence the support point $\mathcal{S}_\mathcal{E}(-q)=\arg\max_{\eta\in\mathcal{E}}(-q)^\intercal\eta$ is well-defined. We next show that $g_\alpha$ is differentiable at zero and that $g_\alpha'(0)=u_3^\intercal\mathcal{S}_\mathcal{E}(-q)$. Let $\varepsilon^*=\mathcal{S}_\mathcal{E}(-q)$. For $s\neq0$, define the (unique) support point $\varepsilon^s \equiv \arg\max_{\eta\in\mathcal{E}}(-q+s u_3)^\intercal\eta$. For $s>0$, optimality of $\varepsilon^*$ at $s=0$ and optimality of $\varepsilon^s$ at $s$ give \begin{align*} (\varepsilon^*)_3 \leq \tfrac{g_\alpha(s)-g_\alpha(0)}{s} \leq (\varepsilon^s)_3 . \end{align*} For $s<0$, the same inequalities hold with the order reversed: \begin{align*} (\varepsilon^s)_3 \leq \tfrac{g_\alpha(s)-g_\alpha(0)}{s} \leq (\varepsilon^*)_3 . \end{align*} The support map is continuous by Berge's maximum theorem, so $\varepsilon^s\to\varepsilon^*$ as $s\to0$. Therefore, $g_\alpha'(0)=(\varepsilon^*)_3=u_3^\intercal\mathcal{S}_\mathcal{E}(-q)$. Since $g_\alpha$ is convex and differentiable at zero, and since zero is an interior point of $[-\bar{s},\bar{s}]$, \begin{align*} g_\alpha(0) = \inf_{s\in[-\bar{s},\bar{s}]}g_\alpha(s) \quad\Longleftrightarrow\quad g_\alpha'(0)=0. \end{align*} Combining this with the derivative calculation yields \begin{align*} h_\mathcal{E}(-q(\alpha)) = \inf_{s\in[-\bar{s},\bar{s}]} \sqrt{1+s^2}\, h_\mathcal{E}\left( \tfrac{-q(\alpha)+s u_3}{\sqrt{1+s^2}} \right) \quad\Longleftrightarrow\quad u_3^\intercal\mathcal{S}_\mathcal{E}(-q(\alpha))=0. \end{align*} The expression inside the supremum in (ref) is nonnegative, as $s=0$ is in the interval $[-\bar{s},\bar{s}]$. Hence, $\Delta_{\mathcal{PF}}^{h}(\bar{s})=0$ if and only if $u_3^\intercal\mathcal{S}_\mathcal{E}(-q(\alpha))=0$ for every $\alpha\in[0,1]$. By Lemma (ref), $\mathcal{PF}=\{\mathcal{S}_\mathcal{E}(-q(\alpha)):\alpha\in[0,1]\}$, so $\Delta_{\mathcal{PF}}^{h}(\bar{s})=0$ if and only if $\mathcal{PF}\subset\{\varepsilon\in\mathcal{E}:\varepsilon_3=0\}$. Under Assumption (ref), Theorem (ref) gives $\mathcal{PF}=\mathcal{F}$ if and only if $\mathcal{PF}\subset\{\varepsilon\in\mathcal{E}:\varepsilon_3=0\}$. Combining the last two equivalences proves $\mathcal{PF}=\mathcal{F}\iff\Delta_{\mathcal{PF}}^{h}(\bar{s})=0$. \end{proof} \subsection{Proofs for Section (ref)} Let $[\cdot]_+=\max\{\cdot,0\}$. Define $\phi: \ell^\infty(\mathbb{S}^2) \times {\mathbb R}^3 \to {\mathbb R}$ as \begin{equation} \phi (f,\varepsilon) \equiv \left[ \max_{q \in \mathbb{S}^2} (q^\intercal \varepsilon -f(q) )\right]_+ + \left[ \min_{q \in \widetilde{\mathbb{S}}^2} (q_1 \varepsilon_1+q_2\varepsilon_2 +|q_3\varepsilon_3| + f(-q) )\right]_+, \end{equation} where $f \in \ell^\infty(\mathbb{S}^2) $ and $\varepsilon \in {\mathbb R}^3 $. Note that $T_n^{\text{LDA}} = \sqrt{n} \left( \phi(\widehat{\boldsymbol{\beta}}^\omega) - \phi(\boldsymbol{\beta}) \right)$. \begin{proof}[Proof of Theorem (ref)] The proof has two parts. In the first part, we use fan:san19 to show that the Bayesian bootstrap is consistent. In the second part, we use romano2012uniform to prove that the test $\varphi_n^{\text{LDA}}$ has asymptotically correct size. \textbf{Part 1:} By kai16 and standard arguments, it follows that $\phi$ defined in (ref) is Hadamard directionally differentiable at $\boldsymbol{\beta}$ tangentially to $\mathbb{D}_0 = C(\mathbb{S}^2) \times {\mathbb R}^3 \subset \mathbb{D} = \ell^\infty(\mathbb{S}^2) \times {\mathbb R}^3 $, where $C(\mathbb{S}^2)$ is the class of continuous functions on $\mathbb{S}^2$. This guarantees Assumption 1 in fan:san19. By Cramer-Wold and Theorems (ref) and (ref), we have $\sqrt{n}(\widehat{\boldsymbol{\beta}}^\omega - \boldsymbol{\beta}) \Rightarrow \mathbb{G}_{\boldsymbol{\beta}}$. Let $K(q,q')$ be the covariance matrix of $\mathbb{G}_{\boldsymbol{\beta}}$ for $q,q' \in \mathbb{S}^2$. It can be show that $K(q,q')$ is continuous in both $q,q' \in \mathbb{S}^2$, which implies that $\mathbb{G}_{\boldsymbol{\beta}} \in \mathbb{D}_0$ with probability 1 by adler2007random. This guarantees Assumption 2 in fan:san19. By fan:san19, we conclude $\sqrt{n} \left( \phi(\widehat{\boldsymbol{\beta}}^\omega) - \phi(\boldsymbol{\beta})\right) \Rightarrow \phi_{\boldsymbol{\beta}}'( \mathbb{G}_{\boldsymbol{\beta}})$. To prove the consistency of the bootstrap, it is sufficient to show that \begin{align} \sup_{f\in\mathcal{BL}_1}\left|\mathbb{E}\big[f\big(\widehat{\phi'}_{\boldsymbol{\beta}}\big(\sqrt{n}\{\widetilde{\boldsymbol{\beta}}^\omega-\widehat{\boldsymbol{\beta}}^\omega\}\big)\big)\,\big|\,\{(Y_i,G_i,X_i,Z_i,\Xi_i)\}_{i=1}^n\big]-\mathbb{E}\big[f\big(\phi_{\boldsymbol{\beta}}'(\mathbb{G}_{\boldsymbol{\beta}})\big)\big]\right|=o_p(1), \end{align} where $\mathcal{BL}_1$ is the set of $1$-Lipschitz functions $f:\mathbb{R}\to\mathbb{R}$ such that $|f|_\infty\leq1$. It is sufficient to verify Assumptions 3 and 4 in fan:san19 to conclude (ref) by fan:san19. Note that Assumption 3 part (i), (iii), and (iv) hold by construction. By HONG2018379, we conclude $\widehat{\phi'}_{\boldsymbol{\beta}}$ defined in Procedure 1 verifies Assumption 4 in fan:san19. Therefore, it is sufficient to verify Assumption 3 part (ii), i.e., the Bayesian bootstrap works for the estimator $\widehat{\boldsymbol{\beta}}^\omega $: \begin{equation} \sup_{f\in\mathcal{BL}_1}\left|\mathbb{E}\left[f\left( \sqrt{n} \left( \widetilde{\boldsymbol{\beta}}^\omega - \widehat{\boldsymbol{\beta}}^\omega \right) \right)\,\big|\,\{(Y_i,G_i,X_i,Z_i,\Xi_i)\}_{i=1}^n\right] -\mathbb{E}\big[f\big( \mathbb{G}_{\boldsymbol{\beta}})\big]\right|=o_p(1) . \end{equation} Define \begin{align} \widehat{\boldsymbol{\beta}}^{\omega,*} &\equiv \left( \tfrac{1}{n} \sum_{i =1}^n \left( \tfrac{\omega_i^h}{\overline{\omega^h}} \right) \zeta_i^*(q;{\mathbf{L}},{\boldsymbol{\eta}}) , \tfrac{1}{n} \sum_{i =1}^n \left( \tfrac{\omega_i^\varepsilon}{\overline{\omega^\varepsilon}} \right) \xi_i^*(a^*;{\mathbf{L}},{\boldsymbol{\eta}})\right)\\ \widetilde{\boldsymbol{\beta}}^{\omega,*} &\equiv \left( \tfrac{1}{n} \sum_{i =1}^n \left( \tfrac{W_i \omega_i^h}{\overline{W^h}} \right) \zeta_i^*(q;{\mathbf{L}},{\boldsymbol{\eta}}) , \tfrac{1}{n} \sum_{i =1}^n \left( \tfrac{W_i \omega_i^\varepsilon}{\overline{W^\varepsilon}} \right) \xi_i^*(a^*;{\mathbf{L}},{\boldsymbol{\eta}})\right) \end{align} By Lemma (ref), we have $\left\|\sqrt{n} \left( \widetilde{\boldsymbol{\beta}}^\omega - \widehat{\boldsymbol{\beta}}^\omega \right) - \sqrt{n} \left( \widetilde{\boldsymbol{\beta}}^{\omega,*} - \widehat{\boldsymbol{\beta}}^{\omega,*} \right) \right\| = o_p(1)$; therefore, we conclude that (ref) holds if the following equation holds $$ \sup_{f\in\mathcal{BL}_1}\left|\mathbb{E}\left[f\left( \sqrt{n} \left( \widetilde{\boldsymbol{\beta}}^{\omega,*} - \widehat{\boldsymbol{\beta}}^{\omega,*} \right) \right)\,\big|\,\{(Y_i,G_i,X_i,Z_i,\Xi_i)\}_{i=1}^n\right] -\mathbb{E}\big[f\big( \mathbb{G}_{\boldsymbol{\beta}})\big]\right|=o(1)~. $$ Finally, the previous equation holds by definition of $ \widetilde{\boldsymbol{\beta}}^{\omega,*}$, $ \widehat{\boldsymbol{\beta}}^{\omega,*}$, by van:wel13, and because $\sqrt{n}\left( \widehat{\boldsymbol{\beta}}^{\omega,*} - \boldsymbol{\beta} \right) \Rightarrow \mathbb{G}_{\boldsymbol{\beta}}$ since $E[\omega^h] = E[\omega^\varepsilon] = 1$. This completes the proof of the bootstrap consistency. \textbf{Part 2:} For an arbitrarily small $\kappa>0$, $\mathbb{P}\left(\sup_{x \in {\mathbb R}} \widetilde{F}_n(x)-F_\infty(x) > \kappa \right)= o(1)$, where \begin{align*} \widetilde{F}_n(x) &= \mathbb{P}\left( \widehat{\phi'}_{\boldsymbol{\beta}}\big(\sqrt{n}\{\widetilde{\boldsymbol{\beta}}^\omega-\widehat{\boldsymbol{\beta}}^\omega\}\big) \le x | \{(Y_i,G_i,X_i,Z_i,\Xi_i)\}_{i=1}^n \right), F_\infty(x) = \mathbb{P}\left( \phi_{\boldsymbol{\beta}}'( \mathbb{G}_{\boldsymbol{\beta}}) \le x \right) . \end{align*} Let $c_{1-\alpha}$ be $(1-\alpha)$-quantile of $\phi_{\boldsymbol{\beta}}'( \mathbb{G}_{\boldsymbol{\beta}})$, i.e., $F_\infty(c_{1-\alpha}) \ge 1-\alpha$. Recall that $ \widehat{c}_{1-\alpha + \kappa}$ is the $(1-\alpha+\kappa)-$quantile of the distribution of $\widehat{\phi'}_{\boldsymbol{\beta}}\big(\sqrt{n}\{\widetilde{\boldsymbol{\beta}}^\omega-\widehat{\boldsymbol{\beta}}^\omega\}\big)$. The following derivation completes the proof of part 2: \begin{align*} \mathbb{P}( T_n^{\text{LDA}} > \widehat{c}_{1-\alpha + \kappa} + \kappa) &\overset{(1)}{\le} \mathbb{P}( T_n^{\text{LDA}} > c_{1-\alpha} + \kappa) + o(1) \overset{(2)}{\le} \mathbb{P}( T_n^{\text{LDA}} \ge c_{1-\alpha} + \kappa/2 ) + o(1) \\ & \overset{(3)}{\le} \mathbb{P}( \phi_{\boldsymbol{\beta}}'( \mathbb{G}_{\boldsymbol{\beta}}) \ge c_{1-\alpha} + \kappa/2 ) + o(1) \\ &\overset{(4)}{\le} \mathbb{P}( \phi_{\boldsymbol{\beta}}'( \mathbb{G}_{\boldsymbol{\beta}}) > c_{1-\alpha} ) + o(1) \overset{(5)}{\le} \alpha + o(1) , \end{align*} where (1) holds since $\mathbb{P}( \widehat{c}_{1-\alpha + \kappa} \ge c_{1-\alpha}) = 1-o(1)$ by romano2012uniform, (2) and (4) hold since $\kappa >0$, (3) holds since $T_n^{\text{LDA}} \overset{d}{\rightarrow} \phi_{\boldsymbol{\beta}}'( \mathbb{G}_{\boldsymbol{\beta}})$, and (5) holds by the definition of $c_{1-\alpha}$. \end{proof} \section{Auxiliary Results } \subsection{Lemmas and Auxiliary Results for Section (ref)} \begin{lemma}[Support function of $\mathcal{C}(\varepsilon^*)$] \textit{For $\varepsilon^*=(\varepsilon_1^*,\varepsilon_2^*,\varepsilon_3^*)\in\mathcal E$ and $q\in{\mathbb R}^3$, the support function of the closed convex set $\mathcal{C}(\varepsilon^*)$ in (ref) is as given in (ref).} \end{lemma} \begin{proof} We have $h_{\mathcal{C}(\varepsilon^*)}(q) = \sup_{\varepsilon \in \mathcal{C}(\varepsilon^*)} q^\intercal{\varepsilon}= \sup_{\substack{\varepsilon_1 \leq \varepsilon_1^*, \varepsilon_2 \leq \varepsilon_2^*\\ |\varepsilon_3| \leq |\varepsilon_3^*|}} (q_1 \varepsilon_1 + q_2 \varepsilon_2 + q_3 \varepsilon_3).$ If $q_1 < 0$, then $\sup_{\varepsilon_1 \leq \varepsilon_1^*} q_1 \varepsilon_1 = +\infty$ (taking $\varepsilon_1 \to -\infty$). Similarly if $q_2 < 0$. For $q_1, q_2 \geq 0$, the supremum over $\varepsilon_j \leq \varepsilon_j^*$ gives $q_j \varepsilon_j^*$, $j=1,2$. For the third coordinate: \begin{align*} \sup_{|\varepsilon_3| \leq |\varepsilon_3^*|} q_3 \varepsilon_3 = \begin{cases} q_3 |\varepsilon_3^*| & \text{if } q_3 \geq 0,\\ -q_3 |\varepsilon_3^*| = |q_3| |\varepsilon_3^*| & \text{if } q_3 < 0. \end{cases} \end{align*} Thus $h_{\mathcal{C}(\varepsilon^*)}(q) = q_1 \varepsilon_1^* + q_2 \varepsilon_2^* + |q_3| |\varepsilon_3^*|$ for $q_1, q_2 \geq 0$. \end{proof} \begin{lemma} \textit{Let $(\Omega,\mathfrak F,\mathbb{P})$ be the probability space on which $(Y,G,X)$ are defined. Then: (i) $\mathcal{E}$ is convex. (ii) Under Assumption (ref), $\mathcal{E}$ is compact. (iii) Under Assumptions (ref)-(ref), $\mathcal{E}$ has a non-empty interior in ${\mathbb R}^3$, and hence $\dim(\mathcal{E})=3$.} \end{lemma} \begin{proof} Recall that $\mathcal A \equiv\{a:\mathcal{X}\to[0,1]\ \text{Borel-measurable}\}$. \textbf{Part (i).} Each coordinate map $a\mapsto e_g^\iota(a)$ is affine in $a$ by (ref). Since $\mathcal A$ is convex and $(e_r^A(a),e_b^A(a),e_r^F(a)-e_b^F(a))$ is an affine image of $a$, its range $\mathcal{E}$ is convex. \textbf{Part (ii).} Define \begin{align*} K\equiv\{u\in L^\infty(\Omega,\mathfrak F,\mathbb{P}): u \text{ is }\sigma(X)\text{-measurable and }0\le u\le 1 \text{ a.s.}\}. \end{align*} As $(\mathcal{X},\mathcal B(\mathcal{X}))$ is standard Borel, the Doob--Dynkin lemma RaoSwift2006 yields $K=\{a(X):a\in\mathcal A\}$ up to a.s. equality. Hence, $\mathcal{E}=\mathbb{E}[\boldsymbol{\theta}_0(X)]+T(K)$, for \begin{align} T(u)\equiv\mathbb{E}[u\Delta\boldsymbol{\theta}(X)]\in{\mathbb R}^3. \end{align} We first show that $K$ is weak-* compact in $L^\infty$, where $L^\infty$ is equipped with the topology $\sigma(L^\infty,L^1)$. Since $0\le u\le 1$ a.s. for every $u\in K$, we have $\|u\|_\infty\le 1$, so $K$ is contained in the closed unit ball of $L^\infty$. Since $L^\infty=(L^1)^*$, the Banach-Alaoglu theorem yields that the closed unit ball of $L^\infty$ is compact for the weak-* topology $\sigma(L^\infty, L^1)$. To prove weak-* compactness of $K$, it suffices to show that $K$ is weak-* closed, since $K$ is contained in the weak-* compact unit ball of $L^\infty$. Accordingly, let $u_\alpha\in K$ and suppose $u_\alpha\to u$ weak-* in $L^\infty$. We show that $u\in K$. First, $u$ is $\sigma(X)$-measurable. Indeed, for every $g\in L^1(\Omega,\mathfrak F,\mathbb{P})$, $\mathbb{E}\left[u_\alpha\bigl(g-\mathbb{E}[g| \sigma(X)]\bigr)\right]=0$, because $u_\alpha$ is $\sigma(X)$-measurable. Passing to the limit using weak-* convergence gives $\mathbb{E}\left[u\left(g-\mathbb{E}[g| \sigma(X)]\right)\right]=0$, for all $g\in L^1$. Taking $g=\mathds{1}_A$ for $A\in\mathfrak F$, we obtain \begin{align*} \int_A ud\mathbb{P}=\int_A \mathbb{E}[u|\sigma(X)]d\mathbb{P},\quad \forall A\in\mathfrak F, \end{align*} hence $u=\mathbb{E}[u|\sigma(X)]$ a.s. Therefore $u$ is $\sigma(X)$-measurable. Next, the bounds $0\le u\le 1$ a.s. are preserved. To show $u\ge 0$ a.s., suppose otherwise. Then $\mathbb{P}(u<0)>0$, so for some $\varepsilon>0$, $A_\varepsilon\equiv\{u\le -\varepsilon\}$ has positive probability. Taking $g=\mathds{1}_{A_\varepsilon}\in L^1$, weak-* convergence yields $\mathbb{E}[u_\alpha g]\to \mathbb{E}[ug]$. Since $u_\alpha\ge 0$ a.s., we have $\mathbb{E}[u_\alpha g]\ge 0$ for all $\alpha$, whereas $\mathbb{E}[ug]=\mathbb{E}[u\mathds{1}_{A_\varepsilon}] \le -\varepsilon\,\mathbb{P}(A_\varepsilon)<0$, a contradiction. Thus $u\ge 0$ a.s. Applying the same argument to $1-u_\alpha$ shows $u\le 1$ a.s. Therefore $u\in K$, so $K$ is weak-* closed. Hence $K$ is weak-* compact. Now consider the map $T:K\to{\mathbb R}^3$ in (ref). This map is well-defined because $u\in L^\infty$ with $|u|\le 1$ a.s. and each component of $\Delta\boldsymbol{\theta}(X)$ is in $L^1$ by Assumption (ref), so each component of $u\Delta\boldsymbol{\theta}(X)$ is in $L^1$. Moreover, for each component $j=1,2,3$, the map from $u$ to the component of $\mathbb{E}[u\Delta\boldsymbol{\theta}(X)]$ is weak-* continuous on $L^\infty$, since each component of $\Delta\boldsymbol{\theta}(X)$ is in $L^1$. Therefore $T:K\to{\mathbb R}^3$ is continuous when $K$ is endowed with the weak-* topology and ${\mathbb R}^3$ with its usual Euclidean topology. Since $K$ is weak-* compact and $T$ is continuous, $T(K)$ is compact in ${\mathbb R}^3$. Therefore, $\mathcal{E}=\mathbb{E}[\boldsymbol{\theta}_0(X)]+T(K)$ is compact. \textbf{Part (iii).} Let $\mathcal A_{\mathrm{disp}}\equiv\{\varepsilon:\mathcal{X}\to{\mathbb R}\ \text{measurable}:\ |\varepsilon(X)|\le 1/2\ \text{$P_X$-a.s.}\}=\{\varepsilon:\mathcal{X}\to{\mathbb R}\ \text{measurable}:\ \|\varepsilon\|_\infty\le 1/2\}$ where $\|\varepsilon\|_\infty\equiv\mathrm{ess}\sup_{x\sim P_X}|\varepsilon(x)|$. For any algorithm $a\in\mathcal A$, let $\varepsilon(X)\equiv a(X)-\tfrac12$ be the corresponding displacement in $\mathcal A_\mathrm{disp}$. Let $e^0=\mathbb{E}\bigl[\boldsymbol{\theta}_0(X)]+\tfrac12\mathbb{E}[\Delta\boldsymbol{\theta}(X)]$. Then for any $a\in\mathcal A$ we have $e(a)-e^0=\mathbb{E}[\varepsilon(X)\Delta\boldsymbol{\theta}(X)]$. Hence, we can write $\mathcal{E}=e^0+\mathcal{E}^0$, where $\mathcal{E}^0$ is the set of feasible displacements: \begin{equation*} \mathcal{E}^0\equiv\left\{\mathbb{E}[\varepsilon(X)\Delta\boldsymbol{\theta}(X)], \varepsilon:\mathcal{X}\to[-\tfrac12,\tfrac12]\ \text{measurable}\right\}, \end{equation*} and the equality holds because every measurable $a:\mathcal{X}\to[0,1]$ corresponds to a unique $\varepsilon\in[-\tfrac12,\tfrac12]$ and vice versa. We next show that $\mathcal{E}^0$ contains a Euclidean ball around $0$, which implies $\mathcal E=e^0+\mathcal{E}^0$ contains a ball around $e^0$. Fix $q\in{\mathbb R}^3$ with $\|q\|=1$. Consider the support function of $\mathcal{E}^0$ in direction $q$: \begin{align*} h_{\mathcal{E}^0}(q)\equiv\sup_{e\in \mathcal{E}^0}q^\intercal e =\sup_{\|\varepsilon\|_\infty\le 1/2}\mathbb{E}\left[\varepsilon(X)q^\intercal\Delta\boldsymbol{\theta}(X)\right]. \end{align*} Pointwise maximization yields the maximizer $\varepsilon_q(X)=\tfrac12\,\mathrm{sign}\big(q^\intercal\Delta\boldsymbol{\theta}(X)\big)$ (with any value in $[-\tfrac12,\tfrac12]$ on the null set where $q^\intercal\Delta\boldsymbol{\theta}(X)=0$), so \begin{equation} h_{\mathcal{E}^0}(q)=\tfrac12\mathbb{E}\left[|q^\intercal\Delta\boldsymbol{\theta}(X)|\right]. \end{equation} By Assumption (ref), there exists a constant $C>0$ such that for all $\delta>0$, $\mathbb{P}(|q^\intercal\Delta\boldsymbol{\theta}(X)|\le \delta)\le C\delta^m$, uniformly in $q\in\mathbb{S}^2\equiv\{q\in{\mathbb R}^3:\Vert q\Vert=1\}$. Choose $\delta_0\equiv (1/(2C))^{1/m}>0$. Then for every $q\in\mathbb{S}^2$, $\mathbb{P}(|q^\intercal\Delta\boldsymbol{\theta}(X)|>\delta_0)\ge 1-C\delta_0^m = \tfrac12$, and hence $\mathbb{E}\left[|q^\intercal\Delta\boldsymbol{\theta}(X)|\right]\ge\delta_0\mathbb{P}(|q^\intercal\Delta\boldsymbol{\theta}(X)|>\delta_0)\ge\tfrac{\delta_0}{2}$. Combining with (ref) yields the uniform bound \begin{equation} h_{\mathcal{E}^0}(q)\ge\tfrac{\delta_0}{4} \qquad\forall q\in\mathbb{S}^2. \end{equation} Hence, (ref) implies that $\mathcal{E}^0$ contains the Euclidean ball $\mathbb{B}(0,r)$ with radius $r=\delta_0/4$ sch93. Therefore, $e^0+\mathbb{B}(0,r)\subseteq \mathcal{E}$ so $\mathcal{E}$ has a nonempty interior in ${\mathbb R}^3$, and hence $\dim(\mathcal{E})=3$. \end{proof} \subsubsection{ Characterization of the Geometry of $\mathcal{F}$ using the Support-Set of $\mathcal{E}$} Let Assumptions (ref)-(ref) hold and let $\mathcal{S}_\mathcal{E}(q)$ be the support-set of $\mathcal{E} \subset {\mathbb R}^3$, \begin{align*} \mathcal{S}_\mathcal{E}(q) =\arg\max_{\varepsilon\in\mathcal{E}}q^\intercal \varepsilon, \qquad q\in\mathbb{S}^2. \end{align*} Since $\mathcal{E}$ is compact and strictly convex, $\mathcal{S}_\mathcal{E}(q)$ is a well-defined continuous function. \begin{lemma} \textit{Let $\mathcal{PF}$ and $\mathcal{F}$ be the Pareto and FA frontiers defined in (ref) and (ref), respectively, and let $\widetilde{\mathbb{S}}^2\equiv\{q\in \mathbb{S}^2:q_1\ge 0,\ q_2\ge 0\}$. Under Assumptions (ref)-(ref), we have \begin{enumerate} • $\mathcal{PF} = \{ \varepsilon = \mathcal{S}_\mathcal{E}(-q): q \in \widetilde{\mathbb{S}}^2,\ q_3 = 0 \}$$\mathcal{F} = \{ \varepsilon = \mathcal{S}_\mathcal{E}(-q): q \in \widetilde{\mathbb{S}}^2,\ q_3 \varepsilon_3 \ge 0\} $ \end{enumerate} In particular, $\mathcal{PF} \subseteq \mathcal{F}$.} \end{lemma} \begin{proof} The proof has two parts. We first characterize the Pareto and FA frontiers using support functions. We then use the characterization to conclude the lemma. \textbf{Part 1:} For $\varepsilon^*=(\varepsilon_1^*,\varepsilon_2^*,\varepsilon_3^*)\in \mathcal{E}$, we define \begin{align} \mathcal{C}^P(\varepsilon^*) \equiv \{\varepsilon=(\varepsilon_1,\varepsilon_2,\varepsilon_3)\in{\mathbb R}^3: \varepsilon_1\le \varepsilon_1^*,\ \varepsilon_2\le \varepsilon_2^*\}, \end{align} the set of group-wise expected accuracy losses that are weakly more accurate than $\varepsilon^*$. It can be verified that the support function of $\mathcal{C}^P(\varepsilon^*)$ equals: \begin{align} h_{\mathcal{C}^P(\varepsilon^*)}(q)\equiv\sup_{\varepsilon\in\mathcal{C}^P(\varepsilon^*)}q^\intercal\varepsilon= \begin{cases} q_1 \varepsilon_1^* + q_2 \varepsilon_2^* & \text{if } q_1\ge 0,\ q_2\ge 0, q_3 = 0\\ +\infty & \text{otherwise.} \end{cases} \end{align} Let $\varepsilon^*=(\varepsilon_1^*,\varepsilon_2^*,\varepsilon_3^*)\in \mathcal{E}$, then $\varepsilon^* \in \mathcal{PF}$ if and only if \begin{equation} \min_{q\in \widetilde{\mathbb{S}}^2,\ q_3 = 0} \left( h_{\mathcal{C}^P(\varepsilon^*)}(q)+h_{\mathcal{E}}(-q) \right)=0 , \end{equation} and $\varepsilon^* \in \mathcal{F}$ if and only if \begin{equation} \min_{q\in \widetilde{\mathbb{S}}^2} \left( h_{\mathcal{C}(\varepsilon^*)}(q)+h_{\mathcal{E}}(-q) \right)=0 , \end{equation} where (ref) follows by similar arguments to those in the proof of Theorem (ref) (ref)-(ref), and (ref) is due to Theorem (ref) (ref). \textbf{Part 2:} We use (ref) and the support set $\mathcal{S}_\mathcal{E}$ to note that (ref) can be written as follows $$ \min_{q\in \widetilde{\mathbb{S}}^2,\ q_3 = 0} \left( q_1 \varepsilon_1^* + q_2 \varepsilon_2^* - q^\intercal \mathcal{S}_\mathcal{E}(-q) \right)=0 ~.$$ Therefore, any $\varepsilon^* = \mathcal{S}_\mathcal{E}(-q)$ for $q\in \widetilde{\mathbb{S}}^2$ and $q_3 =0 $ satisfies the previous equation, implying that $ \{ \varepsilon = \mathcal{S}_\mathcal{E}(-q) : q \in \widetilde{\mathbb{S}}^2,\ q_3 = 0 \} \subset \mathcal{PF}$. Similarly, for any $\varepsilon^*$ that satisfies the previous equation there exists $q \in \widetilde{\mathbb{S}}^2$ such that $q_3 = 0$ and $ q_1 \varepsilon_1^* + q_2 \varepsilon_2^* - q^\intercal \mathcal{S}_\mathcal{E}(-q) = 0$, which implies $\varepsilon^* = \mathcal{S}_\mathcal{E}(-q) $ and $\mathcal{PF} \subset \{ \varepsilon = \mathcal{S}_\mathcal{E}(-q): q \in \widetilde{\mathbb{S}}^2,\ q_3 = 0 \}$. This proves Lemma (ref)(1). We derive a convenient characterization of $\mathcal{F}$ to prove Lemma (ref)(2). Fix $\varepsilon^*\in\mathcal{E}$ and $q\in\tilde{\mathbb{S}}^2$. Since $h_{\mathcal{C}(\varepsilon^*)}(q)=q_1\varepsilon_1^*+q_2\varepsilon_2^*+|q_3||\varepsilon_3^*|$, and $h_{\mathcal{E}}(-q)=-\min_{\varepsilon\in\mathcal{E}}q^\intercal \varepsilon$, we obtain \begin{align*} h_{\mathcal{C}(\varepsilon^*)}(q)+h_{\mathcal{E}}(-q) = q_1\varepsilon_1^*+q_2\varepsilon_2^*+|q_3||\varepsilon_3^*|-\min_{\varepsilon\in\mathcal{E}}q^\intercal \varepsilon. \end{align*} Since $\min_{\varepsilon\in\mathcal{E}}q^\intercal \varepsilon\le q^\intercal \varepsilon^* = q_1\varepsilon_1^*+q_2\varepsilon_2^*+q_3\varepsilon_3^*$, it follows that \begin{align*} h_{\mathcal{C}(\varepsilon^*)}(q)+h_{\mathcal{E}}(-q) \ge |q_3||\varepsilon_3^*|-q_3\varepsilon_3^*\ge 0. \end{align*} Therefore, $\varepsilon^*\in\mathcal{F} \iff \exists\,q\in\tilde{\mathbb{S}}^2 \text{ such that } \varepsilon^*=\mathcal{S}_\mathcal{E}(-q) \text{ and } q_3\varepsilon_3^*\ge 0$, proving Lemma (ref)(2). \end{proof} Lemma (ref) has two immediate implications. First, $\mathcal{PF} \subseteq \mathcal{F}$. Second, the Pareto frontier $\mathcal{PF}$ is a one-dimensional object, while the FA-frontier can be a two-dimensional object. Let $\widetilde{\mathbb{S}}^{2}_{+}\equiv\{q\in\widetilde{\mathbb{S}}^{2}:q_3\geq0\}$. Define the coordinate-edge image \begin{align*} \mathcal{IM}_{ij} \equiv \left\{ \arg\min_{\varepsilon\in\mathcal{E}}\{\alpha\varepsilon_i+(1-\alpha)\varepsilon_j\}: \alpha\in[0,1] \right\},\quad 1\leq i<j\leq3. \end{align*} \begin{corr} \textit{Suppose Assumptions (ref)-(ref) hold and $\mathcal{E}\subset\{\varepsilon\in{\mathbb R}^3:\varepsilon_3\geq0\}$. Then \begin{align} \mathcal{F} = \{\mathcal{S}_\mathcal{E}(-q):q\in\widetilde{\mathbb{S}}^{2}_{+}\}. \end{align} The images under $q\mapsto\mathcal{S}_\mathcal{E}(-q)$ of the three boundary edges of $\widetilde{\mathbb{S}}^{2}_{+}$ are \begin{align*} \mathcal{IM}_{12} = \{\mathcal{S}_\mathcal{E}(-q):q\in\widetilde{\mathbb{S}}^{2}_{+},\ q_3=0\} = \mathcal{PF}, \end{align*} \begin{align*} \mathcal{IM}_{13} = \{\mathcal{S}_\mathcal{E}(-q):q\in\widetilde{\mathbb{S}}^{2}_{+},\ q_2=0\}, \qquad \mathcal{IM}_{23} = \{\mathcal{S}_\mathcal{E}(-q):q\in\widetilde{\mathbb{S}}^{2}_{+},\ q_1=0\}. \end{align*} If Assumption (ref) also holds, then $\mathcal{IM}_{12},\mathcal{IM}_{13},\mathcal{IM}_{23}$ are the relative boundary edge curves of $\mathcal{F}$ and $\mathcal{IM}_{12}=\mathcal{PF}\subsetneq\mathcal{F}$. The case $\mathcal{E}\subset\{\varepsilon_3\leq0\}$ follows by the reflection $(\varepsilon_1,\varepsilon_2,\varepsilon_3)\mapsto(\varepsilon_1,\varepsilon_2,-\varepsilon_3)$.} \end{corr} \begin{proof} If $q\in\widetilde{\mathbb{S}}^{2}_{+}$, then for $\varepsilon=\mathcal{S}_\mathcal{E}(-q)$ we have $\varepsilon_3\geq0$ and $q_3\geq0$, so $q_3\varepsilon_3\geq0$. Lemma (ref) yields $\varepsilon\in\mathcal{F}$. Conversely, let $\varepsilon=\mathcal{S}_\mathcal{E}(-q)\in\mathcal{F}$ for some $q\in\widetilde{\mathbb{S}}^{2}$, so that $q_3\varepsilon_3\geq0$. If $q_3\geq0$, then $q\in\widetilde{\mathbb{S}}^{2}_{+}$. If $q_3<0$, then $\varepsilon_3=0$ and for any $\eta\in\mathcal{E}$, \begin{align*} q_1\varepsilon_1+q_2\varepsilon_2 = q^\intercal\varepsilon \leq q^\intercal\eta = q_1\eta_1+q_2\eta_2+q_3\eta_3 \leq q_1\eta_1+q_2\eta_2, \end{align*} where the last inequality uses $q_3<0$ and $\eta_3\geq0$. The case $q_1=q_2=0$ would force $\eta_3=0$ for all $\eta\in\mathcal{E}$, contradicting that $\mathcal{E}$ has nonempty interior. Hence $q_1+q_2>0$. After normalizing $[q_1,q_2,0]^\intercal$, we obtain a vector in $\widetilde{\mathbb{S}}^{2}_{+}$ exposing the same point $\varepsilon$. This proves (ref). The identities for $\mathcal{IM}_{12}$, $\mathcal{IM}_{13}$, and $\mathcal{IM}_{23}$ follow because positive rescaling of a normal vector does not change the minimizers. For example, $\{q\in\widetilde{\mathbb{S}}^{2}_{+}:q_3=0\}$ consists exactly of positive normalizations of $[\alpha,1-\alpha,0]^\intercal$ for $\alpha\in[0,1]$. Thus its support image is $\mathcal{IM}_{12}$, which equals $\mathcal{PF}$ by Lemma (ref). The other two edges follow by the same argument. Under Assumption (ref), the support map is injective. Hence the map $T:\widetilde{\mathbb{S}}^{2}_{+}\to\mathcal{F}$, $T(q)=\mathcal{S}_\mathcal{E}(-q)$, is injective. It is also continuous by Berge's maximum theorem, and it is surjective by the representation in (ref). Therefore $T$ is a continuous bijection from the compact space $\widetilde{\mathbb{S}}^{2}_{+}$ onto the Hausdorff space $\mathcal{F}\subset{\mathbb R}^3$, and hence $T$ is a homeomorphism. Thus, the images of the three boundary edges of $\widetilde{\mathbb{S}}^{2}_{+}$ are the relative boundary edge curves of $\mathcal{F}$. In particular, $\mathcal{IM}_{12}=T\left(\{q\in\widetilde{\mathbb{S}}^{2}_{+}:q_3=0\}\right)=\mathcal{PF}$. Since $\{q\in\widetilde{\mathbb{S}}^{2}_{+}:q_3=0\}$ is a proper subset of $\widetilde{\mathbb{S}}^{2}_{+}$, injectivity of $T$ implies $\mathcal{IM}_{12}=\mathcal{PF}\subsetneq\mathcal{F}$. \end{proof} \begin{corr} \textit{Let Assumptions (ref)-(ref) hold. Suppose that $\mathcal{E}\cap\{\varepsilon\in{\mathbb R}^3:\varepsilon_3>0\}\neq\emptyset$, $\mathcal{E}\cap\{\varepsilon\in{\mathbb R}^3:\varepsilon_3<0\}\neq\emptyset$, and that either $\mathcal{PF}\subset\{\varepsilon\in{\mathbb R}^3:\varepsilon_3>0\}$ or $\mathcal{PF}\subset\{\varepsilon\in{\mathbb R}^3:\varepsilon_3<0\}$. Recall $\mathcal{E}^A\equiv\pi(\mathcal{E})$, and let \begin{align*} \mathcal{PF}^A \equiv \left\{ c\in\mathcal{E}^A: \not\exists c'\in\mathcal{E}^A \text{ such that } c'_1\leq c_1,\ c'_2\leq c_2, \text{ with at least one strict inequality} \right\}. \end{align*} Then $\mathcal{PF}^A\subsetneq\mathcal{F}^A$ and $\mathcal{PF}\subsetneq\mathcal{F}$.} \end{corr} \begin{proof} It is enough to prove the result when $\mathcal{PF}\subset\{\varepsilon\in{\mathbb R}^3:\varepsilon_3>0\}$; the other case follows by applying the reflection $(\varepsilon_1,\varepsilon_2,\varepsilon_3)\mapsto(\varepsilon_1,\varepsilon_2,-\varepsilon_3)$. Since $\mathcal{E}$ is convex and contains points with positive and negative third coordinate, $\mathcal{E}\cap\{\varepsilon\in{\mathbb R}^3:\varepsilon_3=0\}\neq\emptyset$. Hence, $\pi\left(\mathcal{E}\cap\{\varepsilon_3=0\}\right)$ is a nonempty compact subset of $\mathcal{E}^A$. Choose $c^0\in\arg\min_{c\in\pi(\mathcal{E}\cap\{\varepsilon_3=0\})}(c_1+c_2)$. Then $c^0\in\mathcal{E}^A$ and $d(c^0)=0$ for $d(\cdot)$ defined in (ref), because $c^0$ has a feasible lift with third coordinate equal to zero. We first show that $c^0\in\mathcal{F}^A$. Suppose not. Then there exists $c'\in\mathcal{E}^A$ such that \begin{align*} c'_1\leq c^0_1,\qquad c'_2\leq c^0_2,\qquad d(c')\leq d(c^0)=0, \end{align*} with at least one strict inequality. Since $d(c')\geq0$, we have $d(c')=0$. Therefore $c'\in\pi(\mathcal{E}\cap\{\varepsilon_3=0\})$. Moreover, the strict inequality cannot occur only in the fairness coordinate, because $d(c')=d(c^0)=0$. Hence at least one of $c'_1<c^0_1$ or $c'_2<c^0_2$ holds, so $c'_1+c'_2<c^0_1+c^0_2$, contradicting the definition of $c^0$. Thus $c^0\in\mathcal{F}^A$. Next, $c^0\notin\mathcal{PF}^A$. Indeed, since $c^0\in\pi(\mathcal{E}\cap\{\varepsilon_3=0\})$, the point $\varepsilon^0\equiv(c^0_1,c^0_2,0)$ belongs to $\mathcal{E}$. If $c^0\in\mathcal{PF}^A$, then $\varepsilon^0\in\mathcal{PF}$: otherwise there would exist $\eta\in\mathcal{E}$ with $\eta_1\leq\varepsilon^0_1$, $\eta_2\leq\varepsilon^0_2$, and at least one strict inequality, and then $\pi(\eta)$ would Pareto-dominate $c^0$ in $\mathcal{E}^A$. This contradicts $c^0\in\mathcal{PF}^A$. Therefore $\varepsilon^0\in\mathcal{PF}$, contradicting the maintained case $\mathcal{PF}\subset\{\varepsilon_3>0\}$. Hence $c^0\notin\mathcal{PF}^A$. We have shown that $c^0\in\mathcal{F}^A\setminus\mathcal{PF}^A$. It remains to check that $\mathcal{PF}^A\subseteq\mathcal{F}^A$. If $c\in\mathcal{PF}^A$ but $c\notin\mathcal{F}^A$, then some $c'\in\mathcal{E}^A$ satisfies $c'_1\leq c_1$, $c'_2\leq c_2$, $d(c')\leq d(c)$, with at least one strict inequality. The strict inequality cannot occur only in the fairness coordinate: if $c'_1=c_1$ and $c'_2=c_2$, then $c'=c$, hence $d(c')=d(c)$. Thus at least one of $c'_1<c_1$ or $c'_2<c_2$ holds, contradicting $c\in\mathcal{PF}^A$. Therefore $\mathcal{PF}^A\subseteq\mathcal{F}^A$. Together with $c^0\in\mathcal{F}^A\setminus\mathcal{PF}^A$, this gives \begin{align*} \mathcal{PF}^A\subsetneq\mathcal{F}^A. \end{align*} Finally, because $c^0\in\mathcal{F}^A$, $d(c^0)=0$, and $\varepsilon^0=(c^0_1,c^0_2,0)\in\mathcal{E}$, Theorem (ref) (ref) implies $\varepsilon^0\in\mathcal{F}$. But $\varepsilon^0\notin\mathcal{PF}$, since $\varepsilon^0_3=0$ while $\mathcal{PF}\subset\{\varepsilon_3>0\}$. Since Lemma (ref) gives $\mathcal{PF}\subseteq\mathcal{F}$, we obtain $\mathcal{PF}\subsetneq\mathcal{F}$. \end{proof} \begin{remark}[Implications for Figure (ref)] Lemma (ref) identifies $\mathcal{PF}$ with support directions in $\tilde{\mathbb{S}}^2$ satisfying $q_3=0$, and identifies $\mathcal{F}$ with support directions in $\tilde{\mathbb{S}}^2$ satisfying $q_3\varepsilon_3\geq0$. Combining this characterization with Theorem (ref) gives the classification used in Figure (ref). Under Assumption (ref), $\mathcal{PF}=\mathcal{F}\iff\mathcal{PF}\subset\{\varepsilon\in\mathcal{E}:\varepsilon_3=0\}$. Thus, under Assumption (ref), whenever $\mathcal{PF}\not\subset\{\varepsilon_3=0\}$, one has $\mathcal{PF}\subsetneq\mathcal{F}$, while if $\mathcal{PF}\subset\{\varepsilon_3=0\}$, the two frontiers coincide. In particular, a single intersection of $\mathcal{PF}$ with the plane $\{\varepsilon_3=0\}$, or the fact that only the endpoints $R$ and $B$ lie in that plane, is not sufficient for $\mathcal{PF}=\mathcal{F}$. Corollary (ref) further shows that in the crossing case with $\mathcal{PF}$ strictly on one side of $\{\varepsilon_3=0\}$, the strict inclusions $\mathcal{PF}^A\subsetneq\mathcal{F}^A$ and $\mathcal{PF}\subsetneq\mathcal{F}$ hold already under Assumptions (ref)-(ref) alone. \end{remark} \subsection{Lemmas, Auxiliary Results, and Notation for Section (ref)} Let $\Delta A(x)\equiv A_1(x)-A_0(x)$, so that $\Delta\boldsymbol{\theta}(x;\lambda)\equiv\boldsymbol{\theta}_1(x;\lambda)-\boldsymbol{\theta}_0(x;\lambda)=\Delta A(x)-2B(x)\lambda(x)$, and $ \tau_q(x)\equiv\frac12 q^\intercal\Delta A(x)$. For each $x\in\mathcal{X}$ and for $q\in\mathbb{S}^2$, define \begin{align*} d_q(x) &\equiv\inf_{s\in[0,1]^2} \left|q^\intercal\{ \Delta A(x)-2B(x)s\} \right|,\mathcal{U}(x)\equiv\{B(x)s:s\in[0,1]^2\},\\ K_q^0(x)&\equiv\{A_0(x)+u:u\in\mathcal{U}(x),\ q^\intercal u\le\tau_q(x)\}, K_q^1(x)\equiv\{A_1(x)-u:u\in\mathcal{U}(x),\ q^\intercal u\ge\tau_q(x)\},\\ K_q(x)&\equiv K_q^0(x)\cup K_q^1(x). \end{align*} Define the finite constants $M_0\equiv\operatorname{ess}\sup_x\sup_{u\in\mathcal{U}(x)}\|\Delta A(x)-2u\|<\infty$ and \begin{align*} H_0&\equiv\operatorname{ess}\sup_x\sup_{u\in\mathcal{U}(x)}\max\{\|A_0(x)+u\|,\|A_1(x)-u\|\}<\infty, \end{align*} where the bounds follow from the binary-loss setup and the lower bound on $\mu_g,\,g\in\{r,b\}$. With the convention that $\sup\{\emptyset\}=-\infty$, define \begin{align} \kappa_0(v,q;x)&\equiv\sup_{u\in\mathcal{U}(x):\, q^\intercal u\le\tau_q(x)} v^\intercal(A_0(x)+u),\\ \kappa_1(v,q;x)&\equiv\sup_{u\in\mathcal{U}(x):\, q^\intercal u\ge\tau_q(x)} v^\intercal(A_1(x)-u). \end{align} Recall the fixed-completion feasible set $\mathcal{E}(\lambda)$ in (ref) and its associated support set, \begin{align*} \mathcal{S}_{\mathcal{E}(\lambda)}(v)&=\arg\max_{\eta\in \mathcal{E}(\lambda)}v^\intercal\eta. \end{align*} Thus $\mathcal{S}_{\mathcal{E}(\lambda)}(-q)$ is the set of minimizers of $q^\intercal\eta$ over $\mathcal{E}(\lambda)$. \begin{lemma} \textit{Let Assumptions (ref) and (ref) hold. Then, for every $\lambda\in\Lambda$, the set $\mathcal{E}(\lambda)$ is nonempty, compact, convex, has nonempty interior, and $\mathcal{S}_{\mathcal{E}(\lambda)}(v)$ is a singleton for every $v\in\mathbb{S}^2$. Moreover, $\mathcal{F}(\lambda)=\left\{\varepsilon=\mathcal{S}_{\mathcal{E}(\lambda)}(-q):q\in\widetilde{\mathbb{S}}^2,\ q_3\varepsilon_3\ge0\right\}.$ } \end{lemma} \begin{proof} By Assumption (ref), $\mathbb{P}_X\{|q^\intercal\Delta\boldsymbol{\theta}(X;\lambda)|\le\delta\}\le\mathbb{P}_X\{d_q(X)\le\delta\}\lesssim \delta^m$ and $\mathbb{P}_X\{d_q(X)=0\}=0$ for all $q\in\mathbb{S}^2$. Hence, Assumptions (ref)-(ref) hold for the distribution of $(Y^*,G,X,Z)$ implied by each $\lambda\in\Lambda$. By Lemma (ref) and Theorem (ref), the set $\mathcal{E}(\lambda)$ has nonempty interior, is compact and convex, and $\mathcal{S}_{\mathcal{E}(\lambda)}(p)$ is a singleton for every $p\in\mathbb{S}^2$. By Lemma (ref), for every $\lambda\in\Lambda$, $\mathcal{F}(\lambda)=\left\{\varepsilon=\mathcal{S}_{\mathcal{E}(\lambda)}(-q):q\in\widetilde{\mathbb{S}}^2,\ q_3\varepsilon_3\ge0\right\}$. \end{proof} \begin{lemma} \textit{Let Assumption (ref) hold. Then, for every $q\in\widetilde{\mathbb{S}}^2$, the set $\mathcal{M}_q$ is nonempty, compact, and convex.} \end{lemma} \begin{proof} Since $u(x)\in\mathcal{U}(x)=B(x)[0,1]^2$, for $j=1,2$, $u_j(x)=0$ whenever $B_{jj}(x)=0$, and $0\le u_j(x)\le B_{jj}(x)$ whenever $B_{jj}(x)>0$. Define \begin{align*} \lambda_j(x)=\begin{cases} u_j(x)/B_{jj}(x), & B_{jj}(x)>0,\\ 0, & B_{jj}(x)=0, \end{cases} \qquad j=1,2. \end{align*} Then $\lambda_j(x)\in[0,1]$. Hence, for $u(x)\in\mathcal{U}(x)$, $\mathbb{P}_X$-a.s., with $u$ measurable, there exists a measurable $\lambda\in\Lambda$, such that \begin{align} B(x)\lambda(x)=u(x)\qquad\mathbb{P}_X\text{-a.s.} \end{align} By measurability of $A_0,A_1,B$, $\Delta A$ and $\tau_q(x)=\frac12 q^\intercal\Delta A(x)$ are measurable. Since $B(x)$ has diagonal form, with nonnegative entries $B_{11}(x)$ and $B_{22}(x)$, the graph \begin{align*} \mathrm{Gr}(\mathcal{U}) = \{(x,u)\in\mathcal X\times\mathbb R^3: u_3=0,\ 0\le u_1\le B_{11}(x),\ 0\le u_2\le B_{22}(x)\} \end{align*} is measurable. Hence, for fixed $q$, \begin{align*} \mathrm{Gr}(K_q^0) = \{(x,y):(x,y-A_0(x))\in\mathrm{Gr}(\mathcal{U}),\ q^\intercal(y-A_0(x))\le\tau_q(x)\} \end{align*} is measurable, and, similarly $\mathrm{Gr}(K_q^1)$ is measurable. Therefore $K_q=K_q^0\cup K_q^1$ has a measurable graph. For each $x$, $\mathcal{U}(x)\equiv B(x)[0,1]^2$ is compact. Intersecting $\mathcal{U}(x)$ with a closed half-space preserves compactness, and translating or reflecting preserves compactness. Hence $K_q^0(x)$ and $K_q^1(x)$ are compact, possibly empty sets, and $K_q(x)\equiv K_q^0(x)\cup K_q^1(x)$ is compact. Moreover, $K_q(x)$ is nonempty because $\mathcal{U}(x)\neq\emptyset$, and every $u\in\mathcal{U}(x)$ satisfies either $q^\intercal u\le\tau_q(x)$ or $q^\intercal u\ge\tau_q(x)$. By mol17, $K_q$ is a non-empty random compact set. For each fixed $v\in\mathbb S^2$, the support function of $K_q$ in direction $v$, $\sup_{y\in K_q(x)} v^\intercal y$ is measurable mol17 and admits a measurable maximizer. The binary-loss bounds imply that $K_q$ is integrably bounded. Consequently, the following Aumann integral is well defined: \begin{align*} \left\{\mathbb E[k(X)]:k(x)\in K_q(x)\ \mathbb P_X\text{-a.s.}\right\}. \end{align*} We next show that for every $q\in\mathbb{S}^2$, we can represent $\mathcal{M}_q$ as an Aumann expectation mol:mol18: \begin{align} \mathcal{M}_q= \left\{\mathbb{E}[k(X)]:k(x)\in K_q(x)\ \mathbb{P}_X\text{-a.s.}\right\}. \end{align} To see this, take an element of $\mathcal{M}_q$. Then for some $\lambda\in\Lambda$ and some $q$-optimal rule $a$, let \begin{align*} k(x)=a(x)\boldsymbol{\theta}_1(x;\lambda)+(1-a(x))\boldsymbol{\theta}_0(x;\lambda). \end{align*} Let $u(x)=B(x)\lambda(x)$. If $a(x)=0$, the $q$-optimality condition gives \begin{align*} q^\intercal u(x)\le \frac12 q^\intercal\Delta A(x)=\tau_q(x). \end{align*} Thus $k(x)=A_0(x)+u(x)\in K_q^0(x)$. If $a(x)=1$, the same argument gives $q^\intercal u(x)\ge \tau_q(x)$, so $k(x)=A_1(x)-u(x)\in K_q^1(x)$. Hence $k(x)\in K_q(x)$ a.s., and the first inclusion follows. Conversely, let $k$ be a measurable selector of $K_q$. Since $K_q^0$ and $K_q^1$ have measurable graphs and compact values, a measurable branch-selection argument gives measurable $a:\mathcal{X}\to\{0,1\}$ and measurable $u(x)\in\mathcal{U}(x)$ such that \begin{align*} k(x)=(1-a(x))(A_0(x)+u(x))+a(x)(A_1(x)-u(x)), \end{align*} with $q^\intercal u(x)\le\tau_q(x)$ on $\{a(x)=0\}$, and $q^\intercal u(x)\ge\tau_q(x)$ on $\{a(x)=1\}$. By (ref), there exists $\lambda\in\Lambda$ with $B(x)\lambda(x)=u(x)$ a.s., and $a(x)\in\arg\min_{d\in\{0,1\}}q^\intercal\boldsymbol{\theta}_d(x;\lambda)$, so that $\mathbb{E}[k(X)]\in\mathcal{M}_q$. It follows that $\mathcal{M}_q$ is the Aumann integral of $K_q$. The correspondence $K_q$ is nonempty, measurable, compact-valued, and integrably bounded because $\mathcal{U}(x)$ is compact and because $A_0,A_1,B$ are bounded in the binary-loss setup. Hence, for each $q\in\mathbb{S}^2$, $\mathcal{M}_q$ is convex and compact by mol17, where to apply Thm. 2.1.26 we used that Assumption (ref) implies that $\mathbb{P}_X$ is nonatomic. \end{proof} \begin{lemma} \textit{Let Assumptions (ref) and (ref) hold. Then, for every fixed $\varepsilon\in{\mathbb R}^3$ and $\rho_q(\varepsilon)\equiv\sup_{v\in\mathbb{S}^2}\{v^\intercal\varepsilon-h_{\mathcal{M}_q}(v)\}$, the mapping $q\mapsto \rho_q(\varepsilon)$ is continuous on $\widetilde{\mathbb{S}}^2$. Moreover, the infimum defining $T_{\mathrm{env}}(\varepsilon)$ in (ref) is attained.} \end{lemma} \begin{proof} By Lemma (ref), for all $p,q\in\widetilde{\mathbb{S}}^2$, \begin{align*} \sup_{v\in\mathbb{S}^2}|h_{\mathcal{M}_p}(v)-h_{\mathcal{M}_q}(v)|\le 2H_0 C M_0^m\|p-q\|^m. \end{align*} Therefore, $|\rho_p(\varepsilon)-\rho_q(\varepsilon)| \le \sup_{v\in\mathbb{S}^2} |h_{\mathcal{M}_p}(v)-h_{\mathcal{M}_q}(v)|\le 2H_0 C M_0^m\|p-q\|^m$. Hence $q\mapsto\rho_q(\varepsilon)$ is continuous. Now define $Q(\varepsilon) = \{q\in\widetilde{\mathbb{S}}^2:q_3\varepsilon_3\ge0\}$. This set is nonempty because $(1,0,0)\in Q(\varepsilon)$, and it is compact because it is a closed subset of $\mathbb{S}^2$. Since $q\mapsto[\rho_q(\varepsilon)]_+$ is continuous on $Q(\varepsilon)$, it attains its minimum. Thus, $T_{\mathrm{env}}(\varepsilon)=\min_{q\in Q(\varepsilon)}[\rho_q(\varepsilon)]_+$. \end{proof} \begin{lemma} \textit{Let Assumptions (ref) and (ref) hold. Then, for every $v\in\mathbb{S}^2$, \begin{align} h_{\mathcal{M}_q}(v)=\mathbb{E}\left[\max\{\kappa_0(v,q;X),\kappa_1(v,q;X)\}\right]. \end{align} Moreover, $q_n\to q$ implies $h_{\mathcal{M}_{q_n}}(v)\to h_{\mathcal{M}_q}(v)$, and for all $p,q\in\mathbb{S}^2$, \begin{align} \sup_{v\in\mathbb{S}^2}\left|h_{\mathcal{M}_p}(v)-h_{\mathcal{M}_q}(v)\right|\le 2H_0 C M_0^m\|p-q\|^m. \end{align}} \end{lemma} \begin{proof} By mol17, the support function of $\mathcal{M}_q$ is given by \begin{align*} h_{\mathcal{M}_q}(v)=\mathbb{E}\left[\sup_{y\in K_q(X)}v^\intercal y\right]. \end{align*} Using the definition of $K_q(X),\,\kappa_0(v,q;X),\,\kappa_1(v,q;X)$, we obtain (ref). Next, fix $p,q\in\mathbb{S}^2$. For $u\in\mathcal{U}(x)$, write $z_x(u)=\Delta A(x)-2u$. As $p^\intercal u\le\frac12p^\intercal\Delta A(x)$ if and only if $ p^\intercal(\Delta A(x)-2u)\ge0$, the feasibility inequalities can be rewritten as $p^\intercal u\le\tau_p(x) \iff p^\intercal z_x(u)\ge0$ and $p^\intercal u\ge\tau_p(x)\iff p^\intercal z_x(u)\le0$. Suppose $d_q(x)>M_0\|p-q\|$. Then for every $u\in\mathcal{U}(x)$, \begin{align*} |(p-q)^\intercal z_x(u)|\le\|p-q\|\|z_x(u)\|\le M_0\|p-q\|<d_q(x)\le|q^\intercal z_x(u)|. \end{align*} Therefore $p^\intercal z_x(u)$ and $q^\intercal z_x(u)$ have the same sign for every $u\in\mathcal{U}(x)$. Hence \begin{align*} \{u\in\mathcal{U}(x):p^\intercal u\le\tau_p(x)\}&=\{u\in\mathcal{U}(x):q^\intercal u\le\tau_q(x)\},\\ \{u\in\mathcal{U}(x):p^\intercal u\ge\tau_p(x)\}&=\{u\in\mathcal{U}(x):q^\intercal u\ge\tau_q(x)\}. \end{align*} The objectives in the definitions of $\kappa_0$ and $\kappa_1$ do not depend on the normal vector. Therefore, on the event $\{d_q(x)>M_0\|p-q\|\}$, we have \begin{align*} \max\{\kappa_0(v,p;X),\kappa_1(v,p;X)\}=\max\{\kappa_0(v,q;X),\kappa_1(v,q;X)\}. \end{align*} On the complementary event, the bound defining $H_0$ gives \begin{align*} |\max\{\kappa_0(v,p;X),\kappa_1(v,p;X)\}|\le H_0, \qquad |\max\{\kappa_0(v,q;X),\kappa_1(v,q;X)\}|\le H_0. \end{align*} Thus \begin{align*} |\max\{\kappa_0(v,p;X),\kappa_1(v,p;X)\}-\max\{\kappa_0(v,q;X),\kappa_1(v,q;X)\}|\le 2H_0\mathds{1}\{d_q(x)\le M_0\|p-q\|\}. \end{align*} Taking expectations, $ \left|h_{\mathcal{M}_p}(v)-h_{\mathcal{M}_q}(v)\right|\le 2H_0\mathbb{P}_X\{d_q(X)\le M_0\|p-q\|\}. $ By Assumption (ref), $\mathbb{P}_X\{d_q(X)\le M_0\|p-q\|\}\lesssim M_0^m\|p-q\|^m$. Hence \begin{align*} \left| h_{\mathcal{M}_p}(v)-h_{\mathcal{M}_q}(v) \right| \lesssim 2H_0 M_0^m\|p-q\|^m, \end{align*} establishing (ref). The right-hand side does not depend on $v$, so taking $p=q_n$ establishes continuity of $h_{\mathcal{M}_q}$ in $q$, uniformly $v$. \end{proof} \subsection{Lemmas and Auxiliary Results for Sections (ref)-(ref)} Recall that $\mathfrak{F}_k $ denotes the $\sigma$-algebra generated by all the data except the ones with indices in $\mathcal{I}_k$. \begin{lemma} Let Assumptions (ref) and (ref) hold. Then, there exist $\widehat{\delta}_k: \mathcal{X}_2 \to {\mathbb R}^+$ and an event $\widehat{E}_k \in \mathfrak{F}_k$ such that (i) $\widehat{\delta}_k$ conditional on $\mathfrak{F}_k$ is nonstochastic, (ii) $\mathbb{P}(\widehat{E}_k^c) = o(1)$, (iii) conditional on $\widehat{E}_k$ we have that $\widehat{\boldsymbol{\mu}}_k \in \mathbb{B}_\epsilon(\mu_r,\mu_b)$, $\|\Delta \widehat{\boldsymbol{\theta}}_k(X_i)-\Delta {\boldsymbol{\theta}}(X_i)\| \le \widehat{\delta}_k(X_{2,i})$ and $\|\widehat{\pi}_k(X_i)-{\pi}(X_i)\| \le \widehat{\delta}_k(X_{2,i})$, and (iv) $\mathbb{E}[ \widehat{\delta}_k(X_{2,i})^2 | \mathfrak{F}_k] = o_p(n^{-1/2})$. \end{lemma} \begin{proof} Let $C_F = \sup_{\nu \in {\mathbb R}^{d_\nu}} \sup_{ \gamma\in {\mathbb R}^{d_\gamma}} \sup_{\mu \in \mathbb{B}_\epsilon(\mu_r,\mu_b) }\| D F(\nu, \gamma,\mu)\|$. Consider the event $\widehat{E}_k = \{ \| \widehat{\boldsymbol{\mu}}_k - \boldsymbol{\mu}\|< \epsilon\}$, which holds with probability approaching one by definition of $\widehat{\boldsymbol{\mu}}_k$. By Assumptions (ref) and (ref), the mean value inequality, and conditional on $\widehat{E}_k$, we have $$\left\| \left( \widehat{\pi}(X_i)-{\pi}(X_i)~,~\Delta \widehat{\boldsymbol{\theta}}(X_i)-\Delta {\boldsymbol{\theta}}(X_i) \right) \right\|$$ is lower than or equal to $C_F \left( \| \widehat{\nu}_k(X_{1,i}) - \nu(X_{1,i})\| +\| \widehat{\gamma}_k(X_{2,i}) - \gamma(X_{2,i})\| + \| \widehat{\boldsymbol{\mu}}_k - \boldsymbol{\mu} \|\right)~.$ By Assumption (ref), the previous expression is lower than or equal to $$ \widehat{\delta}_k(X_{2,i}) \equiv C_F \left( \sup_{x_1 \in \mathcal{X}_1} \|\widehat{\nu}_k(x_1) - \nu(x_1)\| +\| \widehat{\gamma}_k(X_{2,i}) - \gamma(X_{2,i})\| + \| \widehat{\boldsymbol{\mu}}_k - \boldsymbol{\mu} \|\right)~.$$ By Assumption (ref) and Jensen's inequality, we have $\mathbb{E}[\widehat{\delta}_k(X_{2,i})^2 | \mathfrak{F}_k] = o_p(n^{-1/2})$. \end{proof} \begin{lemma} Let Assumption (ref) and (ref) hold. Then, $ n_k^{-1/2} \sum_{i \in \mathcal{I}_k } | \widehat{\alpha}_k^h(q; X_i) - \alpha^h(q;X_i)|^2 = o_p(1)$ uniformly on $q \in \mathbb{S}^2$. \end{lemma} \begin{proof} We can write $n_k^{-1/2} \sum_{i \in \mathcal{I}_k } | \widehat{\alpha}_k^h(q; X_i) - \alpha^h(q;X_i)|^2 \le \tilde{A}_k + B_k + C_k$, where \begin{align*} \tilde{A}_k &= n_k^{-1/2} \sum_{i \in \mathcal{I}_k } |q^\intercal \Delta \boldsymbol{\theta} (X_i) | \mathds{1}\{ |q^\intercal \Delta\boldsymbol{\theta}(X_i)| \le |q^\intercal \Delta \widehat{\boldsymbol{\theta}}_k(X_i) - q^\intercal \Delta\boldsymbol{\theta}(X_i)| \} \\ B_k &= n_k^{-1/2} \sum_{i \in \mathcal{I}_k } | q^\intercal \widehat{\boldsymbol{\theta}}_{0,k}(X_i) - q^\intercal \boldsymbol{\theta}_0(X_i) |^2, C_k = n_k^{-1/2} \sum_{i \in \mathcal{I}_k } | q^\intercal \Delta \widehat{\boldsymbol{\theta}}_k (X_i) - q^\intercal \Delta \boldsymbol{\theta} (X_i) |^2. \end{align*} In the proof of Lemma (ref), we conclude $\tilde{A}_k = o_p(1)$ uniformly on $q \in \mathbb{S}^2$, with $\tilde{A}_k$ defined in (ref). Finally, note that $B_k$ and $C_k$ both are $o_p(1)$ uniformly on $q \in \mathbb{S}^2$ by Cauchy-Schwartz and Assumption (ref). \end{proof} \subsubsection{Auxiliary Results for Weighted Estimators for the Support Function} Let $\{ \omega_i: 1 \le i \le n\}$ be a random sample of positive weights that are independent of the data such that $ \mathbb{E}[\omega_i^2] \le C_\omega$. Define \begin{equation} \widehat{h}_\mathcal{E}^{\omega,*}(q; \mathbf{L}, \boldsymbol{\eta}) \equiv \tfrac{1}{n} \sum_{i =1}^n \left( \tfrac{\omega_i}{\overline{\omega}} \right) \zeta_i^*(q;\mathbf{L},{\boldsymbol{\eta}}) , \end{equation} where $\overline{\omega} = n^{-1} \sum_{i=1}^n \omega_i$ and $ \zeta_i^*(q;\mathbf{L},{\boldsymbol{\eta}}) $ is defined in (ref). \begin{lemma} \textit{Let Assumptions (ref)--(ref) hold. Then, \begin{equation} \sup_{q \in \mathbb{S}^2} n^{1/2} \left| \widehat{h}_\mathcal{E}^\omega(q; \widehat{\mathbf{L}}^\omega, \widehat{\boldsymbol{\eta}}) - \widehat{h}_\mathcal{E}^{\omega,*}(q; \mathbf{L}, \boldsymbol{\eta}) \right| = o_p(1) , \end{equation} where $\widehat{h}_\mathcal{E}^\omega(q; \widehat{\mathbf{L}}^\omega, \widehat{\boldsymbol{\eta}})$ and $\widehat{h}_\mathcal{E}^{\omega,*}(q; \mathbf{L}, \boldsymbol{\eta}) $ are defined in (ref) and (ref), respectively. } \end{lemma} \begin{proof} Let $\widehat{\alpha}_k^h$ be the estimator of $\alpha^h$ defined in (ref) using $\widehat{\boldsymbol{\eta}}_k$. Recall $n_k \equiv n/K$ and define $I_{j,n}(q) \equiv \tfrac{1}{K} \sum_{k=1}^K \tfrac{1}{n_k}\sum_{i \in \mathcal{I}_k} \left(\tfrac{\omega_i}{\overline{\omega}} \right) \mathfrak{Z}_{k,i}^{(j)}(q) $ for $j \in \{0,1,2,3\}$, where \begin{align*} \mathfrak{Z}_{k,i}^{(0)}(q) &\equiv \tfrac{q^\intercal \widehat{\mathbf{L}}_{0,i}^\omegaZ_i}{\widehat{\pi}_k (X_i)} + \left(\tfrac{q^\intercal \Delta \widehat{\mathbf{L}}_{i}^\omega Z_i}{\widehat{\pi}_k (X_i)}\right) \mathds{1}\{q^\intercal \Delta\widehat{\boldsymbol{\theta}}_k(X_i)>0\} + \widehat{\alpha}_k^h(q;X_i) \left(1 - \tfrac{Z_i}{\widehat{\pi}_k (X_i)}\right) \\ \mathfrak{Z}_{k,i}^{(1)}(q) &\equiv \tfrac{q^\intercal \widehat{\mathbf{L}}_{0,i}^\omegaZ_i}{\widehat{\pi}_k (X_i)} + \left(\tfrac{q^\intercal \Delta \widehat{\mathbf{L}}_{i}^\omega Z_i}{\widehat{\pi}_k (X_i)}\right) \mathds{1}\{q^\intercal \Delta\boldsymbol{\theta}(X_i)>0\} + \widehat{\alpha}_k^h(q;X_i) \left(1 - \tfrac{Z_i}{\widehat{\pi}_k (X_i)}\right) \\ \mathfrak{Z}_{k,i}^{(2)}(q) &\equiv \tfrac{q^\intercal {\mathbf{L}}_{0,i} Z_i}{\widehat{\pi}_k (X_i)} + \left(\tfrac{q^\intercal \Delta {\mathbf{L}}_{i} Z_i}{\widehat{\pi}_k (X_i)}\right) \mathds{1}\{q^\intercal \Delta\boldsymbol{\theta}(X_i)>0\} + \widehat{\alpha}_k^h(q;X_i) \left(1 - \tfrac{Z_i}{\widehat{\pi}_k (X_i)}\right) \\ &\quad + \sum_{g \in \{r,b\}} \Gamma_g^h(q) \left(1-\tfrac{\mathds{1} \{ G_i = g\}}{\mu_g}\right) \\ \mathfrak{Z}_{k,i}^{(3)}(q) &\equiv \zeta_i^*(q;\mathbf{L}, \boldsymbol{\eta}) . \end{align*} Our goal is to establish (ref). We divide the proof in three claims. \underline{Claim 1}: $\sup_{q \in \mathbb{S}^2 } \sqrt{n} | I_{0,n}(q) - I_{1,n}(q)| = o_p(1)$. We write $I_{0,n}(q) - I_{1,n}(q) = \tfrac{1}{K} \sum_{k=1}^K (A_k + B_k)~,$ where \begin{align} A_k &= \tfrac{1}{n_k} \sum_{i \in \mathcal{I}_k} \left( \tfrac{\omega_i}{\overline{\omega}} \right)\left(\tfrac{q^\intercal \Delta \mathbf{L}_{i} Z_i}{\pi(X_i)}\right) \left( \mathds{1}\{q^\intercal \Delta \widehat{\boldsymbol{\theta}}_k(X_i)>0\} - \mathds{1}\{q^\intercal \Delta\boldsymbol{\theta}(X_i)>0\}\right) \\ B_k &= \tfrac{1}{n_k} \sum_{i \in \mathcal{I}_k} \left( \tfrac{\omega_i}{\overline{\omega}} \right)\left( \tfrac{q^\intercal \Delta \widehat{\mathbf{L}}_{i}^\omega Z_i}{\widehat{\pi}_k(X_i)} - \tfrac{q^\intercal \Delta \mathbf{L}_{i} Z_i}{\pi(X_i)}\right) \left( \mathds{1}\{q^\intercal \Delta \widehat{\boldsymbol{\theta}}_k(X_i)>0\} - \mathds{1}\{q^\intercal \Delta\boldsymbol{\theta}(X_i)>0\}\right) \end{align} Lemma (ref) shows that $A_k = o_p(n^{-1/2})$ uniformly on $q \in \mathbb{S}^2$ by relying on Assumptions (ref) and (ref) and the approach used in the proof of Theorem 4.1 in liu:mol25v2. Using similar ideas, Lemma (ref) shows that $B_k = o_p(n^{-1/2})$ uniformly on $q \in \mathbb{S}^2$. Since $K$ is fixed as $n \to \infty$, we conclude the claim of step 1. \underline{Claim 2}: $\sup_{q \in \mathbb{S}^2 } \sqrt{n} |I_{1,n}(q) - I_{2,n}(q)| = o_p(1)$. Recall $\widehat{\mu}_g^\omega \equiv n^{-1} \sum_{i=1}^n \left( \tfrac{\omega_i}{\overline{\omega}} \right) \mathds{1}\{G_i=g\}$. Note that $n^{-1} \sum_{i=1}^n \left( \tfrac{\omega_i}{\overline{\omega}} \right) \left(1 - \tfrac{\mathds{1}\{G_i = g\} }{ \widehat{\mu}_g^\omega} \right)=0$, which implies \begin{align*} I_{1,n}(q) &= \tfrac{1}{n} \sum_{i=1}^n \left( \tfrac{\omega_i}{\overline{\omega}} \right) \left[\zeta_i(q;\widehat{\mathbf{L}}^\omega, (\Delta {\boldsymbol{\theta}},\widehat{\pi},\widehat{\alpha}^h)) + \sum_{g \in \{r,b\}} \Gamma_g^h(q) \left(1 - \tfrac{\mathds{1}\{G_i = g\} }{ \widehat{\mu}_g^\omega} \right)\right]. \end{align*} Since \begin{align*} I_{2,n}(q) &= \tfrac{1}{n} \sum_{i=1}^n \left( \tfrac{\omega_i}{\overline{\omega}} \right) \left[\zeta_i(q;{\mathbf{L}}, (\Delta{\boldsymbol{\theta}},\widehat{\pi},\widehat{\alpha}^h)) + \sum_{g \in \{r,b\}} \Gamma_g^h(q) \left(1 - \tfrac{\mathds{1}\{G_i = g\} }{ {\mu}_g} \right)\right] , \end{align*} we can write \begin{align*} I_{1,n}(q) - I_{2,n}(q) = \sum_{g \in \{r,b\}} \left( \tfrac{1}{\widehat{\mu}_g^\omega} - \tfrac{1}{\mu_g} \right) \left( \tfrac{1}{K} \sum_{k=1}^K A_k^g + B_k^g \right) , \end{align*} where \begin{align*} A_k^g &= \tfrac{1}{n_k} \sum_{i \in \mathcal{I}_k} \left( \tfrac{\omega_i}{\overline{\omega}} \right) \left\{ q^\intercal \left( \mathbf{L}_{0,i} + \Delta\mathbf{L}_i \mathds{1}\{q^\intercal \Delta\boldsymbol{\theta}(X_i)>0\} \right) \tfrac{Z_i \mu_g}{\pi(X_i)} - \Gamma_g^h(q) \right\} \mathds{1}\{G_i=g\} \\ B_k^g &= \tfrac{1}{n_k} \sum_{i \in \mathcal{I}_k} \left( \tfrac{\omega_i}{\overline{\omega}} \right) \left( \tfrac{Z_i \mu_g}{\widehat{\pi}_k(X_i)} -\tfrac{Z_i \mu_g}{\pi(X_i)} \right) q^\intercal \left( \mathbf{L}_{0,i} + \Delta\mathbf{L}_i \mathds{1}\{q^\intercal \Delta\boldsymbol{\theta}(X_i)>0\} \right) \mathds{1}\{G_i=g\} . \end{align*} We conclude that $A_k^g = O_p(n^{-1/2})$ uniformly on $q \in \mathbb{S}^2$ by using that the weights $\omega_i$ are independent of the data, $\overline{\omega} = \mathbb{E}[\omega_i] + O_p(n^{-1/2})$, Assumptions (ref) and (ref), the definition of $\Gamma_g(q)$, and the Central Limit Theorem. We also conclude that $B_k^g = o_p(1)$ uniformly on $q \in \mathbb{S}^2$ by Cauchy-Schwartz, and Assumptions (ref), (ref), and (ref). Finally, we conclude $ \tfrac{1}{\mu_g} - \tfrac{1}{\widehat{\mu}_g^\omega} = O_p(n^{-1/2})$ uniformly on $q \in \mathbb{S}^2$ by the Central Limit Theorem, Assumption (ref), and the definition of the weights $\omega_i$. \underline{Claim 3}: $\sup_{q \in \mathbb{S}^2 } \sqrt{n} \left| I_{2,n}(q) - I_{3,n}(q)\right| = o_p(1)$. We write ${\sqrt{n}} \left( I_{2,n}(q) - I_{3,n}(q)\right) = \tfrac{1}{\sqrt K} \sum_{k=1}^K (A_k + B_k+ C_k),$ where \begin{align*} A_k &= \tfrac{1}{\sqrt{n_k}} \sum_{i \in \mathcal{I}_k} \left( \tfrac{\omega_i}{\overline{\omega}} \right) \left( \tfrac{Z_i}{\widehat{\pi}_k(X_i)} - \tfrac{Z_i}{\pi(X_i)}\right) \left( q^\intercal\mathbf{L}_{0,i} + \left(q^\intercal \Delta \mathbf{L}_{i} \right) \mathds{1}\{q^\intercal \Delta\boldsymbol{\theta}(X_i)>0\} - \alpha^h(q;X_i) \right), \\ B_k &= \tfrac{1}{\sqrt{n_k}} \sum_{i \in \mathcal{I}_k} \left( \tfrac{\omega_i}{\overline{\omega}} \right) \left( \widehat{\alpha}_k^h(q;X) -\alpha^h(q;X) \right) \left( 1 - \tfrac{Z_i}{\pi(X_i)} \right), \\ C_k &= \tfrac{1}{\sqrt{n_k}} \sum_{i \in \mathcal{I}_k} \left( \tfrac{\omega_i}{\overline{\omega}} \right) \left( \widehat{\alpha}_k^h(q;X) -\alpha^h(q;X) \right) \left( \tfrac{Z_i}{{\pi}(X_i)} - \tfrac{Z_i}{\widehat{\pi}_k(X_i)}\right). \end{align*} Recall that $ (\Delta \widehat{\boldsymbol{\theta}}_k, \widehat{\pi}_k, \widehat{\boldsymbol{\theta}}_{0,k})$ was estimated using all the data except the ones with indices in $\mathcal{I}_k$, and $\mathfrak{F}_k $ is the $\sigma$-algebra generated by all the data except the ones with indices on $\mathcal{I}_k$. The definition of the weights $\omega_i$, Assumptions (ref) and (ref), the definition of $\alpha^h(q;X)$ in (ref), and algebra show $\mathbb{E}[ A_k | \mathfrak{F}_k] = 0$ and $\sup_{q \in \mathbb{S}^2} \mathbb{E}[A_k^2 | \mathfrak{F}_k] = o(1)$. Similarly, we can show $\mathbb{E}[ B_k | \mathfrak{F}_k] = 0$ and $\sup_{q \in \mathbb{S}^2} \mathbb{E}[B_k^2 | \mathfrak{F}_k] = o(1)$. By ChernozhukovDML, we conclude $A_k$ and $B_k$ both are $o_p(1)$ uniformly on $q \in \mathbb{S}^2$. Finally, we conclude that $C_k = o_p(1)$ uniformly on $q \in \mathbb{S}^2$ by Cauchy-Schwartz, Assumptions (ref) and (ref), Lemma (ref), and the definition of the weights $\omega_i$. The triangle inequality and Claims 1, 2, and 3, complete the proof of the lemma. \end{proof} \begin{lemma} Let Assumptions (ref)--(ref) hold. Then, $A_k = o_p(n^{-1/2})$ uniformly on $q \in \mathbb{S}^2$, with $A_k$ defined in (ref). \end{lemma} \begin{proof} It is sufficient to prove that $\mathbb{E}[ A_k | \mathfrak{F}_k, (X_i)_{i \in \mathcal{I}_k} ] = o_p(n^{-1/2})$ to conclude that $A_k$ is $o_p(n^{-1/2})$, by ChernozhukovDML. Define \begin{equation} \widetilde{A}_k \equiv \tfrac{1}{n_k} \sum_{i \in \mathcal{I}_k} \left| q^\intercal \Delta\boldsymbol{\theta}(X_i)\right| \mathds{1}\{ |q^\intercal \Delta\boldsymbol{\theta}(X_i)| \le |q^\intercal \Delta \widehat{\boldsymbol{\theta}}_k(X_i) - q^\intercal \Delta\boldsymbol{\theta}(X_i)| \} . \end{equation} Note that Law of Iterated Expectations, Assumption (ref), and the definitions of the weights $\omega_i$ imply $|\mathbb{E}[ A_k | \mathfrak{F}_k, (X_i)_{i \in \mathcal{I}_k} ]| \le \widetilde{A}_k$; therefore, it is sufficient to show that $\widetilde{A}_k$ is $o_p(n^{-1/2})$ uniformly on $q \in \mathbb{S}^2$. By Cauchy-Schwartz and Lemma (ref), we have $|q^\intercal \Delta \widehat{\boldsymbol{\theta}}(X_i) - q^\intercal \Delta\boldsymbol{\theta}(X_i)| \le \widehat{\delta}_k(X_{2,i})$ conditional on $\widehat{E}_k$, where $\widehat{\delta}_k $ and $\widehat{E}_k$ are as in Lemma (ref). Since $\mathbb{P}(\widehat{E}_k^c) = o(1)$, it is sufficient to show that $\mathbb{E}[ \widetilde{A}_k \mathds{1}\{ \widehat{E}_k\} | \mathfrak{F}_k ] = o_p(n^{-1/2})$ to conclude $\widetilde{A}_k$ is $o_p(n^{-1/2})$ . We verify this next: \begin{align*} \mathbb{E}[ \widetilde{A}_k \mathds{1}\{ \widehat{E}_k\} | \mathfrak{F}_k ] &= \mathbb{E} \left [ \left| q^\intercal \Delta\boldsymbol{\theta}(X_i)\right| \mathds{1}\{ |q^\intercal \Delta\boldsymbol{\theta}(X_i)| \le |q^\intercal \Delta \widehat{\boldsymbol{\theta}}_k(X_i) - q^\intercal \Delta\boldsymbol{\theta}(X_i)| \} \mathds{1}\{ \widehat{E}_k\} | \mathfrak{F}_k \right] \\ &\overset{(1)}{\le} \mathbb{E} \left [ \left| q^\intercal \Delta\boldsymbol{\theta}(X_i)\right| \mathds{1}\{ |q^\intercal \Delta\boldsymbol{\theta}(X_i)| \le \widehat{\delta}_k(X_{2,i}) \} | \mathfrak{F}_k \right] \mathds{1}\{ \widehat{E}_k\} \\ &\overset{(2)}{\le} \mathbb{E} \left [ \widehat{\delta}_k(X_{2,i}) \mathbb{E} \left [ \mathds{1}\{ |q^\intercal \Delta\boldsymbol{\theta}(X_i)| \le \widehat{\delta}_k(X_{2,i}) \} | \mathfrak{F}_k, X_{2,i} \right] | \mathfrak{F}_k \right] \mathds{1}\{ \widehat{E}_k\} \\ &\overset{(3)}\lesssim \mathbb{E} \left [ \widehat{\delta}_k(X_{2,i})^2 | \mathfrak{F}_k \right] \overset{(4)}{=} o_p(n^{-1/2}) \end{align*} where (1) holds by Lemma (ref) and since $\mathds{1}\{ \widehat{E}_k\} \in \mathfrak{F}_k$, (2) holds since $ \widehat{\delta}_k(X_{2,i})$ is nonstochastic conditional on $\mathfrak{F}_k$ and $X_{2,i}$, (3) holds by Assumption (ref) and since $ \mathds{1}\{ \widehat{E}_k\} \le 1$, and (4) holds by Lemma (ref). Finally, note that $\mathbb{P}(\widehat{E}_k^c) =o(1)$ and (4) hold uniformly on $q \in \mathbb{S}^2$ since these objects do not depend on $q$. \end{proof} \begin{lemma} Let Assumptions (ref)--(ref) hold. Then, $B_k = o_p(n^{-1/2})$ uniformly on $q \in \mathbb{S}^2$, with $B_k$ defined in (ref). \end{lemma} \begin{proof} We can write $B_k = B_{k,1} + B_{k,2} $ , where \begin{align*} B_{k,1} &= \tfrac{1}{n_k} \sum_{i \in \mathcal{I}_k} \left( \tfrac{\omega_i}{\overline{\omega} } \right) q^\intercal \Delta \mathbf{L}_{i} Z_i\left( \tfrac{1}{\widehat{\pi}_k(X_i)} - \tfrac{1}{\pi(X_i)}\right) \left( \mathds{1}\{q^\intercal \Delta \widehat{\boldsymbol{\theta}}(X_i)>0\} - \mathds{1}\{q^\intercal \Delta\boldsymbol{\theta}(X_i)>0\}\right) \\ B_{k,2} &= \tfrac{1}{n_k} \sum_{i \in \mathcal{I}_k} \left( \tfrac{\omega_i}{\overline{\omega} } \right) \left( \tfrac{q^\intercal \Delta \widehat{\mathbf{L}}_{i}^\omega Z_i - q^\intercal \Delta \mathbf{L}_{i} Z_i}{\widehat{\pi}_k(X_i)} \right) \left( \mathds{1}\{q^\intercal \Delta \widehat{\boldsymbol{\theta}}(X_i)>0\} - \mathds{1}\{q^\intercal \Delta\boldsymbol{\theta}(X_i)>0\}\right) \end{align*} As in the proof of Lemma (ref), it is sufficient to show that $\mathbb{E}[B_{k,1} | \mathfrak{F}_k, (X_i)_{i\in \mathcal{I}_k}]$ and $B_{k,2}$ are both $o_p(n^{-1/2})$ uniformly on $q \in \mathbb{S}^2$. Define \begin{align*} \widetilde{B}_{k,1} &= \tfrac{1}{n_k} \sum_{i \in \mathcal{I}_k} |\widehat{\pi}_k(X_i) - \pi(X_i)| \mathds{1}\{ |q^\intercal \Delta\boldsymbol{\theta}(X_i)| \le |q^\intercal \Delta \widehat{\boldsymbol{\theta}}_k(X_i) - q^\intercal \Delta\boldsymbol{\theta}(X_i)| \} \end{align*} Note that Law of Iterated Expectations, Assumptions (ref) and (ref), and the definitions of the weights $\omega_i$ imply that $|\mathbb{E}[B_{k,1} | \mathfrak{F}_k, (X_i)_{i\in \mathcal{I}_k}]| \le \left( \tfrac{1}{ \epsilon} \right) \widetilde{B}_{k,1}$; therefore, it is sufficient to show that $\widetilde{B}_{k,1}$ is $o_p(n^{-1/2})$ uniformly on $q \in \mathbb{S}^2$ to conclude the same for $\mathbb{E}[B_{k,1} | \mathfrak{F}_k, (X_i)_{i\in \mathcal{I}_k}]$. By similar arguments used in the proof of Lemma (ref) and by relying on Lemma (ref), $$ \mathbb{E}[ \widetilde{B}_{k,1} \mathds{1}\{\widehat{E}_k \}| \mathfrak{F}_k] \lesssim \mathbb{E} \left [ \widehat{\delta}_k(X_{2,i})^2 | \mathfrak{F}_k \right] = o_p(n^{-1/2})~,$$ where $ \widehat{E}_k$ and $\widehat{\delta}_k$ are as in Lemma (ref). By the next identity $$ q^\intercal \Delta \widehat{\mathbf{L}}_{i}^\omega - q^\intercal \Delta \mathbf{L}_{i} = \sum_{g \in \{r, b\}} \mu_g \left( \tfrac{1}{\widehat{\mu}_g^\omega} - \tfrac{1}{{\mu}_g} \right) \left( q^\intercal \Delta \mathbf{L}_i\right) \mathds{1} \{G_i = g \}~\implies B_{k,2} = \sum_{g \in \{r,b\}} \mu_g \left( \tfrac{1}{\widehat{\mu}_g^\omega} - \tfrac{1}{{\mu}_g} \right) B_{k,2}^g,$$ where \begin{align*} B_{k,2}^g &= \tfrac{1}{n_k} \sum_{i \in \mathcal{I}_k} \left( \tfrac{\omega_i}{\overline{\omega} } \right) \left( \tfrac{ \Delta \mathbf{L}_i \mathds{1} \{G_i = g \} Z_i }{\widehat{\pi}_k(X_i)} \right) \left( \mathds{1}\{q^\intercal \Delta \widehat{\boldsymbol{\theta}}_k(X_i)>0\} - \mathds{1}\{q^\intercal \Delta\boldsymbol{\theta}(X_i)>0\}\right) ,\quad g \in \{r, b\} \end{align*} We conclude that $ \mu_g \left( \tfrac{1}{\widehat{\mu}_g^\omega} - \tfrac{1}{{\mu}_g} \right) =O_p(n^{-1/2})$ uniformly on $q \in \mathbb{S}^2$ by the Central Limit Theorem, Assumption (ref), and the definition of the weights $\omega_i$. Therefore, it is sufficient to show that $B_{k,2}^g = o_p(1)$ uniformly on $q \in \mathbb{S}^2$ to conclude that $B_{k,2} = o_p(n^{-1/2})$ uniformly on $q \in \mathbb{S}^2$. Define $\widetilde{B}_{k,2}^g \equiv \tfrac{1}{n_k} \sum_{i \in \mathcal{I}_k} \mathds{1}\{ |q^\intercal \Delta\boldsymbol{\theta}(X_i)| \le |q^\intercal \Delta \widehat{\boldsymbol{\theta}}_k(X_i) - q^\intercal \Delta\boldsymbol{\theta}(X_i)| \}$. Note that Law of Iterated Expectations, Assumptions (ref) and (ref), and the definitions of the weights imply $|\mathbb{E}[ B_{k,2}^g | \mathfrak{F}_k]| \le C \widetilde{B}_{k,2}^g$, where $C$ is a constant based only on $c_1$, $c_2$, and $\epsilon$ (constants from Assumptions (ref) and (ref)). By similar arguments used in the proof of Lemma (ref) and by relying on Lemma (ref), we conclude $$ \mathbb{E}[ \widetilde{B}_{k,2}^g \mathds{1}\{\widehat{E}_k \}| \mathfrak{F}_k] \lesssim\mathbb{E} \left [ \widehat{\delta}_k(X_{2,i}) | \mathfrak{F}_k \right] = o_p(n^{-1/4})~,$$ which is sufficient to complete the proof of the lemma. \end{proof} \subsubsection{Auxiliary Results for Weighted Estimator of $\varepsilon^*$} Let $\{ \omega_i: 1 \le i \le n\}$ be a random sample of positive weights that are independent of the data such that $ \mathbb{E}[\omega_i^2] \le C_\omega$. Define \begin{align} \widehat{\varepsilon}^\omega(a^*;\widehat{\mathbf{L}}^\omega,\widehat{\boldsymbol{\eta}}) &\equiv \tfrac{1}{n} \sum_{i=1}^n \left( \tfrac{\omega_i}{\overline{\omega}} \right)\xi_i(a^*;\widehat{\mathbf{L}}^\omega,\widehat{\boldsymbol{\eta}}) \\ \widehat{\varepsilon}^{\omega,*}(a^*; \mathbf{L}, \boldsymbol{\eta}) &\equiv \tfrac{1}{n} \sum_{i =1}^n \left( \tfrac{\omega_i}{\overline{\omega}} \right) \xi_i^*(a^*;{\mathbf{L}},{\boldsymbol{\eta}}) , \end{align} where $\overline{\omega} = n^{-1} \sum_{i=1}^n \omega_i$, $\xi_i(a^*;\widehat{\mathbf{L}}^\omega,\widehat{\boldsymbol{\eta}})$, $\widehat{\mathbf{L}}^\omega$ is defined as in (ref)-(ref) with $\mu_g$ replaced by $\widehat{\mu}_g^\omega = n^{-1} \sum_{i=1}^n \left( \tfrac{\omega_i}{\overline{\omega}} \right) \mathds{1}\{G_i=g\}$, and $$\xi_i^*(a^*;\mathbf{L},\boldsymbol{\eta}) = \xi_i(a^*;\mathbf{L},\boldsymbol{\eta}) + \sum_{g \in \{r,b\}} \Gamma_g^{e}\left(1 - \tfrac{\mathds{1}\{G_i=g\}}{\mu_g} \right)$$ and $ \Gamma_g^{e} \equiv \mathbb{E}\big[ \mathds{1}\{G_i=g\} \left( \mathbf{L}_{0,i} + a(X_i)\Delta\mathbf{L}_i\right)\tfrac{Z_i}{\pi(X_i)} \big]$. \begin{lemma} \textit{Let Assumptions (ref)--(ref) hold. Then, \begin{equation} \sup_{q \in \mathbb{S}^2} n^{1/2} \left| \widehat{\varepsilon}^\omega(a^*; \widehat{\mathbf{L}}^\omega, \widehat{\boldsymbol{\eta}}) - \widehat{\varepsilon}^{\omega,*}(a^*; \mathbf{L}, \boldsymbol{\eta}) \right| = o_p(1) \end{equation} } \end{lemma} \begin{proof} Let $\widehat{\boldsymbol{\alpha}}_k^e$ be the estimator of $\boldsymbol{\alpha}^e$ defined in (ref) using $\widehat{\boldsymbol{\eta}}_k$. Define $$\widehat{\varepsilon}^{\omega,*}(a^*; \mathbf{L}, \widehat{\boldsymbol{\eta}}) \equiv \tfrac{1}{n} \sum_{i =1}^n \left( \tfrac{\omega_i}{\overline{\omega}} \right) \xi_i^*(a^*;{\mathbf{L}},\widehat{\boldsymbol{\eta}}) ~.$$ Our goal is to establish (ref). We divide the proofs in two claims. \underline{Claim 1}: $\sup_{q \in \mathbb{S}^2 } \sqrt{n} | \widehat{\varepsilon}^\omega(a^*; \widehat{\mathbf{L}}^\omega, \widehat{\boldsymbol{\eta}}) - \widehat{\varepsilon}^{\omega,*}(a^*; {\mathbf{L}}, \widehat{\boldsymbol{\eta}}) | = o_p(1)$. This claim is similar to Claim 2 in the proof of Lemma (ref) and also its proof; therefore, it is omitted. \underline{Claim 2}: $\sup_{q \in \mathbb{S}^2 } \sqrt{n} \left| \widehat{\varepsilon}^{\omega,*}(a^*; {\mathbf{L}}, \widehat{\boldsymbol{\eta}}) - \widehat{\varepsilon}^{\omega,*}(a^*; {\mathbf{L}},{\boldsymbol{\eta}}) \right| = o_p(1)$. We write ${\sqrt{n}} \left( \widehat{\varepsilon}^{\omega,*}(a^*; {\mathbf{L}}, \widehat{\boldsymbol{\eta}}) - \widehat{\varepsilon}^{\omega,*}(a^*; {\mathbf{L}},{\boldsymbol{\eta}}) \right) = \tfrac{1}{\sqrt K} \sum_{k=1}^K (A_k + B_k+ C_k)$, where \begin{align*} A_k &= \tfrac{1}{\sqrt{n_k}} \sum_{i \in \mathcal{I}_k} \left( \tfrac{\omega_i}{\overline{\omega}} \right) \left( \tfrac{Z_i}{\widehat{\pi}_k(X_i)} - \tfrac{Z_i}{\pi(X_i)}\right) \left( \mathbf{L}_{0,i} + a^*(X_i) \Delta \mathbf{L}_{i} - \boldsymbol{\alpha}^e(a^*;X_i) \right), \\ B_k &= \tfrac{1}{\sqrt{n_k}} \sum_{i \in \mathcal{I}_k} \left( \tfrac{\omega_i}{\overline{\omega}} \right) \left( \widehat{\boldsymbol{\alpha}}_k^e(a^*;X) -\boldsymbol{\alpha}^e(a^*;X) \right) \left( 1 - \tfrac{Z_i}{\pi(X_i)} \right), \\ C_k &= \tfrac{1}{\sqrt{n_k}} \sum_{i \in \mathcal{I}_k} \left( \tfrac{\omega_i}{\overline{\omega}} \right) \left( \widehat{\boldsymbol{\alpha}}_k^e(a^*;X) -\boldsymbol{\alpha}^e(a^*;X) \right) \left( \tfrac{Z_i}{{\pi}(X_i)} - \tfrac{Z_i}{\widehat{\pi}_k(X_i)}\right). \end{align*} Recall that $\widehat{\boldsymbol{\eta}}_k\equiv(\Delta \widehat{\boldsymbol{\theta}}_k, \widehat{\pi}_k, \widehat{\boldsymbol{\theta}}_{0,k})$ was estimated using all the data except the ones with indices in $\mathcal{I}_k$, and $\mathfrak{F}_k $ is the $\sigma$-algebra generated by all the data except the ones with indices on $\mathcal{I}_k$. The definition of the weights $\omega_i$, Assumptions (ref) and (ref), the definition of $\boldsymbol{\alpha}^e(a^*;X)$ in (ref), and algebra show $\mathbb{E}[ A_k | \mathfrak{F}_k] = 0$ and $\sup_{q \in \mathbb{S}^2} \mathbb{E}[A_k^2 | \mathfrak{F}_k] = o(1)$. Similarly, we can show $\mathbb{E}[ B_k | \mathfrak{F}_k] = 0$ and $\sup_{q \in \mathbb{S}^2} \mathbb{E}[B_k^2 | \mathfrak{F}_k] = o(1)$. By ChernozhukovDML, we conclude $A_k$ and $B_k$ both are $o_p(1)$ uniformly on $q \in \mathbb{S}^2$. Finally, we conclude that $C_k = o_p(1)$ uniformly on $q \in \mathbb{S}^2$ by Cauchy-Schwartz, Assumptions (ref) and (ref), and the definition of the weights $\omega_i$. The triangle inequality and Claims 1 and 2 complete the proof of the lemma. \end{proof} \begin{theorem} \textit{Let Assumptions (ref)--(ref) hold and let $\{\omega_i\}_{i=1}^n$ be random positive weights independent of the data such that $E[\omega_i] = 1$. Then, $$ \sqrt{n} \left( \widehat{\varepsilon}^\omega(a^*;\widehat{\mathbf{L}}^\omega,\widehat{\boldsymbol{\eta}}) - \varepsilon^* \right) = \mathbb{G}[\omega_i \{ \xi_i^*(a^*;\mathbf{L},\boldsymbol{\eta}) - \varepsilon^*\}] + o_p(1)~,$$ where $\xi_i^*(a^*;\mathbf{L},\boldsymbol{\eta}) = \xi_i(a^*;\mathbf{L},\boldsymbol{\eta}) + \sum_{g \in \{r,b\}} \Gamma_g^{e}\left(1 - \tfrac{\mathds{1}\{G_i=g\}}{\mu_g} \right)$ and $ \Gamma_g^{e} = \mathbb{E}\big[ \mathds{1}\{G_i=g\} \left( \mathbf{L}_{0,i} + a(X_i)\Delta\mathbf{L}_i\right)\tfrac{Z_i}{\pi(X_i)} \big]$. } \end{theorem} \begin{proof} We use Lemma (ref), the Central Limit Theorem, and similar arguments as in the proof of Theorem (ref); therefore, it is omitted. \end{proof} \begin{lemma} Let Assumptions (ref)--(ref) hold. Then, $$\left\|\sqrt{n} \left( \widetilde{\boldsymbol{\beta}}^\omega - \widehat{\boldsymbol{\beta}}^\omega \right) - \sqrt{n} \left( \widetilde{\boldsymbol{\beta}}^{\omega,*} - \widehat{\boldsymbol{\beta}}^{\omega,*} \right) \right\| = o_p(1)~,$$ where $ \widehat{\boldsymbol{\beta}}^{\omega,*}$ and $\widetilde{\boldsymbol{\beta}}^{\omega,*}$ are defined in (ref)--(ref), respectively. \end{lemma} \begin{proof} The norm $\| \cdot \|$ on $\ell^{\infty}(\mathbb{S}^2) \times \mathbb{R}^3$ is the product norm, i.e., $\| (f,\varepsilon) \| = \|f\|_\infty + \|\varepsilon\|_2$, where $\| \cdot\|_2$ is the euclidean norm on $\mathbb{R}^3$. By the triangular inequality $\left\|\sqrt{n} \left( \widetilde{\boldsymbol{\beta}}^\omega - \widehat{\boldsymbol{\beta}}^\omega \right) - \sqrt{n} \left( \widetilde{\boldsymbol{\beta}}^{\omega,*} - \widehat{\boldsymbol{\beta}}^{\omega,*} \right) \right\|$ is lower or equal than \begin{equation} \left\|\sqrt{n} \left( \widehat{\boldsymbol{\beta}}^\omega - \widehat{\boldsymbol{\beta}}^{\omega,*} \right) \right\| + \left\|\sqrt{n} \left( \widetilde{\boldsymbol{\beta}}^\omega - \widetilde{\boldsymbol{\beta}}^{\omega,*} \right) \right\| . \end{equation} Therefore, it is sufficient to show that each term above is $o_p(1)$. We use Lemmas (ref) and (ref) to conclude $ \left\|\sqrt{n} \left( \widehat{\boldsymbol{\beta}}^{\omega,*} - \widehat{\boldsymbol{\beta}}^\omega \right) \right\| = o_p(1)$. Similarly, we can conclude $ \left\|\sqrt{n} \left( \widetilde{\boldsymbol{\beta}}^{\omega,*} - \widetilde{\boldsymbol{\beta}}^\omega \right) \right\| = o_p(1)~$ by using the weights $W_i \omega_i^h$ and $W_i \omega_i^\varepsilon$ in Lemmas (ref) and (ref), respectively. This completes the proof. \end{proof}