EconBase
← Back to paper

Inference for an Algorithmic Fairness-Accuracy Frontier

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

228,691 characters · 25 sections · 104 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Inference for an Algorithmic Fairness-Accuracy Frontier

frontmatter\begin{aug} \address[id=add1]{ \orgdiv{Department of Economics}, \orgname{Cornell University}} \address[id=add2]{ \orgdiv{Department of Economics}, \orgname{Cornell University}} \end{aug} \support{This draft: June 13, 2025\\ We thank Levon Barseghyan, Gillian Hadfield, Hiroaki Kaido, Nathan Kallus, Jens Ludwig, Chuck Manski, Alice Qi, Chen Qiu, Andres Santos, Vira Semenova, Rahul Singh, Alex Tetenov, Lars Vilhuber, Davide Viviano, reviewers for the EC24 conference, seminar participants at Chicago, Cornell, Geneva, JHU, LSE, MSU, Munich, SciencesPo, Stanford, Toulouse, UCL, Warwick, EC24, ESIF: Economics and AI+ML Meeting, the 2024 Brown University workshop “Using Data to Make Decisions,” and especially Jos{\'e} Montiel-Olea and Thomas Russell for helpful comments. All data and replication files can be accessed at \href{https://github.com/yiqi-liu/TestAlgFair}{github.com/yiqi-liu/TestAlgFair}.} \begin{abstract} Algorithms are increasingly used to aid with high-stakes decision making. Yet, their predictive ability frequently exhibits systematic variation across population subgroups. To assess the trade-off between fairness and accuracy using finite data, we propose a debiased machine learning estimator for the fairness-accuracy frontier introduced by \citet*{lia:lu:mu:oku24}. We derive its asymptotic distribution and propose inference methods to test key hypotheses in the fairness literature, such as (i) whether excluding group identity from use in training the algorithm is optimal and (ii) whether there are less discriminatory alternatives to a given algorithm. In addition, we construct an estimator for the distance between a given algorithm and the fairest point on the frontier, and characterize its asymptotic distribution. Using Monte Carlo simulations, we evaluate the finite-sample performance of our inference methods. We apply our framework to re-evaluate algorithms used in hospital care management and show that our approach yields alternative algorithms that lie on the fairness-accuracy frontier, offering improvements along both dimensions. \end{abstract} \begin{keyword} \kwd{Algorithmic fairness} \kwd{statistical inference} \kwd{support function} \end{keyword}

Introduction

Algorithms are increasingly used in many aspects of life, often to guide or support high-stake decisions, for example by predicting job performance, re-offense risk, loan default, college success, or patient health. These predictions feed, respectively, into the determination of who should be hired; which defendants should receive bail; who should be granted a loan; which students should be admitted to college; and which patients to treat. Yet, a growing body of literature documents that algorithms may exhibit bias against legally protected groups, both in their predictive accuracy and in the decisions they lead to ang:lar:mat:kir16, arn:dob:hul21,obe:pow:vog:mul19,ber:jei:jab:kea:rot21. The bias may arise, for example, due to the choice of labels the algorithm is trained on, the objective function that the algorithm optimizes, the training procedure, and many other factors involved in the design of the algorithm cow:tuc20.

Designing an algorithm often entails a trade-off between making it more fair, i.e., less likely to disproportionately harm a protected class, and more accurate, e.g., better at assigning treatment to those who benefit from it and withholding it from those who do not. As a result, improving fairness often comes at the cost of accuracy. Regulators, policymakers, algorithm designers, and actors affected by algorithmic predictions all have an interest in assessing various aspects of this trade-off.

We provide a set of tools for estimation of and statistical inference on a fairness-accuracy (FA) frontier recently characterized by \citet*[LLMO henceforth]{lia:lu:mu:oku24}, where fairness is measured by the gap between group-specific expected losses. The theoretical analysis in \citetalias{lia:lu:mu:oku24} assumes perfect knowledge of the population distribution of the observable variables and formalizes the trade-off between accuracy and fairness, shading light on how to use properties of the data distribution to determine whether it is optimal for the designer of the algorithm to exclude certain inputs from use. However, in practice, regulators and policymakers typically have access to only finite data. Hence, statistical inference tools are crucial for analyzing properties of algorithms and for their regulation.

We put forward a consistent estimator for \citetalias{lia:lu:mu:oku24}'s FA-frontier and derive its asymptotic distribution. For each point on the FA-frontier, we characterize an algorithm that achieves it. We then develop a method to test hypotheses such as: Is it optimal to fully exclude group identity from use in an algorithm? Does a particular algorithm lead to group-specific expected losses that are on the FA-frontier? How far from the fairest point on the FA-frontier are the group-specific expected losses associated with a given algorithm?

Answers to the first two questions inform the regulation of algorithms and the determination of whether discrimination occurred. The law recognizes two main categories of discrimination: disparate treatment, where individuals are deliberately treated differently based on their membership in a protected class; and disparate impact, where protected classes are adversely affected disproportionately, no matter the intent kle:lud:mul:sun18,bla:spi22. Often, as part of an effort to avoid disparate treatment, algorithms are designed so that they do not take race, gender, or other sensitive attributes as input. Even class-blind algorithms, however, may lead to disparate impact. Our first test informs a fairness-minded policymaker interested in assessing whether banning group identity has the potential to mitigate disparate impact.\footnote{This question is of interest, e.g., when assessing the recent U.S. Supreme Court decision to rule out the use of race in college admissions; for an overview, see \href{https://highered.collegeboard.org/recruitment-admissions/policies-research/access-diversity/2023-scotus-decision/overview}{College Board}.}

Our second test evaluates whether a given algorithm lies on the frontier—and thus whether a less discriminatory alternative (LDA) exists. This test is relevant to both plaintiffs (e.g., job applicants) and defendants (e.g., hiring companies) in disparate impact disputes. For example, if a selection process yields disparate impact, the hiring company may invoke business necessity to justify it. The challenger must then show the existence of an LDA, i.e., a fairer algorithm that is just as accurate. If our test rejects the null that the current algorithm is on the frontier, it supports the plaintiff's claim. Conversely, if the test fails to reject, there is no statistical evidence that the hiring company can build a fairer algorithm without sacrificing accuracy, supporting the business necessity defense. When a given algorithm is not on the frontier, we characterize alternative algorithms that improve upon it in terms of accuracy or fairness (or both).

The third inferential method yields an estimator of the distance from a given algorithm to the fairest point on the frontier and constructs a confidence interval around it. This tool may interest any fairness-minded agent (e.g., a college) willing to trade some accuracy for reduced disparate impact (e.g., via affirmative action), as it provides a measure of the trade-off between promoting equity and achieving accuracy.

The key insight underlying our proposed inference methods is that since the feasible set of group-specific expected losses associated with all possible algorithms is convex, it can be fully represented by its support function. As the FA-frontier is a portion of the boundary of the feasible set, we characterize and estimate it through this support function. We express the hypotheses listed above as restrictions on the support function, yielding easy to understand test statistics that essentially rely on judicious use of the separating hyperplane theorem. Throughout our analysis, the support function serves as a unifying tool for inference on properties of algorithms and of the FA-frontier.

We provide a consistent debiased machine learning (DML) estimator of the support function and establish that it converges to a tight Gaussian process as sample size increases, building on and extending results in ber:mol08, cha:che:mol:sch18 and sem23. We show how to allow for infimum-type test statistics that are directionally-differentiable mappings of the support function, building on results of fan:san19. Earlier uses of the support function for inference in partially identified models ber:mol08, bon:mag:mau12, kai:san14, kai16, cha:che:mol:sch18,mol20,sem23 did not include tests for hypotheses such as the ones we consider. Expressing these hypotheses in terms of restrictions on the support function is one of our main contributions.

We evaluate the finite-sample performance of our inference toolkit using extensive Monte Carlo simulations. We then demonstrate its empirical value by reanalyzing algorithms for high-risk care assignment at a research hospital studied by \citet*{obe:pow:vog:mul19}. We fail to reject the hypothesis that the hospital’s status-quo algorithm admits an LDA, and document fairness and accuracy gains from several alternative algorithms on the frontier that we characterize.

Related Literature. A growing literature in computer science and statistics studies algorithmic fairness; see cho:rot18, bar:har:nar23, and Corbett24 for comprehensive overviews and open questions. Models have been developed to explain algorithmic bias by decomposing disparity sources ram:kle:lud:mul20 or incorporating taste-based discrimination and unobservables in label generation ram:rot20. Fairness has been modeled as a constraint or regularizer in the objective function that maximizes predictive accuracy dwork12,ber:hei:jab:jos:kea:mor:jam:nee:roth17 and incorporated in the preferences of a social planner that uses algorithms in their decision-making process kle:lud:mul:ram18,ram:kle:mul:lud20. In optimal policy targeting, fairness has been set as the criterion to be maximized when choosing a policy from the set of welfare-maximizing rules viv:bra23. When protected class membership is not observed in the data but proxy variables are available, data combination methods have been proposed to partially identify disparity measures kal:mao:zhou22. Yet, tests of hypotheses for properties of the trade-off between fairness and accuracy of algorithms are scant in the literature. aue:lia:tab:oku24 propose a test, which can incorporate exogenous constraints on the algorithm space, using sample splitting for the union null hypothesis that a status quo algorithm can be weakly improved in terms of both fairness and accuracy. Their test is based on finding another algorithm within a user-specified subclass, subject to the constraint that the alternative algorithm is at least as accurate as the status quo algorithm.

In contrast, our test for existence of an LDA is a one-shot test, valid across all algorithms rather than a specific subclass, that reveals if the status quo algorithm is on the FA-frontier and does not require first estimating an alternative algorithm. Characterizing the entire FA-frontier allows us to provide a comprehensive toolkit for statistical inference that can be useful for regulators to determine what algorithm design restrictions and reporting requirements to impose on entities making decisions using algorithms.

Outline. Section (ref) lays out notation and summarizes the derivation of the FA-frontier in \citetalias{lia:lu:mu:oku24}. Section (ref) characterizes the support function of interest and uses it to describe the FA-frontier. Section (ref) derives the asymptotic properties of our DML estimator for the support function. Section (ref) uses these results to obtain a consistent estimator and an asymptotically valid confidence set for the FA-frontier, and characterizes algorithms attaining points on it. Section (ref) formulates hypotheses of interest in the fairness literature as restrictions on the support function and proposes asymptotically valid tests. Section (ref) provides an estimator and inference method for the distance between the expected group-specific losses associated with a given algorithm and the fairest point on the frontier. Section (ref) presents our Monte Carlo simulations and re-evaluation of obe:pow:vog:mul19's study on a research hospital’s use of algorithms for assigning patients to a high-risk care management program. Section (ref) concludes. Our main proofs are in Appendix (ref); Appendix (ref) reports auxiliary results and extensions, and Appendix (ref) includes supplemental empirical results.

Setup

Let a population of individuals be described by an outcome $Y\in\mathcal{Y}\subset\mathbb{R}$, a binary group identity $G\in\{r,b\}$ (red or blue), and a vector of covariates $X\in\mathcal{X}\subset\mathbb{R}^{d_X}$, with the population distribution of $(Y,G,X)$ denoted $\mathbb{P}$. For example, $Y$ may denote an individual's number of active chronic illnesses in the subsequent year, $G$ may denote their race, and $X$ may include age, gender, biomarkers, comorbidity, costs and medication variables. The relation between $G$ and $X$ is left unspecified{, but $G$ is not part of $X$}. Throughout, we assume $G$ is binary, though the results extend to multiple groups. Each individual receives a binary decision $D\in\{0,1\}$, e.g., whether they are automatically enrolled in a high-risk care management program. An algorithm $a:\mathcal{X} \mapsto [0,1]$ assigns a probability distribution to $D$; e.g., the algorithm assigns each patient a health risk score in $[0,1]$, which for simplicity we take to be the only input to the {enrollment decision, and hence to coincide with the enrollment probability}. Let $\mathcal A(\mathcal{X})$ denote the set of all algorithms that map from the input space $\mathcal{X}$ to a probability distribution over $D$, and $\ell:\{0,1\}\times\mathcal{Y}\mapsto\mathbb R$ be a function that measures the loss associated with decision $d\in\{0,1\}$ for an individual with outcome $y\in\mathcal{Y}$. In the example discussed so far, the algorithm designer observes training data consisting of covariates $X$ and a binary outcome $Y$ indicating whether someone has a number of chronic illnesses exceeding a given threshold; the loss function $\ell$ may be the classification loss, $\ell(D,Y)=\mathds{1}\{D\neq Y\}$, which returns the value $1$ if the algorithm mistakenly enrolls a healthy person in the high-risk care program or fails to enroll someone who is very sick. We assume throughout that the training data is drawn from the same distribution as the population that we eventually apply the algorithm to (i.e., the subpopulation for which labels are observed is representative of the entire population).

Given an algorithm $a\in\mathcal A(\mathcal{X})$, let the population expected loss for group $g\in\{r,b\}$ be

align[align omitted — 106 chars of source]

where the expectation is taken with respect to $\mathbb{P}$. We refer to the group-specific expected loss in Eq. (ref) as group risk. Following \citetalias{lia:lu:mu:oku24}, we define a preference ordering over group risk pairs so that $e=(e_r,e_b)$ is preferred to $e'=(e_r',e_b')$, denoted $e>_{FA}e'$, if

align[align omitted — 108 chars of source]

with at least one strict inequality. As shown in \citetalias{lia:lu:mu:oku24}, all of utilitarian, Rawlsian, egalitarian, and various other preferences are consistent with this ordering. One can then define the feasible set of group risk pairs across algorithms from the class $\mathcal{A}(\mathcal{X})$ as

align[align omitted — 184 chars of source]

and the fairness-accuracy (FA) frontier as

align[align omitted — 275 chars of source]

For finite $(\mathcal{X},\mathcal{Y})$, \citetalias{lia:lu:mu:oku24} show that $\mathcal{E}\big(\mathbb{P}, \mathcal{A}(\mathcal{X})\big)$ is a closed convex set {(we extend this convexity result to general $(\mathcal{X},\mathcal{Y})$ in the proof of Proposition (ref))} and $\mathcal{F}\big(\mathbb{P}, \mathcal{A}(\mathcal{X})\big)$ is a specific portion of its boundary connecting three points: the feasible point that minimizes the risk for group $r$, denoted $R$; the feasible point that minimizes the risk for group $b$, denoted $B$, and the feasible point that minimizes the absolute difference in group risks, denoted $F$.\footnote{Ties are broken in favor of the other group's risk for $R$ or $B$. If there are multiple feasible points that minimize the absolute difference in group risks, $F$ is chosen to be the one that has the lowest risk for both groups.} Adapting \citetalias{lia:lu:mu:oku24} nomenclature, call $\big(\mathbb{P}, \mathcal{A}(\mathcal{X})\big)$ group-balanced if $\mathcal{E}\big(\mathbb{P}, \mathcal{A}(\mathcal{X})\big)$ has $R$ and $B$ such that either $R=B=F$, or $e_r < e_b$ at $R$ and $e_r > e_b$ at $B$; call $\big(\mathbb{P}, \mathcal{A}(\mathcal{X})\big)$ $r$-skewed if $e_r<e_b$ at $R$ and $e_r\leq e_b$ at $B$, and $b$-skewed if $e_r\geq e_b$ at $R$ and $e_r > e_b$ at $B$.

To ease notation, we drop the dependence of $\mathcal{E}$ and $\mathcal{F}$ on $\big(\mathbb{P}, \mathcal{A}(\mathcal{X})\big)$ unless explicitly needed. Figure (ref) illustrates these sets and key points on them under a smoothness condition stated in Assumption (ref). \citetalias{lia:lu:mu:oku24} (Theorem 1) show that the shape of $\mathcal{F}$ depends entirely on whether $\big(\mathbb{P}, \mathcal{A}(\mathcal{X})\big)$ is group-balanced or $g$-skewed. If group-balanced, $\mathcal{F}$ is the curve connecting $R$ and $B$, coinciding with the Pareto frontier (panel (a)); if $g$-skewed, $\mathcal{F}$ connects $F$ and the feasible point minimizing risk for group $g$ (panels (b) and (c) for the $r$-skewed case; omitted panels for the $b$-skewed case). \citetalias{lia:lu:mu:oku24} (Proposition 6) further show that excluding group identity as an algorithmic input is uniformly welfare-reducing under strict group balance, where $R$ and $B$ are strictly separated by the 45-degree line.

{

figure[figure omitted — 295 chars of source]

} Notation. We denote by $\|\cdot\|_E$, $\|\cdot\|_{L^2(\mathbb{P})}$, $\|\cdot\|_{\infty}$, respectively, the Euclidean norm, the $L^2$-norm under the probability measure $\mathbb{P}$, and the $L^\infty$-norm (or sup-norm). For a vector $\mathbf{a}$, let $\|\mathbf{a}\|_{L^2(\mathbb{P})}\equiv\bigl\|\|\mathbf{a}\|_E\bigr\|_{L^2(\mathbb{P})}$ and $\|\mathbf{a}\|_{\infty}$ be the supremum over the largest component of $\mathbf{a}$. For a matrix $\mathbf{A}$, let $\|\mathbf{A}\|_{\max}$ denote its max norm (the maximum absolute value among its entries). For two sequences $a_n$ and $b_n$, $a_n\lesssim b_n$ means $a_n\leq c\cdot b_n$ for some constant $c>0$.

Support Function Based Characterizations

We leverage the convexity of the feasible set $\mathcal{E}$ to characterize it by its support function and express the points $R$, $B$, $F$, and the FA-frontier $\mathcal{F}$ in Eq. (ref) through this support function. We begin by observing that Eq. (ref) and the law of iterated expectations yield

align[align omitted — 284 chars of source]

where we denote $\theta_d^g(X)\equiv\mathbb{E}[\ell(d,Y)\mathds{1}\{G=g\}|X]$ the (measurable) conditional expectation of $L_d^g\equiv\ell(d,Y)\mathds{1}\{G=g\}$ given $X$; $\mu_g\equiv \mathbb{P}(G=g)$ the population proportion of group $g\in\{r,b\}$; and the expectation in Eq. (ref) is taken with respect to the population marginal distribution of the covariates, $\mathbb{P}(X)$. To make sure that Eq. (ref) is well defined, we assume:

asm[(Moment Restrictions)] For some constants $0<c_1<1$ and $0<c_2<\infty$, $\mu_g\in(c_1,1-c_1)$ and $\operatorname{ess}\sup_{X\in\mathcal{X}}\mathbb{E}\left[\left(L_d^g\right)^2\big|X\right]<c_2$, for all $d\in\{0,1\}, g\in\{r,b\}$.

Throughout, we let $\boldsymbol{\theta}(X)\equiv[\theta_1^r(X) ~~ \theta_0^r(X) ~~ \theta_1^b(X) ~~ \theta_0^b(X)]^\intercal$ and

align[align omitted — 193 chars of source]

Support Function of the Feasible Set

Given Eqs. (ref)-(ref)-(ref), $\mathcal{E}$ can be written as

align[align omitted — 347 chars of source]

with $\operatorname{conv}(\cdot)$ the convex hull of the set in parentheses, $\Lambda(X)\equiv \operatorname{conv}\left(\{\boldsymbol{\theta}_0(X), \boldsymbol{\theta}_1(X)\}\right)$ a random interval, $\mathcal{M}$ in Eq. (ref), and $\mathbf{E}\left[\mathcal{M}\Lambda(X)\right]$ the Aumann expectation of the scaled random interval $\mathcal{M}\Lambda(X)$ mol:mol18.

As the set $\mathcal{E}$ is non-empty, compact, and convex, its support function in each direction $q=[q_1~q_2]^\intercal\in\mathbb{S}^1\equiv\{v\in\mathbb{R}^2:\, \Vert v\Vert_E=1\}$, defined as

align*[align* omitted — 81 chars of source]

uniquely characterizes $\mathcal{E}$ through the identity roc97

align[align omitted — 144 chars of source]

We next provide a closed-form expression for $h_{\mathcal{E}}(q)$.

propLet Assumption (ref) hold. Then: \begin{align} h_\mathcal{E}(q)&=\mathbb{E}\left[\max\{(\cMq)^\intercal\boldsymbol{\theta}_0(X),(\cMq)^\intercal\boldsymbol{\theta}_1(X)\}\right]\notag\\ &=\mathbb{E}\bigl[(\cMq)^\intercal\mathbf{L}_0+(\cMq)^\intercal(\mathbf{L}_1-\mathbf{L}_0)\mathds{1}\{k(\boldsymbol{\theta}(X),\cMq)>0\}\bigr], \end{align} where $k(\boldsymbol{\theta},\cMq)\equiv(\cMq)^\intercal\bigl(\boldsymbol{\theta}_1(X)-\boldsymbol{\theta}_0(X)\bigr)$, $\mathbf{L}_d\equiv[L_d^r~~L_d^b]^\intercal$, and $L_d^g\equiv\ell(d,Y)\mathds{1}\{G=g\}$.

The support function $h_\mathcal{E}(q)$ is our key inferential tool. One can garner the intuition behind its closed-form expression by rewriting $e_g(a)=\mathbb{E}\left[\tfrac{\theta_0^g(X)}{\mu_g}\right]+\mathbb{E}\left[a(X)\tfrac{\theta_1^g(X)-\theta_0^g(X)}{\mu_g}\right]$ and

align*[align* omitted — 227 chars of source]

The maximum in the above expression is achieved by the algorithm $a^{\texttt{opt}}(X;q)=\mathds{1}\{k(\boldsymbol{\theta}(X),\cMq)>0\}$, yielding Eq. (ref) upon applying the law of iterated expectations.

remarkWe allow for randomized decision rules and for $\mathcal{A}(\mathcal{X})$ to be unrestricted. If instead the family of algorithms is restricted a priori (e.g., by capacity constraints) so that $a(X)=\Pr(D=1|X)\in[\underline{a}(X),\bar{a}(X)]$, $0\le\underline{a}(X)\le\bar{a}(X)\le 1$, with $\underline{a}(\cdot),\bar{a}(\cdot)$ known functions, our analysis continues to apply by replacing $\{\boldsymbol{\theta}_0(X),\boldsymbol{\theta}_1(X)\}$ with $\{\bar{a}(X)\boldsymbol{\theta}_0(X)+(1-\bar{a}(X))\boldsymbol{\theta}_1(X),\underline{a}(X)\boldsymbol{\theta}_0(X)+(1-\underline{a}(X))\boldsymbol{\theta}_1(X)\})$. In Appendix (ref), we also show that our results continue to hold if one restricts attention to threshold rules of the form $D = \mathds{1}\{a(X) \ge 0\}$ for unrestricted $a:\mathcal{X}\mapsto\mathbb{R}$ (and in fact $a^{\texttt{opt}}(X;q)$ is a threshold rule), or to linear threshold rules where $a(X)=[1;~X]^\intercal\beta$ for some $\beta\in\mathbb{R}^{d_X+1}$, provided the space of algorithms is sufficiently rich (see Assumption (ref)).

Best Group-Specific Points on the FA-Frontier

We next define the support set of $\mathcal{E}$ in direction $q\in\mathbb{S}^1$:

align[align omitted — 161 chars of source]

i.e., $\mathcal{S}_\mathcal{E}(q)$ is the intersection between $\mathcal{E}$ and the hyperplane with normal vector $q$ and constant $h_{\mathcal{E}}(q)$, hence collecting the extreme point(s) of $\mathcal{E}$ in direction $q$. To derive a closed-form expression for $\mathcal{S}_\mathcal{E}(q)$, we impose the following assumption:

asm[(Margin Condition)] There exists $0<m\leq1$ such that for any $\delta>0$, $\sup_{q\in \mathbb{S}^1} \mathbb{P}(|k(\boldsymbol{\theta}(X),\cMq)|<\delta)\lesssim \delta^m$, with the probabilities taken with respect to $\mathbb{P}(X)$.

Assumption (ref) is a margin condition that guarantees sufficient smoothness in the distribution of $\boldsymbol{\theta}_1(X)-\boldsymbol{\theta}_0(X)$ for us to show that $h_\mathcal{E}(q)$ is differentiable in $q\in\mathbb{S}^1$ and consequently $\mathcal{S}_\mathcal{E}(q)$ includes a single element in each direction $q$ sch93. We further show that $\mathcal{S}_\mathcal{E}(q)$ equals the gradient of the support function $h_{\mathcal{E}}(\cdot)$ with respect to $q$. We denote by $\mathcal{S}_\mathcal{E}(q)$ both the singleton set and its only element.

propLet Assumptions (ref)-(ref) hold. Then, \begin{align} \nabla_qh_{\mathcal{E}}(q)= \mathbb{E}\bigl[\mathcal{M}\mathbf{L}_0+\mathcal{M}(\mathbf{L}_1-\mathbf{L}_0)\mathds{1}\{k(\boldsymbol{\theta}(X),\cMq)>0\}\bigr]= \mathcal{S}_\mathcal{E}(q), \end{align} uniformly in $q\in\mathbb{S}^1$, where $\mathcal{S}_\mathcal{E}(q)$ is uniformly continuous in $q\in\mathbb{S}^1$.

Let $\mathfrak{u}_1\equiv[-1\,\,\,0]^\intercal$ and $\mathfrak{u}_2\equiv[0\,\,\,-1]^\intercal$. The best group-specific points satisfy:

align[align omitted — 160 chars of source]
remarkAssumption (ref) allows for discrete covariates (see Proposition (ref)), but is violated if $X$ includes discrete covariates only (see Appendix (ref)), in which case we can use data jittering to satisfy Assumption (ref) by adding to one discrete covariate a small amount of smoothly distributed noise, thereby garbling that input. The feasible set constructed with a jittered covariate can be made arbitrarily close to the true feasible set cha:che:mol:sch18.
remarkThe argument in bon:mag:mau12 shows that $\mathcal{E}$ has no kinks (i.e., no support points such that there exist at least two distinct vectors $q$ and $v$ satisfying $\mathcal{S}_\mathcal{E}(q)=\mathcal{S}_\mathcal{E}(v)$) if and only if for any $q,v\in\mathbb{S}^1,q\neqv$, \begin{align} \mathbb{P}\big(k(\boldsymbol{\theta}(X),\cMq)>0,k(\boldsymbol{\theta}(X),\cMv)<0\big)>0. \end{align} If $(\boldsymbol{\theta}_1(X)-\boldsymbol{\theta}_0(X))$ admits a positive density function on a ball of positive radius that includes zero, Eq. (ref) is satisfied. Assumption (ref) in Appendix (ref) is an example of low level conditions yielding this result. The absence of kinks renders simpler limit distributions for the test statistics that we put forward in Sections (ref)-(ref). Nonetheless, Eq. (ref) is not needed for our results to apply and we provide a full treatment allowing for the presence of kinks.

Fairest Point on the FA-Frontier

{

figure[figure omitted — 847 chars of source]

} Determining the coordinates of the fairest point $F$ is more laborious, as they depend on whether $\mathcal{E}$ lies entirely above, entirely below, or on top of the 45-degree line. Figure (ref) illustrates all possible locations of the feasible set $\mathcal{E}$ relative to the 45-degree line. When $\mathcal{E}$ lies entirely on one side of the 45-degree line, as depicted in panels (b) and (d), $F$ is the support set of $\mathcal{E}$, respectively, in directions $\mathfrak{u}_2-\mathfrak{u}_1=[1\,\,-1]^\intercal$ and $\mathfrak{u}_1-\mathfrak{u}_2=[-1\,~~\,1]^\intercal$:

align[align omitted — 338 chars of source]

Complications arise when $\mathcal{E}$ intersects with the 45-degree line, as depicted in panels (a), (c), and (e) of Figure (ref). In this case, the direction at which we can obtain $F$ as the support set of $\mathcal{E}$ is difficult to determine. To circumvent this challenge, we propose a different approach. We focus on the convex set that results when $\mathcal{E}$ intersects the 45-degree line:

align[align omitted — 167 chars of source]

The new set $\Tilde{\mathcal{E}}$ is depicted in panels (a), (c), and (e) of Figure (ref) as an orange line segment. In these cases, $F$ is the support set of $\Tilde{\mathcal{E}}$ in direction $\mathfrak{u}_1$ with identical values for its two coordinates. Hence,

align[align omitted — 185 chars of source]

We are left with providing an expression for $h_{\Tilde{\mathcal{E}}}(q)$. When $\mathcal{E}$ intersects $\mathcal{H}_{45}$,

align[align omitted — 259 chars of source]

where the first equality follows from roc97, and the second follows from the fact that $h_{\mathcal{H}_{45}}(p_2)$ is bounded from above only along the direction $p_2=c[1\,\,-1]^\intercal$ for any scalar $c\in\mathbb{R}$, in which case $h_{\mathcal{H}_{45}}(p_2)=0$. Importantly, the last expression in Eq. (ref) is always well defined, regardless of whether $\mathcal{E}$ intersects with $\mathcal{H}_{45}$ or not; the infimum equals a bounded scalar in the case of intersection and $-\infty$ otherwise.\footnote{To see this, let $q(c)\equivq-c[1~~-1]^\intercal$ and note that $h_\mathcal{E}(q(c))=\|q(c)\|_E\cdoth_\mathcal{E}\big(\frac{q(c)}{\|q(c)\|_E}\big)$. When $\mathcal{E}$ intersects with $\mathcal{H}_{45}$, $h_\mathcal{E}\big(\frac{q(c)}{\|q(c)\|_E}\big)$ is bounded and nonnegative along the sequences $c\to\infty$ and $c\to-\infty$, but when $\mathcal{E}$ and $\mathcal{H}_{45}$ are disjoint, it takes negative value along one of these sequences, yielding $\inf_c h_\mathcal{E}(q(c))=-\infty$.}

Support Function-Based Characterization of the FA-Frontier

We next show that the FA-frontier put forward by \citetalias{lia:lu:mu:oku24} and reproduced in our Eq. (ref) can be characterized using the support function of the feasible set $\mathcal{E}$ and that of an auxiliary set that we introduce in this subsection.

Given an algorithm $a^*\in\mathcal{A}(\mathcal{X})$ that induces the risk pair $e^*=[e_r^*, e_b^*]^\intercal\in\mathcal{E}$, let

align[align omitted — 137 chars of source]

{

figure[figure omitted — 628 chars of source]

} denote the set of risk allocations $e\in\mathbb{R}^2$---whether or not they are feasible---that are both weakly more accurate and weakly fairer than $e^*$. The set $\mathcal{C}(e^*)$, depicted in the two panels of Figure (ref) as the shaded green regions corresponding to two different values of $e^*$, is a closed and convex subset of $\mathbb{R}^2$. Panel (a) depicts a case where $e^*\in\mathcal{F}$, whereas panel (b) depicts a case where $e^*\notin\mathcal{F}$. The key insight from the figure, which we prove can be used to characterize the FA-frontier using the support function of $\mathcal{E}$ and that of $\mathcal{C}(e^*)$, is that for any $e^*\in\mathcal{F}$, the sets $\mathcal{E}$ and $\mathcal{C}(e^*)$ can be properly separated sch93, while in the case where $e^*\notin\mathcal{F}$ they cannot.

propUnder Assumptions (ref)-(ref), $e^*\in\mathcal{F}$ if and only if there exists a hyperplane that properly separates $\mathcal{C}(e^*)$ and $\mathcal{E}$, i.e., there exists $q\in\mathbb{S}^1$ such that $h_{\mathcal{C}^*}(q)=-h_\mathcal{E}(-q)$. Let $\tilde{\mathbb{S}}^1\equiv\mathbb{S}^1\setminus \{q\in\mathbb{S}^1:q_1+q_2<0\}$ and $[\,\,\cdot\,\,]_-\equiv-\min\{\,\cdot\,, 0\}$. We then have that \begin{align} \mathcal{F} &= \left\{e^*\in\mathcal{E}: \left[\max_{q\in\tilde{\mathbb{S}}^1}(-h_{\mathcal{C}(e^*)}(q)-h_\mathcal{E}(-q))\right]_-=0\right\}. \end{align}

Remarkably, this characterization of the FA-frontier does not require knowledge of whether $\mathcal{E}$ intersects the $45^o$ line, or whether there is group balance or skew. As we show in Sections (ref)-(ref), the characterization of the FA-frontier in Eq. (ref) is very helpful for estimation and inference, and for testing whether there exists an LDA to a given algorithm.

Support Function Estimator and Its Asymptotic Distribution

In practice, the policymaker does not have perfect knowledge of $\mathbb{P}$. Hence, $h_{\mathcal{E}}(q)$ can only be estimated from a finite sample. Let a sample of size $n$, $\{(Y_i,G_i,X_i)\}_{i=1}^n$, drawn independently and identically from $\mathbb{P}$, be available. Recall Eq. (ref): $h_\mathcal{E}(q)=\mathbb{E}\bigl[(\cMq)^\intercal\mathbf{L}_0+(\cMq)^\intercal(\mathbf{L}_1-\mathbf{L}_0)\mathds{1}\{k(\boldsymbol{\theta}(X),\cMq)>0\}\bigr]$ with $k(\boldsymbol{\theta},\cMq)\equiv(\cMq)^\intercal\bigl(\boldsymbol{\theta}_1(X)-\boldsymbol{\theta}_0(X)\bigr)$, so that $\boldsymbol{\theta}(X)\equiv[\theta_1^r(X) ~ \theta_0^r(X) ~ \theta_1^b(X) ~ \theta_0^b(X)]^\intercal$ enters the expression for $h_\mathcal{E}(q)$ only through $(\theta_1^r(X)-\theta_0^r(X))$ and $(\theta_1^b(X)- \theta_0^b(X))$. We propose estimating $h_{\mathcal{E}}(q)$ by first estimating the finite dimensional parameters $\mathcal{M}$ by sample averages and the nuisance functions $\Delta\boldsymbol{\theta}(X)\equiv[(\theta_1^r(X)-\theta_0^r(X)) ~~ (\theta_1^b(X)- \theta_0^b(X))]^\intercal$ by flexible machine learning methods (allowing the complexity of the parameter space containing the estimator to grow with sample size), and then plugging their estimators, denoted $\widehat{\mathcal{M}}$ and $\widehat{\Delta\boldsymbol{\theta}}(X)$, into the sample analogue of Eq. (ref). Following the literature on debiased machine learning new94, che:che:dem:duf:han:whi:rob18, sem:che21, che:esc:ich:new:rob22, ich:new22, we show that, at the population $\mathcal{M}$, the moment in Eq. (ref) is Neyman-orthogonal and hence “insensitive” to the errors in the first-stage estimation of $\Delta\boldsymbol{\theta}$ ney79, ney95. We use sample splitting to relax the otherwise needed Donsker condition that limits the complexity of the relevant parameter space bic82, rob:li:tch:van08, rob:li:muk:tch:van17, and we account for the estimation error of $\mathcal{M}$.

Recall that $\mathbf{L}_d\equiv[L_d^r~~L_d^b]^\intercal$, with $L_d^g\equiv \ell(d,Y)\mathds{1}\{G=g\}$, and $\theta_d^g(X)=\mathbb{E}[L_d^g|X]$. Let $\DeltaL^g\equivL_1^g-L_0^g$, so that $\Delta\theta^g(X)=\mathbb{E}[\DeltaL^g|X]$. To learn the nuisance parameter $\Delta\theta^g(X)$, the effective label that, given $X$, we train machine learners to predict is $\DeltaL^g$. Let $\Theta$ denote the convex nuisance parameter space (a subset of a vector space with $L^2(\mathbb{P})$ norm) to which $\Delta\boldsymbol{\theta}=[\Delta\theta^g(X)~~\Delta\theta^b(X)]^\intercal$ belongs and $\Delta\boldsymbol{\vartheta}\equiv[\Delta\vartheta^r(X)~~\Delta\vartheta^b(X)]^\intercal$ be a generic element from $\Theta$ (to simplify notation, we drop the dependence of $\boldsymbol{\theta}$ on $X$ unless explicitly needed). For the $i$-th observation and a given $2\times2$ diagonal matrix $\mathring{\mathcal{M}}$, we note that

align[align omitted — 163 chars of source]

and using Eq. (ref) to recognize that we estimate $\Delta\boldsymbol{\theta}$ while keeping the notation as close as possible to that in Eq. (ref), we define the mapping $\zeta_i(\mathring{\mathcal{M}}q;\,\cdot\,) : \Theta \to \mathbb{R}$ as

align[align omitted — 319 chars of source]

By Proposition (ref), $h_{\mathcal{E}}(q)=\mathbb{E}[\zeta_i(\cMq;\boldsymbol{\theta})]$. We next show that, when $\mathring{\mathcal{M}}$ is fixed at the population $\mathcal{M}$, the score function $\zeta_i(\cMq;\boldsymbol{\vartheta})$ is Neyman-orthogonal at $\Delta\boldsymbol{\vartheta}=\Delta\boldsymbol{\theta}$.

propLet Assumptions (ref)-(ref) hold and $\sup_{\Delta\boldsymbol{\vartheta}\in\Theta}\|\Delta\boldsymbol{\vartheta}-\Delta\boldsymbol{\theta}\|_{L^2(\mathbb{P})}<\infty$. Then the map $\Delta\boldsymbol{\vartheta}\mapsto\mathbb{E}[\zeta_i(\cMq;\boldsymbol{\vartheta})]$ satisfies the Neyman orthogonality condition uniformly in $q\in\mathbb{S}^1$, i.e., for any $\Delta\boldsymbol{\vartheta}\in\Theta$ and scalar $t\in(0,1)$, \begin{align*} \lim_{t\to0}\sup_{q\in\mathbb{S}^1}\biggl|\frac{1}{t}\left(\mathbb{E}\bigl[\zeta_i\bigl(\cMq;\boldsymbol{\theta}+t(\boldsymbol{\vartheta}-\boldsymbol{\theta})\bigr)\bigr]-\mathbb{E}\bigl[\zeta_i(\cMq;\boldsymbol{\theta})\bigr]\right)\biggr|=0. \end{align*}

Intuitively, Proposition (ref) shows that, under its maintained assumptions, the first-order mistake in the sign of $k(\boldsymbol{\theta}(X),\cMq)$ due to the estimation error in $\Delta\boldsymbol{\theta}(X)$ is negligible.

We next show that, provided $n^{1/4}$-consistent first-stage estimators are available, the estimated support function using sample splitting and cross-fitting, as described in Definition (ref) below, converges to a Gaussian process uniformly in $q\in\mathbb{S}^1$. The proof requires showing the residual in estimating the indicator functions is bounded by the quadratic rate of convergence of the nuisance parameter, which we establish under this assumption:

asm[(Nuisance Parameter Structure)] There is a known partition of $X$, $$X=(X_1,X_2,X_{[3:d_X]}),$$ where $X_1,X_2\in\mathbb{R}$ and $X_{[3:d_X]}\in\mathbb{R}^{(d_X-2)}$ are such that $(X_1,X_2)$ has a bounded support, the density of $|k(\boldsymbol{\theta},\cMq)|$ conditional on $X_{[3:d_X]}$ is uniformly bounded in $q\in\mathbb{S}^1$, and \begin{align*} \Delta\theta^g(X)=\alpha^gX_1+\beta^gX_2+\eta^g(X_{[3:d_X]}), \end{align*} for some $\alpha^g, \beta^g\in\mathbb{R}$ satisfying $\alpha^b\cdot\beta^r\neq\alpha^r\cdot\beta^b$ and $\eta^g\in H$ for convex $H\subseteq\Theta$, $g\in\{r, b\}$.

Assumption (ref) requires $\Delta\boldsymbol{\theta}(X)$ to be linearizable in a known set of covariates $(X_1,X_2)$ and that $\alpha^b\cdot\beta^r\neq\alpha^r\cdot\beta^b$. This assures that for each $q\in\mathbb{S}^1$, $k(\boldsymbol{\theta},\cMq)$ depends on at least one of $X_1$ or $X_2$, so that we can employ a proof technique similar to that in sem23 to show that the bias induced by errors in estimating the sign of $k(\boldsymbol{\theta},\cMq)$ is bounded by the desired quadratic $L^2$-rate of the nuisance estimator, accommodating a menu of flexible machine learners in the first stage. For example, one can adapt the method in rob88 to estimate $\Delta\theta^g$. Example machine learning methods that, under suitable conditions, are $n^{1/4}$-consistent in the $L^2$-norm include $\ell_1$-penalized methods, boosting, regression trees and forests, and neural nets che:che:dem:duf:han:whi:rob18. Assumption (ref) can be eliminated at the price of using a more restrictive class of machine learning methods, by deriving rate bounds that depend on the squared $L^\infty$-rate of the first-stage estimation bel:che:fer:han17, sem23. Alternatively, Assumption (ref) can be eliminated by smoothing the indicators che:aus:syr23, par24 at the cost of introducing an additional tuning parameter that controls the degree of smoothing.\footnote{An earlier version of this paper liu:mol24v1 obtains $\sqrt{n}$-Gaussianty of the second-stage estimator without imposing Assumption (ref), but under the classical Donsker condition that limits $\Theta$'s complexity.} Importantly, Assumption (ref) implies the margin condition in Assumption (ref) with $m=1$ whenever $(X_1,X_2)$ is continuously distributed with a bounded density and $X_{[3:d_X]}$ can be either continuous or discrete; see Assumption (ref) and Proposition (ref).

We next define cross-fitting, adapting Definition 3.2 in che:che:dem:duf:han:whi:rob18:

defn[Cross-Fitting] (i) Randomly partition the size-$n$ sample with observations indexed by $i\in[n]\equiv\{1,...,n\}$ to $K\geq2$ subsamples, each of size $n/K$ (assumed to be an integer), where $K$ is a fixed integer. (ii) For each partition $k\in[K]\equiv\{1,...,K\}$ with observations indexed by the set $I_k\subset[n]$, estimate $\Delta\boldsymbol{\theta}$ by $\widehat{\Delta\boldsymbol{\theta}}_k\equiv[(\widehat{\Delta\theta}^r)_k~~(\widehat{\Delta\theta}^b)_k]^\intercal$, where each $(\widehat{\Delta\theta}^g)_k=(\widehat{\alpha}^g)_kX_1+(\widehat{\beta}^g)_kX_2+(\widehat{\eta}^g)_k(X_{[3:d_X]})$ is estimated using only observations from $I_k^c\equiv[n]\backslash I_k$. For $i\in I_k$, let $\widehat{\Delta\boldsymbol{\theta}}(X_i)\equiv\widehat{\Delta\boldsymbol{\theta}}_k(X_i)$. (iii) Let $\widehat{\mathcal{M}}=\operatorname{diag}(1/\widehat{\mu}_r,1/\widehat{\mu}_b)$, with $\widehat{\mu}_g\equiv\frac{1}{n}\sum_{i=1}\mathds{1}\{G_i=g\}$, and construct the second-stage estimator as \begin{align} \widehat{h}_{\mathcal{E}}(q; \widehat{\boldsymbol{\theta}})&=\frac{1}{K}\sum_{k\in[K]}\left(\frac{1}{n/K}\sum_{i\in I_k}\zeta_i(\hcMq;\widehat{\boldsymbol{\theta}}_k)\right)\equiv\frac{1}{n}\sum_{i=1}^n\zeta_i(\hcMq;\widehat{\boldsymbol{\theta}}). \end{align}

In Eq. (ref), to simplify notation and keep it as close as possible to that in Eq. (ref) and $h_{\mathcal{E}}(q)=\mathbb{E}[\zeta_i(\cMq;\boldsymbol{\theta})]$, we use the shorthand notation $\widehat{\Delta\boldsymbol{\theta}}(X_i)\equiv\widehat{\Delta\boldsymbol{\theta}}_k(X_i)$ for $i\in I_k$, suppress the dependence of $\widehat{h}_{\mathcal{E}}(q; \widehat{\boldsymbol{\theta}})$ on $\widehat{\mathcal{M}}$, and adapt Eqs. (ref)-(ref) to let

align[align omitted — 420 chars of source]

Our main asymptotic result shows that $\widehat{h}_{\mathcal{E}}(q; \widehat{\boldsymbol{\theta}})$ converges to a Gaussian process uniformly in the direction $q\in\mathbb{S}^1$, where the score $\zeta_i(\cMq;\boldsymbol{\vartheta})$ in Eq. (ref) evaluated at $\Delta\boldsymbol{\vartheta}=\Delta\boldsymbol{\theta}$ is the influence function that governs the part of the asymptotic distribution of $\widehat{h}_{\mathcal{E}}(q; \widehat{\boldsymbol{\theta}})$ due to $\widehat{\Delta\boldsymbol{\theta}}$, and the the remaining part is attributed to estimating $\mathcal{M}$:

theoremLet Assumptions (ref)-(ref)-(ref) hold and $\{(Y_i,G_i,X_i)\}_{i=1}^n$ be a random sample from $\mathbb{P}$. Define a shrinking neighborhood around $\Delta\boldsymbol{\theta}$ as \begin{align*} &\Theta_n\equiv\bigl\{\Delta\boldsymbol{\vartheta}\in\Theta: \forall g\in\{0,1\}, \Delta\vartheta^g(X)=\tilde{\alpha}^gX_1+\tilde{\beta}^gX_2+\tilde{\eta}^g(X_{3:d_X}),\\ &\max\{|\tilde{\alpha}^g-\alpha^g|, |\tilde{\beta}^g-\beta^g|, \|\tilde{\eta}^g-\eta^g\|_{L^2(\mathbb{P})}\}=o(n^{-1/4})\bigr\}. \end{align*} Let $\widehat{\Delta\boldsymbol{\theta}}_k\in \Theta_n$ with probability approaching 1, $\forall k\in[K]$. Then, for $\widehat{h}_{\mathcal{E}}(q; \widehat{\boldsymbol{\theta}})$ in Eq. (ref), \begin{align*} \sqrt{n}\biggl(\widehat{h}_{\mathcal{E}}(q; \widehat{\boldsymbol{\theta}})-h_{\mathcal{E}}(q)\biggr)=\mathbb{G}[ \zeta_i^*(\cMq;\boldsymbol{\theta})]+o_p(1) \quad in \,\, \ell^\infty(\mathbb{S}^1), \end{align*} where $\mathbb{G}[\zeta_i^*(\cMq;\boldsymbol{\theta})]$ is a Gaussian process in $\ell^\infty(\mathbb{S}^1)$ indexed by \begin{align*} \zeta_i^*(\cMq;\boldsymbol{\theta})\equiv\zeta_i(\cMq;\boldsymbol{\theta})+(\mathcal{M}_i^*q)^\intercal\mathcal{M}^{-1}\mathcal{S}_\mathcal{E}(q), for \mathcal{M}_i^*\equiv\operatorname{diag}\left(\frac{\mathds{1}\{G_i=r\}}{-\mu_r^2}, \frac{\mathds{1}\{G_i=b\}}{-\mu_b^2}\right) \end{align*} with $\zeta_i(\cMq;\boldsymbol{\theta})$ defined in Eq. (ref) and the covariance function of $\mathbb{G}[\zeta_i^*(\cMq;\boldsymbol{\theta})]$ equal to \begin{align} \Omega(q, \Tilde{q})=\mathbb{E}[\zeta_i^*(\cMq;\boldsymbol{\theta})\zeta_i^*(\mathcal{M}\Tilde{q};\boldsymbol{\theta})]-\mathbb{E}[\zeta_i^*(\cMq;\boldsymbol{\theta})]\mathbb{E}[\zeta_i^*(\mathcal{M}\Tilde{q};\boldsymbol{\theta})]. \end{align} If $Var(\mathbf{L}_d|X)$ is positive definite, then $Var(\mathbb{G}[\zeta_i^*(\cMq;\boldsymbol{\theta})])>0$ for each $q\in\mathbb{S}^1$.

Estimation and Inference for the Frontier

Estimation and Inference for the Feasible Set

We propose an estimator of the set $\mathcal{E}$ based on $\widehat{h}_\mathcal{E}(q)$ and Eq. (ref), given by

align[align omitted — 198 chars of source]

which is convex, almost surely compact, and non-empty with probability approaching one if $\mathcal{E}$ has a non-empty interior. As the Hausdorff distance between two non-empty convex and compact sets $A,B\in\mathbb R^d$, denoted $\mathbf{d}_H(A,B)$, equals the uniform distance between their support functions mol:mol18, when $\widehat{\mathcal{E}}$ is non-empty we have

align*[align* omitted — 199 chars of source]

By Theorem (ref) and the continuous mapping theorem, $\mathbf{d}_H(\widehat{\mathcal{E}},\mathcal{E})\xrightarrow[]{p} 0$ and $\sqrt{n}\mathbf{d}_H(\widehat{\mathcal{E}},\mathcal{E})\xrightarrow[]{d}\sup_{q\in\mathbb{S}^1}|\mathbb{G}[ \zeta_i^*(\cMq;\boldsymbol{\theta})]|$. Hence, asymptotically valid tests of hypotheses about $\mathcal{E}$, and confidence sets covering it, can be obtained as in ber:mol08.

Estimation and Inference for the FA-Frontier

As shown in \citetalias{lia:lu:mu:oku24} (Theorem 1), when $\big(\mathbb{P}, \mathcal{A}(\mathcal{X})\big)$ is group-balanced, $\mathcal{F}$ is the curve connecting $R$ and $B$ and coincides with the Pareto frontier (panel (a) in Figure (ref)). However, when $\big(\mathbb{P}, \mathcal{A}(\mathcal{X})\big)$ is $g$-skewed, $\mathcal{F}$ is the curve connecting $F$ with the feasible point that minimizes the risk for group $g$ (panels (b)-(e) in Figure (ref)). A further challenge is that, while it is simple to express $F$ through $\mathcal{S}_\mathcal{E}$ when $\mathcal{E}$ is fully contained in one of the two half-spaces defined by the 45-degree line $\mathcal{H}_{45}$, as shown in Eqs. (ref)-(ref) (panels (b) and (d) in Figure (ref)), characterizing $F$ is not as straightforward when $g$-skew occurs but $\mathcal{E}\cap\mathcal{H}_{45}\neq\emptyset$.

We therefore leverage the characterization of the FA-frontier $\mathcal{F}$ in Proposition (ref), whereby $\mathcal{F} = \left\{e\in\mathcal{E}:~\left[\max_{q\in\tilde\mathbb{S}^1}(-h_{\mathcal{C}(e)}(q)-h_\mathcal{E}(-q))\right]_-=0\right\}$, and the definition of the set $\mathcal{C}(\cdot)$ in Eq. (ref) to sidestep these difficulties. We propose to estimate $\mathcal{F}$ using

align[align omitted — 364 chars of source]

where $[\,\,\cdot\,\,]_+\equiv\max\{\,\cdot\,, 0\}$, $\kappa_n=o(\sqrt{n})$ is a sequence that diverges to infinity, and $\mathbb{B}_C\equiv\{e\in\mathbb{R}^2:\Vert e\Vert_E \le C\}$, with $\mathcal{E}\subset \mathbb{B}_C$ by Assumption (ref) and $C<\infty$ a constant pinned down by $c_1,c_2$ defined in Assumption (ref). The first maximization problem in Eq. (ref) is the sample analog to the requirement that $e\in\mathcal{E}$; the second is the sample analog to the requirement that $\mathcal{E}$ and $\mathcal{C}(e)$ can be properly separated. To implement this estimator, we derive a closed-form expression for the support function of $\mathcal{C}(e)$ in direction $q=[q_1,~q_2]^\intercal$. Only two points are “active” for evaluating $h_{\mathcal{C}(e)}(q)$: $[\min\{e_r,2e_b-e_r\},\,\,e_b]^\intercal$ and $[e_r,\,\,\min\{e_b, 2e_r-e_b\}]^\intercal$, which correspond to the points $e$ and $[e_r,\,\,2e_r-e_b]^\intercal$ (respectively, $[2e_b-e_r,\,\,e_b]^\intercal$ and $e$) when $e$ is above (respectively, below) the 45-degree line as shown in Figure (ref)-panel (a) (respectively, panel (b)). By Proposition (ref), it is without loss of generality to focus on $q\in\tilde{\mathbb{S}}^1$. Hence, the support function of $\mathcal{C}(e)$ at any $q\in\tilde{\mathbb{S}}^1$ equals

align[align omitted — 149 chars of source]

We next propose a test statistic for the null hypothesis $\mathsf{H}_0:~e\in\mathcal{F}$, against the alternative $\mathsf{H}_A:~e\notin\mathcal{F}$. Our test statistic is given by

align[align omitted — 324 chars of source]

As shown in the proof of Proposition (ref), if Eq. (ref) is satisfied and consequently $\mathcal{E}$ has no kinks, the large sample distribution of $T_n^\mathcal{F}(e)$, denoted $\psi^{\mathcal{F}}(e)$, simplifies to

align[align omitted — 157 chars of source]

where $q^*_{\mathbb{S}^1}(e)\equiv\arg\max_{q\in\mathbb{S}^1}q^\intercal e -h_\mathcal{E}(q)$. The quantiles of this distribution can be estimated by standard methods. We build a confidence set for the elements of $\mathcal{F}$ by test inversion:

align[align omitted — 147 chars of source]

where for any $\beta\in(0,1)$, $c_\beta^{\mathcal{F}}$ is the $\beta$-quantile of $\psi^{\mathcal{F}}$. When Eq. (ref) is not assumed and $\mathcal{E}$ has kinks, the expression for $\psi^{\mathcal{F}}(e)$ is more complex and given in Eq. (ref). In this case, restrictions that would guarantee uniform continuity and strict increasing properties for $\psi^{\mathcal{F}}(e)$ are harder to verify, and we follow and:shi13 to replace the confidence set in Eq. (ref) with $\mathcal{CS}_n(\mathcal{F}) = \left\{e\in\mathbb{B}_C:T_n^\mathcal{F}(e)\le c_{1-\alpha+\varsigma}^{\mathcal{F}}(e)+\varsigma\right\}$, for $\varsigma>0$ an arbitrarily small constant. In this case, the critical value can be approximated through bootstrap methods, as in Procedure (ref) in Section (ref).

We next establish consistency of our estimator $\widehat{\mathcal{F}}$ and validity of the confidence set, building respectively on che:hon:tam07 and fan:san19.

propLet the assumptions of Theorem (ref) hold and let $\kappa_n\to\infty$ with $\kappa_n=o(\sqrt{n})$. Then, as $n\to\infty$, \begin{align} \mathbf{d}_H(\widehat{\mathcal{F}},\mathcal{F})&\xrightarrow[]{p} 0\\ \liminf_{n\to\infty}\mathbb{P}\left(e\in\mathcal{CS}_n(\mathcal{F})\right)&\ge 1-\alpha for all e\in\mathcal{F}. \end{align}

Estimation and Inference for the Pareto Frontier

The Pareto Frontier ($\mathcal{PF}$) is the lower boundary of the feasible set $\mathcal{E}$ connecting points $R$ and $B$ (see Figure (ref)). Denoting $\mathbb{Q}\equiv\left\{q\in\mathbb{S}^1:q=[\cos\gamma~~\sin\gamma]^\intercal,~\gamma\in[\pi,3/2\pi]\right\}$, we can express $\mathcal{PF}$ through the support set $\mathcal{S}_\mathcal{E}(\cdot)$, or equivalently through the support function:

subequations\begin{align} \mathcal{PF}&\equiv\left\{\mathcal{S}_\mathcal{E}(q):q\in\mathbb{Q}\right\}\\ &=\left\{ e\in\mathbb{R}^2: \left[\max_{q\in\mathbb{S}^1}(q^\intercal e - h_\mathcal{E}(q))\right]_+= 0,\,\left[\max_{q\in\mathbb{Q}}(q^\intercal e - h_\mathcal{E}(q))\right]_-=0\right\}, \end{align}

where the first condition in Eq. (ref) enforces that $e\in\mathcal{E}$ and the second that it belongs to the supporting hyperplane of $\mathcal{E}$ in a direction $q\in\mathbb{Q}$. Knowing the directions that determine the support points comprising $\mathcal{PF}$ simplifies the expression for it in Eq. (ref) relative to Eq. (ref). It also allows for an estimator, put forward in Proposition (ref) below, based on the support set characterization in Eq. (ref) that is free from the tuning parameter $\kappa_n$, which instead is needed for the estimator of $\mathcal{F}$ in Eq. (ref). Let $\boldsymbol{\zeta}_{\mathcal{S},i}(\hcMq;{\widehat{\boldsymbol{\theta}}})\equiv\widehat{\mathcal{M}}\mathbf{L}_{0_i}+\widehat{\mathcal{M}}(\mathbf{L}_{1_i}-\mathbf{L}_{0_i})\cdot\mathds{1}\bigl\{k\bigl({\widehat{\boldsymbol{\theta}}(X_i)},\hcMq\bigr)>0\bigr\}$, with $k\bigl(\widehat{\boldsymbol{\theta}}(X_i),\hcMq\bigr)=q^\intercal\widehat{\mathcal{M}}\widehat{\Delta\boldsymbol{\theta}}(X_i)$ as in Eq. (ref) and\footnote{As in Eq. (ref), this expression is a shorthand for $\widehat{\mathcal{S}}_\mathcal{E}(q;\widehat{\boldsymbol{\theta}})\equiv\frac{1}{K}\sum_{k\in[K]}\left(\frac{1}{n/K}\sum_{i\in I_k}\boldsymbol{\zeta}_{\mathcal{S},i}({\widehat{\mathcal{M}}}q;\widehat{\boldsymbol{\theta}}_k)\right)$.}

align[align omitted — 215 chars of source]

Proposition (ref) delivers an estimator for $\mathcal{PF}$ based on $\widehat{\mathcal{S}}_\mathcal{E}(q;\widehat{\boldsymbol{\theta}})$ in Eq. (ref) and establishes its Hausdorff-consistency.

propLet the assumptions of Theorem (ref) hold. Let $\widehat{\mathcal{PF}}\equiv\left\{\widehat{\mathcal{S}}_\mathcal{E}(q;\widehat{\boldsymbol{\theta}}):q\in\mathbb{Q}\right\}$. Then, as $n\to\infty$, \begin{align} \max_{q\in\mathbb{Q}}\Vert \widehat{\mathcal{S}}_\mathcal{E}(q;\widehat{\boldsymbol{\theta}})-\mathcal{S}_\mathcal{E}(q)\Vert_E &\xrightarrow[]{p} 0,\\ \mathbf{d}_H(\widehat{\mathcal{PF}},\mathcal{PF})&\xrightarrow[]{p} 0. \end{align}

We next provide a method to test, for a given $e\in\mathbb{R}^2$,

align[align omitted — 117 chars of source]

using the characterization of $\mathcal{PF}$ in Eq. (ref). We first propose a test statistic and derive its asymptotic distribution in Proposition (ref) below, building on results in kai16.

remarkWe use different characterizations of $\mathcal{PF}$ for estimation and for inference due to difficulties with DML estimation of $\mathcal{S}_\mathcal{E}(q)$ that we explain here. Using Eq. (ref) we can write the first (second) coordinate of $\mathcal{S}_\mathcal{E}(q)$ as $\mathbb{E}\big[(\cMv)^\intercal\mathbf{L}_0+(\cMv)^\intercal(\mathbf{L}_1-\mathbf{L}_0)\mathds{1}\{k(\boldsymbol{\theta}(X),\cMq)>0\}\big]$ for $v=[1~~0]^\intercal$ ($v=[0~~1]^\intercal$). As $k(\boldsymbol{\theta}(X),\cMv)=\mathbb{E}[(\cMv)^\intercal(\mathbf{L}_1-\mathbf{L}_0)|X]$, the proof of Theorem (ref) shows that, if $v=q$, the first-stage estimation error in the sign of $k(\boldsymbol{\theta}(X),\cMq)$ is controlled by the size of the error itself (because in case of sign disagreement, $|k(\boldsymbol{\theta}(X),\cMq)|\le|k(\widehat{\boldsymbol{\theta}}(X),\cMq)-k(\boldsymbol{\theta}(X),\cMq)|$), with $k(\widehat{\boldsymbol{\theta}}(X),\cMq)$ as in Eq. (ref). However, when $v\neqq$, which necessarily occurs for at least one coordinate of $\mathcal{S}_\mathcal{E}(q)$, sign errors are not controlled for. Consequently, we switch to the moment inequality characterization in Eq. (ref) that involves $h_\mathcal{E}(q)$ only. If one uses a parametric estimator for $\Delta\boldsymbol{\theta}(X)$, the asymptotic distribution of $\widehat{\mathcal{S}}_\mathcal{E}(\cdot;\widehat{\boldsymbol{\theta}})$ can be obtained and inference is simplified, as a special case of the treatment in liu:mol24v1, who work with a sieve nonparametric estimator of $\Delta\boldsymbol{\theta}(X)$ under Donsker conditions.
propLet $Q_{\mathbb{T}}^*(e)\equiv\arg\max_{q\in\mathbb{T}}q^\intercal e -h_\mathcal{E}(q),~\mathbb{T}\in\{\mathbb{S}^1,\mathbb{Q}\}$. Then, under the null in Eq. (ref) and the assumptions of Theorem (ref), \begin{align} T_n^\mathcal{PF}(e)&\equiv \sqrt{n}\left(\left[\max_{q\in\mathbb{S}^1}q^\intercal e-\widehat{h}_\mathcal{E}(q;\widehat{\boldsymbol{\theta}})\right]_{+} + \left[\max_{q\in\mathbb{Q}}q^\intercal e-\widehat{h}_\mathcal{E}(q;\widehat{\boldsymbol{\theta}})\right]_-\right)\\ &\xrightarrow[]{d}\left[\sup_{q\in Q^*_{\mathbb{S}^1}(e)}\mathbb{G}[-{\zeta_i^*(\cMq;\boldsymbol{\theta})}]\right]_+ + \left[ \sup_{q\in Q^*_{\mathbb{Q}}(e)}\mathbb{G}[-{\zeta_i^*(\cMq;\boldsymbol{\theta})}]\right]_-. \end{align} If $Var(\mathbf{L}_d|X)$ is positive definite for each $d\in\{0,1\},~X-a.s.$, the limit law in Eq. (ref) is absolutely continuous with respect to Lebesgue measure on $\mathbb{R}_{++}$.

If Eq. (ref) is satisfied and the set $\mathcal{E}$ has no kinks (see Remark (ref)), $Q_{\mathbb{S}^1}^*(e)$ and $Q_{\mathbb{Q}}^*(e)$ are singletons that can be consistently estimated through standard methods, and the limit distribution in Eq. (ref) coincides with that in Eq. (ref). If instead kinks are not ruled out for $q\in\mathbb{Q}$, $Q_{\mathbb{S}^1}^*(e)$ and $Q_{\mathbb{Q}}^*(e)$ may not be singletons. In this case we can consistently estimate these sets kai16 as

align[align omitted — 315 chars of source]

where $\mathbb{T}\in\{\mathbb{S}^1,\mathbb{Q}\}$ and $\kappa_n=o(\sqrt{n})$ is a sequence that diverges to infinity. We use Procedure (ref)-Step 1 to obtain a valid bootstrap-based approximation to the Gaussian process in Theorem (ref). Denote $\mathbb{P}^*$ this bootstrap distribution conditional on $\{(Y_i,G_i,X_i)\}_{i=1}^n$. Let

align[align omitted — 344 chars of source]

When $Var(\mathbf{L}_d|X)$ is positive definite, it follows that if $\alpha\in(0,0.5)$,

align*[align* omitted — 222 chars of source]

kai16. A confidence set that covers each $e\in\mathcal{PF}$ with asymptotic probability at least equal to $(1-\alpha)$ can be obtained by test inversion.

remarkThe result in Proposition (ref) can be adapted to testing whether $e^*\in\mathcal{PF}$ when $e^*$ needs to be estimated, by adjusting the covariance function of the limit Gaussian process in Theorem (ref) and using the test statistic in Eq. (ref). If the limit law in Eq. (ref) is not guaranteed to be absolutely continuous on $\mathbb{R}_{++}$, one can use infinitesimal adjustments to the critical value to maintain asymptotic validity, as in kai16.

Algorithms Yielding a Risk Allocation on the Frontier

An algorithm designer or a regulator may wonder if one can characterize the algorithm yielding a specific point on the FA-frontier $\mathcal{F}$ or the Pareto frontier $\mathcal{PF}$. It turns out that, using our support function approach, the answer to this question is affirmative, and particularly simple for points in $\mathcal{PF}$.

Suppose we have two data sets: one for training the algorithm and the other for evaluating it. We denote the training sample as $\{(\tilde{Y}_i,\tilde{G}_i,\tilde{X}_i)\}_{i=1}^{n_1}$ and the evaluation sample as $\{(Y_j,G_j,X_j)\}_{j=1}^{n_2}$, with $\frac{n_1}{n_2}\to c$ for some positive constant $c$ and both samples drawn from the same distribution $\mathbb{P}$. Use the training data to estimate $\Delta\boldsymbol{\theta}$ via machine learning and $\mathcal{M}$ by sample averages as described in Section (ref) and denote the resulting estimators $\widehat{\Delta\boldsymbol{\theta}}_{n_1}$ and $\widehat{\mathcal{M}}_{n_1}$, where the subscript $n_1$ indicates that it is based on the training sample. Let

align[align omitted — 167 chars of source]

with $k\bigl(\widehat{\boldsymbol{\theta}}_{n_1}(X_j),\widehat{\mathcal{M}}_{n_1}q\bigr)$ as in Eq. (ref). Note that $\widehat{a}_{n_1}(X_j;q)$ only takes as input the covariates, $X_j$, and does not depend on group identity $G_j$. Consequently, individuals with the same covariates are assigned the same treatment irrespective of their group. Nonetheless, information on group identity contained in the training data is used in the prediction model (to separately estimate $\Delta\theta^g$, $g\in\{r,b\})$, thereby offering a compromise between a utilitarian perspective as in man:mul:ven23 which advocates for group-aware decisions, and proponents of group-blind decision making.

{Then, for any $q\in\mathbb{S}^1$, algorithm $\widehat{a}_{n_1}(\,\cdot\,;q)$ leads to a loss for each observation $i$ in the evaluation sample equal to $\ell(0,Y_j)+(\ell(1,Y_j)-\ell(0,Y_j))\cdot\mathds{1}\bigl\{k\bigl(\widehat{\boldsymbol{\theta}}_{n_1}(X_j),q\bigr)>0\bigr\}$. Therefore, the average losses for the $r$ group and for the $b$ group in the evaluation sample equal

align[align omitted — 513 chars of source]

} By the same argument as in the proof of Proposition (ref), it follows that

align*[align* omitted — 216 chars of source]

Hence, the algorithm in Eq. (ref) with $q\in\mathbb{Q}$ gives consistent estimators of points on $\mathcal{PF}$.

Algorithms that return consistent estimators of points on $\mathcal{F}$ (other than those on $\mathcal{PF}$) are harder to obtain, as one needs to determine which direction $q$ corresponds to a point $e\in\mathcal{F}$ and plug that direction in the algorithm in Eq. (ref). Suppose Eq. (ref) is satisfied and $\mathcal{E}$ has no kinks. Use the same DML construction in Section (ref) applied to the training sample to obtain an estimator of $h_\mathcal{E}(q)$, denoted $\widehat{h}_{\mathcal{E},n_1}(q;\widehat{\boldsymbol{\theta}}_{n_1})$, as in Eq. (ref), and an estimator of $\mathcal{F}$, denoted $\widehat{\mathcal{F}}_{n_1}$, as in Eq. (ref). {Select $\widehat{e}_{n_1}\in\widehat{\mathcal{F}}_{n_1}$ through a projection method detailed in the proof of Proposition (ref) so that $\Vert\widehat{e}_{n_1}-e\Vert=o_p(1)$ for some $e\in\mathcal{F}$ and let}

align[align omitted — 204 chars of source]

Let $\widehat{\mathcal{S}}_{\mathcal{E},n_2}(\widehat{q}^*_{n_1}(\widehat{e}_{n_1});\widehat{\boldsymbol{\theta}}_{n_1})$ be as in Eq. (ref) but with $\widehat{q}^*_{n_1}(\widehat{e}_{n_1})$ replacing $q$. The next proposition establishes that $\widehat{\mathcal{S}}_{\mathcal{E},n_2}(\widehat{q}^*_{n_1}(\widehat{e}_{n_1});\widehat{\boldsymbol{\theta}}_{n_1})$ is a consistent estimator of $\mathcal{S}_\mathcal{E}(q^*_{\mathbb{S}^1}(e))\in\mathcal{F}$.

propLet the assumptions of Theorem (ref) hold. Then, as $n_1,n_2\to\infty$, \begin{align*} \left\Vert\widehat{\mathcal{S}}_{\mathcal{E},n_2}(\widehat{q}^*_{n_1}(\widehat{e}_{n_1});\widehat{\boldsymbol{\theta}}_{n_1})-\mathcal{S}_\mathcal{E}(q^*_{\mathbb{S}^1}(e))\right\Vert \xrightarrow[]{p} 0. \end{align*}

Regardless of whether one aims at obtaining points in $\mathcal{PF}$ or the entire $\mathcal{F}$, the direction $q$ can be interpreted as the vector of weights that the agent choosing the algorithm puts on each group's risk. In other words, one may think of the agent as evaluating group risks according to the welfare loss function $U(e;q)\equivq_1 e_r(a)+q_2 e_b(a)$. For example, the more the agent cares about group $r$, the closer $q$ is to $\mathfrak{u}_1$. We note that $\widehat{a}_{n_1}(X;q)$ is an empirical success rule, and leave its statistical decision theory analysis to future research.

Hypothesis Testing

In this Section we propose hypothesis tests to answer the following policy questions:

(1) Should the policymaker consider banning group identity as an input to the algorithm?

(2) Is there a less discriminatory alternative (LDA) to an existing algorithm?

We show how to express the first policy question in terms of restrictions on $h_\mathcal{E}(q)$, and we leverage Proposition (ref) to do the same for the second policy question. We then establish asymptotic validity of the corresponding testing procedures.

How to Test Whether Group Identity Should be Banned

\citetalias{lia:lu:mu:oku24} (Proposition 6) show that using $X$ only instead of $(X,G)$ as algorithmic input uniformly worsens the frontier if $R$ and $B$, obtained when only $X$ is used as input, are strictly separated by the 45-degree line.\footnote{“Uniformly worsening the frontier” means that, under the preference relations defined in Eq. (ref), every point on the frontier $\mathcal{F}\big(\mathbb{P},\mathcal{A}(\mathcal{X})\big)$ is dominated by a point on the frontier $\mathcal{F}\big(\mathbb{P},\mathcal{A}(\mathcal{X}\times\{r,b\})\big)$.} We therefore aim at testing the null hypothesis that $R$ and $B$ lie weakly on the same side of the 45 degree line, i.e., that the difference in the two coordinates of $R$ has the same sign as that of $B$, against the alternative that they are strictly separated by the 45-degree line:

align[align omitted — 304 chars of source]

If the null in Eq. ((ref)) is rejected, the policymaker should not ban group identity $G$ as algorithm's input. A Type-I error amounts to the case where one concludes that there is strict group-balance, while instead weak group-skew holds. As a consequence, one does not ban $G$ as input to the algorithm, thinking that banning $G$ is uniformly welfare-reducing, while instead depending on the preferences of the designer it might not be the case. Another interpretation of this test amounts to determining, based on whether $\mathsf{H}_0$ is rejected or not, if one can justify implementing algorithms that lead to Pareto-dominated risks based on the designer's preference over fairness and accuracy. When $\mathsf{H}_0$ holds true, this justification is possible, as the frontier includes an upward-sloping segment (e.g., Panels (b)-(c) of Figure (ref)). On the other hand, when $\mathsf{H}_A$ holds true, such justification is untenable, as the frontier coincides with the Pareto frontier (e.g., Panel (a) of Figure (ref)). Hence, a Type-I error can also be interpreted as a case where one concludes that only Pareto-optimal risks should be implemented, while fairness considerations may justify Pareto-dominated risks.

{As discussed in Remark (ref), carrying out inference if we use DML to directly estimate both coordinates of the points $R$ and $B$ is difficult.} Yet, using Eqs. (ref) and (ref), we can represent these points through moment equalities and inequalities that involve $h_\mathcal{E}(q)$ only:

align[align omitted — 466 chars of source]

where in Eq. (ref), the equality constraint for $R$ restricts it to have horizontal coordinate equal to that of $\mathcal{S}_\mathcal{E}(\mathfrak{u}_1)$, as per Eq. (ref), and the continuum of inequality constraints indexed by $q\in\mathbb{S}^1$ restricts $R$ to be an element of $\mathcal{E}$. The moment constraints that define $B$ are interpreted similarly. Because the support set $\mathcal{S}_\mathcal{E}(\cdot)$ in any direction is a singleton by Proposition (ref), the moments in Eq. (ref) yield points $R$ and $B$ that coincide with Eq. (ref).

We test the null in Eq. (ref) at a given significance level $\alpha\in(0,1)$ based on our procedure to test, for given $e\in\mathbb{R}^2$, whether $e\in\mathcal{PF}$:

proc[Testing Weak Group-Skew] \begin{enumerate} • Build a $(1-\alpha)$-level confidence set for $(R, B)$ by \begin{align} \mathcal{CS}_n(R,B)\equiv\big\{(\Tilde{R},\Tilde{B})\in\mathbb{B}_C\times\mathbb{B}_C: T_n(\Tilde{R},\Tilde{B})\leq \hat{c}_{1-\alpha}(\Tilde{R},\Tilde{B})\big\}, \end{align} where for the support function estimator $\widehat{h}_{\mathcal{E}}(\cdot\,; \widehat{\boldsymbol{\theta}})$ in Theorem (ref), $T_n(\Tilde{R},\Tilde{B})$ adapts the test statistic in Eq. (ref): \begin{align*} T_n(\Tilde{R},\Tilde{B})\equiv\sum_{\substack{(e, \mathfrak{u}_j)\\\in\{(\Tilde{R}, \mathfrak{u}_1),(\Tilde{B}, \mathfrak{u}_2) \}}}\sqrt{n}\left(\left[\max_{q\in\mathbb{S}^1}q^\intercal e-\widehat{h}_\mathcal{E}(q;\widehat{\boldsymbol{\theta}})\right]_{+} + \left[\mathfrak{u}_j^\intercal e-\widehat{h}_\mathcal{E}(\mathfrak{u}_j;\widehat{\boldsymbol{\theta}})\right]_-\right). \end{align*} The critical value $\hat{c}_{1-\alpha}(\Tilde{R},\Tilde{B})$ is obtained similarly to $\hat{c}^\mathcal{PF}_{1-\alpha}(e)$ in Eq. (ref). • Reject $\mathsf{H}_0$ in Eq. (ref) if \begin{align} \varphi_n^{skew}\equiv\mathds{1}\left\{\sup_{(\Tilde{R}, \Tilde{B})\in \mathcal{CS}_n(R,B)}\bigg((\mathfrak{u_1}-\mathfrak{u_2})^\intercal \Tilde{R}\bigg)\bigg((\mathfrak{u_1}-\mathfrak{u_2})^\intercal \Tilde{B}\bigg)<0\right\}=1. \end{align} \end{enumerate}

We note that the test in Eq. (ref) may be conservative as it is based on projection.

propLet the assumptions in Theorem (ref) hold. Then \begin{align} \limsup_{n\to\infty}\mathbb{E}\left[\varphi_n^{skew}\right]\leq \alpha. \end{align}

How to Test for the Existence of an LDA

Given an algorithm $a^*\in\mathcal{A}(\mathcal{X})$ that induces the risk pair $e^*=(e_r^*, e_b^*)\in\mathcal{E}$, call another algorithm that yields a feasible risk pair $e=(e_r,e_b)\in\mathcal{E}$ an LDA if it is at least as accurate as $a^*$ for both groups and at least as fair, with one of these inequalities strict. It follows from the characterization in Proposition (ref) and Eq. (ref) that no LDA to $a^*$ exists if and only if $\mathcal{E}$ can be properly separated from

align*[align* omitted — 127 chars of source]

Recall that the closed form expression for the support function of $\mathcal{C}^*\equiv\mathcal{C}(e^*)$ in direction $q=[q_1,~q_2]^\intercal\in\tilde{\mathbb{S}}^1$ is given in Eq. (ref). We then test the null hypothesis

align[align omitted — 121 chars of source]

against the alternative that $\mathsf{H}_0$ is false. Rejecting the null in Eq. (ref) means that $e^*\notin\mathcal{F}$ and there exists an LDA. We propose estimating $h_{\mathcal{C}^*}(q)$ for $q\in\tilde{\mathbb{S}}^1$ by

align*[align* omitted — 236 chars of source]

where for $g\in\{r,b\}$, $e_g^*$ is estimated by sample means,

align[align omitted — 200 chars of source]

and $\widehat{\mu}_g\equiv\frac{1}{n}\sum_{i=1}\mathds{1}\{G_i=g\}$. We propose the following test statistic:

align[align omitted — 320 chars of source]

$T_n^{\texttt{LDA}}$ differs from $T_n^\mathcal{F}$ in Eq. (ref) only in that $\widehat{e}^*$ is estimated in the former.

propLet the assumptions in Theorem (ref) hold. Then, for any pre-specified significance level $\alpha\in(0,1)$, the test below has asymptotically correct size control: $$ \text{Reject the null in Eq.~\eqref{eqn:null-LDA} if} \quad T_n^{\texttt{LDA}} > c_{1-\alpha+\varsigma}^{\texttt{LDA}}+\varsigma, $$ where $\varsigma>0$ is an arbitrarily small positive constant and for any $\beta\in(0,1)$ the critical value $c_\beta^{\texttt{LDA}}$ is the $\beta$-quantile of $\psi^{\texttt{LDA}}$, for $\psi^{\texttt{LDA}}$ a random variable defined in Eq. (ref).

When Eq. (ref) holds and $Var(\mathbf{L}_d|X)$ is positive definite for $d\in\{0,1\},~X$-a.s., one can take $\varsigma=0$ in Proposition (ref) and the expression for $\psi^{\texttt{LDA}}$ simplifies to that in Eq. (ref). When kinks might be present and the limit distribution is not guaranteed to be continuous and strictly increasing, we take $\varsigma>0$ as in and:shi13. The derivation of $\psi^{\texttt{LDA}}$ uses the fact that $T_n^{\texttt{LDA}}$ is formed by compositions of the $\max$ and $\min$ functions in Eq. (ref) and the $\max$ function in Eq. (ref)---all of which are Hadamard directionally differentiable, as shown in fan:san19 and car:cue:alb20; since compositions preserve directional differentiability Shapiro90, an extension to the functional Delta method can be applied to Theorem (ref). However, standard bootstraps are inconsistent due to the lack of full differentiability fan:san19. As such, we leverage results in fan:san19 to approximate the distribution of $\psi^{\texttt{LDA}}$ and its quantiles via a modified multiplier bootstrap procedure detailed below, similar to that in sem23, where we denote $\phi$ any generic Hadamard directionally differentiable function, $\widehat{he^*}=\widehat{he^*}(\widehat{\boldsymbol{\theta}})\equiv[\widehat{h}_\mathcal{E}(q;\widehat{\boldsymbol{\theta}}),~\widehat{e}_r^*,~\widehat{e}_b^*]^\intercal$ the vector of estimators, and $he^*\equiv[h_\mathcal{E}(q),~e_r^*,~e_b^*]^\intercal$ the vector of truths.

proc[Bootstrap for the Quantiles of $\sqrt{n}\{\phi\big(\widehat{he^*}\big)-\phi\big(he^*\big)\}$] \begin{enumerate} • Draw $\{W_i\}_{i=1}^n$ i.i.d. from the exponential distribution with mean $1$ independent of the sample $\{(Y_i,G_i,X_i)\}_{i=1}^n$ and construct the bootstrap analogue of $\widehat{he^*}$: \begin{align} \widetilde{he^*}=\widetilde{he^*}(\widehat{\boldsymbol{\theta}})\equiv[\widetilde{h}_\mathcal{E}(q;\widehat{\boldsymbol{\theta}}), \widetilde{e}_r^*, \widetilde{e}_b^*]^\intercal, \end{align} \noindentwhere $\widetilde{h}_\mathcal{E}(q;\widehat{\boldsymbol{\theta}})\equiv\frac{1}{n}\sum_{i=1}^n\frac{W_i}{\overline{W}}\zeta_i(\tcMq;\widehat{\boldsymbol{\theta}})$ for $\overline{W}\equiv\frac{1}{n}\sum_{i=1}^n W_i$, $\widetilde{\mathcal{M}}\equiv\operatorname{diag}(1/\widetilde{\mu}_r, 1/\widetilde{\mu}_b)$, $\widetilde{\mu}_g\equiv\frac{1}{n}\sum_{i=1}^n\frac{W_i}{\overline{W}}\mathds{1}\{G_i=g\}$, and $\widetilde{e}_g^*\equiv\frac{1}{n}\sum_{i=1}^n \frac{W_i}{\overline{W}}\frac{Z_i^g}{\widetilde{\mu}_g}$. • Numerically approximate $\phi'_{he^*}(\cdot)$, the directional derivative of $\phi(\cdot)$ at $he^*$, by \begin{align*} \widehat{\phi'}_{he^*}( \Ddot{he})=\frac{1}{s_n}\left(\phi\left(\widehat{he^*}+s_n(\Ddot{he})\right)-\phi\left(\widehat{he^*}\right)\right), \end{align*} where $\Ddot{he}\in\ell^\infty(\mathbb{S}^1)\times\mathbb{R}^2$ is a candidate direction at which we evaluate $\phi'_{he^*}(\cdot)$ and $s_n$ is a vanishing sequence of step sizes such that $\sqrt{n}s_n\to\infty$. • Obtain $\widehat{\phi'}_{he^*}\big(\sqrt{n}\{\widetilde{he^*}-\widehat{he^*}\}\big)$ and \begin{align*} \widehat{c}_\beta\equiv\inf\bigg\{c:\mathbb{P}\bigg(\left.\widehat{\phi'}_{he^*}\big(\sqrt{n}\{\widetilde{he^*}-\widehat{he^*}\}\big)\leq c \,\,\right|\, \{(Y_i,G_i,X_i)\}_{i=1}^n\bigg)\geq \beta\bigg\}, \end{align*} as estimators, respectively, for the limit distribution of $\sqrt{n}\{\phi\big(\widehat{he^*}\big)-\phi\big(he^*\big)\}$ and its $\beta$-quantile, denoted as $c_\beta$. \end{enumerate}

The consistency of the bootstrap outlined in Procedure (ref) is stated in the following result.

propUnder the assumptions of Theorem (ref), \begin{align*} \sup_{f\in\mathcal{BL}_1}\left|\mathbb{E}\big[f\big(\widehat{\phi'}_{he^*}\big(\sqrt{n}\{\widetilde{he^*}-\widehat{he^*}\}\big)\big)\,\big|\,\{(Y_i,G_i,X_i)\}_{i=1}^n\big]-\mathbb{E}\big[f\big(\phi_{he^*}'(\mathbb{G}_{he^*})\big)\big]\right|=o_p(1), \end{align*} where $\mathcal{BL}_1$ is the set of $1$-Lipschitz functions $f:\mathbb{R}\to\mathbb{R}$ such that $|f|_\infty\leq1$ and $\mathbb{G}_{he^*}$ is the Gaussian limit process of $\sqrt{n}(\widehat{he^*}-he^*)$ given in Eq. (ref) in Appendix (ref). If the cdf of $\phi_{he^*}'(\mathbb{G}_{he^*})$ is continuous and increasing at its $\beta$-quantile, denoted $c_\beta$, then $\widehat{c}_\beta=c_\beta+o_p(1)$.

The proof of Proposition (ref) shows that $\psi^{\texttt{LDA}}=\phi_{he^*}'(\mathbb{G}_{he^*})$ for a particular $\phi_{he^*}'(\cdot)$ that is the composition of the directional derivatives of the $\min$, $\max$, and $\inf$ functions that constitute $T_n^{\texttt{LDA}}$. The expression of $\phi_{he^*}'(\cdot)$, given in Eq. (ref), is complex and hence we recommend the numerical approximation approach in Step 2 of Procedure (ref). Under Proposition (ref), we can estimate $c_\beta^{\texttt{LDA}}$, the $\beta$-quantile of $\psi^{\texttt{LDA}}$, by going through the steps in Procedure (ref), where we replace $\phi(\cdot)$ by the composition of the above mentioned directionally differentiable functions accordingly. The same bootstrap procedure can be used to consistently estimate the critical values put forward to carry out inference in Section (ref). The use of the infinitesimal constant $\varsigma$ in Proposition (ref) and Procedure (ref) accounts for the possibility that the cdf of $\phi_{he^*}'(\mathbb{G}_{he^*})$ may not be continuous and increasing at $c_{1-\alpha}$.

Algorithms Yielding an LDA

{

figure[figure omitted — 300 chars of source]

} An algorithm designer or a regulator may wonder if one can characterize the algorithms yielding LDAs to a given algorithm $e^*$. It turns out that it is possible to do so, by combining the algorithm that we put forward in Eq. (ref) with a careful use of our characterization of $\mathcal{F}$ in Proposition (ref). We illustrate the idea for the case that Eq. (ref) holds and $\mathcal{E}$ has no kinks. For a given risk pair $e^*$ induced by an algorithm $a^*\in\mathcal{A}$ such that $e^*\notin\mathcal{F}$, by definition the set of risk pairs $e\in\mathcal{F}$ such that $e>_{FA}e^*$ are:

align[align omitted — 402 chars of source]

The set $\mathcal{F}^*$ is depicted in Figure (ref) as the orange portion of $\mathcal{F}$ that intersects with $\mathcal{C}(e^*)$. Using the same notation as in Section (ref), we can use a training sample $\{(\tilde{Y}_i,\tilde{G}_i,\tilde{X}_i)\}_{i=1}^{n_1}$ and a construction that mimics the procedure to build the estimator of $\mathcal{F}$ in Eq. (ref) to obtain a consistent estimator $\widehat{\mathcal{F}}^*_{n_1}$ of $\mathcal{F}^*$ (the consistency of this estimator can be established through the same steps as in the proof of Proposition (ref)). {Select $\widehat{e}_{n_1}\in\widehat{\mathcal{F}}_{n_1}$ through a projection method detailed in the proof of Proposition (ref) and denote by $e\in\mathcal{F}^*$ the point to which it converges.} Let $\widehat{q}^*_{n_1}(\widehat{e}_{n_1})$ be the consistent estimator of $q^*_{\mathbb{S}^1}(e)$ defined in Eq. (ref). Let $\widehat{\mathcal{S}}_{\mathcal{E},n_2}(\widehat{q}^*_{n_1}(\widehat{e}_{n_1});\widehat{\boldsymbol{\theta}}_{n_1})$ be defined as in Eq. (ref) but with $\widehat{q}^*_{n_1}(\widehat{e}_{n_1})$ replacing $q$. Then $\widehat{\mathcal{S}}_{\mathcal{E},n_2}(\widehat{q}^*_{n_1}(\widehat{e}_{n_1});\widehat{\boldsymbol{\theta}}_{n_1})$ is a consistent estimator of $\mathcal{S}_\mathcal{E}(q^*_{\mathbb{S}^1}(e))\in\mathcal{F}^*$, by the same argument used to establish Proposition (ref).

Distance to the Fairest Point

In this section, we propose a method to build a confidence interval for the distance between the risk $e^*$ induced by a given algorithm and the fairest point $F$ on the frontier, denoted $\rho(e^*,F)$, where $\rho$ is a Hadamard directionally differentiable distance function (e.g., the Euclidean distance, Manhattan distance, Chebyshev distance, etc.). This can inform a decision maker of the relative merits in promoting equity and achieving business efficiency of different algorithms by comparing the confidence intervals on their distance to $F$.

Recall $\mathcal{H}_{45}\equiv\{e\in\mathbb{R}^2:e_r=e_b\}$ denotes the 45-degree line; let $\mathcal{H}_{45}^+\equiv\{e\in\mathbb{R}^2:e_r<e_b\}$ and $\mathcal{H}_{45}^-\equiv\{e\in\mathbb{R}^2:e_r>e_b\}$ denote, respectively, the open halfspace above and below the 45-degree line. As shown in Section (ref), the coordinates of $F$ depend on whether $\mathcal{E}$ intersects with $\mathcal{H}_{45}$. When $\Tilde{\mathcal{E}}\equiv\mathcal{E}\cap\mathcal{H}_{45}\neq\emptyset$, as shown in Eqs. (ref)-(ref) we have $F=\mathfrak{u}\cdoth_{\Tilde{\mathcal{E}}}\left(\mathfrak{u}_1\right)$, with $h_{\Tilde{\mathcal{E}}}\left(\mathfrak{u}_1\right)=\inf_{c\in\mathbb{R}}h_\mathcal{E}\left(\mathfrak{u}_1(c)\right)$, $\mathfrak{u}_1(c)\equiv\mathfrak{u}_1-c[1~~-1]^\intercal$, and $\mathfrak{u}\equiv(\mathfrak{u}_1+\mathfrak{u}_2)$. Note that $h_{\Tilde{\mathcal{E}}}\left(\cdot\right)$ is a Hadamard directionally differentiable function of $h_\mathcal{E}(\cdot)$, and its composition with the distance function $\rho$ is again directionally differentiable. Hence, we use our DML estimator, Theorem (ref), and the results in fan:san19 to directly obtain the limit distribution of an estimator for $\rho(e^*, F)$. However, when $\mathcal{E}\subset \mathcal{H}_{45}^+$ (respectively, $\mathcal{E}\subset \mathcal{H}_{45}^-$), $F$ is given by the support set of $\mathcal{E}$ in direction $(\mathfrak{u}_2-\mathfrak{u}_1)$ (respectively, $(\mathfrak{u}_1-\mathfrak{u}_2)$); see Eqs. (ref)-(ref). {As discussed in Remark (ref), carrying out inference if we use DML to directly estimate both coordinates of a support point is difficult,} and therefore we use Eq. (ref) to represent $F$ through moments that involve $h_\mathcal{E}(\cdot)$ only; see Eqs. (ref)-(ref) below.

Observe that $\mathcal{E}\subset \mathcal{H}_{45}^+$ if and only if $F \in \mathcal{H}_{45}^+$, and $\mathcal{E}\subset \mathcal{H}_{45}^-$ if and only if $F \in \mathcal{H}_{45}^-$, as illustrated in Panels (b) and (d) of Figure (ref). Hence, we partition the parameter space $\mathbb{B}_C$, to which $F$ belongs Assumption (ref), into three sets: $\mathbb{B}_C^+\equiv\mathbb{B}_C\cap\mathcal{H}_{45}^+$, $\mathbb{B}_C^-\equiv\mathbb{B}_C\cap\mathcal{H}_{45}^-$, and $\mathbb{B}_C^{45}\equiv\mathbb{B}_C\cap\mathcal{H}_{45}$. We then have the following expressions for $\rho(e^*, F)$:

enumerate• If $F\in\mathbb{B}_C^{45}$, $\rho(e^*, F)=\rho(e^*,\mathfrak{u}\cdoth_{\Tilde{\mathcal{E}}}\left(\mathfrak{u}_1\right))$. • If $F\in\mathbb{B}_C^+$, $\rho(e^*, F)=\rho(\Tilde{e}, \Tilde{F})$ for $(\Tilde{e}, \Tilde{F})$ satisfying: \begin{align} \left\{\begin{array}{lr} \Tilde{e}-e^*=0,\\ h_{\mathcal{E}}\big((\mathfrak{u}_2-\mathfrak{u}_1)/\sqrt{2}\big)-\big((\mathfrak{u}_2-\mathfrak{u}_1)/\sqrt{2}\big)^\intercal\Tilde{F}=0,\\ h_\mathcal{E}(q)\,-\,q^\intercal \Tilde{F} \geq 0, \forall q\in\mathbb{S}^1. \end{array}\right. \end{align} • If $F\in\mathbb{B}_C^-$, $\rho(e^*, F)=\rho(\Tilde{e}, \Tilde{F})$ for $(\Tilde{e}, \Tilde{F})$ satisfying: \begin{align} \left\{\begin{array}{lr} \Tilde{e}-e^*=0,\\ h_{\mathcal{E}}\big((\mathfrak{u}_1-\mathfrak{u}_2)/\sqrt{2}\big)-\big((\mathfrak{u}_1-\mathfrak{u}_2)/\sqrt{2}\big)^\intercal\Tilde{F}=0,\\ h_\mathcal{E}(q)\,-\,q^\intercal \Tilde{F} \geq 0, \forall q\in\mathbb{S}^1. \end{array}\right. \end{align}

In each of Eqs. (ref)-(ref), the first condition pins down $\Tilde{e}$ to equal $e^*$; the second and third conditions restrict, respectively, $\tilde{F}$ to be a point on the supporting hyperplane of $\mathcal{E}$ in the appropriate direction and $\tilde{F}\in\mathcal{E}$, which together restricts $\tilde{F}$ to be the support set of $\mathcal{E}$ in this direction. We propose the following testing procedure:

proc[Confidence Interval for $\rho(e^*, F)$] \begin{enumerate} • Construct estimators $\widehat{e}^*$ as in Eq. (ref) and $\widehat{h}_\mathcal{E}\big(\cdot;\widehat{\boldsymbol{\theta}}\big)$ as in Theorem (ref). • Emulate the construction in Step 1 of Procedure (ref) to obtain two $(1-\alpha)$-level confidence sets for $(e^*, F)$: let $\mathcal{CS}_n^+(e^*, F)$ denote the one for the moments in Eq. (ref) (in the analog of Eq. (ref), replace $\mathbb{B}_C$ with $\operatorname{cl}{\mathbb{B}_C^+}$), and $\mathcal{CS}_n^-(e^*, F)$ the one for the moments in Eq. (ref) (in the analog of Eq. (ref), replace $\mathbb{B}_C$ with $\operatorname{cl}{\mathbb{B}_C^-}$).\footnote{For a given set $\mathbb{B}$, we denote by $\operatorname{cl}{\mathbb{B}}$ its closure.} • For a given $(\Tilde{e}, \Tilde{F})\in\mathbb{B}_C\times\mathbb{B}_C^{45}$ and $\widehat{h}_{\Tilde{\mathcal{E}}}\left(\mathfrak{u}_1; \widehat{\boldsymbol{\theta}}\right)\equiv\inf_{c\in\mathbb{R}}\widehat{h}_\mathcal{E}\big(\mathfrak{u}_1(c); \widehat{\boldsymbol{\theta}}\big)$, let: \begin{align} T_n^{45}(\rho(\Tilde{e},\Tilde{F}))\equiv\sqrt{n}\left|\rho\bigg(\widehat{e}^*,\mathfrak{u}\cdot\widehat{h}_{\Tilde{\mathcal{E}}}(\mathfrak{u}_1; \widehat{\boldsymbol{\theta}})\bigg)-\rho(\Tilde{e}, \Tilde{F})\right|. \end{align} Let $\psi^{45}$ denote the random variable to which $T_n^{45}(\rho(\Tilde{e},\Tilde{F}))$ converges in distribution for $\Tilde{e}=e^*$ and $\Tilde{F}=F_{45}\equiv\mathfrak{u}\cdoth_{\Tilde{\mathcal{E}}}\left(\mathfrak{u}_1\right)$. Let $c_\beta^{45}$ denote the $\beta$-quantile of $\psi^{45}$ and $\varsigma>0$ an infinitesimal uniformity factor. Use test inversion to construct the confidence set: \begin{align*} \mathcal{CS}_n^{45}(\rho(e^*, F))=\left\{\rho(\tilde{e},\tilde{F}): (\Tilde{e}, \Tilde{F})\in\mathbb{B}_C\times\mathbb{B}_C^{45}, T_n^{45} \le c_{1-\alpha+\varsigma}^{45}+\varsigma\right\}. \end{align*} The expression for $\psi^{45}$ is complex and given in Eq. (ref), with a simpler expression provided in Eq. (ref) for the case that Eq. (ref) is satisfied and $\mathcal{E}$ has no kinks. • Obtain a confidence interval for $\rho(e^*,F)$ as $$ \mathcal{CS}_n^{\rho(e^*,F)}\equiv\left\{\rho(\tilde{e}, \tilde{F}): (\tilde{e}, \tilde{F})\in\mathcal{CS}_n^+(e^*,F)\bigcup\mathcal{CS}_n^-(e^*,F)\right\}\bigcup\left\{\mathcal{CS}_n^{45}(\rho(e^*,F))\right\}. $$ \end{enumerate}

Intuitively, this construction inverts a test that jointly assesses the location of $\mathcal{E}$ relative to $\mathcal{H}_{45}$ and the value of $\rho(e^*,F)$. For example, if $F\in\mathcal{H}_{45}^-$, then both $\mathcal{CS}_n^+(e^*,F)$ and $\mathcal{CS}_n^{45}(\rho(e^*,F))$ are empty with probability approaching one (recall from Section (ref) that $\inf_{c\in\mathbb{R}}h_\mathcal{E}\left(\mathfrak{u}_1(c)\right)$ is unbounded when $F\notin\mathcal{H}^{45}$). Our next result shows that Procedure (ref) delivers an asymptotically valid confidence interval.

propLet $\rho$ be a Hadamard directionally differentiable distance function and the assumptions in Theorem (ref) hold. Then the confidence interval constructed following Procedure (ref) asymptotically covers the true $\rho(e^*,F)$ with probability at least $1-\alpha$.

As we show in the proof, $\psi^{45}$ is again a composition of Hadamard directionally differentiable functions that define $T_n^{\rho(\Tilde{e}, \Tilde{F})}$ in Eq. (ref). We can therefore employ the same bootstrap method detailed in Procedure (ref) to consistently estimate the quantiles $c_\beta^{45}$ for $\beta\in(0,1)$, where we replace $\phi(\cdot)$ by the composition of the directionally differentiable functions, including $\inf$ and $\rho$, that define $T_n^{45}$, whose exact expression we relegate to Appendix (ref).

Monte Carlo Experiments and Empirical Illustration

Monte Carlo Simulations

We evaluate the finite sample properties of the tests introduced in Sections (ref)-(ref) using two distinct data generating processes (DGPs). For both DGPs, the covariates $X\equiv[X_1,\dots,X_{20}]^\intercal\in\mathbb{R}^{20}$ and group identity $G$ are drawn from the following distributions:

align*[align* omitted — 237 chars of source]

We consider two different ways of generating the outcome $Y$:

enumerate• Group-balanced DGP: $Y \,|\, G, X\,\,\stackrel{d}{\sim} Bern\left(\frac{G}{1+e^{-(X_1+X_2+0.5X_3)}}+\frac{(1-G)}{1+e^{-(-X_1-0.5X_2+X_4)}}\right)$; • $r$-skewed DGP: $Y \,|\, G, X\,\,\stackrel{d}{\sim} Bern\left(\frac{G}{1+e^{-2(X_1+X_2+X_3)}}+\frac{(1-G)}{1+e^{-0.7(X_1+0.5X_2+0.6X_4)}}\right)$,

where the group-balanced DGP is such that $X$ is informative about $Y$ in opposite directions for group $r$ ($G=1$) and group $b$ ($G=0$), but its predictive power is similar across groups. In contrast, the $r$-skewed DGP is such that $X$ is systematically more informative about the group-$r$ outcome. We take the loss function to be the classification error, $\ell(d,y)=\mathds{1}\{d\neq y\}$. We also construct a status quo algorithm $a^*$, that we fix in the simulations, by training once a logistic regression on a sample of size $10,000$, with $5,000$ observations from the balanced DGP and $5,000$ observations from the $r$-skewed DGP.\footnote{We train the logistic regression on a mixture of group-balanced and $r$-skewed data so that we test for existence of an LDA to the same algorithm $a^*$ in both DGPs. DGP-specific logistic regressions trained on $10,000$ observations drawn form that DGP for each case yield group risks $e^*$ very close to the ones plotted in Figure (ref).} {

figure[figure omitted — 1,027 chars of source]

}

Figure (ref) depicts the population feasible set $\mathcal{E}$ corresponding to each DGP (pink region), along with a $95\%$ confidence set for the frontier $\mathcal{F}$ (shaded grey region), constructed using a random sample of $10,000$ observations drawn from the respective DGPs, and the risk $e^*$ induced by the status-quo algorithm $a^*$ (green asterisk). Throughout Section (ref), we fit the nuisance parameter $\Delta\boldsymbol{\theta}$ using logit lasso and $5$-fold sample splitting. To assess the finite sample properties of the LDA test in Section (ref) and the distance-to-$F$ test in Section (ref), we test whether $e^*$ is on the frontier and whether it is at a specific distance from $F$.

Table (ref) reports the simulation results. Overall, the Monte Carlo exercise suggests good finite sample properties for our proposed tests, especially as the sample size increases. The top panel corresponds to the test for weak group skew (Section (ref)), where the third column reports the frequency with which the $95\%$ confidence set for the vector $(R, B)$ (Eq. (ref)) based on the available sample fails to cover the true value of $(R,B)$. While the $r$-skewed DGP exhibits some over-rejection at $n=1{,}000$, which is a small sample size relative to the complexity of DML estimation with 20 covariates and estimating the support function across directions, this over-rejection quickly disappears as sample size increases. On the other hand, the weak group skew test (fourth column) rejects the null of weak group skew in the balanced DGP and fails to reject it in the $r$-skewed DGP essentially with probabilities 1 and 0, respectively. This is not surprising because the simulation DGPs are far from the boundary of the null, and because the test is conservative due to the projection step. {

table[table omitted — 4,074 chars of source]

}

The middle panel reports the LDA test results (Section (ref)). The third (fourth) column shows the frequency with which the population point $R$ ($B$), which by definition belongs to $\mathcal{F}$, is rejected by the test of the null that it belongs to $\mathcal{F}$. The fifth (sixth) column reports the frequency with which $(R+B)/2$ (the risk $e^*$ associated with the logit algorithm $a^*$), which by construction does not belong to $\mathcal{F}$, is rejected by the test as an element of $\mathcal{F}$. While at small sample size ($n=1,000$), the test exhibits some over-rejection for the point $B$, this quickly disappears as sample size grows. At small and medium sample sizes ($n=1,000$ and $n=5,000$), the rejection probability for the false null that $(R+B)/2\in\mathcal{F}$ is low for the $r$-skewed DGP, while it is high at all sample sizes for the balanced DGP. This is justifiable in light of Figure (ref), which shows that in the $r$-skewed DGP the chord between $R$ and $B$ is close to the frontier and lies largely inside its 95% confidence set. For $n=10,000$, the test detects the false nulls in columns five and six with substantially higher probability.

The bottom panel reports the results for the distance-to-$F$ test (Section (ref)). This test is more delicate because its implementation requires solving an optimization problem to estimate $F$. Here we observe more substantial over-rejection for $n=1,000$, but the distortion gets markedly reduced as sample size increases.

Empirical Illustration

We revisit the analysis in \citet*[OPVM henceforth]{obe:pow:vog:mul19}, who analyze properties of the algorithm used by a research hospital to determine if a patient should be automatically enrolled in a high-risk care management program. The algorithm used by the research hospital aims at predicting patients' health needs based on total medical expenditures (the label on which the algorithm is trained). It produces a health risk score and automatically enrolls a patient in the high-risk care management program if that patient's risk score exceeds the $97^{\text{th}}$ percentile of all predicted scores. We reassess this algorithm through our testing procedures. To do so, we use the synthetic data made available by \citet*{li:lin:obe19} at \href{https://gitlab.com/labsysmed/dissecting-bias}{GitLab} to replicate all analyses in \citetalias{obe:pow:vog:mul19}. The data include $48,784$ patient observations, of which $5,582$ self report as Black and the others self report as White ($g\in\{\textsf{bl},\textsf{wh}\}$). The data include $149$ covariates, such as age, gender, comorbidity and medication variables, costs, and biomarkers, and provide information about each patient's number of active chronic conditions in the subsequent year, viewed as the true measure of health needs of a patient. Following \citetalias{lia:lu:mu:oku24}, we let $\ell(d,Y)=\mathds{1}\{Y\neq d\}$ be the classification loss and, unless explicitly stated otherwise, $Y_i=\mathds{1}\{$patient $i$ has 6 or more chronic conditions$\}$, with the choice of 6 driven by the fact that it is the $97^{\text{th}}$ percentile of active chronic condition numbers across patients in the sample. We use random forests with $5,000$ trees to estimate $\Delta\boldsymbol{\theta}$ as described in Section (ref).

figure[figure omitted — 566 chars of source]

Feasible Set Estimation and Inference for the Frontier

We report in Figure (ref) an estimate of the feasible set $\mathcal{E}$ based on Eq. (ref) using $1,000$ directions (top-left panel), along with 100 estimated supporting hyperplanes (top-right panel), where across all panels the horizontal (vertical) axis is the risk for Blacks (Whites), denoted as $e_{\textsf{bl}}$ ($e_{\textsf{wh}}$). We zoom in to show the estimated FA-frontier $\widehat\mathcal{F}$ and its $95\%$ confidence set (bottom-left panel), and zoom in further to show the best risk achievable for the Black patients (bottom-right panel, red point labeled $\textsf{BL}$), which coincides with and is overlaid by the best risk for the White patients (bottom-right panel, blue point labeled $\textsf{WH}$).

Hypothesis Testing

We next report the weak group skew test results in Figure (ref)-Panel (a), where both $\textsf{BL}$ and $\textsf{WH}$ are below the $45^\circ$ degree line and the feasible set is $\textsf{wh}$-skewed; the test fails to reject the null of weak group skew, suggesting that implementing Pareto-dominated algorithms could be justified based on the designer's preferences over fairness and accuracy. In Figure (ref)-Panel (b) we plot $\widehat{\mathcal{F}}$, along with its $95\%$ confidence set. Figure (ref)-Panel (b) also plots the estimated group risks associated with the original algorithm used by the hospital (a green asterisk) and three alternative algorithms that \citetalias{obe:pow:vog:mul19} experiment with to assess whether different label choices yield decision rules that are more accurate and fairer than the algorithm currently used by the hospital. One of these algorithms is trained to predict total cost (hollow diamond with a cross in Figure (ref)-Panel (b)), one to predict avoidable costs (filled diamond), and the other one to predict the number of active chronic conditions (hollow diamond). All our tests take into account the finite sample estimation error of the group risks induced by these four algorithms. The top panel of Table (ref) reports the values of the LDA test statistics and the associated critical values that we compute for the four algorithms considered by \citetalias{obe:pow:vog:mul19}. The hypothesis that the original algorithm yields group risks on the frontier is rejected, and so is the same hypothesis for group risks associated with the algorithm trained to predict total costs. On the other hand, we fail to reject that the algorithms trained to predict avoidable costs and the number of active chronic conditions yield group risks on the FA-frontier. Regarding the three algorithms proposed by \citetalias{obe:pow:vog:mul19}, if one were to plot the set $\mathcal{C}^*$ corresponding to the original algorithm, it would be immediate to see that one cannot reject the hypothesis that all other algorithms considered by \citetalias{obe:pow:vog:mul19} improve upon it, both in terms of fairness and accuracy.

figure[figure omitted — 1,088 chars of source]

The bottom panel of Table (ref) shows that the closest risk to $F$ is the one associated with the algorithm predicting total costs. However, all confidence intervals overlap.

{

table[table omitted — 1,609 chars of source]

}

Performance of Algorithms Constructed to Be on a Constrained Frontier

Finally, we compare various outcomes associated with the algorithms considered by \citetalias{obe:pow:vog:mul19} with those of decision rules resulting from the algorithms that we propose in Section (ref). To do so, we randomly split the sample into two halves, and use one half (training data) to estimate the nuisance parameter $\Delta\boldsymbol{\theta}$, and implement Eq. (ref) on the other half (evaluation data) for the following choices of $q$:

enumerate$q=[-1~~~0]^\intercal$, yielding an algorithm that (asymptotically) achieves the best point on $\mathcal{F}$ for Black patients, which we call Rawlsian because $e_\textsf{bl}>e_\textsf{wh}$ at $\textsf{BL}$. • $q=[0~~-1]^\intercal$, yielding an algorithm that (asymptotically) achieves the best point on $\mathcal{F}$ for White patients, which we call Majority as Whites are the majority of the sample. • $\widehat{q}(\widehatF)$, yielding an algorithm based on a direction estimated using the entire sample that (asymptotically) achieves $F$, the fairest point on $\mathcal{F}$, as established in Proposition (ref), which we call Egalitarian. • $q=-\frac{\sqrt{2}}{2}[1~~1]^\intercal$, yielding an algorithm that weighs both groups equally, which we call Utilitarian as it can be expressed as a generalization of the Utilitarian rule.

Throughout, we recognize that to compare algorithms we should enforce a global capacity constraint on the total percentage of patients assigned to the high-risk case management program, so that variation across algorithms is not confounded with a possibly more generous care program. To enforce the capacity constraint, we need to require the algorithms to satisfy $\int a(x)d\mathbb{P}_X\le \bar{a}$ for some known constant $\bar{a}$. For example, $\bar{a}=0.03$ when comparing with the algorithm currently used by the hospital which assigns only patients with risk score above the $97^\text{th}$ percentile to the high-risk care program. Recall from the discussion following Proposition (ref) that without capacity constraints, we have

align[align omitted — 246 chars of source]

When $\mathcal{A}(\mathcal{X})$ is constrained to only include algorithms such that $\int a(x)d\mathbb{P}_X\le \bar{a}$, maximization in Eq. (ref) is achieved by setting

align[align omitted — 182 chars of source]

for $\mathrm{quant}_{k(\boldsymbol{\theta}(X),\cMq)}(\alpha)$ the $\alpha$-quantile of $k(\boldsymbol{\theta}(X),\cMq)$ and $k(\boldsymbol{\theta}(X),\cMq)$ as in Eq. (ref).\footnote{In this case, the constrained feasible set is $\mathcal{E}^\texttt{co}_{\bar{a}}\equiv \left\{\bigl(e_r(a),e_b(a)\bigr)\in\mathbb{R}^2:a\in\mathcal{A}(\mathcal{X})~\text{and}~\int a(x)d\mathbb{P}_X\le \bar{a}\right\}$, a convex subset of $\mathcal{E}$, and the algorithms in Eq. (ref) is on its frontier.} In words, for our example with $\bar{a}=0.03$, high-risk care management is assigned to those patients with positive values of $k(\boldsymbol{\theta}(X),\cMq)$ that exceed its $97^\text{th}$ percentile.

The results of our first exercise are reported in Figure (ref). The top panel replicates Figure 1-(a) in \citetalias{obe:pow:vog:mul19}, and shows that at each percentile of the original algorithm's risk score (the horizontal axis), Black patients have a substantially larger number of active chronic conditions (the vertical axis) than White patients.\footnote{The top panel of Figure (ref) is not identical to Figure 1-(a) of \citetalias{obe:pow:vog:mul19} because it is plotted using the synthetic data, instead of the real data, and because the code used to generate Figure 1-(a) of \citetalias{obe:pow:vog:mul19} had an error that was later corrected after the publication of \citetalias{obe:pow:vog:mul19}; see the documentation of li:lin:obe19 for more details.} Even among patients automatically enrolled in the high-risk care management program (those with a risk score above the $97^{\text{th}}$ percentile), Black patients appear to be in worse health than White patients.

We ask whether the algorithms that we propose are able to select for treatment patients that, in fact, exhibit substantially worse health outcomes one year later. The bottom four panels in Figure (ref) plot, for each of the algorithms described at the beginning of this section, the number of active chronic conditions for (a) Black patients that our algorithm does not assign to the high-risk care program (dash-dotted, light-purple line); (b) White patients that our algorithm does not assign to the high-risk care program (solid salmon line); (c) Black patients that our algorithm assigns to the high-risk care program (dash-dotted, dark-purple line); (b) White patients that our algorithm assigns to the high-risk care program (solid orange line). For comparability with the top panel, on the horizontal axis we continue to report the risk score produced by the algorithm used by the hospital.

figure[figure omitted — 438 chars of source]

The main takeaway from Figure (ref) is that our four proposed algorithms, which use group identity to estimate $\Delta\boldsymbol{\theta}$ but not for treatment assignment, are successful at selecting for treatment patients who, one year later, experience substantially worse health outcomes. Moreover, with the Ralwsian, Majority, and Utilitarian algorithms, Black and White patients assigned to treatment have similar numbers of chronic conditions. Only the Egalitarian algorithm shows notable disparities for patients with hospital risk scores below the $70^{\text{th}}$ percentile. We think this result might be due to two factors: estimation of $F$ is challenging and hence the direction $\widehat{q}(\widehat{F})$ might be imprecisely estimated; and the feasible set is $\textsf{wh}$-skewed and hence the Egalitarian algorithm leads to a Pareto-dominated outcome.

Our last exercise aims at further assessing the extent to which our proposed algorithms may reduce the substantial disparities between Black and White patients in the current program screening practices documented by \citetalias{obe:pow:vog:mul19}. To illustrate the potential for improvement over the hospital's algorithm, \citetalias{obe:pow:vog:mul19} simulate “a counterfactual world with no gap in health conditional on risk” (p.3). They construct an infeasible, couterfactual algorithm that uses group identity and patients' ex-post active chronic health conditions to find the sickest Black patient with health risk score just below a threshold (the “inframarginal Black patient”) and the healthiest White patient with health risk score just above the same threshold (the “supramarginal White patient”). If the number of chronic health conditions of the inframarginal Black patient is larger than that of the supramarginal White patient, they iteratively swap them until the number of chronic health conditions of the inframarginal Black patient equals that of the supramarginal White patient. \citetalias{obe:pow:vog:mul19} find that at all risk thresholds $\alpha$ above the median, the counterfactual algorithm increases the fraction of Black patients treated. Table (ref), columns 2-3, show that across the various thresholds for automatic enrollment in the program that we consider, the fraction of Black patients would rise between 6 and 41 percentage points.\footnote{In columns 2-3, the fractions reported are based on a denominator that equals the total number of patients with risk scores above a certain percentile of risk scores, e.g., the $55^{\text{th}}$ percentile, with this threshold viewed as the threshold above which the patient is automatically enrolled in the high-risk care program, and a numerator that equals the number of Black patients with risk score above that percentile.}

We ask how much of this gap could be filled if one were to use our algorithms, which are feasible and do not rely on ex-post knowledge of the number of active chronic conditions nor on using group identity for assignment to the high-risk care program. The results, reported in Table (ref), columns 4-7,\footnote{For each algorithm, the fraction of Black patients treated at each capacity threshold is computed as the following ratio. The denominator equals the total number of patients treated under a given algorithm, e.g., the Rawlsian algorithm, subject to the constraint that at each threshold $\bar{a}\in[0.55,0.97]$ at most $100(1-\bar{a})\%$ of the evaluation sample is treated, and for each threshold $\bar{a}$ the outcome $Y$ equals the indicator of whether the number of active chronic conditions exceeds the $\bar{a}$-quantile of the distribution of the number of active chronic conditions.} show that each of our four algorithm yields an increase in the fraction of Black patients treated at all capacity thresholds. The Rawlsian and the Utilitarian algorithms yield the largest increases, ranging between 4 and 18 percentage points across the various capacity thresholds. This shows that these two feasible, easy to implement algorithms can close between $26\%-69\%$ of the gap between the algorithm that the hospital uses and the counterfactual, infeasible algorithm simulated by \citetalias{obe:pow:vog:mul19}.

{

table[table omitted — 1,652 chars of source]

}

We conclude by noting that results based on random forests trained using the grf package are not guaranteed to reproduce exactly across platforms, even with the same seed (all simulations and estimation are in R). This is a known feature of grf (see the \href{https://grf-labs.github.io/grf/REFERENCE.html\#forests-predict-different-values-depending-on-the-platform-even-though-the-seed-is-the-same}{reference manual}). Similarly, tests based on optimization via stochastic gradient descent implemented by the torch package are not exactly reproducible, even after seed setting, which affects the $F$ estimator and thus the distance-to-$F$ test (see the \href{https://github.com/mlverse/torch/issues/1311}{repository discussion}). To assess sensitivity of our results to these features, we repeat the entire empirical exercise 20 times with different seeds. Appendix (ref) reports results and includes a robustness check using logit lasso for estimating the nuisance function.\footnote{For the simulations in Section (ref), we expect cross-platform variability to be negligible, as results are averaged over 1000 Monte Carlo replications.} While the results exhibit some nontrivial variation across seeds (although the qualitative results are unchanged), we view this as expected, considering that the nonparametric estimation step involves 149 covariates against a total sample size of $48,784$ and a minority group of size $5,582$.

Conclusion

We provide a consistent nonparametric estimator for a theoretical fairness-accuracy frontier proposed by \citetalias{lia:lu:mu:oku24} and algorithms that attain points on this frontier. We obtain the estimator through judicious use of the separating hyperplane theorem and the support function of the (convex) feasible set of expected losses associated with all possible algorithms, a portion of whose boundary coincides with the FA-frontier. We provide a DML estimator of the support function and show it converges to a tight Gaussian process as sample size increases. We formulate important policy-relevant hypotheses that have received much attention in the fairness literature as restrictions on the support function and construct valid test statistics. We provide an estimator for the distance between a given algorithm and the fairest point on the frontier. We carry out a Monte Carlo exercise that illustrates the good finite sample properties of our method, and demonstrate its practical relevance by revisiting the empirical analysis in \citetalias{obe:pow:vog:mul19}. Our results show that the algorithm that a research hospital employs to screen patients for high-risk care is not on the frontier. Our proposed algorithms substantially improve over this status quo in terms of both fairness and accuracy.

appendix\section{Proofs of Main Results} \subsection{Proofs for Section (ref)} \begin{proof}[Proof of Proposition (ref)] As shown in the discussion leading to Eq. (ref), $\mathcal{E}=\left\{\mathbb{E}[\mathcal{M}\vartheta(X)]: \vartheta(X)\in \Lambda(X) \right\}=\mathbf{E}\left[\mathcal{M}\Lambda(X)\right]$, where $\Lambda(X) \equiv \operatorname{conv}\left(\{\boldsymbol{\theta}_0(X), \boldsymbol{\theta}_1(X)\}\right)$, with $\boldsymbol{\theta}_d(X)$ defined in ((ref)) for $d\in\{0,1\}$. Hence, $\mathcal{M}\Lambda(X)$ is a random compact interval mol:mol18, and $\mathbf{E}\left[\mathcal{M}\Lambda(X)\right]$ is its Aumann expectation mol:mol18, which is well defined because $\Lambda(X)$ is an integrable random convex set owing to $\mathbb{E}[|\theta_d^g(X)|]\leq \mathbb{E}[|\theta_d^g(X)|^2]^{1/2} \leq \mathbb{E}\big[(L_d^g)^2\big]^{1/2}<{\sqrt{c_2}}<\infty$ for any $d\in\{0,1\},g\in\{r,b\}$ by Assumption (ref). It follows that $h_{\mathcal{E}}(q)=h_{\mathbf{E}\left[\mathcal{M}\Lambda(X)\right]}(q)=\mathbb{E}[h_{\mathcal{M}\Lambda(X)}(q)]$ mol:mol18. Observe that \begin{align} h_{\mathcal{M}\Lambda(X)}(q)\equiv\max_{\vartheta(X) \in \Lambda(X)} (\cMq)^\intercal\vartheta(X)=\max\{(\cMq)^\intercal\boldsymbol{\theta}_0(X),(\cMq)^\intercal\boldsymbol{\theta}_1(X)\}, \end{align} where the last equality in Eq. (ref) is well-known in the literature roc97. Taking the expectation with respect to $\mathbb{P}(X)$ yields the first line of Eq. (ref), which we can re-write as \begin{align*} h_\mathcal{E}(q)=\mathbb{E}\bigg[(\cMq)^\intercal\boldsymbol{\theta}_0(X)+(\cMq)^\intercal\big(\boldsymbol{\theta}_1(X)-\boldsymbol{\theta}_0(X)\big)\mathds{1}\big\{(\cMq)^\intercal\big(\boldsymbol{\theta}_1(X)-\boldsymbol{\theta}_0(X)\big)>0\big\}\bigg]. \end{align*} The law of iterated expectations yields the expression in the second line of Eq. (ref). \end{proof} \begin{proof}[Proof of Proposition (ref)] The following proof closely follows the argument from cha:che:mol:sch18. Take any $\|\delta\|_E\to 0$, \begin{align*} &\frac{1}{\|\delta\|_E}\Biggl(\mathbb{E}\big[\big(\mathcal{M}(q+\delta)\big)^\intercal\boldsymbol{\theta}_0 +k\big(\boldsymbol{\theta},\mathcal{M}(q+\delta)\big)\mathds{1}\big\{k\big(\boldsymbol{\theta},\mathcal{M}(q+\delta)\big)>0\big\}\big]\\ &-\mathbb{E}\big[(\cMq)^\intercal\boldsymbol{\theta}_0 +k(\boldsymbol{\theta},\cMq)\mathds{1}\{k(\boldsymbol{\theta},\cMq)>0\}\big]\Biggr) \\ =&\frac{\delta^\intercal}{\|\delta\|_E}\mathbb{E}\big[\mathcal{M}\boldsymbol{\theta}_0 +\mathcal{M}(\boldsymbol{\theta}_1-\boldsymbol{\theta}_0)\mathds{1}\{k(\boldsymbol{\theta},\cMq)>0\}\big]+\frac{1}{\|\delta\|_E}\mathbb{E}[R(q,\delta)], \end{align*} where \begin{align*} R(q,\delta) &\equiv\big(\mathcal{M}(q+\delta)\big)^\intercal(\boldsymbol{\theta}_1-\boldsymbol{\theta}_0)\mathds{1}\big\{k(\boldsymbol{\theta},\cMq)\leq 0< k\big(\boldsymbol{\theta},\mathcal{M}(q+\delta)\big)\big\}\\ &-\big(\mathcal{M}(q+\delta)\big)^\intercal(\boldsymbol{\theta}_1-\boldsymbol{\theta}_0)\mathds{1}\big\{k(\boldsymbol{\theta},\cMq) > 0 \geq k\big(\boldsymbol{\theta},\mathcal{M}(q+\delta)\big)\big\}\\ \implies &\sup_{q\in\mathbb{S}^1} \mathbb{E}[|R(q,\delta)|] \lesssim \|\delta\|_E\cdot \bigl\|\boldsymbol{\theta}_1-\boldsymbol{\theta}_0\bigr\|_{L^2(\mathbb{P})}\cdot\sup_{q\in\mathbb{S}^1} \biggl \|\mathds{1}\bigl\{\left|k(\boldsymbol{\theta},\cMq)\right|<\left|k(\boldsymbol{\theta},\mathcal{M}\delta)\right|\bigr\}\biggr\|_{L^2(\mathbb{P})}\\ &\lesssim\|\delta\|_E\cdot\sup_{q\in\mathbb{S}^1} \mathbb{P}\bigl(\left|k(\boldsymbol{\theta},\cMq)\right|<\left|k(\boldsymbol{\theta},\mathcal{M}\delta)\right|\bigr)^{1/2}\lesssim \|\delta\|_E\big(\|\delta\|_E^{m/2}+\|\delta\|_E^{1/2}\big)^{1/2} \end{align*} where the first inequality follows from Hölder's inequality and that, for each $X$, conditional on the event $\left|k(\boldsymbol{\theta},\cMq)\right|<\left|k(\boldsymbol{\theta},\mathcal{M}\delta)\right|$, we have $|\big(\mathcal{M}(q+\delta)\big)^\intercal(\boldsymbol{\theta}_1-\boldsymbol{\theta}_0)|\leq 2|\mathcal{M}\delta^\intercal(\boldsymbol{\theta}_1-\boldsymbol{\theta}_0)|\leq2\|\mathcal{M}\delta\|_E\|\boldsymbol{\theta}_1-\boldsymbol{\theta}_0\|_E\lesssim\|\delta\|_E\|\boldsymbol{\theta}_1-\boldsymbol{\theta}_0\|_E$; the second inequality follows from $\bigl\|\boldsymbol{\theta}_1-\boldsymbol{\theta}_0\bigr\|_{L^2(\mathbb{P})}<\infty$ by Assumption (ref). The last inequality follows from \begin{align} \sup_{q\in\mathbb{S}^1} \mathbb{P}\bigl(\left|k(\boldsymbol{\theta},\cMq)\right|<\left|k(\boldsymbol{\theta},\mathcal{M}\delta)\right|\bigr)&\leq \sup_{q\in\mathbb{S}^1} \mathbb{P}\bigl(\left|k(\boldsymbol{\theta},\cMq)\right|<\|\delta\|_E^{1/2}\bigr)+\mathbb{P}\bigl(\left|k(\boldsymbol{\theta},\mathcal{M}\delta)\right|\geq\|\delta\|_E^{1/2}\bigr)\notag\\ &\lesssim\|\delta\|_E^{m/2}+\|\delta\|_E^{1/2}, \end{align} where the last line follows because $\sup_{q\in\mathbb{S}^1} \mathbb{P}\bigl(\left|k(\boldsymbol{\theta},\cMq)\right|\bigr)\lesssim\|\delta\|_E^{m/2}$ by Assumption (ref) and $\mathbb{P}\bigl(\left|k(\boldsymbol{\theta},\mathcal{M}\delta)\right|\geq\|\delta\|_E^{1/2}\bigr)\leq\frac{\|\mathcal{M}\delta\|_E\mathbb{E}[\|\boldsymbol{\theta}\|_E]}{\|\delta\|_E^{1/2}}\lesssim\|\delta\|_E^{1/2}$ by Markov's inequality and Assumption (ref). Hence, $\sup_{q\in\mathbb{S}^1}\frac{1}{\|\delta\|_E}\mathbb{E}[R(q,\delta)]\leq \sup_{q\in\mathbb{S}^1}\frac{1}{\|\delta\|_E}\mathbb{E}[|R(q,\delta)|]\lesssim(\|\delta\|_E^{m/2}+\|\delta\|_E^{1/2})^{1/2} \to 0$ and Eq. (ref) follows from applying the law of iterated expectations. Next, the claim that $\nabla_qh_\mathcal{E}(q)=\mathcal{S}_\mathcal{E}(q)$ follows from sch93. Finally, uniform continuity of $\mathcal{S}_\mathcal{E}(q)$ follows from continuity over compact $\mathbb{S}^1$. \end{proof} \begin{proof}[Proof of Proposition (ref)] By definition of $\mathcal{F}$ and $\mathcal{C}(e^*)$, $e^*\in\mathcal{F}$ if and only if $\mathcal{C}(e^*)\cap\mathcal{E}=\{e^*\}$. Suppose $e^*\in\mathcal{F}$. Then $\text{relint}(\mathcal{C}(e^*))\cap\text{relint}(\mathcal{E})=\emptyset$. Since both $\mathcal{C}(e^*)$ and $\mathcal{E}$ are nonempty convex sets, by sch93 $\mathcal{C}(e^*)$ and $\mathcal{E}$ are properly separated. Hence, there exists $q\in\mathbb{S}^1$ and $z\in\mathbb{R}$ such that $$ \forall\,\Tilde{e}\in \mathcal{C}(e^*), \,\,\,\Tilde{e}^\intercalq \leq z\quad\text{and}\quad\forall\, e\in\mathcal{E},\,\,\, e^\intercal(-q)\leq -z. $$ Since $e^*\in \mathcal{C}(e^*)\cap\mathcal{E}$, we have ${e^*}^\intercalq=z$ and $ h_{\mathcal{C}(e^*)}(q)=-h_\mathcal{E}(-q)=z. $ For the other direction, suppose there exists $q\in\mathbb{S}^1$ such that $h_{\mathcal{C}(e^*)}(q)=-h_\mathcal{E}(-q)$. By definition, $e^*$ is feasible and $e^*\in \mathcal{C}(e^*)\cap\mathcal{E}$. If there exists $e'\in \mathcal{C}(e^*)\cap\mathcal{E}$ but $e'\neq e^*$, then $q^\intercal{e'}=q^\intercal {e^*}=-h_\mathcal{E}(-q)$. This means that the support set of $\mathcal{E}$ in direction $-q$ is not a singleton, contradicting Assumption (ref), under which $\mathcal{E}$ has a smooth boundary. To obtain the characterization of $\mathcal{F}$ in Eq. (ref), for $\mathbb{B}^1\equiv\{q:\Vertq\Vert_E\le 1\}$, note that the optimization problem $\max_{q\in\mathbb{B}^1}(-h_{\mathcal{C}(e^*)}(q)-h_\mathcal{E}(-q))$ is dual to the primal problem $\min_{\tilde{e}\in\mathcal{C}(e^*),e\in\mathcal{E}}\Vert\tilde{e}-e\Vert_E$, which measures the distance between $\mathcal{C}(e^*)$ and $\mathcal{E}$. When $e^*\in\mathcal{E}$, this distance is zero, and weak duality along with $\mathbb{S}^1\subset\mathbb{B}^1$ imply $\max_{q\in\mathbb{S}^1}(-h_{\mathcal{C}(e^*)}(q)-h_\mathcal{E}(-q))\leq 0$. Hence, when $e^*\notin\mathcal{F}$ we must have $\sup_{q\in\mathbb{S}^1}(-h_{\mathcal{C}(e^*)}(q)-h_\mathcal{E}(-q))< 0$. Since $h_{\mathcal{C}(e^*)}(q)=\sup_{e\in\mathcal{C}(e^*)}q^\intercal e$ is unbounded for any $q$ such that $q_1+q_2<0$, it is without loss of generality to focus on $q\in\tilde{\mathbb{S}}^1$ and to use the criterion $\left[\max_{q\in\tilde{\mathbb{S}}^1}(-h_{\mathcal{C}(e^*)}(q)-h_\mathcal{E}(-q))\right]_-=0$. \end{proof} \subsection{Proofs for Section (ref)} In the proofs that follow, recall \begin{align*} \zeta_i(\mathring{\mathcal{M}}q;\boldsymbol{\vartheta})&\equiv(\mathring{\mathcal{M}}q)^\intercal\mathbf{L}_{0_i}+(\mathring{\mathcal{M}}q)^\intercal(\mathbf{L}_{1_i}-\mathbf{L}_{0_i})\cdot\mathds{1}\bigl\{k\bigl(\boldsymbol{\vartheta}(X_i),\mathring{\mathcal{M}}q\bigr)>0\bigr\}. \end{align*} defined in Eq. (ref), where $k\bigl(\boldsymbol{\vartheta}(X_i),\mathring{\mathcal{M}}q\bigr)=q^\intercal\mathring{\mathcal{M}}\Delta\boldsymbol{\vartheta}(X_i)$ is given in Eq. (ref) for generic $\Delta\boldsymbol{\vartheta}(X_i)\in\Theta$ and $\mathring{\mathcal{M}}$. For $\widehat{\Delta\boldsymbol{\theta}}$ and $\widehat{\mathcal{M}}$ estimated as per Definition (ref), recall \begin{align*} \zeta_i(\hcMq;\widehat{\boldsymbol{\theta}})= (\hcMq)^\intercal\mathbf{L}_{0_i}+(\hcMq)^\intercal(\mathbf{L}_{1_i}-\mathbf{L}_{0_i})\cdot\mathds{1}\bigl\{k\bigl(\widehat{\boldsymbol{\theta}}(X_i),\hcMq\bigr)>0\bigr\} \end{align*} for $k\bigl(\widehat{\boldsymbol{\theta}}(X_i),\hcMq\bigr)=q^\intercal\widehat{\mathcal{M}}\widehat{\Delta\boldsymbol{\theta}}(X_i)$ as in Eqs. (ref)-(ref). \begin{proof}[Proof of Proposition (ref)] Apply the law of iterated expectations, \begin{align} &\mathbb{E}\bigl[\zeta_i\bigl(\cMq;\boldsymbol{\theta}+t(\boldsymbol{\vartheta}-\boldsymbol{\theta})\bigr)\bigr]-\mathbb{E}\bigl[\zeta_i(\cMq;\boldsymbol{\theta})\bigr]\notag\\ =&\mathbb{E}{\Bigg[\Biggl(\mathds{1}\Bigl\{(\cMq)^\intercal(\Delta\boldsymbol{\theta}+t(\Delta\boldsymbol{\vartheta}-\Delta\boldsymbol{\theta}))\geq0\Bigr\}-\mathds{1}\Bigl\{(\cMq)^\intercal\Delta\boldsymbol{\theta}\geq0\Bigr\}\Biggr)(\cMq)^\intercal\Delta\boldsymbol{\theta}\Bigg]}. \end{align} In what follows, we show Eq. (ref) is bounded in absolute value by $t(t^{m/2}+t^{1/2})^{1/2}$ for $m$ in Assumption (ref), uniformly in $q\in\mathbb{S}^1$. It then follows that, for all $q\in\mathbb{S}^1$, \begin{align*} \lim_{t\to0}\frac{1}{t}\left|\left(\mathbb{E}\bigl[\zeta_i\bigl(\cMq;\boldsymbol{\theta}+t(\boldsymbol{\vartheta}-\boldsymbol{\theta})\bigr)\bigr]-\mathbb{E}\bigl[\zeta_i(\cMq;\boldsymbol{\theta})\bigr]\right)\right|\lesssim \lim_{t\to0}\frac{1}{t}\cdot t(t^{m/2}+t^{1/2})^{1/2}=0. \end{align*} The term in Eq. (ref) is non-zero if and only if the indicator involving $\Delta\boldsymbol{\theta}+t(\Delta\boldsymbol{\vartheta}-\Delta\boldsymbol{\theta})$ equals $1$ and that involving $\Delta\boldsymbol{\theta}$ equals $0$, or vice versa; this happens on a subset of events \begin{align*} \biggl\{(\cMq)^\intercal\Delta\boldsymbol{\theta}\leq0<(\cMq)^\intercal(\Delta\boldsymbol{\theta}+t(\Delta\boldsymbol{\vartheta}-\Delta\boldsymbol{\theta}))\biggr\}\bigcup\biggl\{(\cMq)^\intercal\Delta\boldsymbol{\theta}>0\geq (\cMq)^\intercal(\Delta\boldsymbol{\theta}+t(\Delta\boldsymbol{\vartheta}-\Delta\boldsymbol{\theta}))\biggr\}, \end{align*} which implies the event $\bigl\{|(\cMq)^\intercal\Delta\boldsymbol{\theta}|<t|(\cMq)^\intercal(\Delta\boldsymbol{\vartheta}-\Delta\boldsymbol{\theta})|\bigr\}$. It then follows that \begin{align*} |Eq. (ref)|&\leq \mathbb{E}\left[\mathds{1}\biggl\{|(\cMq)^\intercal\Delta\boldsymbol{\theta}|<t|(\cMq)^\intercal(\Delta\boldsymbol{\vartheta}-\Delta\boldsymbol{\theta})|\biggr\}\big|(\cMq)^\intercal\Delta\boldsymbol{\theta}\big|\right]\\ &\leq \mathbb{E}\left[\mathds{1}\biggl\{|(\cMq)^\intercal\Delta\boldsymbol{\theta}|<t|(\cMq)^\intercal(\Delta\boldsymbol{\vartheta}-\Delta\boldsymbol{\theta})|\biggr\}\cdot t\big|(\cMq)^\intercal(\Delta\boldsymbol{\vartheta}-\Delta\boldsymbol{\theta})\big|\right] \\ &\lesssim t(t^{m/2}+t^{1/2})^{1/2}, \end{align*} where the last line follows from Hölder's inequality, Assumption (ref), and $\sup_{\Delta\boldsymbol{\vartheta}\in\Theta}\|\Delta\boldsymbol{\vartheta}-\Delta\boldsymbol{\theta}\|_{L^2(\mathbb{P})}<\infty$, using a similar argument to that of Eq. (ref). \end{proof} \begin{proof}[Proof of Theorem (ref)] Part 1. We begin by showing that for any fixed $q\in\mathbb{S}^1$, \begin{align} \sqrt{n}\biggl(\widehat{h}_{\mathcal{E}}(q; \widehat{\boldsymbol{\theta}})-h_{\mathcal{E}}(q; \boldsymbol{\theta})\biggr)=\mathbb{G}_n[\zeta_i^*(\cMq;\boldsymbol{\theta})]+o_p(1), \end{align} where for a generic measurable function $t\in\mathcal{T}$ of random variable $O_i$, $\mathbb{G}_n[t(O_i)]\equiv n^{-1/2}\sum_{i=1}^n \bigl(t(O_i)-\mathbb{E}[t(O_i)]\bigr)$ denotes the empirical process indexed by the function class $\mathcal{T}$. For $2\times2$ matrices $\check{\mathcal{M}}$ and $\mathring{\mathcal{M}}$, let \begin{align} \boldsymbol{\xi}_{\mathcal{S},i}(\check{\mathcal{M}};\mathring{\mathcal{M}}q;\boldsymbol{\vartheta})\equiv \check{\mathcal{M}}\mathbf{L}_0+\check{\mathcal{M}}(\mathbf{L}_1-\mathbf{L}_0)\mathds{1}\{k(\boldsymbol{\vartheta},\mathring{\mathcal{M}}q)>0\}. \end{align} To show Eq. (ref), fix $q\in\mathbb{S}^1$ and decompose \begin{align*} \sqrt{n}\biggl(\widehat{h}_{\mathcal{E}}(q; \widehat{\boldsymbol{\theta}})-h_{\mathcal{E}}(q)\biggr) =\underbrace{\sqrt{n}\bigg(\frac{1}{n}\sum_{i=1}^nq^\intercal\big(\widehat{\mathcal{M}}\mathcal{M}^{-1}\big)\big(\boldsymbol{\xi}_{\mathcal{S},i}(\mathcal{M};\hcMq;\widehat{\boldsymbol{\theta}})-\mathbb{E}[\boldsymbol{\xi}_{\mathcal{S},i}(\mathcal{M};\cMq;\boldsymbol{\theta})]\big)\bigg)}_{\equiv A}&\\ +\underbrace{\sqrt{n}q^\intercal\big(\widehat{\mathcal{M}}\mathcal{M}^{-1}-\boldsymbol{I}\big)\mathbb{E}[\boldsymbol{\xi}_{\mathcal{S},i}(\mathcal{M};\cMq;\boldsymbol{\theta})],}_{\equiv B}& \end{align*} where $\boldsymbol{I}$ is the $2\times 2$ identity matrix, and \begin{align} &A=\mathbb{G}_n[\zeta_i(\cMq;\boldsymbol{\theta})] +o_p(1)\notag\\ &+\bigg\{\sqrt{n}\biggl(\frac{1}{n}\sum_{i=1}^n(\cMq)^\intercal\big(\mathbf{L}_{0_i}+(\mathbf{L}_{1_i}-\mathbf{L}_{0_i})\mathds{1}\{k(\widehat{\boldsymbol{\theta}},\hcMq)>0\}\big)\notag\\ &-\mathbb{E}\left[(\cMq)^\intercal\big(\mathbf{L}_{0_i}+(\mathbf{L}_{1_i}-\mathbf{L}_{0_i})\mathds{1}\{k(\widehat{\boldsymbol{\theta}},\hcMq)>0\}\big)\right]\biggr)\notag\\ &\underbrace{-\sqrt{n}\biggl(\frac{1}{n}\sum_{i=1}^n\zeta_i(\cMq;\boldsymbol{\theta})-\mathbb{E}[\zeta_i(\cMq;\boldsymbol{\theta})]\biggr)\biggr\}}_{\equiv R_1}\\ &+\underbrace{\sqrt{n}\biggl(\mathbb{E}\left[(\cMq)^\intercal\big(\mathbf{L}_{0_i}+(\mathbf{L}_{1_i}-\mathbf{L}_{0_i})\mathds{1}\{k(\widehat{\boldsymbol{\theta}},\hcMq)>0\}\big)\right]-\mathbb{E}[\zeta_i(\cMq;\boldsymbol{\theta})]\biggr)}_{\equiv R_2}, \end{align} where the equality follows from $\|\widehat{\mathcal{M}}\mathcal{M}-\boldsymbol{I}\|_{\max}=O_p(n^{-1/2})$ under the theorem's assumptions and adding and subtracting terms and $R_1$ is an empirical process term indexed by $\Delta\zeta_i\big(q;\widehat{\boldsymbol{\theta}},\widehat{\mathcal{M}}\big)\equiv(\cMq)^\intercal(\mathbf{L}_{1_i}-\mathbf{L}_{0_i})\mathds{1}\big\{k\big(\widehat{\boldsymbol{\theta}},\hcMq\big)>0\big\}-(\cMq)^\intercal(\mathbf{L}_{1_i}-\mathbf{L}_{0_i})\mathds{1}\big\{k\big(\boldsymbol{\theta},\cMq\big)>0\big\}$. Recall that $k(\widehat{\boldsymbol{\theta}} ,\hcMq)=q^\intercal\widehat{\mathcal{M}}\widehat{\Delta\boldsymbol{\theta}}$, where $\widehat{\Delta\boldsymbol{\theta}}(X_i)=\widehat{\Delta\boldsymbol{\theta}}_k(X_i)$ for $i\in I_k$. Fixing any $k\in[K]$ and conditional on the $I_k^c$ sample, $\Delta\widehat{\boldsymbol{\theta}}_k$ is non-stochastic and \begin{align} \mathbb{E}[R_1^2 | I_k^c] &\lesssim\sup_{\Delta\boldsymbol{\vartheta}\in\Theta_n}\mathbb{E}\biggl[\Delta\zeta_i\big(q;\widehat{\boldsymbol{\theta}},\widehat{\mathcal{M}}\big)^2\biggr]\notag\\ &=\sup_{\Delta\boldsymbol{\vartheta}\in\Theta_n}\mathbb{E}\left[\big((\cMq)^\intercal(\mathbf{L}_{1_i}-\mathbf{L}_{0_i})\big)^2\left|\mathds{1}\{q^\intercal\widehat{\mathcal{M}}\Delta\boldsymbol{\vartheta}>0\}-\mathds{1}\{q^\intercal\mathcal{M}\Delta\boldsymbol{\theta}>0\}\right|\right]\notag\\ &\lesssim\sup_{\Delta\boldsymbol{\vartheta}\in\Theta_n}\mathbb{E}\left[\big((\cMq)^\intercal(\mathbf{L}_{1_i}-\mathbf{L}_{0_i})\big)^2\mathds{1}\big\{|q^\intercal\mathcal{M}\Delta\boldsymbol{\theta}|<|q^\intercal(\widehat{\mathcal{M}}\Delta\boldsymbol{\vartheta}-\mathcal{M}\Delta\boldsymbol{\theta})|\big\}\right]\notag\\ &\lesssim\sup_{\Delta\boldsymbol{\vartheta}\in\Theta_n}\mathbb{P}\left(|q^\intercal\mathcal{M}\Delta\boldsymbol{\theta}|<|q^\intercal(\widehat{\mathcal{M}}\Delta\boldsymbol{\vartheta}-\mathcal{M}\Delta\boldsymbol{\theta})|\right)=o_p(1), \end{align} where the first inequality follows from che:che:dem:duf:han:whi:rob18, the third inequality follows by Assumption (ref), and the last equality follows from the definition of $\Theta_n$, $\|\widehat{\mathcal{M}}-\mathcal{M}\|_{\max}=O_p(n^{-1/2})$, and a similar argument to that of Eq. (ref). Hence $|R_1|=o_p(1)$. In addition, by Proposition (ref), $R_2$ encapsulates higher-order Gateaux derivatives and can be bounded by \begin{align*} |R_2|\lesssim\sqrt{n} \underbrace{\mathbb{E}\left[\mathds{1}\biggl\{|q^\intercal\mathcal{M}\Delta\boldsymbol{\theta}|<|q^\intercal(\widehat{\mathcal{M}}\widehat{\Delta\boldsymbol{\theta}}-\mathcal{M}\Delta\boldsymbol{\theta})|\biggr\}|q^\intercal\mathcal{M}\Delta\boldsymbol{\theta}|\right]}_{R_3}, \end{align*} and we show $|R_3|=o_p(n^{-1/2})$ by leveraging the proof technique from sem23 under Assumption (ref). It then follows that $|R_2|=o_p(1)$. Under Assumption (ref), for $g\in\{r,b\}$ and the $k$-th fold, \begin{align*} &\bigg|\frac{(\widehat{\Delta\theta}^g)_k(X)}{\widehat{\mu}_g}-\frac{\Delta\theta^g(X)}{\mu_g}\bigg|\\ \lesssim &\bigg|\frac{(\widehat{\alpha}^g)_k}{\widehat{\mu}_g}-\frac{\alpha^g}{\mu_g}\bigg|+\bigg|\frac{(\widehat{\beta}^g)_k}{\widehat{\mu}_g}-\frac{\beta^g}{\mu_g}\bigg|+\bigg|\frac{(\widehat{\eta}^g)_k(X_{[3:d_X]})}{\widehat{\mu}_g}-\frac{\eta^g(X_{[3:d_X]})}{\mu_g}\bigg|\equiv\delta_k^g(X_{[3:d_X]}), \end{align*} where $\lesssim$ follows from the assumption that $(X_1, X_2)$ has bounded support. Let $\delta_k(X_{[3:d_X]})\equiv\sum_{g\in\{r,b\}}\delta_k^g(X_{[3:d_X]})$. Denote the distribution of $|q^\intercal\mathcal{M}\Delta\boldsymbol{\theta}(X)|$ conditional on $X_{[3:d_X]}$ by $\mathbb{P}_{\bigl(|q^\intercal\mathcal{M}\Delta\boldsymbol{\theta}|~\bigl|X_{[3:d_X]}\bigr)}$. Conditional on $X$ and the sample $I_k^c$, $|q^\intercal(\widehat{\mathcal{M}}\widehat{\Delta\boldsymbol{\theta}}_k-\mathcal{M}\Delta\boldsymbol{\theta})|\leq\|\widehat{\mathcal{M}}\widehat{\Delta\boldsymbol{\theta}}_k-\mathcal{M}\Delta\boldsymbol{\theta}\|_E\lesssim\delta_k(X_{[3:d_X]})$, and \begin{align*} |R_3|\lesssim\mathbb{E}_{X_{[3:d_X]}}\Bigg[\int_{-\delta_k(X_{[3:d_X]})}^{\delta_k(X_{[3:d_X]})}\delta_k(X_{[3:d_X]})\mathbb{P}_{\bigl(|q^\intercal\mathcal{M}\Delta\boldsymbol{\theta}| \bigl|X_{[3:d_X]}\bigr)}\Biggr]\lesssim\mathbb{E}_{X_{[3:d_X]}}\big[\delta_k(X_{[3:d_X]})^2\big]=o_p(n^{-1/2}), \end{align*} where the second inequality follows from Assumption (ref) that $\mathbb{P}_{\bigl(|q^\intercal\mathcal{M}\Delta\boldsymbol{\theta}|~\bigl|X_{[3:d_X]}\bigr)}$ is bounded, and the last equality follows from $\widehat{\Delta\boldsymbol{\theta}}_k\in\Theta_n$ as $n\to\infty$, $\|\widehat{\mathcal{M}}-\mathcal{M}\|_{\max}=O_p(n^{-1/2})$, and repeated application of triangle inequality. Therefore, \begin{align*} A=\mathbb{G}_n[\zeta_i(\cMq;\boldsymbol{\theta})]+o_p(1). \end{align*} In addition, \begin{align*} B=\sqrt{n}\big((\widehat{\mathcal{M}}-\mathcal{M})q\big)^\intercal\mathcal{M}^{-1}\mathcal{S}_\mathcal{E}(q)=\mathbb{G}_n[(\mathcal{M}_i^*q)^\intercal\mathcal{M}^{-1}\mathcal{S}_\mathcal{E}(q)]+o_p(1) \end{align*} for $\mathcal{M}_i^*\equiv\operatorname{diag}\left(\frac{\mathds{1}\{G_i=r\}}{-\mu_r^2}, \frac{\mathds{1}\{G_i=b\}}{-\mu_b^2}\right)$ by the Delta method. We hence conclude \begin{align*} \sqrt{n}\biggl(\widehat{h}_{\mathcal{E}}(q; \widehat{\boldsymbol{\theta}})-h_{\mathcal{E}}(q; \boldsymbol{\theta})\biggr)=\mathbb{G}_n[\zeta_i^*(\cMq;\boldsymbol{\theta})]+o_p(1) \end{align*} for $\zeta_i^*(\cMq;\boldsymbol{\theta})\equiv\zeta_i(\cMq;\boldsymbol{\theta})+(\mathcal{M}_i^*q)^\intercal\mathcal{M}^{-1}\mathcal{S}_\mathcal{E}(q)$. \noindentPart 2. Since the class of functions over the random variables $(\mathbf{L}_{1_i}, \mathbf{L}_{0_i}, X_i)$, \begin{align*} \mathcal{Z}\equiv\bigl\{\zeta_i^*(\cMq;\boldsymbol{\theta}): q\in\mathbb{S}^1\bigr\}, \end{align*} is composed of linear functions (linear in $q\in\mathbb{S}^1$) and their indicators, $\mathcal{Z}$ is VC-subgraph and94. An envelope function of $\mathcal{Z}$ is \begin{align} \sum_{d\in\{0,1\},g\in\{r,b\}}(|L_{d_i}^g|+C)/c_1. \end{align} As $C$ and $c_1$ are constant and $L_{d_i}^g$ is square-integrable, Eq. (ref) is square-integrable under $\mathbb{P}$. By van:wel13, $\mathcal{Z}$ is $\mathbb{P}$-Donsker, and hence \begin{align*} \sqrt{n}\biggl(\widehat{h}_{\mathcal{E}}(q; \widehat{\boldsymbol{\theta}})-h_{\mathcal{E}}(q)\biggr)=\mathbb{G}[ \zeta_i^*(\cMq;\boldsymbol{\theta})]+o_p(1) \quad in \,\, \ell^\infty(\mathbb{S}^1). \end{align*} \noindentPart 3. The fact that $Var\left(\mathbb{G}[\zeta_i(q;\boldsymbol{\theta})]\right)>0$ for each $q\in\mathbb{S}^1$ follows using the variance decomposition formula and the same argument as in ber:mol08. \end{proof} \subsection{Proofs for Section (ref)} \begin{proof}[Proof of Proposition (ref)] Recalling the definition of the set $Q_{\mathbb{T}}^*(e),~\mathbb{T}\in\{\mathbb{S}^1,\tilde{\mathbb{S}}^1\}$ in Proposition (ref), by the same argument as in the proof of Proposition (ref) below, \begin{align*} \left[\sqrt{n}\max_{q\in\mathbb{S}^1}(q^\intercal e-\widehat{h}_\mathcal{E}(q;\widehat{\boldsymbol{\theta}}))\right]_{+}\xrightarrow[]{d}\left[\sup_{q\in Q^*_{\mathbb{S}^1}(e)}\mathbb{G}[-{\zeta_i^*(\cMq;\boldsymbol{\theta})}]\right]_+. \end{align*} Applying the argument in the proof of Proposition (ref) to the second maximization problem in the definition of $T_n^\mathcal{F}$ in Eq. (ref), we have \begin{align} \psi^\mathcal{F}(e)=\left[\sup_{q\in Q^*_{\mathbb{S}^1}(e)}\mathbb{G}[-{\zeta_i^*(\cMq;\boldsymbol{\theta})}]\right]_+ + \left[\sup_{q\in Q^*_{\tilde{\mathbb{S}}^1}(e)}\mathbb{G}[-{\zeta_i^*(-\cMq;\boldsymbol{\theta})}]\right]_-. \end{align} When Eq. (ref) holds, $\mathcal{E}$ has no kinks or flat faces (by Assumption (ref)); hence under the null: \begin{align*} Q^*_{\mathbb{S}^1}(e)=\arg\max_{q\in\mathbb{S}^1}\left(q^\intercal e-h_\mathcal{E}(q)\right)=\left\{q^*_{\mathbb{S}^1}(e)\right\}, Q^*_{\tilde{\mathbb{S}}^1}(e)=\arg\max_{q\in\tilde{\mathbb{S}}^1}\left(-h_{\mathcal{C}(e)}(q)-h_\mathcal{E}(-q)\right)\equiv\{q^*_{\tilde{\mathbb{S}}^1}(e)\}. \end{align*} Under the null $e\in\mathcal{F}$, $(q^*_{\mathbb{S}^1}(e))^\intercal e-h_\mathcal{E}(q^*_{\mathbb{S}^1}(e))=0$ and $e=\mathcal{S}_\mathcal{E}(q^*_{\mathbb{S}^1}(e))$; since $e\in\mathcal{F}$, $-(q^*_{\tilde{\mathbb{S}}^1}(e))^\intercal e-h_\mathcal{E}(-q^*_{\tilde{\mathbb{S}}^1}(e))=0$, so that $e=\mathcal{S}_\mathcal{E}(-q^*_{\tilde{\mathbb{S}}^1}(e))$. By Eq. (ref) kinks are absent, and hence $q^*_{\mathbb{S}^1}(e)=-q^*_{\tilde{\mathbb{S}}^1}(e)$. We therefore obtain the expression in Eq. (ref), as \begin{align*} \psi^\mathcal{F}(e)=\left[\mathbb{G}[-{\zeta_i^*(\cMq^*_{\mathbb{S}^1}(e);\boldsymbol{\theta})}]\right]_+ + \mathbb{G}\left[-{\zeta_i^*(\cMq^*_{\mathbb{S}^1}(e);\boldsymbol{\theta})}\right]_- =\left|\mathbb{G}\left[{\zeta_i^*(\cMq^*_{\mathbb{S}^1}(e);\boldsymbol{\theta})}\right]\right|. \end{align*} Eq. (ref) follows by standard arguments. We establish Eq. (ref) by verifying Condition C.1 in che:hon:tam07, under which Eq. (ref) follows from their Theorem 3.1-(1). By definition, the parameter space $\mathbb{B}_C$ is a compact (and convex) set. Our criterion function is \begin{align*} \mathsf{f}(e)\equiv\left[\max_{q\in\mathbb{S}^1}(q^\intercal e-h_\mathcal{E}(q))\right]_{+}+\left[\max_{q\in\Tilde{\mathbb{S}}^1}(-h_{\mathcal{C}(e)}(q)-h_\mathcal{E}(-q))\right]_-, \end{align*} and the population set is $\mathcal{F}=\{e\in\mathbb{B}_C:~\mathsf{f}(e)=0\}$. The criterion function $\mathsf{f}(e)$ is lower-semicontinuous by Berge's Maximum Theorem and composition with a continuous function. The sample criterion function \begin{align*} \widehat{\mathsf{f}}(e)\equiv \left[\max_{q\in\mathbb{S}^1}(q^\intercal e-\widehat{h}_\mathcal{E}(q;\widehat{\boldsymbol{\theta}}))\right]_{+}+\left[\max_{q\in\tilde\mathbb{S}^1}(-h_{\mathcal{C}(e)}(q)-\widehat{h}_\mathcal{E}(-q;\widehat{\boldsymbol{\theta}}))\right]_- \end{align*} takes values in $\mathbb{R}_+$ and is jointly measurable in the parameter $e\in\mathcal{E}$ and the data by standard arguments. Finally, by Proposition (ref), $\sup_{e\in\mathbb{B}_C}\left(\mathsf{f}(e)-\widehat{\mathsf{f}}(e)\right)_+=O_p(1/\sqrt{n})$ and $\sup_{e\in\mathcal{F}}~\widehat{\mathsf{f}}(e)=O_p(1/\sqrt{n})$. \end{proof} \begin{proof}[Proof of Proposition (ref)] Recall that by Proposition (ref), $\mathcal{S}_\mathcal{E}(q)=\mathbb{E}\bigl[\boldsymbol{\zeta}_{\mathcal{S}}(\cMq;\boldsymbol{\theta})\bigr]$ for $\boldsymbol{\zeta}_{\mathcal{S}}(\cMq;\boldsymbol{\theta})\equiv\mathcal{M}\mathbf{L}_0+\mathcal{M}(\mathbf{L}_1-\mathbf{L}_0)\mathds{1}\{k(\boldsymbol{\theta},\cMq)>0\}$. Using the notation $\boldsymbol{\xi}_{\mathcal{S},i}(\check{\mathcal{M}};\mathring{\mathcal{M}}q;\boldsymbol{\vartheta})$ defined in Eq. (ref), we have: \begin{subequations} \begin{align} &\left\Vert \widehat{\mathcal{S}}_\mathcal{E}(q;\widehat{\boldsymbol{\theta}})-\mathcal{S}_\mathcal{E}(q)\right\Vert= \left\Vert\frac{1}{n}\sum_{i=1}^n\boldsymbol{\zeta}_{\mathcal{S},i}(\hcMq;\widehat{\boldsymbol{\theta}})-\mathcal{S}_\mathcal{E}(q)\right\Vert\notag\\ \le&\left\Vert \widehat{\mathcal{M}}\mathcal{M}^{-1}\right\Vert \Bigg\Vert \frac{1}{n}\sum_{i=1}^n\boldsymbol{\xi}_{\mathcal{S},i}(\mathcal{M};\hcMq;\widehat{\boldsymbol{\theta}})-\mathbb{E}\left[\boldsymbol{\xi}_{\mathcal{S},i}(\mathcal{M};\cMq;\boldsymbol{\theta})\right]\Bigg\Vert\notag\\ &+\left\Vert \left(\widehat{\mathcal{M}}\mathcal{M}^{-1}-\boldsymbol{I}\right) \mathbb{E}[\boldsymbol{\zeta}_{\mathcal{S},i}(\cMq;\boldsymbol{\theta})]\right\Vert\notag\\ \le&\Bigg\Vert \frac{1}{n}\sum_{i=1}^n\boldsymbol{\xi}_{\mathcal{S},i}(\mathcal{M};\cMq;\boldsymbol{\theta})-\mathbb{E}\left[\boldsymbol{\xi}_{\mathcal{S},i}(\mathcal{M};\cMq;\boldsymbol{\theta})\right]\Bigg\Vert\\ &+\Bigg\Vert \left(\frac{1}{n}\sum_{i=1}^n\boldsymbol{\xi}_{\mathcal{S},i}(\mathcal{M};\hcMq;\widehat{\boldsymbol{\theta}})-\mathbb{E}\left[\boldsymbol{\xi}_{\mathcal{S},i}(\mathcal{M};\hcMq;\widehat{\boldsymbol{\theta}})\right]\right)\notag\\ &-\left(\frac{1}{n}\sum_{i=1}^n\boldsymbol{\xi}_{\mathcal{S},i}(\mathcal{M};\cMq;\boldsymbol{\theta})-\mathbb{E}\left[\boldsymbol{\xi}_{\mathcal{S},i}(\mathcal{M};\cMq;\boldsymbol{\theta})\right]\right)\Bigg\Vert\\ &+\left\Vert \frac{1}{n}\sum_{i=1}^n\mathbb{E}\left[\boldsymbol{\xi}_{\mathcal{S},i}(\mathcal{M};\hcMq;\widehat{\boldsymbol{\theta}})\right]-\mathbb{E}\left[\boldsymbol{\xi}_{\mathcal{S},i}(\mathcal{M};\cMq;\boldsymbol{\theta})\right] \right\Vert+o_p(1), \end{align} \end{subequations} where the $o_p(1)$ term in line (ref) follows as $\Vert \widehat{\mathcal{M}}\mathcal{M}^{-1}-\boldsymbol{I}\Vert_{\max}=O_p(n^{-1/2})$. The term in Eq. (ref) equals $o_p(1)$ by the Law of Large Numbers. Using that $\mathbb{E}\left[\mathcal{M}(\mathbf{L}_1-\mathbf{L}_0)|X\right]=\mathcal{M}(\boldsymbol{\theta}_1(X)-\boldsymbol{\theta}_0(X))=[-\mathfrak{u}_1^\intercal\mathcal{M} \Delta\boldsymbol{\theta}(X),~-\mathfrak{u}_2^\intercal\mathcal{M} \Delta\boldsymbol{\theta}(X)]^\intercal$ and omitting dependence on $X$ to shorten notation, we have that the first term of Eq. (ref) is upper bounded by \begin{align*} &\sup_{\Delta\boldsymbol{\vartheta}\in\Theta_n}\Big\Vert\mathbb{E}\Bigl[\mathcal{M}(\mathbf{L}_{1_i}-\mathbf{L}_{0_i})\left(\mathds{1}\bigl\{q^\intercal\widehat{\mathcal{M}}\Delta\boldsymbol{\vartheta}>0\bigr\}-\mathds{1}\bigl\{q^\intercal\mathcal{M}\Delta\boldsymbol{\theta}>0\bigr\}\right)\Bigr]\Big\Vert\\ =&\sup_{\Delta\boldsymbol{\vartheta}\in\Theta_n}\Big\Vert\mathbb{E}\Bigl[[-\mathfrak{u}_1^\intercal\mathcal{M} \Delta\boldsymbol{\theta}, -\mathfrak{u}_2^\intercal\mathcal{M} \Delta\boldsymbol{\theta}]^\intercal\left(\mathds{1}\bigl\{q^\intercal\widehat{\mathcal{M}}\Delta\boldsymbol{\vartheta}>0\bigr\}-\mathds{1}\bigl\{q^\intercal\mathcal{M}\Delta\boldsymbol{\theta}>0\bigr\}\right)\Bigr]\Big\Vert\\ \lesssim &\sup_{\Delta\boldsymbol{\vartheta}\in\Theta_n}\max_{v\in\{-\mathfrak{u}_1,-\mathfrak{u}_2\}}\Big|\mathbb{E}\Bigl[v^\intercal\mathcal{M} \Delta\boldsymbol{\theta}\left(\mathds{1}\bigl\{q^\intercal\widehat{\mathcal{M}}\Delta\boldsymbol{\vartheta}>0\bigr\}-\mathds{1}\bigl\{q^\intercal\mathcal{M}\Delta\boldsymbol{\theta}>0\bigr\}\right)\Bigr]\Big|\\ \le &\sup_{\Delta\boldsymbol{\vartheta}\in\Theta_n}\max_{v\in\{-\mathfrak{u}_1,-\mathfrak{u}_2\}}\mathbb{E}\Bigl[\big|v^\intercal\mathcal{M}\Delta\boldsymbol{\theta}\big|\mathds{1}\bigl\{\big|q^\intercal\mathcal{M}\Delta\boldsymbol{\theta}\big|<\big|q^\intercal\widehat{\mathcal{M}}\Delta\boldsymbol{\vartheta}-q^\intercal\mathcal{M}\Delta\boldsymbol{\theta}\big|\bigr\}\Bigr]\\ \lesssim &\sup_{\Delta\boldsymbol{\vartheta}\in\Theta_n}\mathbb{E}\left[\mathds{1}\bigl\{\big|q^\intercal\mathcal{M}\Delta\boldsymbol{\theta}\big|<\big|q^\intercal\widehat{\mathcal{M}}\Delta\boldsymbol{\vartheta}-q^\intercal\mathcal{M}\Delta\boldsymbol{\theta}\big|\bigr\}\right]^{1/2}=o_p(1) \end{align*} where the last inequality follows by Cauchy-Schwartz and Assumption (ref), and the last equality follows from Assumption (ref), the fact that $\Delta\boldsymbol{\vartheta}\in\Theta_n$, and that under the proposition's assumptions $\Vert \widehat{\mathcal{M}}\mathcal{M}^{-1}-\boldsymbol{I}\Vert_{\max}=O_p(n^{-1/2})$; one can then show Eq. (ref) is $o_p(1)$ by the same argument as that used to establish Eq. (ref) in the proof of Theorem (ref). We conclude by establishing Eq. (ref). Using the definition of Hausdorff distance, \begin{align*} \mathbf{d}_H(\widehat{\mathcal{PF}},\mathcal{PF})=\max\left\{\sup_{\hat{e}\in\widehat{\mathcal{PF}}}\inf_{e\in\mathcal{PF}}\left\Vert\hat{e}-e\right\Vert, \sup_{e\in\mathcal{PF}}\inf_{\hat{e}\in\widehat{\mathcal{PF}}}\left\Vert\hat{e}-e\right\Vert\right\} \end{align*} Suppose by contradiction that there is $e^*\in\mathcal{PF}$ such that for all $n\ge 1$, $\inf_{\hat{e}\in\widehat{\mathcal{PF}}}\left\Vert\hat{e}-e^*\right\Vert>c$ for some constant $c>0$. By definition, there exists $q^*\in\mathbb{Q}$ such that $\mathcal{S}_\mathcal{E}(q^*)=e^*$. By Eq. (ref), $\Vert\widehat{\mathcal{S}}_\mathcal{E}(q^*;\widehat{\boldsymbol{\theta}})-e^*\Vert=o_p(1)$, and by definition $\widehat{\mathcal{S}}_\mathcal{E}(q^*;\widehat{\boldsymbol{\theta}})\in\widehat{\mathcal{PF}}$, yielding a contradiction. The same argument holds by swapping the role of $\widehat{\mathcal{PF}}$ and $\mathcal{PF}$. \end{proof} \begin{proof}[Proof of Proposition (ref)] Our proof follows arguments in kai16. We first observe that for $e\in\mathcal{PF}$, $\max_{q\in\mathbb{S}^1}(q^\intercal e - h_\mathcal{E}(q)) = 0$ and $\max_{q\in\mathbb{Q}}(q^\intercal e - h_\mathcal{E}(q))=0$. Hence, for $\mathbb{T}\in\{\mathbb{S}^1,\mathbb{Q}\}$ and $\phi_{e,\mathbb{T}}(f)\equiv\sup_{q\in\mathbb{T}}(q^\intercal e-f(q))$, we can write \begin{multline*} T_n^\mathcal{PF}(e) =\max\left\{\sqrt{n}\left(\phi_{e,\mathbb{S}^1}\big(\widehat{h}_\mathcal{E}(q;\widehat{\boldsymbol{\theta}})\big)-\phi_{e,\mathbb{S}^1}\big(h_\mathcal{E}(q)\big)\right),0\right\} \\ - \min\left\{\sqrt{n}\left(\phi_{e,\mathbb{Q}}\big(\widehat{h}_\mathcal{E}(q;\widehat{\boldsymbol{\theta}})\big)-\phi_{e,\mathbb{Q}}\big(h_\mathcal{E}(q)\right),0\right\}. \end{multline*} kai16 establishes that $\phi_{e,\mathbb{T}}$ is Hadamard directionally differentiable at $h_\mathcal{E}(\cdot)$ with Hadamard directional derivative equal to $\phi^\prime_{e,\mathbb{T}}(y)=\sup_{q\in Q^*_\mathbb{T}(e)}-y(q)$. By Theorem (ref), the assumptions in kai16 are satisfied, and the result follows by the Continuous Mapping Theorem and as argued in kai16. Absolute continuity of the limit law in Eq. (ref) follows using the same argument as in ber:mol08. \end{proof} \begin{proof}[Proof of Proposition (ref)] Take a point $e^0$ in the set $\{\tilde{e}: \Vert \tilde{e}\Vert=2C~\text{and}~\tilde{e}_r+\tilde{e}_b\le 0\}$. Select $\widehat{e}_{n_1}$ as the metric projection of $e^0$ onto $\widehat{\mathcal{F}}$ and let $e$ be the metric projection of $e^0$ onto $\mathcal{F}$. By the consistency result in Proposition (ref) and by mol17, $\Vert \widehat{e}_{n_1} - e\Vert \xrightarrow[]{p} 0$ as $n_1\to\infty$. Hence, given that by Theorem (ref) the support function estimator converges to the population support function uniformly in $q\in\mathbb{S}^1$, and given that $\widehat{q}^*_{n_1}(\widehat{e}_{n_1})$ is an extremum estimator and all the conditions for its consistency are satisfied, $\Vert \widehat{q}^*_{n_1}(\widehat{e}_{n_1})-q^*_{\mathbb{S}^1}(e)\Vert \xrightarrow[]{p} 0$ as $n_1\to\infty$. Next, using the same argument as in the proof of Proposition (ref) leading to Eqs. (ref)-(ref), and $\boldsymbol{\xi}_{\mathcal{S},i}(\check{\mathcal{M}};\mathring{\mathcal{M}}q;\boldsymbol{\vartheta})$ defined in Eq. (ref), \begin{subequations} \begin{align} \Big\Vert\widehat{\mathcal{S}}_\mathcal{E}(\widehat{q}^*_{n_1}(\widehat{e}_{n_1});\widehat{\boldsymbol{\theta}}_{n_1}) - \mathcal{S}_\mathcal{E}(q^*_{\mathbb{S}^1}(e))\Big\Vert& \notag\\ = \left\Vert \frac{1}{n_2}\sum_{i=1}^{n_2}\boldsymbol{\zeta}_{\mathcal{S},i}(\widehat{\mathcal{M}}\widehat{q}^*_{n_1}(\widehat{e}_{n_1});\widehat{\boldsymbol{\theta}}_{n_1})-\mathbb{E}[\boldsymbol{\zeta}_{\mathcal{S},i}(\cMq^*_{\mathbb{S}^1}(e);\boldsymbol{\theta})]\right\Vert&\notag\\ \le o_p(1)+\Bigg\Vert \frac{1}{n_2}\sum_{i=1}^{n_2}\boldsymbol{\xi}_{\mathcal{S},i}(\mathcal{M};\cMq^*_{\mathbb{S}^1}(e);\boldsymbol{\theta})-\mathbb{E}\left[\boldsymbol{\xi}_{\mathcal{S},i}(\mathcal{M};\cMq^*_{\mathbb{S}^1}(e);\boldsymbol{\theta})\right]\Bigg\Vert& \\ +\left\Vert \left(\frac{1}{n_2}\sum_{i=1}^{n_2}\boldsymbol{\xi}_{\mathcal{S},i}(\mathcal{M};\widehat{\mathcal{M}}\widehat{q}^*_{n_1}(\widehat{e}_{n_1});\widehat{\boldsymbol{\theta}}_{n_1})-\mathbb{E}\left[\boldsymbol{\xi}_{\mathcal{S},i}(\mathcal{M};\widehat{\mathcal{M}}\widehat{q}^*_{n_1}(\widehat{e}_{n_1});\widehat{\boldsymbol{\theta}}_{n_1})\right]\right)\right. &\notag\\ \left.-\left(\frac{1}{n_2}\sum_{i=1}^{n_2}\boldsymbol{\xi}_{\mathcal{S},i}(\mathcal{M};\cMq^*_{\mathbb{S}^1}(e);\boldsymbol{\theta})-\mathbb{E}\left[\boldsymbol{\xi}_{\mathcal{S},i}(\mathcal{M};\cMq^*_{\mathbb{S}^1}(e);\boldsymbol{\theta})\right]\right)\right\Vert& \\ +\left\Vert \mathbb{E}\left[\boldsymbol{\xi}_{\mathcal{S},i}(\mathcal{M};\widehat{\mathcal{M}}\widehat{q}^*_{n_1}(\widehat{e}_{n_1});\widehat{\boldsymbol{\theta}}_{n_1})\right]-\mathbb{E}\left[\boldsymbol{\xi}_{\mathcal{S},i}(\mathcal{M};\cMq^*_{\mathbb{S}^1}(e);\boldsymbol{\theta})\right] \right\Vert&. \end{align} \end{subequations} By the same argument as in the proof of Proposition (ref), the terms in Eqs. (ref)-(ref) are $o_p(1)$. We next show that the same holds for the term in Eq. (ref). Take any $\delta>0$, \begin{align*} &\Big\Vert \mathbb{E}\left[\boldsymbol{\xi}_{\mathcal{S},i}(\mathcal{M};\widehat{\mathcal{M}}\widehat{q}^*_{n_1}(\widehat{e}_{n_1});\widehat{\boldsymbol{\theta}}_{n_1})\right]-\mathbb{E}\left[\boldsymbol{\xi}_{\mathcal{S},i}(\mathcal{M};\cMq^*_{\mathbb{S}^1}(e);\boldsymbol{\theta})\right]\Big\Vert \\ \lesssim & \sup_{\Delta\boldsymbol{\vartheta}\in\Theta_n}\max_{v\in\{-\mathfrak{u}_1,\mathfrak{u}_2\}}\mathbb{E}\Bigl[\big|v^\intercal\mathcal{M}\Delta\boldsymbol{\theta}\big|\mathds{1}\bigl\{\big| q^*_{\mathbb{S}^1}(e)^\intercal\mathcal{M}\Delta\boldsymbol{\theta} \big|<\big| \widehat{q}^*_{n_1}(\widehat{e}_{n_1})^\intercal\widehat{\mathcal{M}}\Delta\boldsymbol{\vartheta}-q^*_{\mathbb{S}^1}(e)^\intercal\mathcal{M}\Delta\boldsymbol{\theta} \big|\bigr\}\Bigr]\\ \lesssim & \sup_{\Delta\boldsymbol{\vartheta}\in\Theta_n}\mathbb{E}\Bigl[\mathds{1}\bigl\{\big|q^*_{\mathbb{S}^1}(e)^\intercal\mathcal{M}\Delta\boldsymbol{\theta}\big|<\big|\widehat{q}^*_{n_1}(\widehat{e}_{n_1})^\intercal\widehat{\mathcal{M}}\Delta\boldsymbol{\vartheta}-q^*_{\mathbb{S}^1}(e)^\intercal\mathcal{M}\Delta\boldsymbol{\theta}\big|\bigr\}\Bigr]^{1/2}\\ \leq&\left(\mathbb{E}\Bigl[\mathds{1}\bigl\{\big|q^*_{\mathbb{S}^1}(e)^\intercal\mathcal{M}\Delta\boldsymbol{\theta}\big|<\delta\big\}\Bigr]+\mathbb{E}\Bigl[\mathds{1}\bigl\{\Vert\widehat{\mathcal{M}}\widehat{q}^*_{n_1}(\widehat{e}_{n_1})-\cMq^*_{\mathbb{S}^1}(e)\Vert>\delta\big\}\Bigr]\right)^{1/2}=o_p(1), \end{align*} using Assumption (ref), $\Vert \widehat{\mathcal{M}}\mathcal{M}^{-1}-\boldsymbol{I}\Vert_{\max}=O_p(n^{-1/2})$, Markov inequality, and that $\delta>0$ is arbitrary. \end{proof} \subsection{Proofs for Section (ref)} \begin{proof}[Proof of Proposition (ref)] By the same argument as in the proof of Proposition (ref) and the discussion in Section (ref), \begin{align} \liminf_{n\to\infty}\mathbb{P}\big\{(R,B)\in \mathcal{CS}_n(R,B)\big\}\geq 1-\alpha. \end{align} To obtain the result in Eq. (ref), observe that under the null in Eq. (ref), by Eq. (ref), \begin{align*} \mathbb{E}\left[\varphi_n^{\texttt{skew}}\right]&=1-\mathbb{\mathbb{P}}\bigg\{\sup_{(\Tilde{R}, \Tilde{B})\in \mathcal{CS}_n(R,B)}\big((\mathfrak{u_1}-\mathfrak{u_2})^\intercal \Tilde{R}\big)\big((\mathfrak{u_1}-\mathfrak{u_2})^\intercal \Tilde{B}\big)\geq0\bigg\}\\ &\leq 1-\mathbb{\mathbb{P}}\big\{(R,B)\in \mathcal{CS}_n(R,B)\big\}\leq \alpha. \end{align*} \end{proof} \begin{proof}[Proof of Proposition (ref)] For $g\in\{r,b\}$, recall $Z_i^g$ defined in Eq. (ref) and note that \begin{align*} \sqrt{n}\left(\widehat{e}_g^*-e_g^*\right)=&\sqrt{n}\left(\frac{1}{n}\sum_{i=1}^n \frac{Z_i^g}{\widehat{\mu}_g}-\frac{\mathbb{E}[Z_i^g]}{\mu_g}\right)\\ =&\frac{1}{\mu_g}\mathbb{G}_n[Z_i^g]-\frac{\mathbb{E}[Z_i^g]}{\mu_g^2}\mathbb{G}_n[\mathds{1}\{G_i=g\}]+ o_p(1) =\mathbb{G}_n[Z_i^{g,*}]+o_p(1), \end{align*} where \begin{align} Z_i^{g,*}\equiv\frac{Z_i^g}{\mu_g}-\frac{\mathbb{E}[Z_i^g]}{\mu_g^2}\mathds{1}\{G_i=g\}. \end{align} Since $\max_{d\in\{0,1\},g\in\{r,b\}}\mathbb{E}[(L_d^g)^2]<\infty$, $Z_i^{g,*}$ has finite first and second moments. By the Lindeberg–Lévy central limit theorem, we have $\sqrt{n}\left(\widehat{e}_g^*-e_g^*\right)=\mathbb{G}[Z_i^{g,*}]+o_p(1)$, where $\mathbb{G}[Z_i^{g,*}]$ is a mean-zero Gaussian random variable with variance $\mathbb{E}\bigl[(Z_i^{g,*})^2\bigr]$. By Theorem (ref) and the Cramér–Wold theorem, we have that jointly, \begin{align} \sqrt{n}\begin{bmatrix} \widehat{h}_\mathcal{E}(q; \widehat{\boldsymbol{\theta}})-h_\mathcal{E}(q)\\ \widehat{e}_r^*-e_r^*\\ \widehat{e}_b^*-e_b^* \end{bmatrix}\xrightarrow[]{d}\begin{bmatrix} \mathbb{G}[\zeta_i^*(\cMq;\boldsymbol{\theta})]\\ \mathbb{G}[Z_i^{r,*}]\\ \mathbb{G}[Z_i^{b,*}] \end{bmatrix}\equiv\mathbb{G}_{he^*} \quad \text{ in }\ell^\infty(\mathbb{B}^1), \end{align} Next, we analyze the two parts of the test statistic in Eq. (ref), beginning with the first part. By the same argument as in the proof of Proposition (ref), under the null that $e^*\in\mathcal{F}$ we have $\max_{q\in\mathbb{S}^1}(q^\intercal e^* - h_\mathcal{E}(q)) = 0$. Hence, for $\phi_{\mathbb{S}^1}\big(e,f(\cdot)\big)\equiv\sup_{q\in\mathbb{S}^1}\big(q^\intercal e-f(q)\big)$, we can write $\sqrt{n}\left[\max_{q\in\mathbb{S}^1}(q^\intercal \widehat{e}^*-\widehat{h}_\mathcal{E}(q;\widehat{\boldsymbol{\theta}}))\right]_{+}=\max\left\{\sqrt{n}\left(\phi_{\mathbb{S}^1}\big(\widehat{e}^*,\widehat{h}_\mathcal{E}(q;\widehat{\boldsymbol{\theta}})\big)-\phi_{\mathbb{S}^1}\big(e^*,h_\mathcal{E}(q)\big)\right),0\right\}$. kai16 shows that $\phi_{\mathbb{S}^1}$ is Hadamard directionally differentiable at $\big(e^*,h_\mathcal{E}(\cdot)\big)$ with derivative $\phi^\prime_{\mathbb{S}^1,(e^*,h_\mathcal{E}(\cdot))}\big(e,f(\cdot)\big)=\sup_{q\in Q^*_{\mathbb{S}^1}(e^*)}\big(q^\intercal e-f(q)\big)$. For the second part of the test statistic in Eq. (ref), note that for any $s,t\in\mathbb{R}$, $\min\{s,2t-s\}=(2t-s)-\max\{2(t-s),0\}$ and $\min\{t,2s-t\}=t-\max\{2(t-s),0\}$, so we can plug in $t=e_r^*$ and $s=e_b^*$ to rewrite Eq. (ref) as \begin{align*} &h_{\mathcal{C}^*}(q)=\max\biggl\{q_1(e_r^*-\max\{2(e_r^*-e_b^*),0\})+q_2 e_b^*,\,\,\, q_1 e_r^*+q_2\bigl(2e_r^*-e_b^*-\max\{2(e_r^*-e_b^*),0\}\bigr)\biggr\}\\ &=\max\biggl\{2q_2\bigr(e_r^*-e_b^*\bigr)+(q_1-q_2)\max\{2(e_r^*-e_b^*),0\},\,\,\,0 \biggr\}+q_1 e_r^* + q_2 e_b^* - q_1\max\{2(e_r^*-e_b^*),0\}. \end{align*} It follows that we can write \begin{align} \mathfrak{h}_1\big(h_\mathcal{E}(\cdot), e_r^*, e_b^*; q\big) \equiv&-h_{\mathcal{C}^*}(q)-h_\mathcal{E}(-q)\notag\\ =&-\bigg(\max\biggl\{2q_2\bigr(e_r^*-e_b^*\bigr)+(q_1-q_2)\max\{2(e_r^*-e_b^*),0\},\,\,\,0 \biggr\}\notag\\ &+q_1 e_r^* + q_2 e_b^* - q_1\max\{2(e_r^*-e_b^*),0\}\bigg)-h_\mathcal{E}(-q). \end{align} Hence, the second part of the test statistic in Eq. (ref) is the composition of two mappings applied to $\{\widehat{h}_\mathcal{E}(\cdot; \widehat{\boldsymbol{\theta}}),\widehat{e}_r^*,\widehat{e}_b^*\}$: $\mathfrak{h}_1(\,\cdot\,;q)$ in Eq. (ref) and $\mathfrak{h}_2(\cdot)\equiv\max_q(\cdot)$. Each of these mappings is Hadamard directionally differentiable at $\{h_\mathcal{E}(\cdot),e_r^*,e_b^*\}$ tangentially to $\ell^\infty(\tilde{\mathbb{S}}^1)\times\mathbb{R}^2$ (fan:san19, fan:san19, Example 2.1; car:cue:alb20, car:cue:alb20, Theorem 2.1). By Shapiro90, \begin{align} \max_{q\in\tilde{\mathbb{S}}^1}\left(-h_{\mathcal{C}^*}(q)-h_\mathcal{E}(-q)\right)=\mathfrak{h}_2\circ\mathfrak{h}_1\big(h_\mathcal{E}(\cdot), e_r^*, e_b^*;\,q\big)\equiv\mathfrak{h}\big( h_\mathcal{E}\big(\cdot), e_r^*, e_b^*\big)\equiv\mathfrak{h} \end{align} is directionally differentiable at $(h_\mathcal{E}(\cdot),e_r^*,e_b^*)$ tangentially to $\ell^\infty(\tilde{\mathbb{S}}^1)\times\mathbb{R}^2$, with \begin{align} \mathfrak{h}'(\cdot)=\mathfrak{h}'_{2,\mathfrak{h}_1^*(q)}\circ\mathfrak{h}'_{1,s^*}(\,\cdot ;q), \end{align} where $\mathfrak{h}_1^*(q)\equiv\mathfrak{h}_1\big(h_\mathcal{E}(\cdot),e_r^*,e_b^*\,;q\big)$ and $s^*\equiv(q_1-q_2)\max\{2(e_r^*-e_b^*),0\}+2q_2\bigr(e_r^*-e_b^*\bigr)$; for any $s_r,s_b,s\in\mathbb{R}$, $s_h\in\ell^\infty(\tilde{\mathbb{S}}^1)$ (the space of bounded functions over the compact set $\tilde{\mathbb{S}}^1$) and continuous $f\in\ell^\infty(\tilde{\mathbb{S}}^1)$, \begin{align} \mathfrak{h}'_{1,s^*}\big(s_h(\cdot), s_r, s_b\,;q\big)&=-\phi_{1,s^*}^\prime(2q_2[s_r-s_b]+2(q_1-q_2)\phi_{1,e_r^*-e_b^*}^\prime[s_r-s_b])\notag\\ &-q_1 s_r-q_2 s_b+2q_1\phi_{1,e_r^*-e_b^*}^\prime[s_r-s_b]-s_h(-q),\\ \mathfrak{h}'_{2,\mathfrak{h}_1^*(q)}(f)&= \max_{\bigl\{q\in\tilde{\mathbb{S}}^1:\mathfrak{h}_1^*(q)=\mathfrak{h}\bigr\}} f(q). \end{align} In Eq. (ref), for any $t\in\mathbb{R}$ \begin{align} \phi_{1,s^*}'(t)\equiv\begin{cases} t, &\text{ if $s^*>0$}\\ \max\{t,0\}, &\text{ if $s^*=0$} \\ 0, &\text{ if $s^*<0$} \end{cases} \end{align} is the Hadamard directional derivative of $\phi_1(s)\equiv\max\{s,0\}$ at $s^*$ fan:san19. Eq. (ref) results from car:cue:alb20, as $\mathfrak{h}_1(\cdot~;q)$ is a continuous function over compact support. By Shapiro90, $T_n^{\texttt{LDA}}$ is Hadamard directionally differentiable at $(h_\mathcal{E}(\cdot),e_r^*,e_b^*)$ tangentially to $\ell^\infty(\tilde{\mathbb{S}}^1)\times\mathbb{R}^2$, with directional derivative given by the sum of the two directional derivatives derived above. By kai16, fan:san19, and car:cue:alb20, for $\mathbb{G}[\mathbf{Z}_i^{*}]\equiv\left[\mathbb{G}[Z_i^{r,*}]~~~~~\mathbb{G}[Z_i^{b,*}]\right]^\intercal$, $\psi^{\texttt{LDA}}$ takes the form \begin{multline} \psi^{\texttt{LDA}}=\left[\sup_{q\in Q^*_{\mathbb{S}^1}(e^*)}q^\intercal\mathbb{G}[\mathbf{Z}_i^{*}]-\mathbb{G}[\zeta_i^*(\cMq;\boldsymbol{\theta})]\right]_+\\ +\left[\mathfrak{h}'_{2,\mathfrak{h}_1^*(q)} \left\{ \mathfrak{h}_{1,s^*}'\big(\mathbb{G}[\zeta_i^*(-\mathcal{M}(\,\cdot\,);\boldsymbol{\theta})], \mathbb{G}[Z_i^{r,*}], \mathbb{G}[Z_i^{b,*}]\,;q\big)\right\}\right]_-. \end{multline} If $c^{\texttt{LDA}}_{1-\alpha+\varsigma}+\varsigma$ is a continuity point of the distribution of $\psi^{\texttt{LDA}}$, \begin{align*} \lim_{n\to\infty}\mathbb{P}\bigl(T_n^{\texttt{LDA}}>c^{\texttt{LDA}}_{1-\alpha+\varsigma}+\varsigma\bigr) =\mathbb{P}\bigl(\psi^{\texttt{LDA}}>c^{\texttt{LDA}}_{1-\alpha+\varsigma}+\varsigma\bigr)\le\alpha. \end{align*} For $\varsigma$ small enough, if $c^{\texttt{LDA}}_{1-\alpha+\varsigma}+\varsigma$ is a discontinuity point of $\psi^{\texttt{LDA}}$, then $\psi^{\texttt{LDA}}$ is continuous at $c^{\texttt{LDA}}_{1-\alpha}$, and since $\mathbb{P}(T_n^{\texttt{LDA}}>c^{\texttt{LDA}}_{1-\alpha+\varsigma}+\varsigma)\le\mathbb{P}(T_n^{\texttt{LDA}}>c^{\texttt{LDA}}_{1-\alpha})$, the result follows. If Eq. (ref) is satisfied, $\mathcal{E}$ has no kinks or flat faces (by Assumption (ref)). Under the null, since $e^*\in\mathcal{F}$, $\left\{q\in\tilde{\mathbb{S}}^1:\mathfrak{h}_1^*(q)=\mathfrak{h}\right\}=\arg\max_{q\in\tilde{\mathbb{S}}^1}\left(-h_{\mathcal{C}^*}(q)-h_\mathcal{E}(-q)\right)=\{q^*_{\tilde{\mathbb{S}}^1}\}$ and $-(q^*_{\tilde{\mathbb{S}}^1})^\intercal e^*-h_\mathcal{E}(-q^*_{\tilde{\mathbb{S}}^1})=0$, so that $e^*=\mathcal{S}_\mathcal{E}(-q^*_{\tilde{\mathbb{S}}^1})$. It then follows that under Eq. (ref), \begin{multline} \psi^{\texttt{LDA}}=\left[q^{*\intercal}_{\mathbb{S}^1}\mathbb{G}[\mathbf{Z}_i^{*}]-\mathbb{G}[\zeta_i^*(\cMq^*_{\mathbb{S}^1};\boldsymbol{\theta})]\right]_+ \\ + \left[\mathfrak{h}_{1,s^*}'\big(\mathbb{G}[\zeta_i^*(-\mathcal{M}(\cdot);\boldsymbol{\theta})], \mathbb{G}[Z_i^{r,*}], \mathbb{G}[Z_i^{b,*}]\,;q^*_{\mathbb{S}^1}\big)\right]_-. \end{multline} \end{proof} \begin{proof}[Proof of Proposition (ref)] We verify the assumptions in Theorem 3.2 of fan:san19, from which it follows that Proposition (ref) holds and that $\widehat{c}_\beta=c_\beta+o_p(1)$ fan:san19 when $c_\beta$ is a point at which the cdf of $\phi_{he^*}'(\mathbb{G}_{he^*})$ is continuous and increasing. Assumptions 1, 3(i), 3(iii), and 3(iv) of fan:san19 hold by construction; their Assumption 2 holds by Eq. (ref), and Assumption 4 holds by Lemma S.3.8 of fan:san19. Lastly, to show Assumption 3(ii) holds, let $\mathcal{O}_n\equiv\{(Y_i,G_i,X_i)\}_{i=1}^n$. In the next paragraph, we establish that \begin{align} &\sup_{f\in\mathcal{BL}_1}\left|\mathbb{E}\left[f\left(\sqrt{n}\{\widetilde{he^*}(\widehat{\boldsymbol{\theta}})-\widehat{he^*}(\widehat{\boldsymbol{\theta}})\}\right)\,\bigg|\,\mathcal{O}_n\right]-\mathbb{E}[f(\mathbb{G}_{he^*})]\right|\notag\\ &= \sup_{f\in\mathcal{BL}_1}\left|\mathbb{E}\left[f\left(\begin{bmatrix} \mathbb{G}_n[(W_i-1)\zeta_i^*(\cMq;\boldsymbol{\theta})]\\ \mathbb{G}_n[(W_i-1)Z_i^{r,*}]\\ \mathbb{G}_n[(W_i-1)Z_i^{b,*}] \end{bmatrix}\right)\,\bigg|\,\mathcal{O}_n\right]-\mathbb{E}[f(\mathbb{G}_{he^*})]\right|+o_p(1), \end{align} where $\mathbb{G}_{he^*}$ is defined in Eq. (ref), and by Theorem 3.6.13 of van:wel13, Eq. $\eqref{eq:bs-nice}=o_p(1)$, and therefore Assumption 3(ii) of fan:san19 holds. To show that the equality in Eq. (ref) holds, we let $\widetilde{\mu}_g^W\equiv\overline{W}\widetilde{\mu}_g$, $\widetilde{\mathcal{M}}^W\equiv\operatorname{diag}(1/\widetilde{\mu}_r^W, 1/\widetilde{\mu}_b^W)$. Recall $\boldsymbol{\xi}_{\mathcal{S},i}(\check{\mathcal{M}};\mathring{\mathcal{M}}q;\widehat{\boldsymbol{\theta}})$ defined in Eq. (ref). Decompose the bootstrapped process by \begin{align*} &\sqrt{n}\big\{\widetilde{he^*}(\widehat{\boldsymbol{\theta}})-\widehat{he^*}(\widehat{\boldsymbol{\theta}})\big\}=\sqrt{n}\big\{\widetilde{he^*}(\widehat{\boldsymbol{\theta}})-he^*\big\}-\sqrt{n}\big\{\widehat{he^*}(\widehat{\boldsymbol{\theta}})-he^*\big\}\\ &=\sqrt{n}\begin{bmatrix} \frac{1}{n}\sum_{i=1}^nW_iq^\intercal\boldsymbol{\xi}_{\mathcal{S},i}(\widetilde{\mathcal{M}}^W;\tcMq;\widehat{\boldsymbol{\theta}})-\mathbb{E}[\zeta_i(\cMq;\boldsymbol{\theta})]\\ \frac{1}{n}\sum_{i=1}^nW_i\frac{Z_i^r}{\widetilde{\mu}_r^W}-\frac{\mathbb{E}[Z_i^r]}{\mu_r}\\ \frac{1}{n}\sum_{i=1}^nW_i\frac{Z_i^b}{\widetilde{\mu}_b^W}-\frac{\mathbb{E}[Z_i^b]}{\mu_b} \end{bmatrix}-\begin{bmatrix} \mathbb{G}_n[\zeta_i^*(\cMq;\boldsymbol{\theta})]\\ \mathbb{G}_n[Z_i^{r,*}]\\ \mathbb{G}_n[Z_i^{b,*}] \end{bmatrix}+o_p(1), \end{align*} where the second equality follows from Theorem (ref) and the proof of Proposition (ref). Next, observe that \begin{align*} &\sqrt{n}\left(\frac{1}{n}\sum_{i=1}^nW_iq^\intercal\boldsymbol{\xi}_{\mathcal{S},i}(\widetilde{\mathcal{M}}^W;\tcMq;\widehat{\boldsymbol{\theta}})-\mathbb{E}[\zeta_i(\cMq;\boldsymbol{\theta})]\right)\\ &=\underbrace{\sqrt{n}\left(\frac{1}{n}\sum_{i=1}^nq^\intercal(\widetilde{\mathcal{M}}^W\mathcal{M}^{-1})\big(W_i\boldsymbol{\xi}_{\mathcal{S},i}(\mathcal{M};\tcMq;\widehat{\boldsymbol{\theta}})-\mathbb{E}[W_i\boldsymbol{\xi}_{\mathcal{S},i}(\mathcal{M};\cMq;\boldsymbol{\theta})]\big)\right)}_{\equiv\widetilde{A}}\\ &+\underbrace{\sqrt{n}q^\intercal\left(\widetilde{\mathcal{M}}^W\mathcal{M}^{-1}-\boldsymbol{I}\right)\mathbb{E}[\boldsymbol{\xi}_{\mathcal{S},i}(\mathcal{M};\cMq;\boldsymbol{\theta})]}_{\equiv\widetilde{B}}\\ &=\mathbb{G}_n[W_i\zeta_i^*(\cMq;\boldsymbol{\theta})]+o_p(1), \end{align*} where $\mathbb{E}[W_i\boldsymbol{\xi}_{\mathcal{S},i}(\mathcal{M};\cMq;\boldsymbol{\theta})]=\mathbb{E}[\boldsymbol{\xi}_{\mathcal{S},i}(\mathcal{M};\cMq;\boldsymbol{\theta})]$ by independence of $W_i$ and $\mathbb{E}[W_i]=1$, $\widetilde{A}$ (resp., $\widetilde{B}$) is the bootstrapped analogue of $A$ (resp., $B$) in the proof of Theorem (ref), and the last equality follows from a similar argument used in the proof of Theorem (ref). In addition, for $g\in\{r,b\}$, by the Delta method, \begin{align*} &\sqrt{n}\left(\frac{1}{n}\sum_{i=1}^nW_i\frac{Z_i^g}{\widetilde{\mu}_g^W}-\frac{\mathbb{E}[Z_i^g]}{\mu_g}\right)\\ &=\sqrt{n}\left(\frac{1}{n}\sum_{i=1}^nW_i\frac{Z_i^g}{\mu_g}-\frac{\mathbb{E}[Z_i^g]}{\mu_g}\right)+\sqrt{n}\left(\frac{1}{\widetilde{\mu}_g^W}-\frac{1}{\mu_g}\right)\left(\frac{1}{n}\sum_{i=1}^nW_iZ_i^g\right)\\ &=\mathbb{G}_n[W_iZ_i]\frac{1}{\mu_g}-\mathbb{G}_n[W_i\mathds{1}\{G_i=g\}]\frac{\mathbb{E}[Z_i^g]}{\mu_g^2}+o_p(1)=\mathbb{G}_n[W_iZ_i^{g,*}]+o_p(1). \end{align*} Therefore, \begin{align*} \sqrt{n}\big\{\widetilde{he^*}(\widehat{\boldsymbol{\theta}})-\widehat{he^*}(\widehat{\boldsymbol{\theta}})\big\}=\begin{bmatrix} \mathbb{G}_n[(W_i-1)\zeta_i^*(\cMq;\boldsymbol{\theta})]\\ \mathbb{G}_n[(W_i-1)Z_i^{r,*}]\\ \mathbb{G}_n[(W_i-1)Z_i^{b,*}] \end{bmatrix}+o_p(1), \end{align*} yielding Eq. (ref). \end{proof} \subsection{Proofs for Section (ref)} \begin{proof}[Proof of Proposition (ref)] Consider the null hypothesis $\mathsf{H}_0:\rho(e^*,F)=\delta$ against $\mathsf{H}_A:\rho(e^*,F)\neq\delta$ for some $\delta>0$, and view our confidence interval as the result of inverting this hypothesis test. Let $\varphi_n^{\texttt{dist}}\equiv\mathds{1}\left\{\inf_{\rho(\Tilde{e}, \Tilde{F})\in \mathcal{CS}_n^{\rho(e^*,F)}}\left|\rho(\Tilde{e}, \Tilde{F}) - \delta\right|>0\right\}$ and partition the parameter space of the location of $\mathcal{E}$ relative to $\mathcal{H}_{45}$ such that \begin{align} \varphi_n^{\texttt{dist}} = & \varphi_n^{\texttt{dist}}\mathds{1}\big\{\mathcal{E}\cap\mathcal{H}_{45}^-=\emptyset\big\}\\ & +\varphi_n^{\texttt{dist}}\mathds{1}\big\{\mathcal{E}\cap\mathcal{H}_{45}^+=\emptyset\big\}\\ & +\varphi_n^{\texttt{dist}}\mathds{1}\big\{\mathcal{E}\cap\mathcal{H}_{45}^+\neq\emptyset, \mathcal{E}\cap\mathcal{H}_{45}^-\neq\emptyset\big\}. \end{align} Note that whenever $\mathcal{CS}_n^+(e^*,F)\neq\emptyset$, \begin{multline} \inf_{\rho(\Tilde{e}, \Tilde{F})\in \mathcal{CS}_n^{\rho(e^*,F)}}|\rho(\Tilde{e}, \Tilde{F})-\delta|\le \inf_{(\Tilde{e}, \Tilde{F})\in \mathcal{CS}_n^+(e^*,F)}|\rho(\Tilde{e}, \Tilde{F})-\delta|,\\ \Rightarrow \mathbb{E}[\varphi_n^{\texttt{dist}}]\le\mathbb{E}\left[\mathds{1}\left\{\inf_{(\Tilde{e}, \Tilde{F})\in \mathcal{CS}_n^+(e^*,F)}|\rho(\Tilde{e}, \Tilde{F})-\delta| > 0\right\}\right], \end{multline} and similarly when $(\Tilde{e}, \Tilde{F})\in\mathcal{CS}_n^+(e^*,F)$ is replaced with $(\Tilde{e}, \Tilde{F})\in\mathcal{CS}_n^-(e^*,F)$ or $\rho(\Tilde{e}, \Tilde{F})\in\mathcal{CS}_n^{45}(\rho(e^*,F))$ (and under the case where these sets are non-empty). When $\mathcal{E}\cap\mathcal{H}_{45}^-=\emptyset$, we have $ F\in\mathcal{H}_{45}^+\cup\mathcal{H}_{45}$ and $\lim_{n\to\infty}\mathbb{P}\big((e^*,F)\in \mathcal{CS}_n^+(e^*,F)\big)\geq1-\alpha$ by a similar argument to the proof of Proposition (ref). It then follows that under the null $\rho(e^*,F)=\delta$, $\mathbb{E}[\eqref{eq:tstat-above}]\leq \alpha\mathds{1}\{\mathcal{E}\cap\mathcal{H}_{45}^-=\emptyset\}$ as $n\to\infty$, using Eq. (ref) and the fact that $\mathbb{P}\left(\inf_{(\Tilde{e},\Tilde{F})\in \mathcal{CS}_n^+(e^*,F)}\left|\rho(\Tilde{e}, \Tilde{F})-\delta\right|>0 \right)\leq 1-\mathbb{P}\big((e^*,F)\in \mathcal{CS}_n^+(e^*,F)\big)$. Similarly, $\mathbb{E}[\eqref{eq:tstat-below}]\leq\alpha\mathds{1}\{\mathcal{E}\cap\mathcal{H}_{45}^+=\emptyset\}$ as $n\to\infty$. We complete the proof by showing $\mathbb{E}[\eqref{eq:tstat-cross}]\leq\alpha\mathds{1}\{\mathcal{E}\cap\mathcal{H}_{45}^+\neq\emptyset, \mathcal{E}\cap\mathcal{H}_{45}^-\neq\emptyset\}$ if $\rho(e^*,F)=\delta$. In the event $\mathcal{E}\cap\mathcal{H}_{45}^+\neq\emptyset$ and $ \mathcal{E}\cap\mathcal{H}_{45}^-\neq\emptyset$, $F\in\mathcal{H}_{45}$ and---since $\mathcal{E}$ is \textit{not} tangent to $\mathcal{H}_{45}$ in this case---the direction in which $F$ is the support set of $\mathcal{E}$ does \textit{not} live in the span $\{c[1~ -1]^\intercal: c\in\mathbb{R}\}$. Hence, $c^*\equiv\arg\inf_{c\in\mathbb{R}}h_\mathcal{E}(\mathfrak{u}_1(c))$ for $\mathfrak{u}_1(c)\equiv\mathfrak{u}_1-c[1~-1]^\intercal$ is bounded, since otherwise it would require $\mathcal{E}$ to be tangent to $\mathcal{H}_{45}$. In addition, as explained in Footnote \hyperref[ftnt:bound_inf_E_tilde]{3}, $h_\mathcal{E}(\mathfrak{u}_1(c))$ is bounded from below when $F\in\mathcal{H}_{45}$. We can therefore restrict attention to $c\in[-c_3, c_3]\equiv\mathcal{C}_3\subset\mathbb{R}$ for some bounded constant $c_3>0$ in the characterization of $F$ in Eq. (ref) so that $h_\mathcal{E}(\mathfrak{u}_1(\cdot))$ is a bounded function over compact support $\mathcal{C}_3$. Under the maintained assumptions, Eq. (ref) holds and implies that \begin{align*} \sqrt{n}\begin{bmatrix} \widehat{h}_\mathcal{E}\big(\mathfrak{u}_1(c); \widehat{\boldsymbol{\theta}}\big)-h_\mathcal{E}\big(\mathfrak{u}_1(c)\big)\\ \widehat{e}_r^*-e_r^*\\ \widehat{e}_b^*-e_b^* \end{bmatrix}\xrightarrow[]{d}\begin{bmatrix} \|\mathfrak{u}_1(c)\|_E\mathbb{G}[\zeta_i^*(\mathcal{M}\mathfrak{u}_1(c)/\|\mathfrak{u}_1(c)\|_E;\boldsymbol{\theta})]\\ \mathbb{G}[Z_i^{r,*}]\\ \mathbb{G}[Z_i^{b,*}] \end{bmatrix} \text{in }\ell^\infty(\mathcal{C}_3), \end{align*} where we use the property of support functions that for any constant $\tilde{c}>0$, $h_\mathcal{E}(\tilde{c}\cdotq)=\tilde{c}\cdoth_\mathcal{E}(q)$ sch93. Recall $F_{45}\equiv\mathfrak{u}\cdoth_{\Tilde{\mathcal{E}}}(\mathfrak{u}_1)$ and let $\widehat{F}_{45}$ be its estimator with $h_\mathcal{E}(\cdot)$ replaced by $\widehat{h}_\mathcal{E}(\,\cdot\,;\widehat{\boldsymbol{\theta}})$ in the expression of $h_{\Tilde{\mathcal{E}}}(\cdot)$ given in Eq. (ref). Observe that $\rho(e^*,F_{45})$ is a composition of two Hadamard directionally differentiable functions: let $\mathfrak{h}_4:\ell^\infty(\mathcal{C}_3)\to\mathbb{R}$ be the $\inf_{c\in\mathcal{C}_3}(\cdot)$ function, \begin{align*} \rho(e^*,F_{45})= \rho\bigg(e^*, \mathfrak{u}\cdot\mathfrak{h}_4\big\{h_\mathcal{E}\big(\mathfrak{u}_1(c)\big)\big\}\bigg), \end{align*} where by car:cue:alb20, $\mathfrak{h}_4$ is directionally differentiable at $h_\mathcal{E}\big(\mathfrak{u}_1(\cdot)\big)$ tangentially to $\ell^\infty(\mathcal{C}_3)$, with \begin{align*} \mathfrak{h}'_{4,h_\mathcal{E}(\mathfrak{u}_1(\cdot))}(f)=\inf_{\bigl\{c\in\mathcal{C}_3:h_\mathcal{E}\big(\mathfrak{u}_1(c)\big)=h_{\Tilde{\mathcal{E}}}(\mathfrak{u}_1)\bigr\}} f(c), \text{for continuous } f\in\ell^\infty(\mathcal{C}_3). \end{align*} Denote $\rho_{(e^*,F_{45})}' : \mathbb{R}^2\times\mathbb{R}^2\to\mathbb{R}$ the directional derivative of $\rho$ at $(e^*,F_{45})$. Then by Shapiro90, fan:san19, and car:cue:alb20, \begin{align*} \sqrt{n}\begin{bmatrix} \rho(\widehat{e}^*,\widehat{F}_{45})-\rho(e^*,F_{45}) \end{bmatrix}\xrightarrow[]{d}\psi^{\rho(e^*,F_{45})}, \end{align*} where, for $\mathbb{G}[\mathbf{Z}_i^{*}]\equiv\left[\mathbb{G}[Z_i^{r,*}]~~~~~\mathbb{G}[Z_i^{b,*}]\right]^\intercal$, \begin{align*} \psi^{\rho(e^*,F_{45})}\equiv\rho_{(e^*,F_{45})}'\left( \mathbb{G}[\mathbf{Z}_i^{*}],\mathfrak{u}\cdot\mathfrak{h}'_{4,h_\mathcal{E}(\mathfrak{u}_1(\cdot))}\big(\|\mathfrak{u}_1(c)\|_E\mathbb{G}[\zeta_i^*(\mathcal{M}\mathfrak{u}_1(c)/\|\mathfrak{u}_1(c)\|_E;\boldsymbol{\theta})]\big)\right), \end{align*} so by the continuous mapping theorem, \begin{align} \psi^{45}=\left|\psi^{\rho(e^*,F_{45})}\right|, \end{align} Next, if Eq. (ref) holds and $\mathcal{E}$ has no kinks, we show that the set $\left\{c\in\mathcal{C}_3:h_\mathcal{E}\big(\mathfrak{u}_1(c)\big)=h_{\Tilde{\mathcal{E}}}(\mathfrak{u}_1)\right\}$ is a singleton and the expression of $\psi^{\rho(e^*, F_{45})}$ simplifies. Recall $c^*\equiv\arg\inf_{c\in\mathcal{C}_3}h_\mathcal{E}\big(\mathfrak{u}_1(c)\big)$. By contradiction, assume there exists $\Tilde{c}\neq c^*$ and $h_\mathcal{E}\big(\mathfrak{u}_1(\Tilde{c})\big)=h_\mathcal{E}\big(\mathfrak{u}_1(c^*)\big)=h_{\Tilde{\mathcal{E}}}(\mathfrak{u}_1)$. This implies that the two linear equations $(-1-\Tilde{c})e_r+\Tilde{c}e_b=h_{\Tilde{\mathcal{E}}}(\mathfrak{u}_1)$ and $(-1-c^*)e_r+c^*e_b=h_{\Tilde{\mathcal{E}}}(\mathfrak{u}_1)$ intersect at some point $(e_r^\star,e_b^\star)$. Replacing this value in the equations, we find $\Tilde{c}(e_b^\star-e_r^\star)=c^*(e_b^\star-e_r^\star)$. If $c^*=0$, for $\Tilde{c}$ not to equal $c^*$ it must be the case that $\Tilde{c}\neq 0$, in which case $e_b^\star=e_r^\star$, implying that $(e_r^\star,e_b^\star)=R=F$ and in turn by Eq. (ref) it must be the case that $\Tilde{c}=c^*$. Similarly, if $c^*\neq 0$, then either $e_b^\star=e_r^\star$, which implies $(e_r^\star,e_b^\star)=F$ and hence $\Tilde{c}=c^*$, or $\Tilde{c}/c^*=1$ and the claim follows. Let $\Tilde{\mathfrak{u}}_1(c^*)\equiv\frac{\mathfrak{u}_1(c^*)}{\|\mathfrak{u}_1(c^*)\|_E}$, we get: \begin{align} \psi^{45}=\left|\rho_{(e^*,F_{45})}'\bigg( \mathbb{G}[\mathbf{Z}_i^{*}],\mathfrak{u}\cdot\|\mathfrak{u}_1(c^*)\|_E\cdot\mathbb{G}\left[\zeta_i^*\left(\mathcal{M}\Tilde{\mathfrak{u}}_1(c^*);\boldsymbol{\theta}\right]\right)\bigg)\right|. \end{align} If $F\in\mathcal{H}_{45}$, $\lim_{n\to\infty}\mathbb{P}\left(\mathcal{CS}_n^{45}(\rho(e^*,F))=\emptyset\right)=0$. Under the null $\rho(e^*,F)=\delta$, and using again Eq. (ref), if $c_{1-\alpha+\varsigma}^{\rho(\Tilde{e}, \Tilde{F})}+\varsigma$ is a continuity point of the distribution of $\psi^{\rho(\Tilde{e}, \Tilde{F})}$, \begin{align*} \lim_{n\to\infty}\mathbb{E}[\varphi_n^{\texttt{dist}}]\le&\lim_{n\to\infty}\mathbb{P}\big(T_n^{\rho(\Tilde{e}, \Tilde{F})}>c_{1-\alpha+\varsigma}^{\rho(\Tilde{e}, \Tilde{F})}+\varsigma \big)\leq\alpha \end{align*} and if it is a discontinuity point for an infinitesimal $\varsigma$, then $c_{1-\alpha}^{\rho(\Tilde{e}, \Tilde{F})}$ is a continuity point and \begin{align*} \lim_{n\to\infty}\mathbb{E}[\varphi_n^{\texttt{dist}}]\leq&\lim_{n\to\infty}\mathbb{P}\big(T_n^{\rho(\Tilde{e}, \Tilde{F})}>c_{1-\alpha}^{\rho(\Tilde{e}, \Tilde{F})}\big)=\alpha. \end{align*} Therefore, under the null $\rho(e^*,F)=\delta$, we conclude \begin{align*} \lim_{n\to\infty}\mathbb{E}[\varphi_n^{\texttt{dist}}]\leq\alpha. \end{align*} Test inversion yields the coverage result. \end{proof} \numberwithin{asm}{section} \section{Auxiliary Results} \subsection{Threshold Rules} In the paper, we allow $\mathcal A(\mathcal{X})$, the set of all algorithms that map from the input space $\mathcal{X}$ to $[0,1]$, to be completely unrestricted. This includes randomized rules where the event $D=1$ occurs with probability $a(X)$. Here instead we consider \emph{threshold rules}. Let $\mathcal{A}^\texttt{th}(\mathcal{X})$ denote a space of algorithms $\{a:\mathcal{X} \mapsto \mathbb{R}\}$ such that $a\in\mathcal{A}^\texttt{th}(\mathcal{X})$ induces the decision rule \begin{align*} D_{a} = \mathds{1}\{a(X) \ge 0\}. \end{align*} One could alternatively pick a constant $\kappa$ a priori and use a decision rule of the form $\mathds{1}\{a(X) \ge \kappa\}$, but to ease notation we absorb the threshold $\kappa$ in $a$. We maintain the following richness assumption on $\mathcal{A}^\texttt{th}(\mathcal{X})$: \begin{asm} \textit{The set of algorithms $\mathcal{A}^\texttt{th}(\mathcal{X})$ is (i) convex; and (ii) sufficiently rich, in the sense that, $X$-a.s., \begin{align*} \exists a'\in\mathcal{A}^\texttt{th}(\mathcal{X}): \mathds{1}\{a'(X) \ge 0\}&=1,\\ \exists a”\in\mathcal{A}^\texttt{th}(\mathcal{X}): \mathds{1}\{a”(X) \ge 0\}&=0. \end{align*} } \end{asm} \begin{remark}[Linear Threshold Rules] Assumption (ref) is satisfied by linear threshold rules with $\mathcal A^\texttt{th}(\mathcal{X})=\{[1;X]^\intercal\beta:~\beta\in\mathbb{R}^{d_X+1}\}$ and $D_{\beta}=\mathds{1}\{[1;X]^\intercal\beta \ge 0\}$. \end{remark} When using threshold rules, similar to Eq. (ref), the group risks can be expressed as \begin{align} e_g(D_{a})&\equiv\frac{1}{\mu_g}\mathbb{E}_{X}\left[D_{a}\theta_1^g(X)+(1-D_{a})\theta_0^g(X)\right]\notag\\ &= \frac{1}{\mu_g}\mathbb{E}\left[L_0^g+(L_1^g-L_0^g)\mathds{1}\{a(X) \ge 0\}\right]. \end{align} Compare Eq. (ref) with Eq. (ref): it follows immediately that threshold rules given by $D_{k(\boldsymbol{\theta}(X),\cMq)}=\mathds{1}\{k(\boldsymbol{\theta}(X), \cMq)\ge 0\}$ yield the extreme points of the set $\mathcal{E}$ in Eq. (ref). As we show next, under Assumption (ref), the feasible set associated with threshold rules, denoted $\mathcal{E}^\texttt{th}$, is convex. To see this, note that \begin{align} \mathcal{E}^\texttt{th}&=\left\{\mathbb{E}[\mathcal{M}\vartheta(X)]: \vartheta(X)\in \left\{\boldsymbol{\theta}_0(X), \boldsymbol{\theta}_1(X)\right\} \right\}\equiv \mathbf{E}\left[\mathcal{M}\tilde\Lambda(X)\right], \end{align} with $\boldsymbol{\theta}_d(X)$ defined in Eq. (ref) and $\tilde\Lambda(X)\equiv\left\{\boldsymbol{\theta}_0(X), \boldsymbol{\theta}_1(X)\right\}$. Relative to Eq. (ref), the fundamental difference here is that $\left\{\boldsymbol{\theta}_0(X), \boldsymbol{\theta}_1(X)\right\}$ is a two-point set instead of an interval. Nonetheless, the set on the right-hand-side of Eq. (ref) is by definition mol:mol18 the \emph{Aumann expectation} of the two-point set $\tilde\Lambda(X)$. Next, observe that under Assumption (ref) the probability space on which $(Y,G,X)$ are defined is non-atomic (non-atomicity follows as long as one of the variables in $X$ has a continuous distribution, and if all variables in $X$ had a discrete distribution, Assumption (ref) would fail).\footnote{To see this, let $X$ have countable support, take $x\in\mathcal{X}$ such that $\mathbb{P}(X=x)=\varsigma>0$. Then for $q=\frac{[\{\theta_1^b(x)-\theta_0^b(x)\}/\mu_b \quad \{\theta_0^r(x)-\theta_1^r(x)\}/\mu_r]^\intercal}{\Vert \mathcal{M}\{\boldsymbol{\theta}_1(x)-\boldsymbol{\theta}_0(x) \}\Vert}$, $\mathbb{P}(|q_1\{\theta_1^r(x)-\theta_0^r(x)\}/\mu_r+q_2\{\theta_1^b(x)-\theta_0^b(x)\}/\mu_b|=0)\geq\varsigma>0$, and hence $\mathbb{P}(|q_1\{\theta_1^r(x)-\theta_0^r(x)\}/\mu_r+q_2\{\theta_1^b(x)-\theta_0^b(x)\}/\mu_b|<\delta)\geq\varsigma>0$ for any $\delta>0$.} As both $\boldsymbol{\theta}_0(X)$ and $\boldsymbol{\theta}_1(X)$ are absolutely integrable, all conditions required for Theorem 3.4 in mol:mol18 are satisfied, yielding: \begin{align} \mathbf{E}\left[\mathcal{M}\tilde\Lambda(X)\right]=\mathbf{E}\big[\mathcal{M}\operatorname{conv}\left(\{\boldsymbol{\theta}_0(X),\boldsymbol{\theta}_1(X)\}\right)\big]=\mathbf{E}\left[\mathcal{M}\Lambda(X)\right]. \end{align} Theorem 3.11 in mol:mol18 also applies, and $h_{\mathcal{E}}(q)=h_{\mathbf{E}\left[\mathcal{M}\Lambda(X)\right]}(q)=h_{\mathbf{E}\left[\mathcal{M}\tilde\Lambda(X)\right]}(q)=\mathbb{E}[h_{\mathcal{M}\Lambda(X)}(q)]$. Hence, under Assumption (ref), the feasible set associated with threshold rules is convex and the support function fully characterizes it. \begin{remark}[Richness of Threshold Rules] The result in Eq. (ref) shows that threshold rules corresponding to a rich algorithm space can replicate any unconstrained algorithm. Inspecting further the linear case is instructive. If one allows for any $\beta\in\mathbb{R}^{d_X+1}$, Assumption (ref) is satisfied. By the same argument given above, the feasible set associated with linear threshold rules, denoted $\mathcal{E}^\texttt{lin}$, is convex. Indeed, for each $\beta$ one can write \begin{multline*} e_g(\beta)\equiv\frac{1}{\mu_g}\mathbb{E}\left[L_0^g\big|[1;X]^\intercal\beta< 0\right]\mathbb{P}\left([1;X]^\intercal\beta< 0\right) +\frac{1}{\mu_g}\mathbb{E}\left[L_1^g\big|[1;X]^\intercal\beta\ge 0\right]\mathbb{P}\left([1;X]^\intercal\beta\ge 0\right). \end{multline*} Let $\mathcal{X}_\beta^+\equiv\{X\in\mathcal{X}:[1;X]^\intercal\beta\ge 0\}$ and $\mathcal{X}_\beta^-\equiv\{X\in\mathcal{X}:[1;X]^\intercal\beta< 0\}$ denote the two sets in which a linear threshold rule with parameter $\beta$ partitions $\mathcal{X}$. As at least one component of $X$ has continuous distribution and $\beta$ has support $\mathbb{R}^{d_X+1}$, each realization of $X$ can be allocated either in $\mathcal{X}_\beta^+$ or in $\mathcal{X}_\beta^-$ for some $\beta$. Hence, convexification occurs. The support function of $\mathcal{E}^\texttt{lin}$ can be expressed as \begin{align*} h_{\mathcal{E}^\texttt{lin}}(q)=\max_{\beta\in\mathbb{R}^{d_X+1}}q^\intercal e_g(\beta)=\max_{\beta\in\mathbb{R}^{d_X+1}}(\cMq)^\intercal \mathbb{E}\left[\mathbf{L}_0\mathds{1}([1;X]^\intercal\beta< 0)+\mathbf{L}_1\mathds{1}([1;X]^\intercal\beta\ge 0)\right]. \end{align*} \end{remark} \subsection{Sufficient Conditions Yielding Strict Convexity and No Kinks} Assumption (ref) plays multiple roles in our analysis. It assures that $\mathcal{E}$ is strictly convex, hence its support set in any direction $q\in\mathbb{S}^1$ is a singleton, and it assures Neyman orthogonality of the moment condition defining $h_\mathcal{E}(\cdot)$. sem23 provides sufficient conditions for this assumption, based on joint Gaussianity of $\boldsymbol{\theta}_1(X)-\boldsymbol{\theta}_0(X)$. A mild strengthening of Assumption (ref) is sufficient both for Assumption (ref) to hold with $m=1$ and for Eq. (ref) to be satisfied, guaranteeing the absence of kinks in $\mathcal{E}$, as we show next. \begin{asm} \textit{(i) The distribution of $(X_1,X_2)|X_{[3:d_X]}$ is continuous with a bounded density and $\mathbb{E}\big[\big|\eta^g(X_{[3:d_X]})\big|\big]<\infty$ for $g\in\{r,b\}$; (ii) for a set $\tilde{\mathcal{X}}_{[3:d_X]}$ of realizations of $X_{[3:d_X]}$ with positive probability, the density of $(X_1,X_2)|X_{[3:d_X]}$ is positive on a ball of radius $c>0$ that includes $\mathbf{0}$ and the image of $\tilde{\mathcal{X}}_{[3:d_X]}$ under $\eta^g(\cdot)$ includes $0$.} \end{asm} \begin{prop} \textit{If Assumptions (ref) and (ref)(i) hold, Assumption (ref) is implied. If Assumption (ref)(ii) also holds, Eq. (ref) holds and the set $\mathcal{E}$ has no kinks.} \end{prop} \begin{proof} Take $\delta>0$, and note that $\sup_{q\in \mathbb{S}^1} \mathbb{P}(|k(\boldsymbol{\theta}(X),\cMq)|<\delta)$ can be written as \begin{align*} &\sup_{q\in \mathbb{S}^1}\mathbb{E}_{X_{[3:d_X]}}\left[\int_0^\delta \left|q_1\Delta\theta^r(X)/\mu_r+q_2\Delta\theta^b(X)/\mu_b \right|d\mathbb{P}_{(X_1,X_2)|X_{[3:d_X]}}\right]\\ \lesssim &\max_{g\in\{r,b\}}\mathbb{E}_{X_{[3:d_X]}}\left[\int_0^\delta \left(|\alpha^g|+|\beta^g|+|\eta^g(X_{[3:d_X]})|\right)d\mathbb{P}_{(X_1,X_2)|X_{[3:d_X]}}\right]\\ \lesssim &\max_{g\in\{r,b\}}\mathbb{E}_{X_{[3:d_X]}}\left[\delta \left(|\alpha^g|+|\beta^g|+|\eta^g(X_{[3:d_X]})|\right)\right]\lesssim\delta, \end{align*} where the first inequality follows from $\sup_{q\in \mathbb{S}^1}\|q\|_E=1$, $\mu_g\in(0,1)$, and that $(X_1,X_2)$ has bounded support. The second equality follows from continuity and boundedness of $d\mathbb{P}_{(X_1,X_2)|X_{[3:d_X]}}$. The last inequality follows from $\mathbb{E}\big[\big|\eta^g(X_{[3:d_X]})\big|\big]<\infty$. For the second result, observe that $\boldsymbol{\theta}_1(X)-\boldsymbol{\theta}_0(X)=\Delta\boldsymbol{\theta}(X)$ equals: \begin{align*} \underbrace{\begin{bmatrix} \alpha^r & \beta^r \\ \alpha^b & \beta^b \end{bmatrix}}_{\equiv A_1}\begin{bmatrix} X_1\\ X_2 \end{bmatrix}+\underbrace{\begin{bmatrix} \eta^r(X_{[3:d_X]})\\ \eta^b(X_{[3:d_X]}) \end{bmatrix}}_{\equiv A_2}, \end{align*} where the matrix $A_1$ is invertible by $\alpha^b\beta^r\neq\alpha^r\beta^b$. Under Assumption (ref)(ii), on the set $\tilde{\mathcal{X}}_{[3:d_X]}$, $A_1[X_1~X_2]^\intercal+A_2$ realizes in a set containing $0$. Hence, Eq. (ref) holds by applying the law of iterated expectations. \end{proof} \numberwithin{asm}{section} \section{Variability of the Empirical Results to Nuisance Parameter Estimation} \subsection{Estimating the Nuisance Parameters Using Lasso} In this subsection, we report the analogs of Figures (ref)-(ref)-(ref) and Tables (ref)-(ref) using multinational logit lasso from the \href{https://glmnet.stanford.edu/index.html}{\texttt{glmnet}} package to estimate the nuisance parameter $\Delta\boldsymbol{\theta}$. That is, the only difference between the results reported in this subsection and those in Section (ref) lies in the choice of what machine learner is used to estimate nuisance parameters. Respectively, these are Figures (ref)-(ref)-(ref) and Tables (ref)-(ref). \begin{figure}[H]\captionsetup[subfigure]{font=footnotesize} \caption{Top-left panel: $\widehat{\mathcal{E}}$; top-right panel: $\widehat{\mathcal{E}}$ along with one hundred supporting hyperplanes; bottom-left panel: zoom-in to $\widehat{\mathcal{F}}$ and the $95\%$ confidence set around this frontier; bottom-right panel: further zoom-in to the best group-specific points $\textsf{BL}$ and $\textsf{WH}$, and the fairest point $F$. $\Delta\boldsymbol{\theta}$ is estimated by logit lasso.} \end{figure} \begin{figure}[H]\captionsetup[subfigure]{font=footnotesize} \subfigure[Candidate values for $(\textsf{BL}, \textsf{WH})$] \subfigure[$3$ experimental algorithms] \caption{Panel (a): Plum-colored (respectively, light-blue colored) circles correspond to candidate values for $e_\textsf{BL}$ ($e_\textsf{WH}$) sampled from a normal distribution centered at $\widehat{e}_\textsf{BL}$ ($\widehat{e}_\textsf{WH}$), and red (blue) diamonds correspond to non-rejected values. Panel (b): $\widehat{\mathcal{F}}$ along with its $95\%$ confidence set and the estimated group risks for four algorithms considered by \citetalias{obe:pow:vog:mul19}: the original algorithm used by the hospital (asterisk); one that predicts total cost (hollow diamond with a cross); one that predicts avoidable costs (filled diamond); and one that predicts the number of active chronic conditions (hollow diamond). $\Delta\boldsymbol{\theta}$ is estimated by logit lasso.} \end{figure} \begin{figure}[H]\captionsetup[subfigure]{font=footnotesize} \caption{Average number of active chronic conditions within each risk-score percentile bin by treatment group under the alternative algorithms on the FA frontier subject to $3\%$ capacity constraint, averaged across $20$ replications of the $50$-$50$ split. $\Delta\boldsymbol{\theta}$ is estimated by logit lasso.} \end{figure} \subsection{Variability of Empirical Results due to the Randomness in the Nuisance Estimation} Both random forests and lasso involve randomness in their respective construction: for forests implemented by the \href{https://grf-labs.github.io/grf/index.html}{\texttt{grf}} package, randomness comes from subsampling in the construction of individual trees and randomly splitting features at each tree node, whereas for lasso implemented by the \href{https://glmnet.stanford.edu/index.html}{\texttt{glmnet}} package, randomness comes from choosing the optimal penalty parameter via cross-validation. Therefore, even if the same seed is set for reproducibility whenever possible, results will vary across different seeds. For this reason, we provide an assessment of how the empirical results reported in Section (ref) vary across different seeds by repeating the empirical exercises $20$ times, with results for forests and lasso reported respectively in Table (ref) and Table (ref). { \begin{table}[H] \scriptsize \caption{Results for the LDA Test and Confidence Sets for the Distance to $F$ (for $\alpha=0.05$)} \resizebox{0.85\columnwidth}{!}{ \begin{tabular} {@{\extracolsep{3pt}}c@c@@c@@c@@c@@c@@} \multicolumn{5}{c}{\cellcolor{pink!20}\textbf{Test \textit{$\mathsf{H}_0:$ There Is No LDA}}}\\ \cline{1-5} \multicolumn{1}{c} & \multicolumn{1}{c}{Original} & \multicolumn{1}{c}{Total Costs} & \multicolumn{1}{c}{Avoid. Costs} & \multicolumn{1}{c}{Act. Chr. Cond.} \\ \cline{1-5} Estimated Risks & $ (0.065, 0.034)$ & $(0.059, 0.033)$ & $(0.054, 0.023)$ & $(0.037, 0.015)$ \\ Test Statistic & $2.043$ & $0.787$ & $0.440$ & $2.774$\\ Critical Value & $2.014$ & $1.865$ & $1.937$ & $1.886$ \\ Conclusion & Rejected & Not Rejected & Not Rejected & Rejected \\ \cline{1-5} \multicolumn{5}{c}{\cellcolor{pink!20}\textbf{Distance to $F=(0.063, 0.063)$}}\\ \cline{1-5} Estimated Distance & $0.0009$ & $0.0009$ & $0.0017$ & $0.0029$ \\ Confidence Set & $(0.000, 0.002)$ & $(0.000, 0.002)$ & $(0.000, 0.003)$ & $(0.001, 0.006)$\\ \cline{1-5} \end{tabular} } \begin{tablenotes} • Top panel: LDA test statistics and $0.05$-level critical values associated with the original algorithm and the three experimental algorithms (predicting, respectively, total costs; avoidable costs; number of active chronic conditions) analyzed by \citetalias{obe:pow:vog:mul19}. Bottom panel: estimated squared-Euclidean distance to the $F$ point and corresponding confidence set for this distance. $\Delta\boldsymbol{\theta}$ is estimated by logit lasso \end{tablenotes} \end{table} } { \begin{table}[H] \scriptsize \caption{Fraction of Black Patients Treated among All Treated } \resizebox{0.95\columnwidth}{!}{ \begin{tabular}{c|c@c@|cccc} \cline{1-7} \multicolumn{1}{c}{\cellcolor{pink!20}} & \multicolumn{2}{c}{\cellcolor{pink!20}\textbf{Algorithms from obe:pow:vog:mul19}} & \multicolumn{4}{c}{\cellcolor{pink!20}\textbf{Algorithms on the FA-Frontier}}\\ \cline{1-7} Capacity Threshold & $\hspace{.5cm}$Original & Counterfactual & Rawlsian & Majority & Egalitarian & Utilitarian \\ \cline{1-7} $55$ & $\hspace{.5cm}0.120$ & $0.184$ & $0.173$ & $0.173$ & $0.174$ & $0.172$\\ $69$ & $\hspace{.5cm}0.128$ & $0.255$ & $0.217$ & $0.200$ & $0.171$ & $0.202$\\ $82$ & $\hspace{.5cm}0.138$ & $0.327$ & $0.241$ & $0.224$ & $0.130$ & $0.223$\\ $89$ & $\hspace{.5cm}0.151$ & $0.407$ & $0.264$ & $0.247$ & $0.118$ & $0.249$\\ $94$ & $\hspace{.5cm}0.167$ & $0.498$ & $0.324$ & $0.284$ & $0.124$ & $0.294$\\ $97$ & $\hspace{.5cm}0.184$ & $0.592$ & $0.369$ & $0.318$ & $0.143$ & $0.339$\\ \cline{1-7} \end{tabular} } \begin{tablenotes} • The distribution of the number of active chronic conditions is such that the $55^\text{th}$ to the $68^\text{th}$ percentiles all correspond to $1$ active chronic condition, the $69^\text{th}$-$81^\text{st}$ correspond to $2$, the $82^\text{nd}$-$88^\text{th}$ correspond to $3$, the $89^\text{th}$-$92^\text{nd}$ correspond to $4$, the $94^\text{th}$-$95^\text{th}$ correspond to $5$, and the $96^\text{th}$-$97^\text{th}$ correspond to $6$. $\Delta\boldsymbol{\theta}$ is estimated by logit lasso. \end{tablenotes} \end{table} } { \begin{table}[H] \scriptsize \caption{Variability of Empirical Results: Random Forests} \resizebox{0.8\columnwidth}{!}{ \begin{tabular} {cccccccc} \cline{1-8} {\cellcolor{pink!20}} & {\cellcolor{pink!20}Mean} & {\cellcolor{pink!20}SD} & {\cellcolor{pink!20}Min} & {\cellcolor{pink!20}$25$-th} &{\cellcolor{pink!20}$50$-th} & {\cellcolor{pink!20}$75$-th} & {\cellcolor{pink!20}Max} \\ \cline{1-8} \multicolumn{8}{c}{Weak Skew Test}\\ Conclusion & $0$ & $0$ & $0$ & $0$ & $0$ & $0$ & $0$ \\ \cline{1-8} \multicolumn{8}{c}{LDA Test} \\ \multicolumn{8}{l}{\textit{Original Algorithm:}}\\ Test Statistic & $3.338$ & $0.318$ & $2.666$ & $3.173$ & $3.301$ & $3.562$ & $3.823$\\ Critical Value & $1.894$ & $0.051$ & $1.810$ & $1.866$ & $1.896$ & $1.919$ & $2.004$ \\ Conclusion & $1$ & $0$ & $1$ & $1$ & $1$ & $1$ & $1$\\ \multicolumn{8}{l}{\textit{Algorithm that Predicts Total Costs:}}\\ Test Statistic & $2.149$ & $0.327$ & $1.475$ & $1.977$ & $2.111$ & $2.361$ & $2.761$\\ Critical Value & $1.844$ & $0.066$ & $1.736$ & $1.804$ & $1.829$ & $1.876$ & $1.998$ \\ Conclusion & $0.8$ & $0.410$ & $0$ & $1$ & $1$ & $1$ & $1$\\ \multicolumn{8}{l}{\textit{Algorithm that Predicts Avoidable Costs:}} \\ Test Statistic & $1.488$ & $0.054$ & $1.398$ & $1.440$ & $1.493$ & $1.521$ & $1.577$\\ Critical Value & $1.721$ & $0.060$ & $1.612$ & $1.682$ & $1.724$ & $1.754$ & $1.836$ \\ Conclusion & $0$ & $0$ & $0$ & $0$ & $0$ & $0$ & $0$\\ \multicolumn{8}{l}{\textit{Algorithm that Predicts the Number of Active Chronic Conditions:}} \\ Test Statistic & $1.194$ & $0.346$ & $0.674$ & $1.007$ & $1.135$ & $1.395$ & $1.812$ \\ Critical Value & $1.632$ & $0.049$ & $1.516$ & $1.600$ & $1.634$ & $1.654$ & $1.725$ \\ Conclusion & $0.15$ & $0.366$ & $0$ & $0$ & $0$ & $0$ & $1$\\ \cline{1-8} \multicolumn{8}{c}{Confidence Set for the Distance to $F$} \\ Estimated $F$ & $0.052$ & $0.003$ & $0.049$ & $0.050$ & $0.052$ & $0.054$ & $0.058$\\ \multicolumn{8}{l}{\textit{Original Algorithm:}}\\ Estimated Distance & $0.0005$ & $0.0000$ & $0.0005$ & $0.0005$ & $0.0005$ & $0.0005$ & $0.0006$\\ Lower $95\%$ CI & $0.0001$ & $0.0001$ & $0.0000$ & $0.0000$ & $0.0000$ & $0.0001$ & $0.0002$\\ Upper $95\%$ CI & $0.0012$ & $0.0003$ & $0.0008$ & $0.0009$ & $0.0011$ & $0.0014$ & $0.0018$\\ \multicolumn{8}{l}{\textit{Algorithm that Predicts Total Costs:}}\\ Estimated Distance & $0.0004$ & $0.0001$ & $0.0003$ & $0.0004$ & $0.0004$ & $0.0004$ & $0.0006$ \\ Lower $95\%$ CI & $0.0000$ & $0.0000$ & $0.0000$ & $0.0000$ & $0.0000$ & $0.0000$ & $0.0001$ \\ Upper $95\%$ CI & $0.0013$ & $0.0004$ & $0.0007$ & $0.0010$ & $0.0013$ & $0.0016$ & $0.0022$ \\ \multicolumn{8}{l}{\textit{Algorithm that Predicts Avoidable Costs:}} \\ Estimated Distance & $0.0009$ & $0.0002$ & $0.0007$ & $0.0007$ & $0.0008$ & $0.0009$ & $0.0013$ \\ Lower $95\%$ CI & $0.0000$ & $0.0000$ & $0.0000$ & $0.0000$ & $0.0000$ & $0.0000$ & $0.0002$ \\ Upper $95\%$ CI & $0.0024$ & $0.0005$ & $0.0018$ & $0.0020$ & $0.0023$ & $0.0027$ & $0.0033$ \\ \multicolumn{8}{l}{\textit{Algorithm that Predicts the Number of Active Chronic Conditions:}} \\ Estimated Distance & $0.0016$ & $0.0003$ & $0.0012$ & $0.0013$ & $0.0016$ & $0.0017$ & $0.0023$ \\ Lower $95\%$ CI & $0.0003$ & $0.0002$ & $0.0000$ & $0.0002$ & $0.0003$ & $0.0004$ & $0.0006$ \\ Upper $95\%$ CI & $0.0041$ & $0.0006$ & $0.0030$ & $0.0035$ & $0.0042$ & $0.0046$ & $0.0051$\\ \cline{1-8} \end{tabular} } \begin{tablenotes} • Table (ref) reports the variability of the empirical results in Section (ref) due to randomness in the estimation of $\Delta\boldsymbol{\theta}$ using random forests, reported as the mean, standard deviation, minimum, the 25-th percentile, 50-th percentile, 75-th percentile, and the maximum across 20 replications. Test conclusions are recorded as $1$ if the conclusion is rejection, and $0$ otherwise. \end{tablenotes} \end{table} } { \begin{table}[H] \scriptsize \caption{Variability of Empirical Results: Logit Lasso} \resizebox{0.8\columnwidth}{!}{ \begin{tabular} {cccccccc} \cline{1-8} {\cellcolor{pink!20}} & {\cellcolor{pink!20}Mean} & {\cellcolor{pink!20}SD} & {\cellcolor{pink!20}Min} & {\cellcolor{pink!20}$25$-th} &{\cellcolor{pink!20}$50$-th} & {\cellcolor{pink!20}$75$-th} & {\cellcolor{pink!20}Max} \\ \cline{1-8} \multicolumn{8}{c}{Weak Skew Test}\\ Conclusion & $0$ & $0$ & $0$ & $0$ & $0$ & $0$ & $0$ \\ \cline{1-8} \multicolumn{8}{c}{LDA Test} \\ \multicolumn{8}{l}{\textit{Original Algorithm:}}\\ Test Statistic & $2.037$ & $0.182$ & $1.784$ & $1.874$ & $2.001$ & $2.166$ & $2.390$ \\ Critical Value & $1.976$ & $0.074$ & $1.762$ & $1.937$ & $1.989$ & $2.024$ & $2.066$ \\ Conclusion & $0.550$ & $0.510$ & $0$ & $0$ & $1$ & $1$ & $1$ \\ \multicolumn{8}{l}{\textit{Algorithm that Predicts Total Costs:}}\\ Test Statistic & $0.804$ & $0.173$ & $0.562$ & $0.651$ & $0.788$ & $0.908$ & $1.154$ \\ Critical Value & $1.915$ & $0.055$ & $1.801$ & $1.890$ & $1.914$ & $1.961$ & $1.995$ \\ Conclusion & $0$ & $0$ & $0$ & $0$ & $0$ & $0$ & $0$ \\ \multicolumn{8}{l}{\textit{Algorithm that Predicts Avoidable Costs:}} \\ Test Statistic & $0.478$ & $0.171$ & $0.100$ & $0.375$ & $0.491$ & $0.591$ & $0.821$ \\ Critical Value & $1.913$ & $0.057$ & $1.783$ & $1.888$ & $1.923$ & $1.941$ & $2.006$ \\ Conclusion & $0$ & $0$ & $0$ & $0$ & $0$ & $0$ & $0$\\ \multicolumn{8}{l}{\textit{Algorithm that Predicts the Number of Active Chronic Conditions:}} \\ Test Statistic & $2.769$ & $0.174$ & $2.470$ & $2.668$ & $2.754$ & $2.889$ & $3.184$ \\ Critical Value & $1.873$ & $0.058$ & $1.768$ & $1.841$ & $1.880$ & $1.905$ & $2.020$ \\ Conclusion & $1$ & $0$ & $1$ & $1$ & $1$ & $1$ & $1$\\ \cline{1-8} \multicolumn{8}{c}{Confidence Set for the Distance to $F$} \\ Estimated $F$ & $0.063$ & $0.001$ & $0.061$ & $0.062$ & $0.062$ & $0.063$ & $0.065$ \\ \multicolumn{8}{l}{\textit{Original Algorithm:}}\\ Estimated Distance & $0.0008$ & $0.0001$ & $0.0007$ & $0.0008$ & $0.0008$ & $0.0009$ & $0.0010$ \\ Lower $95\%$ CI & $0.0000$ & $0.0000$ & $0.0000$ & $0.0000$ & $0.0000$ & $0.0000$ & $0.0001$\\ Upper $95\%$ CI & $0.0020$ & $0.0003$ & $0.0014$ & $0.0017$ & $0.0019$ & $0.0021$ & $0.0026$ \\ \multicolumn{8}{l}{\textit{Algorithm that Predicts Total Costs:}}\\ Estimated Distance & $0.0009$ & $0.0001$ & $0.0008$ & $0.0008$ & $0.0009$ & $0.0009$ & $0.0010$ \\ Lower $95\%$ CI & $0.0000$ & $0.0000$ & $0.0000$ & $0.0000$ & $0.0000$ & $0.0000$ & $0.0001$ \\ Upper $95\%$ CI & $0.0022$ & $0.0003$ & $0.0017$ & $0.0021$ & $0.0022$ & $0.0025$ & $0.0025$ \\ \multicolumn{8}{l}{\textit{Algorithm that Predicts Avoidable Costs:}} \\ Estimated Distance & $0.0016$ & $0.0001$ & $0.0014$ & $0.0015$ & $0.0016$ & $0.0017$ & $0.0018$ \\ Lower $95\%$ CI & $0.0003$ & $0.0001$ & $0.0000$ & $0.0002$ & $0.0003$ & $0.0004$ & $0.0005$ \\ Upper $95\%$ CI & $0.0035$ & $0.0004$ & $0.0030$ & $0.0032$ & $0.0035$ & $0.0038$ & $0.0042$ \\ \multicolumn{8}{l}{\textit{Algorithm that Predicts the Number of Active Chronic Conditions:}} \\ Estimated Distance & $0.0028$ & $0.0002$ & $0.0026$ & $0.0027$ & $0.0028$ & $0.0029$ & $0.0032$ \\ Lower $95\%$ CI & $0.0009$ & $0.0003$ & $0.0005$ & $0.0008$ & $0.0009$ & $0.0011$ & $0.0016$\\ Upper $95\%$ CI & $0.0054$ & $0.0004$ & $0.0047$ & $0.0051$ & $0.0054$ & $0.0057$ & $0.0062$ \\ \cline{1-8} \end{tabular} } \begin{tablenotes} • Table (ref) reports the variability of the empirical results in Section (ref) due to randomness in the estimation of $\Delta\boldsymbol{\theta}$ using logit lasso, reported as the mean, standard deviation, minimum, the 25-th percentile, 50-th percentile, 75-th percentile, and the maximum across 20 replications. Test conclusions are recorded as $1$ if the conclusion is rejection, and $0$ otherwise. \end{tablenotes} \end{table} }