EconBase
← Back to paper

On Statistical Discrimination as a Failure of Social Learning: A Multi-Armed Bandit Approach

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

81,024 characters · 20 sections · 54 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

On Statistical Discrimination as a Failure of Social Learning: A Multi-Armed Bandit Approach

abstractWe analyze statistical discrimination in hiring markets using a multi-armed bandit model. Myopic firms face workers arriving with heterogeneous observable characteristics. The association between the worker's skill and characteristics is unknown ex ante; thus, firms need to learn it. Laissez-faire causes perpetual underestimation: minority workers are rarely hired, and therefore, the underestimation tends to persist. Even a marginal imbalance in the population ratio frequently results in perpetual underestimation. We propose two policy solutions: a novel subsidy rule (the hybrid mechanism) and the Rooney Rule. Our results indicate that temporary affirmative actions effectively alleviate discrimination stemming from insufficient data.

Introduction

Statistical discrimination refers to discrimination against minority people, taken by fully rational and non-prejudiced agents. Previous studies have shown that, even in the absence of prejudice, discrimination can occur persistently because of various reasons, including the discouragement of human capital investment arrow1973theory,Foster1992AnEA,coate1993will,moro_general_2004, information friction phelps1972statistical,cornell1996culture,Bardhi2019Spiraling, and search friction mailath2000endogenous,Che2019RatingsGuided. The literature has proposed various affirmative-action policies to solve statistical discrimination, with many having been implemented in practice.

This paper demonstrates that statistical discrimination may appear as a failure of social learning. We endogenize the evolution of biased beliefs and analyze their consequences. Our model assumes that (i) all firms (decision-makers) are fully rational and non-prejudiced (i.e., attempt to hire the most productive worker), and (ii) all workers are ex ante symmetric. In such an environment, an unbiased decision policy---hiring workers with superior skills---satisfies numerous fairness notions articulated in scholarly literature, including equalized odds and demographic parity. It also achieves efficiency by maximizing each firm's payoff. However, the long-term persistence of biased beliefs could still occur. This paper underscores that temporary affirmative actions can effectively enhance both welfare and equality.

Although our model applies more broadly, we use the terminology of hiring markets to describe our model. We develop a multi-armed bandit model of social learning, in which many myopic and short-lived firms sequentially make hiring decisions. In each round, a firm hires one worker from a set of candidates. Each firm's utility is determined by the hired worker's skill, which cannot be observed directly until employment. However, as in the standard statistical discrimination model, each worker also has observable characteristics associated with their unobservable skills. Firms learn the statistical association between characteristics and skills using data pertaining to past hiring cases (shared through, e.g., private communication, social media, and recommendation letters) and use the estimators to predict the skills of candidates.

Each worker belongs to a group that represents, for example, their gender, race, and ethnicity. We assume that the characteristics of workers who belong to different groups should be interpreted differently. This assumption is realistic. First, previous studies have revealed that underrepresented groups receive unfairly low evaluations.\footnote{For instance, trix2003exploring analyze letters of recommendation for medical faculty, finding systematic differences between those written for female and male applicants. hanna2012discrimination postulate that students belonging to lower castes in India tend to receive unjustifiably lower exam scores. In the context of teaching evaluations, macnell2015s and mitchell2018gender illustrate that students rate male identities significantly higher than female ones. In a study of online freelance marketplaces, hannak2017bias establish that gender and race significantly correlate with worker evaluations.} When these evaluations are used as the observable characteristics, firms should be aware of the potential bias. Second, evaluations may reflect differences in cultures, living environments and social systems precht1998cross,al2004get. For instance, firms need to be conversant with the norms of drafting recommendation letters to interpret them accurately. Therefore, observable characteristics, such as curriculum vitae, exam scores, grading reports, recommendation letters, and so forth, might convey starkly different implications despite their similar presentations. If firms are unbiased and cognizant of these potential biases, they should adapt their interpretation methods for these characteristics, applying varied statistical models to different groups.

When firms learn the statistical association from data, with some probability, the minority group is underestimated because of a large estimation error raised by insufficient data. Once the minority group is underestimated, it is difficult for a minority worker to appear to be the best candidate---even if he has the greatest skill among the candidates, the firm often dismisses this fact and tends to hire a majority worker. As long as firms only hire majority workers, society cannot learn about the minority group; thus, the imbalance persists even in the long run. We call this phenomenon perpetual underestimation.

We use a linear contextual bandit model to analyze the consequence of social learning. To gauge policy performance, we utilize regret, a widely adopted measure in machine-learning literature that assesses welfare loss relative to the optimal decision rule. Regret arises if proficient minority workers are overlooked due to biased estimates by firms; hence, regret not only signifies efficiency but also encapsulates fairness.\footnote{In Appendix (ref), we formally prove that a decision rule has sublinear regret only if it aligns with equalized odds. This notion stipulates that society's hiring policy performs equitably across groups. Moreover, with symmetric groups, sublinear regret also harmonizes with demographic parity, which ensures hiring decisions are irrespective of membership in a minority group.}

We focus on how regret grows as the total number of firms (denoted by $N$) increases. When regret is sublinear in $N$, firms make fair and efficient decisions in the long run. We first analyze the equilibrium consequence of laissez-faire (no policy intervention). When the groups are ex ante symmetric and the population ratio is equal, laissez-faire results in $\tilde{O}(\sqrt{N})$ regret.\footnote{$\tilde{O}, \tilde{\Omega}$, and $\tilde{\Theta}$ are a Landau notations that ignore polylogarithmic factors. We often treat polylogarithmic factors as if they were constant because these factors grow very slowly ($o(N^\epsilon)$ for any exponent $\epsilon > 0$).} However, when the population ratio is unbalanced, this no longer holds, and expected regret is linear: $\tilde{\Omega}(N)$.

We study two policy interventions toward fair and efficient social learning. The first policy is a subsidy rule, based on the idea of upper confidence bound (UCB). UCB is an effective solution for balancing exploration and exploitation lai1985,auer2002. By incentivizing firms to take actions that are consistent with the recommendations of the UCB, social learning can promote sublinear regret in the long run. The subsidy is adjusted to the degree of information externality. We demonstrate that the UCB mechanism has the expected regret of $\tilde{O}(\sqrt{N})$. The subsidy required to implement the UCB mechanism is also $\tilde{O}(\sqrt{N})$.

Improving the UCB mechanism, this paper proposes a hybrid mechanism, which lifts affirmative actions upon the collection of a sufficiently rich data set. The hybrid mechanism takes advantage of spontaneous exploration: Once firms obtain a certain amount of data, the diversity of workers' characteristics naturally promotes learning about the minority group. The hybrid mechanism achieves $\tilde{O}(\sqrt{N})$ regret with $\tilde{O}(1)$ subsidy.

The second policy is the Rooney Rule, which requires each firm to interview at least one minority candidate as a finalist for each job opening. We analyze the effect of the Rooney Rule using a two-stage model in which firms observe additional signals of each finalist. The Rooney Rule enables minority workers to reveal the additional signal to the firm, which leaves a chance of breaking down the underestimation. However, our assessment of the Rooney Rule is mixed. The imposed interviewing quota could unjustly deprive skilled majority workers of employment opportunities, suggesting reverse discrimination. This drawback is lessened if the Rooney Rule is implemented temporarily.

This paper is framed as a positive analysis elucidating how discrimination arises from social learning conducted by small, rational, and unbiased firms. Alternatively, our study could be viewed as a normative analysis showcasing an efficient hiring policy targeting the long-term average skill of workers hired by a large firm Bergman2020. For this latter scenario, our results for the hybrid mechanism indicate that a firm can cease affirmative action once it has accumulated reasonably comprehensive information about minority workers.

Related Literature

\paragraph{Statistical Discrimination}

Various studies have analyzed statistical discrimination both theoretically phelps1972statistical,arrow1973theory,Foster1992AnEA,coate1993will,cornell1996culture,mailath2000endogenous and experimentally neumark2018. We contribute to this literature by articulating a new channel of discrimination: endogenous data imbalance and insufficiency. Similar to previous studies, we assume otherwise ex ante identical individuals from different groups to demonstrate how discrimination evolves and persists. Meanwhile, our results provide further indication that demographic minorities suffer from discrimination as an inevitable consequence of laissez-faire.

Hu2018Short examines a dynamic reputation model in a labor market, where workers can endogenously select their skill level. As highlighted by Foster1992AnEA and coate1993will, statistical discrimination can potentially discourage minorities from enhancing their skills. Implementing a fairness constraint through affirmative action at the entry-level may rectify inequalities within the entire labor market. Our study adds to this body of literature by demonstrating that short-term affirmative action successfully tackles inefficiency and inequality, even when the skill level is fixed.

Kannan2019Downstream study how a college can design an admission and grading policy to achieve fair employment, assuming employers form a Bayesian belief about students' skills based on the information provided by the college. We also consider a government that introduces an affirmative-action policy taking into account stakeholders' (firms') endogenous response. Che2019RatingsGuided examine a rating-guided market, demonstrating that feedback loops can cause discriminatory inferences concerning social groups. We identify endogenously created informational disparities due to feedback loops within a distinct model inspired by a hiring market, deliberate on the underlying causes (demographic imbalance) that instigate discrimination, and propose policy solutions.

Bohren2019inaccurate,bohren2019dynamics and monachou2019discrimination have demonstrated how misspecified beliefs about groups generate discrimination. Thus far, this literature has attributed belief misspecification to psychological biases and bounded rationality. In contrast, we demonstrate that misspecified beliefs may evolve and persist endogenously, even in the long run. Through a laboratory experiment, dianat2020lab reveal that affirmative action's impact becomes fleeting if the measure is discontinued before beliefs undergo transformation. Our hybrid mechanism offers a resolution by optimally choosing the timing to terminate the program, thereby preventing the persistence of underestimation.

\paragraph{Social Learning} The economics literature has extensively studied herding, information cascade, and social learning bikhchandani1992theory,banerjee1992simple,smith2000pathological. Additionally, various papers have studied improvements to social welfare through subsidy for exploration frazier2014incentivizing,kannan2017fairness and selective information disclosure kremer2014implementing,papanastasiou2018crowdsourcing,Immorlica2020incentivizing,mansour2020bayesian. We propose novel policy interventions to improve social learning in fairness and efficiency.

\paragraph{Multi-armed Bandit} A multi-armed bandit problem stems from the literature of statistics thompson1933,robbins1952. This problem is driven by the question of how a single long-lived decision-maker can maximize his payoff by balancing exploration and exploitation. More recently, the machine-learning community has proposed the contextual bandit framework, in which payoffs associated with “arms” (actions) depend not only on the hidden state but also on additional information, referred to as “contexts” abe1999,langford2007. We adopt the contextual bandit framework because context enables us to capture the diversity of worker characteristics.\footnote{The trade-off between exploration and exploitation presents itself in a wider context. For example, Owen2020Tie propose “tie-breaker designs” which are hybrids of randomized controlled trials and regression discontinuity designs, and solve the optimal tradeoff between information gain (exploration) and efficiency in the treatment allocation (exploitation).}

Several previous studies have considered a linear contextual bandit problem and studied the performance of a “greedy” algorithm, which makes decisions myopically in accordance with the current information. Because firms take greedy actions under laissez-faire, their results are also relevant to our model. Bastani17 and kannan2018 have shown that a greedy algorithm leads to sublinear regret in the long run, if the contexts are diverse enough.\footnote{Our simulation, included in Appendix (ref), shows that our hybrid mechanism can be interpreted as an efficient approach to collecting initial samples.} We characterize the relationship between the diversity of contexts and the rate of learning. Moreover, we show that the population ratio is crucial to the regret rate (Section (ref)). As an efficient intervention, kannan2017fairness consider a contextually fair UCB-based subsidy rule. Although our subsidy policy also originates from the idea of UCB (Section (ref)), we establish a novel mechanism (the hybrid mechanism, Section (ref)) that reduces budget expenditure by utilizing spontaneous exploration.

The multi-armed bandit approach has recently found applications in labor market analyses. Bardhi2019Spiraling demonstrate that a minor difference in initial beliefs about each worker's type can ultimately yield a substantial disparity in workers' payoffs. Johari2018Exploration examine how a labor platform can discern workers' skills to attain an optimal worker assignment when the platform can only observe the outcomes generated by teams, not individual workers.

Bergman2020 portray a large firm's hiring process as a multi-armed bandit problem and empirically compare the performance of the status quo (screening via manual work), a greedy policy, and a UCB method. They reveal that a UCB method not only screens job applicants efficiently but also preserves diversity. Their findings suggest that a UCB method is both fairer and more efficient when implemented by a large firm. Interpreting this paper as a study of an efficient hiring policy by a large firm, our results enhance Bergman2020 by providing theoretical foundations that outline the performances of a greedy policy (corresponding to laissez-faire) and a UCB method. Moreover, we illustrate that affirmative action can be discontinued shortly by characterizing the performance of a hybrid mechanism.

\paragraph{Algorithmic Fairness} The literature on algorithmic fairness is growing. This literature has implicitly assumed exogenous asymmetry in worker skills and pursued the approaches to correct between-group inequality. To this end, “discrimination-aware” constraints such as equalized odds HardtPNS16 and demographic parity PedreschiRT08,CaldersV10 have been proposed, with several papers applying these constraints in the context of multi-armed bandit problems joseph_fairness_2016 or more general sequential learning RaghavanSVW18,BechavodL0WW19,ChenAAMAS2020. While these fairness goals are conflicting in general, we analyze an environment in which many fairness goals are aligned and demonstrate how affirmative action improves them.

\paragraph{Rooney Rule} The Rooney Rule was originally introduced in the context of the hiring of National Football League senior staff Eddo-Lodge2017Rooney. While it is widely used in practice, theoretical analyses of the Rooney Rule are scarce. DBLP:conf/innovations/KleinbergR18 show that, when a recruiter is unconsciously biased against a group, the Rooney Rule not only improves the representation of that group but also leads to a higher payoff for the recruiter. To the best of our knowledge, this study (Section (ref)) constitutes the first attempt to demonstrate the advantage of the Rooney Rule by modeling unbiased agents.

Model

\paragraph{Basic Setting} We develop a linear contextual bandit problem with myopic agents (firms). We consider a situation where $N$ firms (indexed by $n = 1, \ldots, N$) sequentially hire one worker for each.\footnote{While real-world firms are long-lived and hire multiple workers, the number of workers hired by one firm is typically much smaller than the total number of workers hired in a hiring market. Accordingly, even if we allowed firms to hire multiple (but a small number of) workers, the conclusion would not change qualitatively. Note also that various seminal papers within the social learning literature banerjee1992simple,bikhchandani1992theory,smith2000pathological have made the same assumption.} In each round $n$, a set of workers $I(n)$ (i.e., arms) arrives. Each worker $i \in I(n)$ takes no action, and firm $n$ hires only one worker $\iota(n) \in I(n)$. Both firms and workers are short-lived. Upon round $n$ ending, firm $n$'s payoff is finalized, and all rejected workers leave the market.\footnote{This assumption is for the sake of simplicity. Since firms have no private information, the fact that a worker was previously rejected by another firm does not influence the worker's evaluation (given that the current firm can also observe the worker's characteristics); thus, entrant workers and incumbent workers have no informational difference. Accordingly, even if workers stay in the hiring market for multiple periods, our conclusion will not be changed qualitatively.}

Each worker $i \in I$ belongs to a group $g \in G$. We assume that the population ratio is fixed: for every round $n$, the number of workers belonging to group $g$ is $K_g \in \mathbb{N}$ and $K = \sum_{g \in G} K_g$. Slightly abusing the notation, we denote the group worker $i$ belongs to by $g(i)$. Each worker $i$ also has observable characteristics $\bm{x}_i \in \mathbb{R}^d$, with $d \in \mathbb{N}$ as their dimension. Finally, each worker $i$ also has a skill $y_i \in \mathbb{R}$ that is not observable until worker $i$ is hired. The characteristics and skills are random variables.

Because each firm's payoff is equal to the hired worker's skill $y_i$ (plus the subsidy assigned to worker $i$ as an affirmative action, if any), firms want to predict the skill $y_i$ based on the characteristics $\bm{x}_i$. We assume that characteristics and skills are associated as $y_i = \bm{x}_i' \bm{\theta}_{g(i)} + \epsilon_i$, where $\bm{\theta}_g \in \mathbb{R}^d$ is a coefficient parameter, and $\epsilon_i \sim \mathcal{N}(0, \sigma^2_\epsilon)$ i.i.d.\ is an unpredictable error term. We assume $||\bm{\theta}_g|| \le S$ for some $S \in \mathbb{R}_{+}$, where $||\cdot||$ is the standard L2-norm. Since $\epsilon_i$ is unpredictable, $q_i \coloneqq \bm{x}_i' \bm{\theta}_{g(i)}$ is the best predictor of worker $i$'s skill $y_i$.

The coefficient parameters $(\bm{\theta}_{g})_{g\in G}$ are initially unknown. Hence, unless firms share information about past hires, firms are unable to predict each worker's skill $y_i$. We assume that firms share information about past hiring cases.\footnote{Alternatively, we can assume that firms only share information about a certain fraction of workers. We expect that, under this assumption, (i) the results would not change qualitatively, and (ii) the statistical discrimination would become severer because it becomes more difficult to accumulate information about the minority group.} Accordingly, when firm $n$ makes a decision, in addition to the characteristics and groups of current workers $(\bm{x}_i, g(i))_{i \in I(n)}$, firm $n$ observes the characteristics, groups, and skills of previously hired workers $(x_{\iota(n')}, g(\iota(n')), y_{\iota(n')})_{n' = 1}^{n - 1}$. We refer to all realizations of these variables as the history in round $n$, and denote it by $h(n)$. Formally, $h(n)$ is given by

equation[equation omitted — 135 chars of source]

Note that, $h(n)$ does not include information about (i) the worker hired by firm $n$, or (ii) that worker's actual skill. This is because the notation $h(n)$ represents the information set firm $n$ faces when it makes a hiring decision. We denote the set of all the possible histories in round $n$ by $H(n)$. The firm's decision rule for hiring and the government's subsidy rule are defined as a function that maps a history to a hiring decision and the subsidy amount (described later). For notational convenience, we often omit $h(n)$.

\paragraph{Prediction} We assume that firms are not Bayesian but frequentists. Hence, firms do not have a prior belief about the parameter $\bm{\theta}$ but estimate it only using the available data set. We expect that essentially the same results will be obtained with Bayesian firms (see Appendix (ref)).

We assume that each firm predicts skill using ridge regression (L2-regularized least square).\footnote{For the properties of the ridge estimator, see Kennedy2008, for example.} Let $N_g(n)$ be the number of rounds at which group-$g$ workers are hired before round $n$. Let $\bm{X}_g(n) \in \mathbb{R}^{N_g(n)\times d}$ be a matrix that lists the characteristics of group-$g$ workers hired by round $n$: each row of $\bm{X}_g(n)$ corresponds to $\{\bm{x}_{\iota(n')}: \iota(n') = g\}_{n' = 1}^{n - 1}$. Likewise, let $Y_g(n) \in \mathbb{R}^{N_g(n)}$ be a vector that lists the skills of group-$g$ workers hired by round $n$: each element of $Y_g(n)$ corresponds to $\{y_{\iota(n')}: \iota(n') = g\}_{n'=1}^{n-1}$. We define $\bm{V}_g(n) \coloneqq (\bm{X}_g(n))'\bm{X}_g(n)$. For a parameter $\lambda > 0$, we define $\bar{\bm{V}}_g(n) = \bm{V}_g(n) + \lambda \bm{I}_d$, where $\bm{I}_d$ denotes the $d\times d$ identity matrix. Firm $n$ estimates the parameter as follows:

equation[equation omitted — 123 chars of source]

Firm $n$ predicts worker $i$'s skill $q_i$, while substituting the true predicted skill $\bm{\theta}_g$ with estimated skill $\hat{\bm{\theta}}_g(n)$: $\hat{q}_i(n) \coloneqq \bm{x}_i' \hat{\bm{\theta}}_{g(i)}(n)$. Hence, $\hat{q}_i(n)$ and $\hat{\bm{\theta}}_{g}(n)$ depend on the history $h(n)$. The ordinary least squares (OLS) estimator corresponds to the ridge estimator with $\lambda = 0$. We use the ridge estimator instead of the OLS estimator to stabilize the small-sample inference. For example, for some history, $\bm{V}_g(n)$ may not have full rank, and the OLS estimator may not be well-defined. Even for such histories, the ridge estimator is always well-defined.

For analytical tractability, we assume that for the first $N^{(0)}$ rounds, each firm $n$ must hire from a pre-specified group, $g_n$. We refer to the first $N^{(0)}$ rounds as the initial sampling phase. We assume $N^{(0)}$ to be small and deal $N^{(0)}$ as a constant.\footnote{The required size of $N^{(0)}$ is specified by Eq. (ref) in Appendix.} Let $N^{(0)}_g \coloneqq \sum_{n=1}^{N^{(0)}}\textbf{1}[g_n = g]$ as the data size of initial sampling for group $g$, where $\textbf{1}[\mathcal{A}] = 1$ if event $\mathcal{A}$ holds or $0$ otherwise. The initial sampling phase is exogenous. That is, we ignore the incentives and payoffs of firms and assume that the characteristics $\bm{x}$ of the hired candidate constitute an i.i.d.\ sample of the corresponding group. We analyze mechanism, social welfare, and budget after round $n > N^{(0)}$. The initial sampling phase can be interpreted as data that has already been produced in history. The welfare cost is already sunk, and the government can no longer make policy interventions for the event that has already occurred in the past.

\paragraph{Mechanism}

In addition to worker skills, firms are also concerned about subsidies. We assume that firm preferences are risk-neutral and quasi-linear. Hence, if firm $n$ hires worker $i$, its payoff (von-Neumann--Morgenstern utility) is given by $y_i + s_i$, where $s_i \in \mathbb{R}_+$ denotes the amount of the subsidy assigned to worker $i$.

In the beginning of the game, the government commits to a subsidy rule $s_i(n,\cdot): H(n) \to \mathbb{R}_+$, which maps a history to a subsidy amount. Hence, once a history $h(n)$ is specified, firm $n$ can identify the subsidy assigned to each worker $i\in I(n)$. Firm $n$ attempts to maximize

equation[equation omitted — 115 chars of source]

Firm $n$'s decision rule $\iota(n, \cdot): H(n) \to I(n)$ specifies the worker that firm $n$ hires after history $h(n)$. We say that, a decision rule $\iota$ is implemented by a subsidy rule $s_i$ if for all $n$ and $h(n)$, we have

equation[equation omitted — 158 chars of source]

Throughout this paper, any ties are broken arbitrarily. We call a pair of a decision rule and subsidy rule a mechanism. We often drop $h(n)$ from the input of decision rule $\iota$ when it does not cause confusion.

\paragraph{Regret}

Regret is a standard measure for evaluating the performance of algorithms in multi-armed bandit models:

equation[equation omitted — 149 chars of source]

Since $\epsilon_i$ is unpredictable, it is natural to evaluate the performance of the algorithm (or the equilibrium consequence of the policy intervention) by comparing it with $q_i$. If the parameter $(\bm{\theta}_g)_{g\in G}$ were known, each firm could easily calculate $q_i$ for each worker $i$ and hire the best worker, $i^*(n) \coloneqq \operatorname*{arg\,max}_{i\in I(n)}q_i$. In this case, regret would be zero. The goal of the policy design is to establish a mechanism that minimizes the expected regret $\mathbb{E}[\mathrm{Reg}(N)]$, where the expectation is taken on a random draw of workers. This aim is equivalent to maximizing the sum of the skill of workers hired.

Following the literature, we often evaluate the performance by the limiting behavior (order) of expected regrets. A decision rule $\iota$ is said to have sublinear regret if $\mathbb{E}[\mathrm{Reg}(N)]=O(N^a)$ for some $a < 1$. Small regret implies not only efficiency but also fairness. Regret measures the disparate impact that is not justified by skill disparity, and sublinear regret is achieved if firms hire the most skillful workers without regard to the group of workers. In Appendix (ref), we demonstrate that a sublinear-regret decision rule asymptotically aligns with a fairness notion called equalized odds, which requires that candidate workers in the majority and minority groups have an equal true positive rate (hired when they have the highest skill predictor $q_i$) and equal false negative rate (not hired when they have the highest $q_i$).

\paragraph{Budget}

Some of the policies we study incentivize exploration through subsidies. The total budget required by a subsidy rule is also an important policy concern. The total amount of the subsidy is given by $\mathrm{Sub}(N) \coloneqq \sum_{n = N^{(0)}+1}^N s_{\iota(n)}(n)$.

Laissez-Faire

This section analyzes the equilibrium under laissez-faire; that is, the consequence of social learning in the absence of policy intervention.

definition[Laissez-Faire] The laissez-faire decision rule always selects the worker who has the greatest estimated skill, i.e., $\iota(n) = \operatorname*{arg\,max}_{i\in I(n)}\hat{q}_i(n)$. This decision rule is implemented by the laissez-faire subsidy rule, which provides no subsidy $s_i(n) = 0$ after any history.

Laissez-faire makes no intervention. Each firm hires the worker with the greatest estimated skill, as predicted by the current data set. The multi-armed bandit literature refers to the laissez-faire decision rule as the greedy algorithm.

Symmetry and Diverse Characteristics

To illustrate a failure of social learning, we make three assumptions. First, as a minimal environment to analyze discrimination, we focus on the two-group case.

assp[Two Groups] The population comprises two groups $G = \{1,2\}$.

When we consider asymmetric equilibria, we refer to group $1$ as the majority (dominant) group and group $2$ as the minority (discriminated-against) group. The two-group assumption enables the elucidation of how the minority group is discriminated against.

Second, we assume that groups are symmetric.

assp[Symmetric Groups] The characteristics of all groups are identical, and the coefficient parameters are the same across the groups. That is, a probability distribution $F$ such that for all $i \in I$, $\bm{x}_i \sim F$, and there exists $\bm{\theta} \in \mathbb{R}^d$ such that, for all $g \in G$, $\bm{\theta}_g = \bm{\theta}$.

Note that although we assume that groups are symmetric, firms do not know the true parameters, and therefore, apply different statistical models to different groups. That is, even though the true coefficients are identical ($\bm{\theta}_g = \bm{\theta}_g'$ for all $g,g' \in G$), firms estimate them separately; thus, the values of the estimated coefficients are typically different ($\hat{\bm{\theta}}_g(n) \neq \hat{\bm{\theta}}_{g'}(n)$ for $g\neq g'$).

Although Assumption (ref) is unrealistic (because the characteristics should evidently be interpreted differently), it is useful for elucidating how laissez-faire nourishes statistical discrimination. Under Assumption (ref), agents are ex ante identical arrow1973theory,Foster1992AnEA,coate1993will,moro_general_2004, and therefore the differences we observe in the equilibrium are entirely attributed to social learning.

Furthermore, when groups are symmetric, disparate impact is unambiguously unfair. It is well-known that popular fairness notions aim at different goals and are compatible with each other only in highly constrained special cases KleinbergMR17. The symmetric environment specified by Assumption (ref) one of such exceptions: In this environment, sublinear regret implies not only equalized odds but also demographic parity, i.e., the probability of a worker to be hired is independent of his group (see Appendix (ref)). Since this paper's focus is not to debate which of the various types of fairness notions should be respected, we will concentrate only on the symmetric environment.\footnote{We confirmed through simulations that the proposed mechanisms are effective in a broad class of asymmetric environments. See Appendices (ref) and (ref).}

Third, we assume that characteristics are normally distributed, and therefore, the distribution is non-degenerate. This assumption captures the diversity of workers.

assp[Normally Distributed Characteristics] For every candidate $i$, \begin{equation} \bm{x}_i \sim \mathcal{N}(\bm{\mu}_{xg(i)}, \sigma_{xg(i)}^2 \bm{I}_d), \end{equation} where $\bm{\mu}_{xg} \in \mathbb{R}^d$ and $\sigma_{xg} \in \mathbb{R}_{++}$ for every $g \in G$. We also denote $\bm{x}_i = \bm{\mu}_{xg(i)} + \bm{e}_{xi}$ to highlight the noise term $\bm{e}_{xi}$.

We consider essentially the same results to hold more generally as long as the characteristics are sufficiently diverse. Note that when we have both Assumptions (ref) and (ref), then there exist $\bm{\mu}_x, \sigma_x$ such that $\bm{\mu}_{xg} = \bm{\mu}_x$ and $\sigma_{xg} = \sigma_x$ for all $g \in G$. Hence, $\bm{x}_i \sim \mathcal{N}(\bm{\mu}_x,\sigma_x^2 \bm{I}_d)$ for all $i$.

Perpetual Underestimation

To determine whether social learning incurs linear expected regret, it is useful to check whether it results in perpetual underestimation with a significant probability.

definition[Perpetual Underestimation] A group $g_0$ is perpetually underestimated if, for all $n > N^{(0)}$, we have $g(\iota(n)) \neq g_0$.

When group $g_0$ is perpetually underestimated, no worker from group $g_0$ is hired after the initial sampling phase. If social learning generates perpetual underestimation with a significant probability, then linear expected regret often results. In particular, under Assumption (ref), perpetual underestimation against any group $g \in G$ implies that firms fail to hire at least $(K_g/K)\left(N - N^{(0)} \right)$ best candidate, which is linear in $N$. Hence, the constant probability of perpetual underestimation (independent of $N$) precipitates linear expected regret.

Perpetual underestimation is not only inefficient but also unfair in the sense of various fairness notions (formally defined in Appendix (ref)); it results in a candidate belonging to an underestimated group not being hired, implying a violation of demographic parity. Furthermore, under a symmetric environment, such a hiring policy cannot be justified by workers' underlying skills, implying a violation of equalized odds. Hence, perpetual underestimation is an extreme form of discrimination that persists for a long time.

Sublinear Regret with Balanced Population

This section analyzes the case of only one candidate arriving from each group during each period. The contextual variation implicitly urges firms to explore all the groups with some frequency. Consequently, laissez-faire has sublinear regret, implying that statistical discrimination is eventually resolved.

thm[Sublinear Regret with a Balanced Population] Suppose Assumptions (ref), (ref), and (ref). Suppose also that $K_g = 1$ for $g = 1,2$. Then, expected regret $\mathrm{Reg}^{\text{LF}}(N)$ under the laissez-faire policy is bounded as \begin{equation} \mathbb{E}[\mathrm{Reg}^{LF}(N)] = \tilde{O}(\sqrt{N}). \end{equation}

Let $\mu_x = ||\bm{\mu}_x||$ and $\Phi$ be the cumulative distribution function of the standard normal distribution. The constant on the top of $\mathbb{E}[\mathrm{Reg}^{\text{LF}}(N)]$ is inverse proportional to $1 - \Phi(\mu_x/\sigma_x)$, which approximately scales as $\exp(-(\mu_x/\sigma_x)^2/2)$.

\paragraph*{Proof.} See Appendix (ref).

To prove Theorem (ref), we characterize the condition with which underestimation is spontaneously resolved. Let indices $i_1$ and $i_2$ denote the majority candidate and the minority candidate. With a constant (i.e., independent of $N$) probability, the minority group is underestimated (i.e, $\hat{\bm{\theta}}_2(n)$ is misestimated in such that $\bm{x}_{i_2}\hat{\bm{\theta}}_2(n) \ll \bm{x}_{i_2}\bm{\theta}_2$ often occurs) in early rounds due to a bad realization of the error term. Even in such a case, there is some probability of the minority candidate being hired. Since characteristics are diverse (i.e., $\sigma_x > 0$), with some probability, the majority candidate $i_1$ is not very good (i.e., $x_{i_1}\hat{\bm{\theta}}_1(n) \approx x_{i_1}\bm{\theta}_1$ is small). In such a round, $x_{i_1} \hat{\bm{\theta}}_1(n) < x_{i_2} \hat{\bm{\theta}}_2(n)$ holds despite group $2$ being underestimated, and the minority candidate $i_2$ is hired. In such a case, firms update their belief about the minority, leading to a resolution of underestimation. Such events occur more frequently when workers have more diverse characteristics, i.e., $\mu_x/\sigma_x$ is small.

As anticipated by the theory of least squares, the standard deviation of $\hat{\bm{\theta}}_g(n)$ is proportional to $(\bar{\bm{V}}_g(n))^{-1/2}$, and we demonstrate that its diameter $(\lambda_{\mathrm{min}}(\bar{\bm{V}}_g(n)))^{-1/2}$ shrinks as $\tilde{O}(1/\sqrt{n})$, where $\lambda_{\mathrm{min}}$ is the minimum eigenvalue of a matrix. The regret per error is defined by this quantity, with the total regret being $\tilde{O}(\sum_{n\le N} (1/\sqrt{n}))= \tilde{O}(\sqrt{N})$.

Theorem (ref) indicates that statistical discrimination is resolved spontaneously when candidate variation is large. At a glance, this appears to contradict widely known results that state laissez-faire (greedy) may lead to suboptimal results in bandit problems due to underexploration. However, the variation in characteristics naturally incentivizes selfish agents to explore the underestimated group, and therefore, with some additional conditions, the probability of perpetual underestimation is bounded.

remarkIn Theorem (ref), we assumed that there is one candidate for each group, $K_1 = K_2 = 1$, for tractability. If we assume a larger but balanced population, $K_1 = K_2 = K/2$, then the analysis would become significantly more challenging because the maximum of normally distributed variables is not normally distributed. However, we conjecture that a similar result would hold more generally because the variance of the expected skill of the best candidate in each group decreases only slowly as $K$ increases.\footnote{Lemma (ref) in the Appendix implies that the variance is in the order of $O(1/\log K)$.}
remarkTheorem (ref) shares certain intuitions with the previous research kannan2018,Bastani17 demonstrating that the variation in contexts (characteristics) improves the performance of the greedy algorithm (laissez-faire) in contextual multi-armed bandit problems. However, in contrast to kannan2018, our theorem makes no assumptions regarding the length of the initial sampling phase. Theorem 1 in Bastani17 corresponds to our paper's Theorem (ref), and we further characterize the factor of the regret as a function of $\mu_x/\sigma_x$ rather than the diameter of the characteristics.

Large Regret with Unbalanced Population

While Theorem (ref) implies that statistical discrimination is spontaneously resolved in the long run, it crucially relies on one unrealistic assumption---the balanced population ratio. In many real-world problems, the population ratio is unbalanced, and the discriminated group is often a demographic minority in the relevant market. We indeed find that the population ratio crucially impacts the equilibrium consequence under laissez-faire.

thm[Substantial Regret with Unbalanced Populations] Suppose Assumptions (ref), (ref), and (ref). Suppose also that $K_2 = 1$ and $d=1$. Let $K_1 > \log_2 N$. Then, under the laissez-faire decision rule, group $2$ is perpetually underestimated with a probability of at least $C_{\text{imb}} = \tilde{\Theta}(1)$. Accordingly, the expected regret associated with the laissez-faire decision rule is \begin{equation} \mathbb{E}\left[\mathrm{Reg}^{LF}(N)\right] \ge \frac{C_{imb}(N-N^{(0)})}{K} = \tilde{\Omega}(N). \end{equation}

\paragraph*{Proof.} See Appendix (ref). The explicit form of $C_{\text{imb}}$ is shown in Eq. (ref).

In the proof of Theorem (ref), we evaluate the probability that the following two events occur: (i) $\hat{\theta}_2$ is underestimated, and (ii) the characteristics and skills of the hired majority workers are not very bad throughout rounds (i.e., $\max_{i: g(i)=1} x_i \hat{\theta}_1 \ge c \mu_x \theta$ for some constant $c>0$). The probability of (i) is polylogarithmic to $N$ (i.e., $\tilde{\Theta}(1)$) and the probability that (ii) consistently holds for all the rounds $n=N^{(0)}+1,\dots,N$ is polylogarithmic if $K_1 > \log_2 N$. When both (i) and (ii) occur, we always have $\max_{i \in I(n)\setminus\{i_2(n)\}} x_i \hat{\theta}_1 > x_{i_2(n)} \hat{\theta}_2$ (where $i_2(n)$ is the unique minority candidate of round $n$); thus, the minority worker is never hired. Note that the majority group does not suffer from perpetual underestimation (with a significant probability) because the event that all the majority workers are bad occasionally occurs.

Theorem (ref) indicates that we should not be too optimistic about the consequence of laissez-faire. A small imbalance in the population ratio (the ratio of majority to minority is just $\log_2 N$ to $1$) could lead to a substantially unfair job allocation. Once the minority group is underestimated and the majority candidate pool is reasonably large, then the minority group is afforded no hiring opportunity, perpetuating underestimation. This insight applies to many real-world problems because unbalanced populations are commonplace.

We conjecture a substantial probability under a broader environment than the premise of Theorem (ref). Specifically, the assumptions of $d = 1$ and $K_1 > \log_2 N$ are made only for analytical tractability, and (approximately) linear regret should be obtained under a weaker set of assumptions. Theorem (ref) (i) focuses on perpetual underestimation, which is an extreme form of statistical discrimination, and (ii) evaluates the probability of perpetual estimation occurring loosely. In Section (ref), we demonstrate that perpetual underestimation occurs with a significant probability even under the assumptions of $d = 5$ and $(K_1, K_2) = (10, 2)$, where the premise of Theorem (ref) does not hold.

The Upper Confidence Bound Mechanism

Section (ref) has discussed the equilibrium consequences of laissez-faire. We observed that an unbalanced population ratio leads to a substantial probability of underestimation being perpetuated. Policy intervention is demanded to improve social welfare and the fairness of the hiring market.

This section proposes a subsidy rule to resolve underestimation. We employ the idea of the upper confidence bound (UCB) algorithm lai1985,auer2002, which has widely been used in the literature on the bandit problem. The UCB algorithm balances exploration and exploitation by developing a confidence interval for the true reward and evaluating each arm's performance according to its upper confidence bound to achieve this balance. Firms are generally unwilling to follow the UCB decision rule voluntarily; therefore, the government needs to provide a subsidy to incentivize firms to hire a candidate with the greatest UCB index. This section establishes a UCB-based subsidy rule and evaluates its performance.

The adaptive selection of candidates based on history can induce some bias, meaning the standard confidence bound no longer applies. To overcome this issue, we use martingale inequalities selfnormalized,rusmevichientong2010,abbasi2011. We here introduce the confidence interval for the true coefficient parameter, $(\bm{\theta}_g)_{g\in G}$.

definition[Confidence Interval] Given the group $g$'s collected data matrix $\bar{\bm{V}}_g(n)$, the confidence interval of group $g$'s coefficient parameter $\bm{\theta}_g$ is given by \begin{equation} \mathcal{C}_g(n; \delta) \coloneqq \left\{ \bar{\bm{\theta}}_g \in \mathbb{R}^d: \left\lVert\bar{\bm{\theta}}_g - \hat{\bm{\theta}}_g(n)\right\rVert_{\bar{\bm{V}}_g(n)} \le \sigma_\epsilon \sqrt{d \log\left(\frac{\det(\bar{\bm{V}}_g(n))^{1/2}\det(\lambda \bm{I}_d)^{-1/2}}{\delta}\right)} + \lambda^{1/2} S \right\}, \end{equation} where $||\bm{v}||_{\bm{A}} = \sqrt{\bm{v}' \bm{A} \bm{v}}$ for a $d$-dimensional vector $\bm{v}$ and $d \times d$ matrix $\bm{A}$.

abbasi2011 study the property of this confidence interval, and they prove that the true parameter $\bm{\theta}_g$ lies in $\mathcal{C}_g(n;\delta)$ with probability $1-\delta$ (Lemma (ref)). By choosing a sufficiently small $\delta$,\footnote{We typically choose $\delta = 1/N$ to make the confidence interval asymptotically correct in the limit of $N \to \infty$.} it is “safe” to assess that worker $i$'s skill is at most

equation[equation omitted — 144 chars of source]

We call $\tilde{q}_i(n)$ the UCB index of worker $i$'s skill. Intuitively, $\tilde{q}_i(n)$ is worker $i$'s skill in the most optimistic scenario. The confidence interval $\mathcal{C}_g(n;\delta)$ shrinks as we obtain more data about group $g$. Hence, the UCB index $\tilde{q}_i(n)$ converges to true predicted skill $q_i(n)$ as the size of the data set increases.

definition[UCB Decision Rule] The UCB decision rule selects the worker with the greatest UCB index; i.e., \begin{equation} \iota(n) = \operatorname*{arg\,max}_{i \in I(n)}\tilde{q}_i(n). \end{equation}

The UCB index $\tilde{q}_i(n)$ is close to the pointwise estimate $\hat{q}_i(n)$ when society has rich data about group $g(i)$, because $\mathcal{C}_{g(i)}(n;\delta)$ is small in such cases. However, when information about group $g(i)$ is insufficient, $\tilde{q}_i(n)$ is much larger than $\hat{q}_i(n)$, because the firm is unsure about the true skill of worker $i$ and $\mathcal{C}_{g(i)}(n;\delta)$ is large. In this sense, the UCB decision rule offers affirmative actions toward underexplored groups.

The subsidy amount is proportional to the uncertainty surrounding the candidate's characteristics, which is represented by the confidence interval $\mathcal{C}_g(n)$ for $g=g(i)$. The magnitude of the confidence interval $\mathcal{C}_g(n)$ is inverse proportional to ${\bar{\bm{V}}_g(n)} = \bm{V}_g(n) + \lambda \bm{I}_d$.\footnote{The standard OLS has a confidence bound of the form $\bm{\theta}_g - \hat{\bm{\theta}}_g(n) \sim \mathcal{N}(0, \sigma_\epsilon^2 \bm{V}^{-1}_g(n))$ and thus $|\bm{\theta}_g - \hat{\bm{\theta}}_g(n)| \sim \sigma_\epsilon \bm{V}^{-1/2}_g(n)$. The price of adaptivity causes the martingale confidence bound $\mathcal{C}_g(n)$ to be larger than the OLS confidence bound for two factors: (i) $\sqrt{d}$ factor, and (ii) $\sqrt{\log(\det(\bar{\bm{V}}_g(n)))}$ factor. As discussed in Xu2018AFA, the $\sqrt{d}$ factor unnecessarily overestimates the confidence bound in most cases.} Hence, if the data $\bm{V}_g(n)$ do not vary substantially for a particular dimension of $\bm{x}_i$, then that dimension's prediction can be inaccurate. In such cases, the UCB decision rule recommends hiring a candidate that contributes to increasing that dimension's data. For example, when a candidate possesses skills previous hires do not, then the candidate's UCB index tends to become large.

The UCB decision rule efficiently balances exploration and experimentation. Accordingly, it has sublinear regret in general environments.

thm[Sublinear Regret of UCB] Suppose Assumption (ref). Let $\mathrm{Reg}^{\text{UCB}}$ be the regret from the UCB decision rule. Let $\lambda \ge \max(1,(L_{1/N})^2)$, where $L_{1/N}$ is an $O(\sqrt{d\log{KN}})$ value defined in Lemma (ref) in Appendix. Then, by choosing $\delta=1/N$, regret under the UCB decision rule is bounded as \begin{equation} \mathbb{E}[\mathrm{Reg}^{\mathrm{UCB}}(N)] = \tilde{O}(\sqrt{N}). \end{equation}

\paragraph*{Proof.} See Appendix (ref).

There are three remarks. First, $\tilde{O}(\sqrt{N})$ regret is the optimal rate for these sequential optimization problems under partial feedback pmlr-v15-chu11a. Hence, Theorem (ref) states that the UCB decision rule effectively prevents perpetual underestimation and is asymptotically efficient. Second, Theorem (ref) relies only on Assumption (ref), and therefore, the regret under UCB is sublinear even when groups have a fundamental disparity besides their group sizes. Accordingly, even when the groups are asymmetric, the UCB decision rule satisfies several fairness notions (see Appendix (ref) for details). Third, differing from the case of laissez-faire, where the factor depends on the variation of the context (Theorem (ref)), Theorem (ref) provides a reasonably small regret bound even when $\sigma_x$ is very small.

To implement the UCB decision rule, we need to satisfy the firms' obedience condition (ref) in conjunction with the UCB decision rule (ref). In the following, we propose one of the most straightforward subsidy rules.

definition[UCB Index Subsidy Rule] The UCB index subsidy rule $s$ subsidizes firm $n$ to hire worker $i$ who arrives by \begin{equation} s_i(n;h(n)) = \tilde{q}_i(n;h(n)) - \hat{q}_i(n;h(n)). \end{equation}

The UCB index subsidy rule aligns each firm's incentive with the maximization of the UCB index, thereby incentivizing firms to follow the UCB decision rule.

thm[Sublinear Subsidy of the UCB Index Subsidy Rule] Under the same assumptions as Theorem (ref), the amount of the subsidy required by the UCB index subsidy rule is bounded as \begin{equation} \mathbb{E}[\mathrm{Sub}^{UCB-I}(N)] = \tilde{O}(\sqrt{N}). \end{equation}

\paragraph*{Proof.} See Appendix (ref).

remarkThe UCB index subsidy rule is an index policy in the sense that the subsidy amount is independent of the information about rejected workers. The UCB index subsidy rule demands the smallest budget among all index policies implementing the UCB decision rule. In Appendix (ref), we consider a non-index subsidy rule that implements the UCB decision rule with a smaller budget.

The Hybrid Mechanism

Although the UCB mechanism effectively prevents perpetual underestimation and achieves sublinear regret in general environments, it has one drawback: it continues subsidies in perpetuity. Even for a large $n$, there remains a gap between estimated skill $\hat{q}_i(n)$ and the UCB index $\tilde{q}_i(n)$. This is undesirable for several reasons. First, introducing a permanent policy is often more politically difficult than introducing a temporary policy. Second, a long-term distribution of subsidies tends to increase the required budget. Third, in addition to the subsidy itself, the permanent allocation of the subsidy features (unmodeled) administrative costs.

To overcome these limitations, we propose the hybrid mechanism, which initially uses the UCB mechanism but switches to laissez-faire by terminating the subsidy at some point. We abandon the UCB phase upon receiving sufficient minority-group data to induce spontaneous exploration. Similar to the UCB mechanism, our hybrid mechanism has $\tilde{O}(\sqrt{N})$ regret. Furthermore, its expected total subsidy amount is $\tilde{O}(1)$, while the UCB mechanism needs $\tilde{O}(\sqrt{N})$ subsidy.

The construction of the hybrid mechanism is as follows. Let $s^{\text{U-I}}_i(n) = \tilde{q}_i(n) - \hat{q}_i(n)$ be the size of the confidence bound. Note that, $s^{\text{U-I}}_i(n)$ corresponds to the amount of the subsidy allocated by the UCB index subsidy rule (Definition (ref)). The hybrid index $\tilde{q}^{\mathrm{H}}_i$ is defined as

equation[equation omitted — 257 chars of source]

where $a \ge 0$ is the mechanism's parameter.

The hybrid index is literally a “hybrid” of estimated skill $\hat{q}_i(n)$ and the UCB index $\tilde{q}_i(n)$. If the difference between the UCB index and estimated skill surpasses the threshold (i.e., $s^{\text{U-I}}_i(n) > a ||\hat{\bm{\theta}}_{g(i)}(n)||$), then the hybrid index is equal to the UCB index $\tilde{q}_i(n)$. The confidence bound $|\tilde{q}_i(n) - \hat{q}_i(n)|$ is large when society has insufficient knowledge about group $g(i)$, which is typically the case during early stages of the game. Once this gap falls below the threshold (i.e., $s^{\text{U-I}}_i(n) \le a ||\hat{\bm{\theta}}_{g(i)}(n)||$), then the hybrid index switches to the estimated skill $\hat{q}_i(n)$.

The hybrid decision rule is defined as the rule that hires the greatest hybrid index.

definition[Hybrid Decision Rule] The hybrid decision rule selects the worker who has the greatest hybrid index; i.e., \begin{equation} \iota^{H}(n;h(n)) = \operatorname*{arg\,max}_{i\in I(n)} \tilde{q}^{\mathrm{H}}_i(n;h(n)). \end{equation}

Since the hybrid decision rule is a hybrid of the UCB decision rule and the laissez-faire decision rule, it can be implemented by mixing the laissez-faire subsidy rule and the UCB index subsidy rule.

definition[Hybrid Index Subsidy Rule] Let $s^{\text{U-I}}_i$ be the UCB index subsidy rule. The hybrid index subsidy rule $s^{\text{H-I}}$ is defined by \begin{equation} s^{H-I}_i(n;h(n)) \coloneqq \begin{cases} s^{U-I}_i(n;h(n)) &if s^{U-I}_i(n;h(n)) > a ||\hat{\bm{\theta}}_{g(i)}(n;h(n))||, \\ 0 & otherwise. \end{cases} \end{equation}

The following theorems characterize the regret and the total subsidies associated with the hybrid mechanism.

thm[Performance of the Hybrid Mechanism] Suppose Assumptions (ref), (ref), and (ref). Then, by choosing $\delta=1/N$, regret associated with the hybrid decision rule $\iota^{\text{H}}$ is bounded as \begin{equation} \mathbb{E}[\mathrm{Reg}^{H}(N)] = \tilde{O}(\sqrt{N}). \end{equation} Furthermore, for any $a > 0$, the total amount of the subsidy under the hybrid index subsidy rule ($\mathrm{Sub}^{\text{H-I}}$) is bounded as \begin{align} \mathbb{E}[\mathrm{Sub}^{H-I}(N)] = \tilde{\Theta}(1). \end{align}

\paragraph*{Proof.} See Appendix (ref).

Theorem (ref) states that (i) the order of the regret under the hybrid decision rule is the same as the original UCB, and (ii) the subsidy amount is reduced to $\tilde{O}(1)$ (with respect to $N$). This is a substantial improvement from the UCB mechanism, which requires the $\tilde{O}(\sqrt{N})$ subsidy.

The threshold for switching from the UCB mechanism to laissez-faire is crucial for guaranteeing the performance of the hybrid mechanism. Our threshold, $a ||\hat{\bm{\theta}}(n)||$, is determined such that the hybrid decision rule $\iota^{\text{H}}$ satisfies proportionality, a new concept that this paper establishes. We prove that the amount of exploration exerted by the hybrid decision rule is proportional to the UCB decision rule. This property guarantees that the hybrid rule resolves underestimation and secures the expected regret of $\tilde{O}(\sqrt{N})$. The formal statement of the proportionality appears in Lemma (ref) in Appendix (ref).

Interviews and the Rooney Rule

Although subsidy rules effectively resolve statistical discrimination, they are often difficult to implement in practice. This section articulates the advantages and disadvantages of the Rooney Rule, a regulation that requires each firm to invite at least one candidate from each group to an on-site interview. The Rooney Rule is easier to implement because it requires neither a subsidy nor meeting a hiring quota.

To incorporate the additional information firms acquire through the interview, we modify the model as follows. In the modified model, each round $n$ comprises two stages. At the first stage, firm $n$ observes the characteristics $\bm{x}_i$ of each arriving agent $i \in I(n)$. Based on $\bm{x}_i$, firm $n$ selects a shortlist of finalists $I^F(n) \subseteq I(n)$, where $|I^F(n)| = K^F$ for some constant $K^F\in \mathbb{N}$. At the second stage, by interviewing finalists, firm $n$ observes an additional signal $\eta_i$ for each finalist $i$ DBLP:conf/innovations/KleinbergR18. Firm $n$ predicts each finalist $i$'s skill from the characteristics $\bm{x}_i$ and the additional signal $\eta_i$, and hires one worker from the set of finalists, $\iota(n) \in I^F(n)$. Firms are not allowed to hire a worker not selected as a finalist. After the firm's decision, the skill of the hired worker $y_{\iota(n)}$ is publicly disclosed.

We assume the following linear relationship between skill $y_i$ and observable variables $\bm{x}_i$: $ y_i = \bm{x}_i' \bm{\theta}_{g(i)} + \eta_i + \epsilon_i $ The “noise” term comprises two variables: $\eta_i$ and $\epsilon_i$. $\eta_i$ is revealed as an additional signal when the firm chooses $i$ as a finalist. However, $\epsilon_i$ remains unpredictable even after the interview. For analytical tractability, we make the following two assumptions.

assp[Two Finalists] Each firm can invite only two finalists; i.e., $K^F = 2$.
assp[Normal Additional Signals] Each additional signal that a finalist reveals follows a normal distribution, $\eta_i \sim \mathcal{N}(0, \sigma_\eta^2)$, i.i.d.
remarkIf $\sigma_\eta = 0$, then the two-stage model is equivalent to the one-stage model that we have considered in the previous sections.

Failure of Laissez-Faire in the Two-Stage Model

This subsection analyzes the performance of laissez-faire in this two-stage setting. The result is analogous to the one-stage case (Theorem (ref)): laissez-faire often falls in perpetual underestimation, and therefore, has linear regret.

First, we define regret. As in the one-stage model, the benchmark is the first-best decision rule, which is the rule firms would apply if the coefficient parameter $\bm{\theta}$ were known. Clearly, the first-best decision rule would greedily invite top-$K^F$ workers in terms of $q_i$ to the final interview. We denote this set of finalists chosen by the first-best decision rule in round $n$ by $\bar{I}^F(n)$. Formally, $\bar{I}^F(n)$ is obtained by solving the following problem:

equation[equation omitted — 160 chars of source]

After that, the first-best decision rule would observe the realization of $\eta_i$ for $i\in \bar{I}^F(n)$, and then hire the worker $i$ who has the greatest skill predictor: $q_i + \eta_i$. Unconstrained two-stage regret (U2S-Reg) is defined as the loss compared with this first-best decision rule. (This type of regret is named “unconstrained” because we later introduce an alternative definition.)

definition[Unconstrained Two-Stage Regret] In the two-stage hiring model, the unconstrained two-stage regret $\text{U2S-Reg}$ of decision rule $\iota$ is defined as follows: \begin{align} U2S-Reg(N) &= \sum_{n=1}^N \left\{\max_{i \in \bar{I}^F(n)} \left(q_i + \eta_i \right) - \left( q_{\iota(n)} + \eta_{\iota(n)} \right)\right\}. \end{align}

Under laissez-faire, the optimal strategy of firm $n$ is to choose candidates greedily based on their estimated skills, i.e.,

equation[equation omitted — 133 chars of source]

After observing the realization of the additional signals $\eta_i$, firm $n$ selects the candidate who has the greatest estimated skill: $\iota(n) = \operatorname*{arg\,max}_{i \in I^F(n)}\left\{\hat{q}_i(n) + \eta_i\right\}$.

Even in the two-stage model, laissez-faire has linear regret when the population ratio is unbalanced.

thm[Failure of Laissez-Faire in the Two-Stage model] Suppose Assumptions (ref), (ref), (ref), (ref), and (ref). Suppose also that $K_2 = 1$ and $d=1$. Let $K_1 - \log_2(K_1 + 1) > \log_2 N$. Then, under the laissez-faire decision rule, group $2$ is perpetually underestimated with the probability $\tilde{\Omega}(1)$. Accordingly, the expected regret associated with the laissez-faire decision rule is \begin{equation} \mathbb{E}\left[U2S-Reg^{\mathrm{LF}}(N)\right] = \tilde{\Omega}(N). \end{equation}

\paragraph*{Proof.} See Appendix (ref).

The proof idea of Theorem (ref) is as follows. Under laissez-faire, each firm $n$ interviews the two finalists with the greatest estimated skills, $\hat{q}_i(n)$. If both finalists belong to the majority group, then minority candidates are never hired, regardless of the $\eta_i$ for each finalist. By evaluating the probability that both finalists are majority candidates, we derive the probability of perpetual underestimation. Thus, even in a two-stage setting, the laissez-faire decision has linear regret under an imbalanced population.

The Rooney Rule and Exploration

Given laissez-faire does not mitigate perpetual underestimation, desirable policy intervention is necessary.

definition[Rooney Rule] In the two-stage hiring model, the Rooney Rule requires each firm $n$ to select at least one finalist from every group $g \in G$; i.e., for every $n$ and every $g\in G$, $I^F(n)$ must satisfy \begin{equation} \left|\left\{i\in I^F(n) \mid g(i) = g\right\}\right| \ge 1. \end{equation}

Under Assumption (ref) and (ref), each firm interviews one majority candidate and one minority candidate. To analyze how the Rooney Rule resolves statistical discrimination, we introduce a weaker notion of regret, constrained two-stage regret.

definition[Constrained Two-Stage Regret] In the two-stage hiring model, the constrained two-stage regret ($\text{C2S-Reg}$) of decision rule $\iota$ is defined as follows: \begin{align} C2S-Reg(N) &= \sum_{n=1}^N \left\{\max_{i \in \breve{I}^F(n)} \left(q_i + \eta_i \right) - \left( q_{\iota(n)} + \eta_{\iota(n)} \right)\right\}, \end{align} where $\breve{I}^F(n)$ is given by \begin{align} &\breve{I}^F(n) = \operatorname*{arg\,max}_{I'\subseteq I(n)} \sum_{i\in I}q_i\\ s.t. &|I'| = K^F, \\ &\forall g\in G, \ \left|\left\{i\in I' \mid g(i) = g\right\}\right| \ge 1. \end{align}

In plain words, $\breve{I}^F(n)$ is the best set of finalists who satisfy the constraint (ref). If Eq. (ref) is imposed as an “exogenous constraint” (rather than a policy), the first-best decision rule would interview $\breve{I}^F(n)$ to maximize social welfare. Constrained regret enables us to identify whether the Rooney Rule prevents perpetual underestimation: if perpetual underestimation occurs under the Rooney Rule, then the constrained regret is linear in $N$.

Under the Rooney Rule, myopic firm $n$ greedily chooses candidates based on estimator $\hat{q}_i(n)$ subject to the following constraints:

align[align omitted — 236 chars of source]

and $\iota(n) = \operatorname*{arg\,max}_{i \in I^F(n)} \left\{\hat{q}_i(n) + \eta_i\right\}$.

The following theorem states that the Rooney Rule resolves underestimation.

thm[Sublinear Constrained Regret under the Rooney Rule] Suppose Assumptions (ref), (ref), (ref), (ref), and (ref). Then, regret under the Rooney Rule is bounded as \begin{equation} \mathbb{E}\left[C2S-Reg^{\mathrm{Rooney}}(N)\right] = \tilde{O}(\sqrt{N}). \end{equation}

\paragraph*{Proof.} See Appendix (ref).

In the proof of Theorem (ref), we show that the factor of (ref) exhibits an exponential dependency\footnote{See definition of $C_6$ in the proof.} on signal variance $\sigma_\eta$, which implies that a sufficiently large $\sigma_\eta$ is required for a reasonable bound.

The Rooney Rule and Exploitation

Although the Rooney Rule prevents statistical discrimination (Theorem (ref)), it may worsen social welfare in terms of the original unconstrained regret. The intuition is as follows. An unbalanced population ratio produces a significant probability that more than one majority candidate is highly skilled. In that case, the true predicted skill of the second-best majority candidate ($q_i$) is likely to be greater than that of the best minority candidate. This feature raises constant regret per round: when $\eta_i$ is normally distributed, any finalist has a positive probability of being hired. Hence, the skill level of all finalists matters, and therefore, firms prefer to interview top-$K^F$ candidates who have the greatest skill. The Rooney Rule prevents this outcome. This effect would present even when firms had perfect information about coefficients $\bm{\theta}$. Consequently, the loss from the constraint (ref) is constant per round, and the Rooney Rule results in $\Omega(N)$ unconstrained regret for $N$ rounds.

thm[Linear Unconstrained Regret under the Rooney Rule] Suppose Assumptions (ref), (ref), (ref), (ref), and (ref). Then, regret under the Rooney Rule is bounded as \begin{equation} \mathbb{E}\left[U2S-Reg^{\mathrm{Rooney}}(N)\right] = \Omega(N). \end{equation}

The proof is straightforward from the argument above, and therefore, is omitted.

Although the laissez-faire and the Rooney Rule have linear unconstrained regret, these two results have different causes for the outcome in each case: laissez-faire produces linear regret due to underexploration, whereas the Rooney Rule produces linear regret due to underexploitation. One way to resolve this is by combining the two. By starting with the Rooney Rule and abolishing it after obtaining sufficiently rich data, we could mitigate the approach's disadvantage. Section (ref) demonstrates the performance of such a mechanism.

Simulation

This section presents the outcomes of our simulations. Unless specified, model parameters are set as $d = 5, \bm{\theta} = (1,1,1,1,1), \bm{\mu}_x = (1.5, \dots, 1.5), \sigma_x = 1$, $\sigma_\epsilon = 0.5$, $\lambda = 1$, and $N= 1,000$. Group sizes are set to be $(K_1,K_2) = (10, 2)$. The initial sample size is $N^{(0)} = K_1 + K_2$, and the sample size for each group is equal to its population ratio: $N^{(0)}_1 = K_1, N^{(0)}_2 = K_2$. We draw $4,000$ paths independently for each simulation scenario. The value of $\delta$ in the confidence bound is set to $0.1$.

The Effects of Population Ratio

figure[figure omitted — 935 chars of source]

We test how the population ratio impacts the frequency of perpetual underestimation. The decision rule is fixed to laissez-faire (LF). We fix the number of minority candidates in each round to two (i.e., $K_2 = 2$) and vary the number of majority candidates ($K_1 = 2, 10, 30, 100$).

Figure (ref) exhibits the simulation result. Consistent with our theoretical analyses, we observe that (i) as indicated by Theorem (ref), laissez-faire rarely produces perpetual underestimation if the population is balanced (i.e., $K_1$ is close to $K_2 = 2$), and (ii) as indicated by Theorem (ref), the larger the population of majority workers (i.e., $K_1$ increases), the more frequently perpetual underestimation occurs. With $K_1 = 10$, perpetual underestimation occurs more than 2% of runs, which is large enough to ensure that laissez-faire produces (approximately) linear regret.

Laissez-Faire vs the UCB Mechanism

figure[figure omitted — 815 chars of source]

Figure (ref) compares the regret associated with the laissez-faire (LF) decision rule and the UCB decision rule. As indicated by Theorem (ref), our simulation shows that laissez-faire has a significant probability of underestimating the minority group. Consequently, laissez-faire sometimes causes perpetual underestimation, and regret grows (approximately) linearly to $n$. Furthermore, due to the possibility of perpetual underestimation, the confidence intervals of the sample paths (denoted by the red area) are very large, indicating the highly uncertain performance of laissez-faire. In contrast, consistent with Theorem (ref), the UCB decision rule performs much more stably. Since the UCB rule avoids underexploration, it does not cause perpetual underestimation.

The UCB Mechanism vs the Hybrid Mechanism

Next, we compare the performance of the UCB and hybrid mechanisms. The parameter of the hybrid mechanism is set to be $a= 0.5$. Figure (ref) shows the associated regret. As Theorems (ref) and (ref) anticipated, the regret associated with the two decision rules are similar (these two decision rules have the same order: $\tilde{O}(\sqrt{N})$). Figure (ref) compares the subsidy rules. As Theorems (ref) and (ref) predicted, the subsidy required for the UCB index rule grows at the rate of $\tilde{O}(\sqrt{N})$, whereas the hybrid index subsidy rule only requires only a constant subsidy, implying that the policy intervention can be terminated at some point. Furthermore, the hybrid index subsidy rule requires a much smaller budget than the UCB index subsidy rule. To summarize, the hybrid mechanism produces similar regret as the UCB mechanism with a much smaller budget.

In Appendix (ref), we demonstrate that while the budget required by UCB is improved substantially if the subsidy rule does not have to be an index policy, whereas its total subsidy cannot be bounded by a constant and requires a large subsidy in the long run.

The Rooney Rule

figure[figure omitted — 959 chars of source]

This subsection compares the performance of the Rooney Rule with that of the laissez-faire decision rule. Figure (ref) depicts the relationship between the frequency of perpetual underestimation and the informativeness of the signal obtained at the second stage (measured by $\sigma_\eta^2$, the variance of $\eta_i$) under both rules. When $\sigma_\eta^2$ is large, the Rooney Rule effectively resolves underestimation.

Figure (ref) compares U2S-Reg associated with each rule. We set $\sigma_\eta = 6$. While both rules produce linear regret, the Rooney Rule suffers from more regret due to underexploitation. This shortcoming can be overcome by using the Rooney Rule as a temporary policy. the “Rooney-LF” decision rule begins with the Rooney Rule and shifts to laissez-faire after $50$ rounds. This approach achieves both less regret and fairer hiring.

Conclusion

We have studied statistical discrimination using a contextual multi-armed bandit model. Our dynamic model articulates how a failure of social learning produces statistical discrimination. In our model, the insufficiency of data about minority groups is endogenously generated. This data shortage prevents firms from accurately estimating the skill of minority candidates. Consequently, firms tend to prefer hiring majority candidates, leading the data sufficiency to persist. This form of statistical discrimination is not only unfair but also inefficient. We have demonstrated that an unbalanced population ratio leads laissez-faire to tend toward perpetual underestimation, an unfair and inefficient consequence.

We analyzed two possible policy interventions. One is subsidy rules that incentivize firms to hire minority candidates. Our hybrid mechanism achieves $\tilde{O}(\sqrt{N})$ regret with $\tilde{O}(1)$ subsidy. Another intervention is the Rooney Rule, which requires firms to interview at least one minority candidate. Our result indicates that terminating the Rooney Rule at an appropriate point would resolve statistical discrimination while maintaining the social welfare level. These results contrast with some of the previous studies Foster1992AnEA,coate1993will,moro_general_2004 demonstrating the possible counterproductivity of affirmative-action policies.

Our analyses of the two interventions provide a consistent policy implication: Affirmative actions effectively resolve statistical discrimination caused by data insufficiency, but such actions should be lifted upon acquiring sufficient information. Accordingly, a temporary affirmative action constitutes the best approach to resolving statistical discrimination as a social learning failure.