EconBase
← Back to paper

A Job I Like or a Job I Can Get: Designing Job Recommender Systems Using Field Experiments

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

128,517 characters · 41 sections · 49 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

A Job I Like or a Job I Can Get: Designing Job Recommender Systems Using Field Experiments

abstractRecommendation systems (RSs) are increasingly used to guide job seekers on online platforms, yet the algorithms currently deployed are typically optimized for predictive objectives such as clicks, applications, or hires, rather than job seekers' welfare. We develop a job-search model with an application stage in which the value of a vacancy depends on two dimensions: the utility it delivers to the worker and the probability that an application succeeds. The model implies that welfare-optimal RSs rank vacancies by an expected-surplus index combining both, and shows why rankings based solely on utility, hiring probabilities, or observed application behavior are generically suboptimal, an instance of the inversion problem between behavior and welfare. We test these predictions and quantify their practical importance through two randomized field experiments conducted with the French public employment service. The first experiment, comparing existing algorithms and their combinations, provides behavioral evidence that both dimensions shape application decisions. Guided by the model and these results, the second experiment extends the comparison to an RS designed to approximate the welfare-optimal ranking. The experiments generate exogenous variation in the vacancies shown to job seekers, allowing us to estimate the model, validate its behavioral predictions, and construct a welfare metric. Algorithms informed by the model-implied optimal ranking substantially outperform existing approaches and perform close to the welfare-optimal benchmark. Our results show that embedding predictive tools within a simple job-search framework and combining it with experimental evidence yields recommendation rules with substantial welfare gains in practice.\\\ JEL Classification: J64, J68, L86, C78, C55, C61 Keywords: Job Recommender Systems, Matching, Experiments, Machine Learning.

\newgeometry{top=1in, bottom=1in, left=1in, right=1in}

\onehalfspacing \doparttoc \faketableofcontents

Introduction

Recommendation systems (RSs) are transforming how job seekers interact with online labor-market platforms. Their adoption is accelerating, and many public employment services (PES) are considering integrating such tools organisation2023artificial. Yet designs vary widely, and evidence on their relative effectiveness or their ability to generate meaningful labor-market improvements at scale remains limited, even if they may benefit specific subgroups. This paper develops a framework providing a welfare criterion for comparing designs, and uses two field experiments to show that algorithms informed by this criterion substantially outperform existing approaches.

A large data-science literature develops job RSs by training algorithms on historical data to predict user behavior such as clicks, applications, or hires, and evaluating performance using predictive metrics freire2021recruitment, de2021job, mashayekhi2022challenge. While highly effective at forecasting observed behavior, their normative interpretation is less clear: predicting observed interactions does not imply that recommended opportunities maximize job seekers' expected welfare.

At the same time, a growing empirical literature in economics evaluates job recommendations in the field, often through randomized experiments conducted in partnership with PESs in several countries. Many of these interventions provide occupational or employer-level recommendations designed to broaden search or redirect job seekers toward labor markets with better prospects belot2019providing, altmann2022direct, behaghel2024potential, belot2025advising, bachli2025helping. These studies provide valuable causal evidence, but they typically evaluate a small number of pre-specified recommendation rules against a limited set of outcomes, and offer little guidance on how to compare such rules against a common welfare criterion.

These two literatures have largely developed in parallel, though recent contributions connect them. roland2022 evaluate a collaborative filtering RS based on clicking behavior; SuBayoumiJoachims2022 consider social welfare maximization. Other work questions the practice of treating observed behavior as a proxy for welfare agan2023automating, kleinberg2022challenge, kleinberg2024inversion, mullainathan2025economics. kleinberg2024inversion formalize the inversion problem: when behavior is generated under frictions, predicting behavior is not equivalent to recovering the latent objective guiding optimal decisions.

This paper develops a job-search framework to discipline the design of RSs. The framework incorporates the application stage and highlights two objects: the utility a job seeker associates with a vacancy and the probability that an application results in a hire. The welfare-optimal rule ranks vacancies by an expected-surplus index combining both. A central implication is that neither utility or hiring outcomes alone provide a sufficient basis for optimal recommendations: the optimal algorithm requires a specific combination of utility and hiring probability that reflects the expected gains from applying. This is an instance of the inversion problem: recommendation rules that optimize predictive objectives such as predicted utility, observed application behavior, or observed hiring outcomes alone will generally miss it, even in a simple and frictionless job-search environment. An important empirical question is whether this inversion problem is economically significant in practice, and whether welfare-oriented rankings deliver substantial improvements over robust predictive rules in realistic environments.

We bring this framework to the data through close collaboration with the French PES. Starting from two existing RSs, one reflecting stated preferences, another a state-of-the-art ML system vadoreijcai predicting hiring outcomes, we conduct a sequence of two randomized field experiments. These experiments are not designed to evaluate the large-scale causal impact of deploying recommendations per se. Rather, they are conceived as beta-tests in a learning cycle athey2018impact, aimed at comparing alternative algorithmic designs and iteratively improving them.

A first experiment compares recommendations based on existing algorithms separately and in combination. We find that job seekers respond more favorably to recommendations combining information about preferences and hiring prospects, suggesting that neither dimension alone is sufficient to shape search behavior.

Guided by these findings and the model's structure, we design a new class of algorithms that better approximate the model-implied optimal ranking by incorporating explicit proxies for application behavior and preference-related signals into hiring predictions. This yields a family of recommendation rules including the original algorithms, an application-based algorithm, and an approximation of the welfare-optimal algorithm that combines information on preferences, applications, and hiring.

We evaluate these alternative designs in a second randomized experiment. The newly introduced algorithms, particularly the application-based algorithm and the approximation of the welfare-optimal rule, substantially outperform the initial approaches, especially in terms of clicks and applications. While they remain distinct from the theoretically optimal rule, their strong empirical performance highlights the practical gains from incorporating application behavior and preference-related signals into recommendation design.

A central goal of the paper is to use the conceptual structure of the job-search model to formally compare alternative RSs. Because optimality is defined within a behavioral model and a welfare criterion, this comparison requires that the model’s implications for application behavior be empirically plausible. We therefore exploit the exogenous variation generated by the random assignment of recommendation algorithms to estimate a structural model of application behavior. The results validate the model's core behavioral predictions: both the utility score and the inverse hiring probability are highly significant predictors of application decisions, with coefficients stable across specifications. This supports using the model as a conceptual basis for welfare comparisons.

The experimental data allow us to identify the predictive components needed to reconstruct the model-implied optimal recommendation rule. Using observed outcomes under different algorithms, we estimate predictions for both applications and hiring, and combine them to construct an estimate of the optimal recommendation score. This score is then used to identify the vacancies that would have been recommended under the optimal algorithm and to rank the tested algorithms against a common welfare benchmark. The application-based algorithm and our approximation of the welfare-optimal rule strongly outperform the initial algorithms. The latter further improves upon the application-based algorithm, albeit by a smaller margin, and performs remarkably close to the model-implied optimum.

Taken together, our results point to an important implication: RSs that target predictive objectives such as clicks, applications, or hiring outcomes are generally not aligned with job seekers' welfare. At the same time, the predictive tasks underlying these systems provide essential building blocks for welfare-relevant recommendation design. By combining structural modeling with experimental evidence that disciplines its behavioral assumptions, our approach illustrates how predictive tools can be used to construct, evaluate, and iteratively improve recommendation rules.

Our paper speaks to a growing empirical literature in economics that studies job RSs and related forms of automated advice, often through randomized field experiments conducted in partnership with public employment services belot2019providing, altmann2022direct, behaghel2024potential, belot2025advising, bachli2025helping, roland2022. A central feature of this literature is the diversity of objectives implicitly targeted by recommendation rules. Some interventions rely on observed transitions or predicted hiring probabilities to steer job seekers toward occupations or employers with higher employment prospects belot2019providing, altmann2022direct, behaghel2024potential, belot2025advising; others emphasize measures of fit or expressed interest inferred from search behavior or skills profiles bachli2025helping, roland2022. While these approaches provide valuable causal evidence on the effects of specific recommendation designs, they do not offer a general framework to compare alternative algorithmic objectives or relate them to a common welfare criterion.

We contribute to this literature by providing such a framework and comparing different RSs using experiments. Within a simple job-search model that explicitly incorporates the application stage, we show that two dimensions are central to recommendation design: the utility that a job seeker associates with a vacancy and the probability that an application results in a hire. Existing approaches can be interpreted as emphasizing one of these dimensions in isolation, but neither is sufficient on its own. An economically meaningful ranking must combine both into a single expected-surplus index, which allows us to place diverse recommendation designs on a common footing and to evaluate their relative performance.

Our analysis also relates to a growing literature at the intersection of machine learning and economics that questions the normative interpretation of observed behavior on digital platforms agan2023automating, kleinberg2022challenge, kleinberg2024inversion, mullainathan2025economics. In our setting, the inversion problem arises because application decisions only reveal whether applying is privately profitable, hiring outcomes capture only part of the expected gains, whereas welfare-relevant ranking of vacancies depends on the magnitude of expected gains.

Our contribution is to make this insight operational in a labor-market environment. By combining experimental variation with a structural model of application behavior, we characterize job seekers' welfare-relevant objective, distinguish it from commonly used behavior-based rankings, and quantify the importance of the inversion problem. While the model implies that welfare-relevant rankings differ from application-based recommendations, our results show that the gap is positive but quantitatively modest, an assessment that would be difficult to obtain without jointly leveraging experimental evidence and structural modeling.

Our methodology relates closely to work emphasizing experimentation and economic modeling in the design of algorithmic decision rules athey2018impact. Rather than evaluating the large-scale causal impact of a fixed RS, we use sequential beta-tests to compare designs, feed the results back into the model, and construct improved algorithms. The structural model disciplines which signals to incorporate and how to combine them, while the experiments in turn discipline the model’s behavioral assumptions and provide the variation needed to estimate the welfare metric. Empirically, RSs informed by both utility and hiring probabilities substantially outperform approaches based on either alone, yielding sizeable gains relative to algorithms currently used in practice, while the inversion problem has modest quantitative implications in this setting.

The paper proceeds as follows. Section (ref) presents the job-search model and derives the optimal recommendation rule. Section (ref) describes two representative RSs and their underlying scores. Section (ref) presents the design and results of the randomized experiments. Section (ref) estimates the structural model and compares alternative algorithms using the proposed metric. Section (ref) concludes with implications for the design of job RSs in practice.

Job search model with recommender systems

\paragraph{Overview.} We develop a job-search model in which job seekers encounter vacancies sequentially and decide whether to apply. Vacancies are lotteries characterized by (i) the utility they would deliver to the worker and (ii) the probability that an application succeeds. We then introduce a RS as a technology that (a) restricts the set of vacancies a job seeker is exposed to based on a score and (b) potentially increases the rate at which vacancies can be processed. The model delivers (i) a model-implied value index for vacancies, (ii) conditions under which an RS improves welfare, and (iii) an optimal ranking rule in the benchmark case of myopic job seekers. Finally, we discuss how these results inform RS design and clarify why ranking by utility alone, by hiring probability alone, or by predicted application behavior is generally suboptimal.

Environment and primitives

Our job search model builds on the following environment and primitives. \paragraph{Job seekers and vacancies.} Job seekers are indexed by characteristics $x$ and vacancies by characteristics $y$. An unemployed job seeker receives flow utility $u(b)$. \paragraph{Payoffs.} A vacancy $y$ yields utility $U(x,y)+\varepsilon_{i,y}$ to job seeker $i$ of type $x$, where $\varepsilon$ is an idiosyncratic taste shock observed by the job seeker. We assume $\varepsilon$ follows a logistic distribution with scale parameter $\sigma$ and cumulative distribution function $F_{\varepsilon}(\cdot)=F(\cdot/\sigma)$.

\paragraph{Hiring probability.} Conditional on applying, the job seeker is hired with probability $p(x,y)$, which is known to the job seeker at the application decision stage. \paragraph{Vacancy distribution and reparametrization.} Vacancies are drawn from $F_0(y)$. For a given type $x$, the induced distribution of $(p(x,y),U(x,y))$ is denoted $F_0(p,U)$.\footnote{with $x$ suppressed in the notation.} \paragraph{Other primitives.} Applications entail a cost $k$ and rejection entails a psychological cost $R$. Matches separate at rate $q$ and future utility is discounted at rate $r$.

Baseline sequential search without an RS

This section characterizes job seekers’ behavior and the value of unemployment in the baseline sequential search environment without an RS. The key friction is that job seekers explore vacancies sequentially, drawing from the distribution $F_0(p,U)$ at rate $\alpha_0$, and cannot simultaneously compare all available opportunities. Upon encountering a vacancy, the job seeker observes its characteristics $(p,U,\varepsilon)$ and decides whether to apply. All formal derivations and proofs are relegated to Appendix (ref).

Reservation utility and surplus

Let $V_0(x)$ denote the discounted value of unemployment for a job seeker of type $x$ in the absence of a RS. Each time the job seeker encounters a vacancy with characteristics $(p,U,\varepsilon)$, she faces a binary decision: apply to it or continue searching.

If the job seeker does not apply, she remains unemployed and retains continuation value $V_0(x)$. If she applies, she pays the application cost $k$; with probability $p$ the application succeeds and she transitions into employment with discounted present value of utility $V_e(x,U+\varepsilon)$, while with probability $1-p$ she is rejected and continues searching, incurring the rejection cost $R$. Conditional on applying, the discounted expected payoff is therefore \[ p\,V_e(x,U+\varepsilon) + (1-p)\big(V_0(x)-R\big) - k. \] To characterize this apply-or-not decision, it is useful to introduce a reservation utility:

equation[equation omitted — 106 chars of source]

where $\overline{R}=(r+q)R$ and $\overline{k}=(r+q)k$. The quantity $U_0^*(x,p)$ represents the minimum utility level that makes a vacancy with hiring probability $p$ worth applying to.

We then define the surplus associated with a vacancy as

equation[equation omitted — 70 chars of source]

This surplus compares the utility provided by the vacancy to the reservation utility associated with its hiring probability. Vacancies with lower hiring probabilities must offer higher utility in order to be attractive, reflecting the costs associated with unsuccessful applications. To streamline notation, we henceforth suppress the dependence on $x$ whenever this does not create ambiguity.

Application behavior and vacancy value

The job seeker applies to a vacancy whenever the realized surplus is positive:

equation[equation omitted — 58 chars of source]

This rule admits a natural interpretation: a job seeker applies if the realized utility $U+\varepsilon$ exceeds the reservation utility $U_0^*(p)$.

Under the assumed logistic distribution for $\varepsilon$, this decision rule implies that the probability of applying to a vacancy $(p,U)$ is $ p_a(p,U) = F\!\left(\Delta(p,U)/\sigma\right)$.

The discounted value of unemployment satisfies

equation[equation omitted — 114 chars of source]

where the expectation is taken over the distribution of vacancies and

equation[equation omitted — 180 chars of source]

The index $\Gamma(p,U)$ represents the expected contribution of encountering a vacancy with characteristics $(p,U)$ to the job seeker’s continuation value. It aggregates the probability of applying, the probability of being hired conditional on application, and the surplus generated by a successful match conditional on application.

Under the logistic assumption on $\varepsilon$, $\Gamma(p,U)$ admits the closed-form expression

equation[equation omitted — 105 chars of source]
prop[Application rule and vacancy value] In the absence of a RS, application behavior is governed by the surplus $\Delta(p,U)$ through the rule (ref), while the welfare-relevant value of a vacancy is summarized by the index $\Gamma(p,U)$.
proofSee Appendix (ref).

Proposition (ref) makes explicit the distinction between application behavior and vacancy value. This non-equivalence arises because application decisions are governed by a binary profitability condition, whereas vacancy values depend on the expected magnitude of the gains. A vacancy with a high hiring probability $p$ contributes more to the job seeker's continuation value than one with low $p$, even if both exceed the application threshold. Rankings based on observed applications therefore do not generally coincide with rankings that maximize job seekers' welfare. This distinction is central to the design of RSs, and Section (ref) characterizes it precisely.

Recommender systems: selection and exposure

A recommender system (RS) scores the pool of available vacancies and pre-selects a subset to be shown to the job seeker, potentially also increasing the rate at which vacancies can be processed. While decisions remain sequential on the worker side, the RS transforms sequential search over vacancies into sequential applications over a curated set.

In our framework, RSs affect job search through two channels:

enumerate[label=(\roman*)] • Selection. The RS restricts the job seeker’s consideration set to vacancies whose score exceeds a threshold, i.e., the top $s$ fraction according to a score. • Exposure. The RS may change the effective rate at which vacancies can be processed, from $\alpha_0$ to $\alpha_1$.

Scores and the induced distribution of considered vacancies

For a job seeker with characteristics $x$ and a vacancy described by characteristics $y$, the RS computes a score $S(x,y)$. To keep the notation light, we suppress the dependence on $x$ and $y$ whenever it is unambiguous.

For a given $x$, the joint distribution of vacancy characteristics and scores is the distribution of $\big(p(x,y),\,U(x,y),\,S(x,y)\big)$, when $y \sim F_0$, which we denote by $F_0(p,U,S)$.\footnote{As in Section (ref), this notation suppresses the dependence on $x$. Formally, $F_0(p,U,S)$ is the distribution of $(p(x,y),U(x,y),S(x,y))$ induced by $F_0(y)$ conditional on $x$.}

We model RS selection as recommending the vacancies whose score lies above the $(1-s)$-quantile of the score distribution. Let $\overline{q}_S(s)$ denote the quantile of order $1-s$ of $S$ under $F_0(p,U,S)$. Then the RS induces the truncated distribution

equation[equation omitted — 114 chars of source]

which is the distribution of $(p,U,S)$ among recommended vacancies.

Exposure: arrival/processing rate

In addition to changing the composition of vacancies considered, RSs may also change the intensity of exposure to vacancies. We capture this by allowing the effective vacancy-processing rate under recommendations to be $\alpha_1$ rather than $\alpha_0$. This reduced-form formulation accommodates multiple mechanisms: recommendations may lower cognitive and search costs, facilitate navigation, or simply deliver a fixed number of suggestions over a given period.

Myopic benchmark: value under recommendations

We first consider the case of myopic job seekers, who do not adjust their reservation utility in response to the introduction of the RS. They continue to evaluate vacancies using the baseline reservation utility $U_0^*(p)$ derived in Section (ref). Accordingly, $\Gamma^m(p,U)$ is defined by the same expression as $\Gamma(p,U)$, but evaluated at the baseline continuation value $V_0$. In particular, $\Gamma^m(p,U)=\Gamma(p,U)$ in the absence of an RS.

Under an RS characterized by $(S,s,\alpha_1)$, the discounted value for a myopic unemployed job seeker is

equation[equation omitted — 179 chars of source]

where the expectation is taken with respect to $F_0(p,U,S)$.

Optimal RS and improvement conditions

Objective and definition of an optimal RS

We now define what it means for a RS to be optimal. Since the RS is designed and implemented by the PES, we must specify the objective it pursues. Throughout the paper, we assume that the PES aims to maximize the welfare of job seekers. We abstract from firms’ outcomes and from broader social objectives.

RSs rank vacancies using a score $S$ and recommend the top fraction $s$ of vacancies according to that score. An optimal RS is therefore one that, for any recommendation intensity $s$, maximizes the expected value of the recommended vacancies.

definitionFor myopic job seekers, an optimal RS is a measurable scoring rule $S^*$ of $(p,U)$ that solves, for each $s\in[0,1]$, \[ S^* \in \operatorname*{\arg\!\max}_{S } \; \mathbb{E}\!\left[ \Gamma^{m}(p,U)\; {\rm {\large 1}\hspace{-2.3pt}{\large l}}\big\{ S(p,U)>\overline{q}_{S}(s) \big\} \right]. \]

The requirement that optimality holds for all $s$ emphasizes that the RS defines a global ranking, rather than being tailored to a specific cutoff or number of recommendations.

Characterization and sufficient conditions

Characterizing the optimal RS in the myopic case is straightforward and yields sharp implications for practice.

propConsider a RS based on a score $S(p,U)$ that selects the top fraction $s$ of vacancies. \begin{enumerate} • The optimal ranking is obtained by $S(p,U)=\Gamma^m(p,U)$. • If $\alpha_1\geq\alpha_0$, a sufficient condition for an $S$-based RS to improve job seekers’ welfare relative to baseline search is that $z\mapsto \mathbb{E}\!\left(\Gamma^m(p,U)\mid S=z\right)$ is increasing. In particular, the RS based on $S(p,U)=\Gamma^m(p,U)$ strictly improves welfare whenever $\Gamma^m$ is non-degenerate. \end{enumerate}
proofSee Appendix (ref).

\paragraph{Interpretation.} Proposition (ref) establishes that the optimal RS solves a global ranking problem over the entire distribution of vacancies, whereas application behavior reflects a binary profitability condition. Together with Proposition (ref), this creates an inversion problem: algorithms that target observed applications or hiring outcomes are not generally aligned with job seekers' welfare-relevant ranking. Section (ref) characterizes the nature and magnitude of this gap for several natural recommendation rules.

Implications for the design of recommender systems

This section studies the gap between the optimal RS characterized in Proposition (ref) and alternative recommendation rules that may appear reasonable a priori. We consider rankings based on perceived utility $U$, on the probability of hire $p$, on observed application behavior $p_a$, and on observed hires. The latter are of special interest, since applications are the observable outcome of job seekers’ decisions and hiring is a key outcome of the search process, making it natural to ask whether recommendation rules that replicate observed applications or hires can be normatively justified.

Building on the model and welfare criterion introduced above, we compare these alternative rules to the welfare-relevant objective. The key result of this section is a unifying decomposition of the welfare-relevant score, from which the non-optimality of several natural recommendation rules follows directly.

Decomposing the welfare-relevant score

In the myopic case, the welfare-relevant index $\Gamma(p,U)$ admits the decomposition:

equation[equation omitted — 171 chars of source]

This expression makes clear that the optimal ranking is the product of three distinct components: (i) the probability of being hired conditional on applying ($p$), (ii) the probability of applying ($p_a(p,U)$), and (iii) the expected surplus generated by a successful application, conditional on applying.

Since $p_a(p,U)=\mathbb{P}(\Delta(p,U)+\varepsilon>0 \mid p, U)$ is a monotone function of $\Delta$ under mild conditions, the conditional expectation in (ref) can be expressed as a function of the application probability. Let $m(p_a(p,U)) := \mathbb{E}\!\left[\Delta(p,U)+\varepsilon \mid \Delta(p,U)+\varepsilon>0, p, U\right]$, which allows us to rewrite the optimal score as $ \Gamma(p,U) = p \times p_a(p,U) \times m(p_a(p,U))$.

This decomposition is generic and does not rely on a specific parametric assumption on the distribution of the idiosyncratic shock $\varepsilon$. Under the logistic assumption adopted in the baseline model, $\Gamma(p,U)$ admits the closed-form expression given in Equation (ref), which can be written as

equation[equation omitted — 129 chars of source]

The last term, $m(p_a)=-\log(1-p_a)/p_a$, is a convex and increasing function of $p_a$, with $\lim_{p_a\to 0}m(p_a)=1$ and $\lim_{p_a\to 1}m(p_a)=+\infty$. Thus, for small $p_a$, a second-order Taylor expansion gives $m(p_a)\approx 1+p_a/2$, so that $\Gamma(p,U)\approx p_h(p,U)\times(1+p_a(p,U)/2)$, where $p_h$ is the unconditional probability of being hired,

equation[equation omitted — 63 chars of source]

i.e., the product of the hiring probability conditional on application and the application probability. When application probabilities are empirically small, $\Gamma$ is therefore well approximated by $p_h$, which rationalizes the strong empirical performance of hiring-based rankings documented in Section (ref).

More generally, one can show that if the distribution of $\varepsilon$ is log-concave, then the conditional surplus $m(p_a)$ is an increasing function of the application probability. The precise shape of this function depends on the distributional assumption. Appendix Figure (ref) illustrates this additional term for three standard cases: logistic, Gumbel (EV1), and normal distributions, all normalized to have unit variance.

The decomposition (ref) immediately implies that neither $U$, $p$, $p_a$, nor $p_h$ alone is sufficient to recover the optimal ranking: each captures only a subset of the dimensions jointly determining $\Gamma(p,U)$.

Why ranking by application probability is not optimal

Ranking vacancies according to observed application behavior is a natural benchmark, since applications are directly observed and summarize job seekers’ choices. Within the model, however, the Bellman equation makes clear that application behavior and welfare-relevant vacancy values rely on fundamentally different objects.

When a job seeker applies to a vacancy $(p,U)$, the contribution of this vacancy to her continuation value is proportional to $p\big(U - U_0^*(x,p) + \varepsilon\big)$, where $p$ is the probability of being hired conditional on application. By contrast, the application decision itself is governed by a threshold rule: a job seeker applies whenever this expression is positive.

As a result, application behavior only reveals whether applying is privately profitable, but abstracts from the probability $p$ that scales the contribution of the vacancy to expected welfare. Two vacancies may generate the same surplus $U - U_0^*(p) + \varepsilon$ and therefore induce the same application decision, yet differ substantially in their welfare contribution because they are associated with different hiring probabilities $p$.

This wedge follows directly from the dynamic structure of the search problem and provides a structural explanation for the inversion problem emphasized in the algorithmic fairness literature: a score that reproduces application behavior generally fails to recover the welfare-relevant ranking of vacancies.

Why ranking by hiring probability is not optimal

A closely related benchmark is to rank vacancies according to the probability of hire $p_h(p,U)$ (see Equation (ref)). Using the decomposition in (ref), the optimal score can be written as $$\Gamma(p,U) = p_h(p,U) \times m\!\big(p_a(p,U)\big).$$ This expression shows that $p_h$ is still an incomplete measure of vacancy value. Under log-concavity of the distribution of $\varepsilon$, the function $m(p_a)$ is increasing, implying that $p_h$ does not place sufficient weight on perceived utility and application behavior. Ranking by $p_h$ underweights vacancies that generate large conditional surpluses.

That said, for standard distributions such as the logistic or Gumbel (EV1), the function $m(p_a)$ varies relatively slowly when $p_a$ is small. As a result, in environments where application probabilities are low, rankings based on $p_h$ may perform reasonably well in practice, even though they are not theoretically optimal.

Two implementable routes to welfare-oriented recommendation

Proposition (ref) characterizes the welfare-optimal ranking through the score $\Gamma(p,U)$. In practice, two routes can be followed to implement such a ranking.

\paragraph{Route A: structural implementation.} A first approach is to recover the primitives entering $\Gamma(p,U)$ by separately estimating a utility component $U_{i,j}$ and the probability of recruitment $p_{i,j}$, and then computing the welfare score implied by the model. While conceptually direct, this route requires specifying how observed platform signals map into the latent utility object $U_{i,j}$.

\paragraph{Route B: reduced-form welfare index.} An alternative is to construct the welfare score directly from observable transition probabilities. Let \[ p_{a,i,j} = \mathbb{P}(C_{i,j}=1 \mid x_i,y_j), \qquad p_{i,j} = \mathbb{P}(H_{i,j}=1 \mid C_{i,j}=1, x_i,y_j) \] denote the probability of applying and the probability of recruitment conditional on application. Under the logistic specification discussed above, the welfare score can be written as (ref)-(ref), which yields $\Gamma_{i,j} = p_{i,j}\,[-\log(1-p_{a,i,j})]$, up to a positive scale normalization.

This second route is directly implementable since it relies only on predicting applications and hires. The empirical strategy developed below follows this approach: we estimate $\widehat p_{a,i,j}$ and $\widehat p_{i,j}$ using the available signals, construct $\widehat{\Gamma}_{i,j}$, and use it to form welfare-oriented recommendation sets.

Implementing such a welfare-oriented ranking relies on the behavioral structure linking applications to both the surplus component $U_{i,j}$ and the probability of recruitment $p_{i,j}$. Establishing the empirical relevance of this structure is therefore a natural prerequisite for algorithm design.\footnote{While the experiments reported below provide evidence consistent with this behavioral structure, they do not identify the exact distribution of taste shocks.} Random assignment of recommendation algorithms generates exogenous variation in the signals observed by job seekers. This variation allows us to identify how application decisions respond to both utility-related signals and recruitment probabilities.

Identification of the objective $\Gamma$ from application data

Identification of $\Delta(p,U)$ and reconstruction of $\Gamma(p,U)$

Under the logit taste-shock specification, application decisions identify the surplus index $\Delta(p,U)$ that governs job seekers’ application behavior. As Equation (ref) shows, the welfare-relevant vacancy value $\Gamma(p,U)$ is a function of this surplus index. Given identification of $\Delta(p,U)$, the structural assumptions of the model allow us to reconstruct the welfare-relevant objective $\Gamma(p,U)$ and thus the optimal ranking. Identification of $\Gamma(p,U)$ in this framework therefore primarily hinges on the ability to identify $\Delta(p,U)$ from application data.

We denote by $C_{i,j}=1$ if job seeker $i$ applies to vacancy $j$, and $0$ otherwise. Let $\tilde U_{i,j} \equiv U_{i,j}-U_{0,i}^*(1)$ denote utility net of the continuation value of unemployment.

A sufficient parametric identification result

We provide a sufficient parametric identification result under a logistic specification.

propLet $\varepsilon$ be distributed as a logistic random variable with scale parameter $\sigma$. Assume that sequences of individual $i$'s application decisions on vacancies $j$ are observed, together with $p_{i,j}$ and $\tilde U_{i,j}$. Then, the parameters $\alpha$,\footnote{This $\alpha$ is not to be confused with the vacancy-arrival rates $\alpha_0, \alpha_1$.} $\beta$, and $\gamma$ in the binary choice model \begin{equation} \mathbb{P}(C_{i,j}=1 \mid p_{i,j},\tilde U_{i,j}) = F\!\left( \alpha \tilde U_{i,j} -\beta/p_{i,j} +\gamma \right) \end{equation} are identified. Moreover, for generic values $(p,U)$, the structural objects $\sigma$, $\overline{k}+\overline{R}$, $\Delta(p,U)$, and $\Gamma(p,U)$ in equations (ref) and (ref) are identified as follows: \begin{itemize}[label=--] • $1/\alpha$ identifies $\sigma$; • $\beta/\alpha $ identifies $\overline{k}+\overline{R}$; • $U-(\beta/\alpha)/p+\gamma/\alpha$ identifies $\Delta(p,U)$; • $p\log\!\left(1+e^{\alpha U+\gamma-\beta/p}\right)/\alpha$ identifies $\Gamma(p,U)$. \end{itemize}
proofSee Appendix (ref).

In Proposition (ref), the assumption that $\tilde U_{i,j}$ is directly observed can be relaxed by allowing for an individual-specific error term $\nu_i$, so that utilities take the form $\tilde U_{i,j}+\nu_i$. In this case, identification of $\sigma$, $\overline{k}+\overline{R}$, $\Delta(p,U)$, and $\Gamma(p,U)$ follows from a fixed-effects panel logit specification adapted from Equation (ref).

Extensions and scope

This section discusses extensions of the baseline framework and clarifies the scope of the analysis. Our objective is not to provide a full treatment of these extensions in the main text, but rather to explain how the core insights extend beyond the baseline model and to indicate where additional complexities arise.

Forward-looking job seekers

The analysis in the main text focuses on myopic job seekers. In Online Appendix (ref), we extend the framework to allow for forward-looking behavior. We address two issues.

First, we examine how to evaluate the gains associated with the use of a RS when job seekers are not myopic. We show that, using observable data, it is possible to adjust the welfare evaluation derived under the myopia assumption to account for the endogenous adjustment of the reservation utility implied by forward-looking behavior.

Second, we consider the implications for RS design. While the optimal rule under myopia is no longer exactly optimal in this setting, recommendation sets constructed under this assumption can still be improved. Although deriving a fully optimal rule with forward-looking job seekers is challenging, the RS remains effective in steering job seekers toward vacancies with higher expected value.

Imperfect knowledge about the distribution of opportunities

The baseline model abstracts from potential misperceptions about the distribution of available vacancies. In practice, job seekers may hold biased beliefs about the types of opportunities they are likely to encounter and form expectations accordingly.

To fix ideas, consider the case of myopic job seekers. Absent recommendations, job seekers base their search decisions on a subjective distribution of vacancies, which gives rise to a continuation value denoted $V_0^{JS}$. By contrast, the true distribution of opportunities would imply a value $V_0^{True}$. Introducing a RS provides access to vacancies drawn from the true distribution and yields a continuation value $V_1^m$, without modifying the model-implied optimal ranking rule $\Gamma^m$ for recommended vacancies.

The overall effect of the RS can therefore be decomposed as: \[ \underbrace{V_1^m - V_0^{JS}}_{\text{Full RS effect}} = \underbrace{V_1^m - V_0^{True}}_{\text{Pure RS effect}} + \underbrace{V_0^{True} - V_0^{JS}}_{\text{Information effect}}. \]

This decomposition highlights that, in addition to alleviating search frictions, RSs may generate value by correcting job seekers’ misperceptions about the distribution of opportunities.\footnote{Formally, subjective beliefs can be represented by assuming that job seekers draw vacancies from a subjective distribution $dF_0^{JS}$, while recommended vacancies are drawn from the true distribution $dF_0^{True}$. The relation between the two can be expressed, for example, as $dF_0^{True} = \phi(p,U)\, dF_0^{JS}$, where $\phi(\cdot)$ captures systematic belief distortions. Under this formulation, recommendations shift job seekers from draws based on $dF_0^{JS}$ to draws based on $dF_0^{True}$, without altering the model-implied optimal ranking. For forward-looking job seekers, the same logic applies, although the interaction with the endogenous reservation utility makes the analysis more involved without altering the qualitative insight.}

Competition and congestion

The model abstracts from competition among job seekers and from congestion effects. At the recommendation stage, vacancies are ranked independently for each job seeker, so the RS effectively solves a collection of individual optimization problems.

In environments with many job seekers, this approach may lead multiple individuals to be recommended the same vacancy, potentially exacerbating congestion and altering effective hiring probabilities. Accounting for such interactions would require formulating a global assignment problem that imposes constraints on how often a vacancy can be recommended.

From an algorithmic perspective, this can be implemented as a post-processing step on top of a proximity matrix between job seekers and vacancies, for instance using optimal transport methods that explicitly trade off match quality and congestion bied2021congestion,mashayekhi2023recon. In this paper, we deliberately abstract from congestion and focus on recommendation rules that ignore these interactions.

Two representative job RSs

Section (ref) shows that welfare-optimal recommendations must combine two primitives: the utility $U(x,y)$ a vacancy delivers to a job seeker and the probability $p(x,y)$ that an application succeeds. This section presents the two RSs at the core of the paper's experimental design, chosen precisely because each emphasizes one of these two components. The first, a state-of-the-art ML algorithm trained on realized hires, primarily targets $p$; the second, a knowledge-based matching algorithm derived from job seekers' stated search criteria, primarily targets $U$. Neither is welfare-optimal in isolation (Proposition (ref)), but they provide the building blocks for the welfare-approximating algorithms constructed in Section (ref). Before describing these two systems, we briefly situate them within the broader RS landscape.

Many forms of RSs

All RSs operate according to a common principle: they rely on the computation of a matching score that summarizes information about the expected value of a job seeker–vacancy match. Specifically, for an individual $i$ described by a set of characteristics $x_i$ and a vacancy $j$ described by characteristics $y_j$, the system computes a score $S_{i,j}$ depending on $x_i$ and $y_j$. Higher values of $S_{i,j}$ indicate a stronger match and are therefore preferred. Once these scores are computed, generating recommendations is straightforward. For a given job seeker $i_0$, vacancies are ranked according to their scores $S_{i_0,j}$ from the most to the least desirable. To make $k$ recommendations to $i_0$, an intuitive solution is to pick the $k$ vacancies with the highest scores.\footnote{Throughout the paper, we define the rank of a vacancy as its position in this ordering, with rank $1$ corresponding to the highest score.}

Although this underlying principle is shared across systems, job recommendation approaches vary widely in both the computer science literature and real-world applications. As surveyed by freire2021recruitment, de2021job, mashayekhi2022challenge, this diversity reflects a multitude of application contexts, data availability, and algorithmic strategies. Importantly, these approaches also differ in the type of information they exploit and in the outcomes they are designed to predict. Table (ref) summarizes the main families of approaches and highlights their key characteristics.

Knowledge-based RSs leverage expert ontologies of occupations, skills, and locations to match workers to vacancies based on assessed fit. A prominent example is WCC ELISE, used by several PESs and private entities.\footnote{See the dedicated websites \href{https://www.roberthalf.com/us/en/find-jobs}{\textcolor{blue}{Robert Half}} and \href{https://www.wcc-group.com/employment/products/elise-job-matching-search-and-match/}{\textcolor{blue}{ELISE}}.} An alternative class of approaches leverages machine learning techniques. Collaborative filtering relies on interaction histories to infer similarity patterns, as in the click-based algorithm studied by roland2022 at the Swedish PES. By contrast, content-based RSs exploit observable characteristics (e.g. occupation, education and skills) to predict interaction probabilities; CareerBuilder provides an example zhao2021embedding.

Hybrid RSs combine these approaches. Examples include the RecSys 2017 winner Volkovs2017 or LinkedIn's system, which predicts matches based on user and vacancy characteristics, incorporating individual and recruiter-level fixed effects when sufficient interaction data are available shi2022generalized. Finally, Indeed's RS ma2022jobs combines collaborative filtering and content-based methods, with a final hybrid stage involving deep learning and a rule-based engine.

Beyond algorithmic design, an important dimension of heterogeneity across RSs concerns the choice of target variable. Table (ref) summarizes selected contributions from data science and economics, as well as the algorithms studied in this paper. The table highlights substantial variation in outcome measures, ranging from clicks and applications to hires. Notably, data science contributions predominantly focus on intermediate outcomes such as clicks or applications, whereas economic studies typically emphasize hires. Our paper implements an experimental comparison of RSs with different designs within the same setting, evaluated against a common welfare metric derived from the model. This contrasts with the existing literature, where studies typically evaluate one or two pre-specified algorithms targeting different outcomes in different populations, making cross-study comparisons difficult.

table[table omitted — 1,825 chars of source]

A state-of-the-art RS based on hiring predictions

We present the baseline ML-based RS we initially developed bied2021congestion,vadoreijcai. This is a state-of-the-art content-based RS building on the insights of the winning algorithm of the RecSys 2017 challenge Volkovs2017. In the framework developed in Section (ref), this algorithm can be interpreted as primarily targeting the probability of successful matching $p(x,y)$. It is trained on data from job seekers who found employment, using the vacancy that led to a hire as the positive example for each job seeker, and a set of vacancies that did not result in a hire as negative examples. The training objective seeks to rank the realized match above the alternatives for each job seeker, it does not directly maximize reemployment rates, but rather learns a score that separates successful from unsuccessful job seeker, vacancy pairs.

Model performance is naturally evaluated using recall@$k$: the share of job seekers $i$ hired a given week for whom the algorithm ranks the realized vacancy among the top $k$ recommendations available that week, where $k$ is usually 10, 20, 50, or 100. While the recall@$k$ provides a meaningful evaluation metric, it is intractable for direct optimization. We therefore follow the learning to rank literature, and learn a similarity score $S_{i,j}$ between job seeker $i$ and vacancy $j$. The objective is to ensure that, for any job seeker $i$, the score associated with the realized match $j^*(i)$ exceeds that of any alternative vacancy $j'$. This leads to minimizing the triplet margin loss corresponding to the following objective:

equation[equation omitted — 113 chars of source]

where $\eta > 0$ is a scalar hyperparameter, $[x]_+ = \max (x, 0)$, the outer sum ranges over all job seekers with matches, the inner one over all vacancies weinberger2009distance. The objective enforces a separation of at least $\eta$ between the scores of matched and unmatched vacancies for each job seeker.

Given job seeker and vacancies characteristics, respectively denoted by $X_i$ and $Y_j$, the score $S_{ij}$ is defined as: \[ S_{i,j}(X_{i,j}) = \phi(X_i)^{\top} A \psi(Y_j), \] where $X_{i,j} = (X_i,Y_j)$, $\phi, \psi$ are feed-forward neural networks with several layers, and $A$ is an affinity matrix. Feed-forward neural networks provide flexible, differentiable representations well suited to high-dimensional inputs and large datasets.\footnote{The interested reader may consult goodfellow2016deeplearningbook for a textbook treatment.}

In this context, $\phi(X_i)$ and $\psi(Y_j)$ can be interpreted as latent representations of job seeker $i$ and vacancy $j$. The matrix $A$ captures cross-dimensional affinities: the parameter $A_{k,l}$ measures the complementarity between dimension $k$ of the job seekers' latent space and dimension $l$ of the vacancy’s latent space. Both latent spaces have dimension 872. Table (ref) lists the observable characteristics used on the job seeker and vacancy sides to predict hires.

Importantly, $\phi$, $\psi$ and $A$ are given a structure which incorporates three main blocks corresponding to geography, skills, and all remaining features. This design explicitly incorporates key dimensions emphasized in the job recommendation and labor market literature, most notably location and (soft) skills belot2019providing,belot2025advising,altmann2022direct,bachli2025helping, while leveraging the power of ML methods to detect the most promising interactions and transitions.

Formally (see Figure (ref)), the similarity score decomposes as \[ S_{i,j}(X_{i,j}) = \sum_{b \in \{``geography", ``skills" , ``other features" \} }\phi_{b}(X_{i})^{\top} A_b \psi_b(Y_{j}), \] where $\phi_b$, $\psi_b$ and $A_b$ are block-specific transformations and affinity matrices.

The parameters that are optimized during training include the neural network weights defining the mappings to the latent spaces, as well as the affinity matrices $A_b$. The resulting non-convex objective is minimized using mini-batch stochastic gradient descent. For computational tractability, non-matching job seeker–vacancy pairs are heavily and uniformly subsampled. Finally, for a given job seeker $i_0$, we define the $\mathcal{P}$-ranking as the ordering of vacancies induced by the score $S_{i_0,j}$.

\paragraph{Data used to train algorithms.}

Three features of the data are particularly relevant for this study. First, administrative records on job seekers and vacancies are matched to behavioral data (clicks, applications, and hiring outcomes) on the same platform, providing a joint view of both sides of the market and the interactions between them. Second, we observe click data alongside application data, allowing us to distinguish early-stage interest from actual application decisions; the use of click data to measure job seekers' expressed interest at this early stage of online search is novel in this context. Third, the scale of the data, over 1.1 million job-seeker search sessions, 516,776 unique vacancies, and 75,744 observed hires, enables reliable training of high-dimensional ML models and estimation of the structural application model.

We use rich historical administrative data from the Public Employment Service (PES) to train and evaluate several job recommendation systems (RSs). Table (ref) summarizes the type of data used in this paper. This includes descriptions of vacancies posted on the PES, job seekers' characteristics and search parameters, as well as user interactions on the PES website, such as clicks, applications, and subsequent hiring outcomes. Our analysis focuses on the former French region of Rhône-Alpes, which offers substantial economic and geographic diversity while remaining sufficiently contained for detailed empirical analysis.

The PES website is open to all employers and job seekers and constitutes the largest platform for vacancy postings in the French labor market.\footnote{Using the same application data, le2021gender estimate that vacancies posted on this website represented 60% of all vacancies in 2010 (see their Section VI.A.)} The data provide extensive information on job postings, including publication date, the occupation at several levels of granularity, the posted wage, required experience; contract type (permanent, temporary, or fixed-term); workplace location; weekly working hours; and required qualifications. Importantly, the data contain information on desired hard and soft skills, textual descriptions of both the vacancy and the firm, firm size, the number of applications received by the vacancy and by the firm over the previous six months, and the time elapsed since the vacancy was posted.

Information on job seekers is drawn from administrative records on unemployment spells (for example, the fichier historique, FH, of the PES). These records include demographic characteristics and detailed job search histories, such as date of registration, geographic location, experience, skills, unemployment duration, applications in the last six months. They also contain various individual and postal code level socio-demographic characteristics. Importantly, we observe job search parameters declared at registration, including reservation wage, maximum commuting time, desired occupation, desired type of contract (temporary vs. long-term), and working hours (full-time vs. part-time).

This comprehensive information on both sides of the market is complemented by detailed data on user behavior on the PES website, notably clicks on vacancies and subsequent applications. While application data from the PES have been used in previous work marinescu2021unemployment,glover2019job,algan2020active, the use of click data to measure job seekers’ interest at the early stages of online search appears to be novel in this context. Applications are observed through three of the PES channels: applications submitted directly by job seekers, potential matches initiated by firms, and applications suggested by PES caseworkers. Finally, we also exploit the final outcome of these interactions, whether a hire occurs, which is recorded by caseworkers.\footnote{As noted in algan2020active, these hiring data are subject to measurement error, notably because hires occurring outside the PES may not be observed. To mitigate this limitation, we complement the PES records with comprehensive administrative data on all hires (Déclarations préalable à l'embauche) whenever a hire can be linked to an identifiable PES posted vacancy.}

An RS based on search criteria

The French PES relies on a matching algorithm based on WCC Elise to recommend relevant vacancies to job seekers. Like most knowledge-based RSs, this algorithm starts from a comparison between characteristics of the job desired by job seekers and those of the available vacancies and can be interpreted, in the framework developed in Section (ref), as primarily capturing the utility component of a job seeker–vacancy match. For each characteristic considered, a sub-score is determined, ranging from zero to one, reflecting the degree of compatibility between the job seeker’s preferences or profile and the vacancy’s requirements. The sub-scores are then aggregated to form a global score. Aggregation is primarily based on a weighted average.

For the purpose of this study, we construct a RS inspired by the PES algorithm and based on the same set of criteria. Each criterion is matched exactly with its counterpart on the recruiter’s side (i.e., the requirements specified in the vacancy and the characteristics of the job offered). For each characteristic $k$, we define a consistency measure $c_{k,i,j}\in [0,1]$, which captures the degree to which characteristic $k$ of job seeker $i$ is compatible with that of vacancy $j$.

The characteristics entering the definition of the score and their associated weights are as follows: Occupation (0.332), Skills in occupation (0.332), Geographic mobility (0.1), Reservation wage (0.066), Diploma (0.033), Working hours (0.033), Driving license (0.033), Languages (0.033), Years of experience in occupation (0.033), Duration and type of contract (0.003). The resulting matching score is defined as:\footnote{Throughout the paper, we use the formula from Equation (ref) together with the weights used by the PES. The score $\mathcal{U}^{PES}$ actually used by the PES and the score $\mathcal{U}$-rec \ used in our experiments are not identical. The exact criteria used at the PES share the same principles but allow for smoother definitions of several sub-criteria and incorporates additional nonlinearities, such as censoring based on geographic fit. We abstract from these features here in order to preserve transparency and interpretability.}

equation[equation omitted — 80 chars of source]

We refer to the ordering of vacancies for a given job seeker $i_0$ according to the criterion $\mathcal{U}_{i_0,j}$ as the $\mathcal{U}$ ranking. As for the $\mathcal{P}$-based RS, recommendations are the vacancies ranked highest (i.e., with the largest score) under this ordering.

table[table omitted — 2,069 chars of source]

Understanding the two RSs

The $\mathcal{U}$ score as a utility signal. Our previous analysis shows that any recommendation algorithm score can be interpreted as a combination of two components: a signal about utility, $U_{i,j}-U_{0,i}^*(1)$ and a signal about the probability of recruitment, $p_{i,j}$. Consequently, the scores associated with the different RSs generally mix information about both dimensions. We interpret the $\mathcal{U}$ score primarily as a signal of surplus utility $U_{i,j}-U_{0,i}^*(1)$. Salary is the job attribute most directly linked to utility, but the other job characteristics included in the definition of $\mathcal{U}$ such as occupation, required skills, geographic location, working hours, and contract type also plausibly affect job seekers’ utility. Similarly, requirements related to diplomas, experience, driving licenses, and languages describe the set of skills that a job seeker would deploy in the position. Skill matching is not only relevant for productivity but is also associated with job satisfaction, personal fulfillment, and the accumulation of human capital, all of which are valued by job seekers. The weights used to aggregate these characteristics are those set by the PES. Nevertheless, as shown in Appendix (ref), these weights can alternatively be estimated from the data at our disposal using application behavior.

The $\mathcal{P}$ score as a match probability signal. Analogously, we interpret the ranking induced by the $\mathcal{P}$ score as primarily reflecting the probability of recruitment, $p_{i,j}$. Appendix (ref) provides empirical support for this interpretation. Following the approach of chernozhukov2018generic, we use the predictive content of the ML algorithm's score to build a generic best logistic predictor of the matching probability. Specifically, we exploit the history of sequential job applications to estimate a model linking the probability of a successful application to the matching score between job seeker $i$ and the vacancies $j(i)$ to which they applied. The results in Table (ref) strongly validate the association between the score $S_{i,j}$ and recruitment outcomes. This procedure also allows us to map the score $S_{i,j}$ into an estimated probability of success, $p_{i,j}$. As a result, the rankings induced by the $S$-score and $\mathcal{P}$-score, which we call $\mathcal{P}$-ranking, can be interpreted as rankings based on recruitment probabilities.\footnote{This approach provides us with an estimate of $p_{i,j}$ for each application–vacancy pair. This quantity is typically unobserved, but is revealed here through the machine-learning-based estimation.}

Two different scores. Figure (ref) compares the sets of top-ranked vacancies recommended to a given job seeker under the $\mathcal{U}$ and $\mathcal{P}$ rankings. The overlap between the vacancies ranked highest according to each criterion is very limited. On average, the vacancy that maximizes $\mathcal{P}$ is ranked 2{,}027th in the $\mathcal{U}$ ordering, while the vacancy that maximizes $\mathcal{U}$ is ranked 4{,}403rd in the $\mathcal{P}$ ordering. These large rank reversals highlight that the two scores emphasize markedly different dimensions of job seeker–vacancy matches.

Appendix (ref) further documents these differences by comparing the distributions of $\mathcal{U}$ and $\mathcal{P}$ among the top-ranked vacancies under each criterion (see Figure (ref)). For example, the median probability of success of a vacancy recommended under the $\mathcal{U}$ ranking is 0.02, compared with 0.06 for a vacancy recommended under the $\mathcal{P}$ ranking.

Taken together, these results support the intuitive interpretation that the $\mathcal{U}$-based RS primarily captures the utility dimension, while the $\mathcal{P}$-based RS primarily captures the probability of successful matching. This distinction is useful for expositional clarity. However, neither score was originally designed to isolate a single dimension. There is no guarantee that $\mathcal{U}_{i,j}$ perfectly identifies $U_{i,j}-U_{0,i}^*(1)$ and that $\mathcal{P}_{i,j}$ perfectly identifies $p_{i,j}$. More plausibly, both scores are functions of the two underlying components, such that $\mathcal{U}_{i,j}=\mathcal{U}(U_{i,j}-U_{0,i}^*,p_{i,j})$ and $\mathcal{P}_{i,j}=\mathcal{P}(U_{i,j}-U_{0,i}^*,p_{i,j})$. Even if $\mathcal{U}$ predominantly reflects utility and $\mathcal{P}$ predominantly reflects recruitment probabilities, each score is likely to contain information about both dimensions. The framework of Section (ref) identifies the theoretically correct combination, and the experiments of Section (ref) are designed to test whether algorithms that move in this direction, by enriching hiring-based predictions with preference-related signals, deliver welfare gains in practice. Importantly, $\mathcal{U}$ and $\mathcal{P}$ were developed independently of this interpretation and were not designed to be combined; the model provides the basis for doing so.

\FloatBarrier

Using field experiments to design recommender systems

This section reports two randomized field experiments that are distinctive in three respects. First, six algorithms are compared head-to-head in the same setting, using the same population and the same outcome measures, a direct experimental comparison of this breadth does not exist in the prior literature. Second, the set of algorithms spans the full spectrum from welfare-misaligned (pure $\mathcal{P}$ or pure $\mathcal{U}$) to welfare-approximating ({\sc Vadore}.2), so the experiments trace the welfare gains from progressively better-aligned designs. Third, the design of the second experiment is derived from the theoretical predictions of the model and guided by the results of the first experiment, as part of an iterative learning cycle.

The first field experiment (beta-test 1) tests whether job seekers respond more favorably to recommendations that combine hiring-based and preference-based rankings than to recommendations based on either dimension alone, the core prediction of Proposition (ref). Section (ref) and Appendix (ref) establish that the $\mathcal{P}$ and $\mathcal{U}$ rankings differ substantially, so this comparison has empirical bite. The results of this experiment motivate a redesign of the algorithm that more directly incorporates preference information while maintaining hiring prediction as the primary objective.

The second field experiment (beta-test 2) evaluates a new family of algorithms designed in response to these results, including an application-based algorithm (Application) and an enriched hiring-prediction algorithm ({\sc Vadore}.2) that incorporates both preference-related and application-related signals. Application plays a dual role: it is a theoretically meaningful benchmark closely related to job seekers’ decision rules, and a building block used to improve hiring predictions in {\sc Vadore}.2.

The two field experiments we conducted follow the same protocol, which is described in detail in Appendix (ref). Table (ref) summarizes these experiments. The eligible population consists of job seekers registered at France Travail\ in the Auvergne-Rhône-Alpes region who were actively seeking employment. In each experiment, a single email was sent to a randomly selected sample of job seekers among this population, respectively 102,314 for the March 2022 experiment and 150,000 for the June 2023 one, providing access to a list of job recommendations. Job seekers who clicked the consent link and viewed the list were enrolled, resulting in 18,947 and 30,973 participants, respectively. There is no control group; instead, participants were randomly assigned to recommendation lists generated by different algorithms. The first experiment focuses on combining the two generic algorithms presented in Section (ref). The second experiment emphasizes new algorithms developed based on insights from the first experiment. In both experiments, we analyze the $\mathcal{U}$ and $\mathcal{P}$ scores of recommended vacancies, as well as clicks, applications, and hires.

Beta-test 1: testing ordinal combinations of $\mathcal{U}$ and $\mathcal{P}$

Experimental design and treatments

The experiment randomized two dimensions of the intervention: (i) the algorithm used to generate job recommendations, and (ii) the display of additional information. Job seekers were randomly assigned to one of ten treatment arms, corresponding to five algorithmic variants crossed with two display conditions.

\paragraph{Recommendation algorithms}

All recommendations are drawn from a consideration set, a pool of vacancies that includes job vacancies highly ranked by at least one of the two base algorithms ({\sc Vadore}.0\ or $\mathcal{U}$-rec). This ensures that the recommended vacancies are relevant according to either match likelihood or stated preferences.

Each job seeker was randomly assigned to one of five algorithms, which differed in how they selected and ranked job vacancies from within this consideration set:\footnote{Independently of the recommendation algorithm, job seekers were also randomly assigned to different information display conditions. Some participants were shown additional performance indicators (star ratings summarizing preference match and predicted hiring probability) for the first two recommended vacancies. In the main analysis, we pool all display conditions. Appendix Table (ref) shows that accounting for this variation does not affect our conclusions.}

itemize• {\sc Vadore}.0 (ML-based recommendations): The RS ranking vacancies using the $\mathcal{P}$ score (detailed in Section (ref)), designed to predict successful matches. • $\mathcal{U}$-rec \ (Preference-based recommendations): The RS using the $\mathcal{U}$ score (see Section (ref)), which ranks vacancies based on the job seeker's stated preferences. On top of this baseline score, it also includes a final censoring step based on the geographic fit (see footnote (ref)). • Mix\ algorithms (ordinal combinations of {\sc Vadore}.0 and $\mathcal{U}$-rec): Three hybrid variants combine the rankings from {\sc Vadore}.0\ and $\mathcal{U}$-\textsc{rec}, relying solely on ordinal information. The steps are as follows (see Appendix (ref) for details): \begin{enumerate} • Rank vacancies in the consideration set according to {\sc Vadore}.0 ($\mathcal{P}$); • Filter vacancies in the consideration set: \begin{itemize} • \textbf{\textsc{Mix}-$\sfrac{1}{4}$:} Retains the top 25% according to $\mathcal{P}$ ranking; • \textbf{\textsc{Mix}-$\sfrac{1}{2}$:} Retains the top 50% according to $\mathcal{P}$ ranking; • \textbf{\textsc{Mix}-$\sfrac{3}{4}$:} Retains the top 75% according to $\mathcal{P}$ ranking; \end{itemize} • Re-rank the filtered vacancies using the $\mathcal{U}$ ranking; • Select the 10 first vacancies in this ranking. \end{enumerate} We refer to these three variants collectively as the \textit{\textsc{Mix}} group.

Clearly, the corresponding algorithms, ordered as {\sc Vadore}.0, Mix-$\sfrac{1}{4}$, Mix-$\sfrac{1}{2}$, Mix-$\sfrac{3}{4}$, and $\mathcal{U}$-rec, progressively shift the weight from $\mathcal{P}$ to $\mathcal{U}$ ranking.

\paragraph{Survey protocol and data description.}

Each of the 102,314 job seekers randomly selected for invitation to the experiment was first assigned to one of ten randomization groups. They received an email containing a link to a Qualtrics survey in which job ads were listed. Of the 102,314 individuals invited, 100,879 successfully received the email (the remainder were affected by technical issues), and 18,947 (18.6%) opened the survey, thereby enrolling in the experiment. Each job seeker was shown the top 10 ads corresponding to their profile according to the algorithm to which they were assigned.

Reduced-form estimates

We estimate the following specification at the job seeker--vacancy pair level using ordinary least squares (OLS):

equation[equation omitted — 127 chars of source]

where \( Y_{ij} \) denotes one of the following outcomes for job seeker \( i \) and vacancy \( j \): the hiring score (\( \mathcal{P} \)), the matching score (\( \mathcal{U} \)), whether the job seeker clicked or applied to the vacancy, or the subjective rating given to the vacancy. The indicator \( G_{a,i} \) denotes assignment of individual \( i \) to algorithm \( a \in \mathcal{A}_1 = \{\text{\textsc{Vadore}.0}, \text{\textsc{Mix}}-\sfrac{1}{4}, \text{\textsc{Mix}}-\sfrac{1}{2}, \text{\textsc{Mix}}-\sfrac{3}{4}, \text{$\mathcal{U}$-\textsc{rec}}\} \). The vector \( Z_i \) includes a set of indicators for the position of the vacancy in the displayed list (slots 1 to 10). Standard errors are clustered at the job seeker level.

Figure (ref) (see also Table (ref), taking $\mathcal{U}$-rec\ as reference) presents the results using data from the 18,947 job seekers randomly assigned to receive job recommendations from one of the five algorithms.\footnote{Preregistered outcomes for this beta-test are: ratings, clicks and applications on recommended vacancies. We also preregistered broader labor market outcomes related to job search but do not use them here.}

Figures (ref)-(a) and (ref)-(b) use the full set of 10 job recommendations provided to each participant (for a total of 189,470 observations) to examine how the assigned algorithm affects the distribution of vacancies by predicted hiring probability and matching score. As intended by the experimental design, assignment to algorithms with higher weight on {\sc Vadore}.0\ results in higher average hiring scores (\( \mathcal{P} \)) but lower matching scores (\( \mathcal{U} \)). Specifically, compared to vacancies recommended by $\mathcal{U}$-rec, recommendations from {\sc Vadore}.0\ have an average hiring score that is 0.046 points higher (against a baseline mean of 0.054), representing nearly a doubling in expected hiring probability. All estimated coefficients are statistically significant, and we observe a large jump in hiring scores between Mix-$\sfrac{3}{4}$\ and Mix-$\sfrac{1}{2}$. The three algorithms with greater weight on {\sc Vadore}.0\ ({\sc Vadore}.0, Mix-$\sfrac{1}{4}$, and Mix-$\sfrac{1}{2}$) yield similar outcomes on this dimension.

Conversely, matching scores (\( \mathcal{U} \)) decline markedly as the weight on {\sc Vadore}.0\ increases. Vacancies recommended by {\sc Vadore}.0\ have an average matching score 0.19 points lower than those recommended by $\mathcal{U}$-rec\ (baseline: 0.773). The pattern of decreasing \( \mathcal{U} \) scores mirrors the increase in \( \mathcal{P} \) scores, and again, the first three algorithms ({\sc Vadore}.0, Mix-$\sfrac{1}{4}$, Mix-$\sfrac{1}{2}$) yield relatively close estimates.

Figures (ref)-(c) uses data on subjective ratings provided by job seekers for the first two recommended vacancies (36,668 observations). These ratings range from 0 to 10 and were elicited via the question: “Overall, what rating out of 10 would you give this job vacancy?”. In Table (ref), all treatment coefficients are positive and statistically significant, indicating that vacancies recommended by algorithms in \( \mathcal{A}_1\setminus\{\text{$\mathcal{U}$-\textsc{rec}}\} \) are rated more favorably than those recommended by $\mathcal{U}$-rec. Importantly, the highest coefficient is not associated with {\sc Vadore}.0, but—as anticipated from the model presented in the previous section, with hybrid algorithms that combine the rankings of $\mathcal{U}$-rec\ and {\sc Vadore}.0. The Mix-$\sfrac{1}{2}$\ algorithm yields the highest average rating.

Figures (ref)-(d) and (ref)-(e) present results for clicks and applications using again the full set of ten recommendations per job seeker. Both outcomes are relatively rare, particularly applications. The lowest rates are observed under $\mathcal{U}$-rec: 4.2% for clicks and 0.45% for applications. Click behavior follows a pattern consistent with the subjective ratings: algorithms ({\sc Vadore}.0, Mix-$\sfrac{1}{4}$, Mix-$\sfrac{1}{2}$) yield relatively close and higher estimates, the difference with $\mathcal{U}$-rec \ being positive and statistically significant. The Mix-$\sfrac{1}{2}$\ algorithm again produces the largest increase, with a click rate 0.64 percentage points higher than $\mathcal{U}$-rec, an increase of approximately 15%. Application rates are also higher under {\sc Vadore}.0\ and \textsc{Mix}-$\sfrac{1}{2}$, with the former effect being significant at the 10% level and representing roughly 16% increases relative to the $\mathcal{U}$-\textsc{rec}\ baseline. Since we only observed three hires based on the recommendations made in this first experiment, the results regarding the effects on hiring are not significant and are therefore not reported.

figure[figure omitted — 2,215 chars of source]

Designing new algorithms based on the experimental results

The key finding of beta-test 1 is that the hybrid algorithm Mix-$\sfrac{1}{2}$, which combines both $\mathcal{P}$ and $\mathcal{U}$ rankings, outperforms the two pure strategies on ratings and clicks. This is the empirical counterpart of Proposition (ref): since $\Delta(p,U)$ and $\Gamma(p,U)$ are not ordinally equivalent, neither utility-based nor hiring-based rankings alone suffice for welfare-optimal recommendations. These results therefore motivate the exploration of additional recommendation principles consistent with the model.

Building on this insight, we consider two complementary directions for extending the initial recommendation scores. First, the model highlights application behavior as a key behavioral object. As shown in Section (ref), application decisions identify the surplus index $\Delta(p,U)$ that governs job seekers’ choices. A recommendation score based on predicted applications therefore constitutes a theoretically meaningful benchmark, even though it does not coincide with the welfare-relevant score $\Gamma$. We accordingly train a RS that predicts applications rather than hirings, which we denote Application. Beyond its role as a benchmark, this score provides an empirical proxy for perceived job utility, which is not directly observed.

Second, the model implies that welfare-relevant rankings depend not only on hiring probabilities but also on job seekers’ preferences. To operationalize this dimension within a hiring-based RS, we enrich the original hiring-prediction algorithm by incorporating the preference components $c_{k,i,j}$ used in the $\mathcal{U}$ score (see Equation (ref)) as additional predictors. This modification yields an intermediate version of the algorithm, denoted {\sc Vadore}.1. We then combine these two extensions by introducing the application-based score as an additional input into the hiring-prediction architecture used for {\sc Vadore}.1. This results in the final algorithm, {\sc Vadore}.2, which augments hiring predictions with information on both job seekers’ preferences and application behavior. Figure (ref) schematically illustrates how these components are integrated within the algorithm.

Importantly, the algorithms evaluated in the second field experiment also allow us to assess empirically a natural approximation of the welfare-relevant score characterized in Section (ref). As shown in the model, the optimal score can be written as $\Gamma(p,U)=p_h(p,U)\times m\!\big(p_a(p,U)\big)$, where the function $m(\cdot)$ is increasing in the application probability under fairly general conditions. While $p_a(p,U)$ is not directly observed, both application-based predictions and preference-based scores provide empirical proxies that are positively related to job seekers’ propensity to apply.

From this perspective, combining a hiring-based score with additional information on job utility constitutes a conceptually grounded way of enriching $p_h(p,U)$ in the direction suggested by the model. In particular, mixing the hiring score produced by {\sc Vadore}.2\ with the $\mathcal{U}$ score amounts to testing whether reinforcing the utility content of a hiring-based RS improves performance, in line with the monotonicity properties implied by the model. Although alternative combinations, such as directly mixing $p_h$ with predicted application probabilities, would also be consistent with this logic, the mixtures considered here allow us to assess whether incorporating job utility into hiring-based recommendations moves the algorithm in the direction predicted by the theory.

Beta-test 2: evaluating the model-guided algorithm family

Experimental design and treatments

The second field experiment is designed to evaluate the relative performance of the new algorithms developed in response to the results of the first beta-test. Its objective is twofold: first, to compare these newly designed algorithms against one another; second, to benchmark them against the two reference algorithms based on hiring predictions ($\mathcal{P}$) and stated preferences ($\mathcal{U}$) that motivated the initial analysis.

Among the six algorithms tested in this second experiment, two were already included in the March 2022 study and serve as benchmarks: $\mathcal{U}$-rec\ and {\sc Vadore}.0.\footnote{Although their architecture is unchanged, both algorithms were retrained using more recent data to reflect labor market conditions prevailing at the time of the June 2023 experiment.} They are complemented by four additional algorithms that reflect different ways of incorporating preference-related information into hiring-based recommendations. These include the enhanced hiring-based algorithm {\sc Vadore}.2, which integrates the design improvements described above; the Application algorithm, which predicts application behavior and serves both as a theoretically meaningful benchmark and as an input into {\sc Vadore}.2; and a hybrid ranking that combines {\sc Vadore}.2\ and $\mathcal{U}$-rec, denoted $\text{\textsc{Mix}}\sfrac{1}{2}(\text{\textsc{Vadore}.2})$, constructed in the same spirit as the mixtures tested in the first experiment. Finally, we include XGBoost, a benchmark algorithm based on gradient boosting that mirrors the objective of {\sc Vadore}.0\ (predicting hires) but replaces neural networks with a standard gradient-boosting architecture. Its purpose is to disentangle the effect of the recommendation objective from the effect of the architecture: since {\sc Vadore}.0\ and \textsc{XGBoost} target the same outcome but differ in their ML method, any performance difference between them reflects architectural choices rather than the welfare alignment of the objective. Conversely, differences between \textsc{XGBoost} and the welfare-enriched algorithms ({\sc Vadore}.2, \textsc{Application}) reflect objective alignment rather than architecture.

The full set of algorithms evaluated in this second experiment is therefore:\footnote{ In addition to algorithmic variation, the experiment also randomized the type of information displayed alongside recommendations for a subset of algorithms (including {\sc Vadore}.2, Application, and $\text{\textsc{Mix}}\sfrac{1}{2}(\text{\textsc{Vadore}.2})$), resulting in multiple display conditions. As in the first experiment, our baseline analysis abstracts from this dimension and aggregates all display variants for a given algorithm. Appendix Table (ref) shows that accounting explicitly for display variation does not affect our main conclusions. } $$\mathcal{A}_2 = \{ \text{\textsc{Vadore}.0},\ \text{\textsc{Vadore}.2},\ \text{\textsc{Mix}}\sfrac{1}{2}(\text{\textsc{Vadore}.2}),\ \text{\textsc{Application}},\ \text{\textsc{XGBoost}},\ \mathcal{U}\text{-}\textsc{rec} \}.$$ The experiment closely replicated the design of the 2022 study. As in the earlier experiment, job seekers were invited by email and enrollment was conditional on clicking a consent button. The survey templates were kept very similar to those previously used. Conducted in June 2023, the experiment invited 150,000 job seekers, of whom 30,973 were enrolled.

The reduced-form analysis follows the same methodology as in the first experiment, as specified in Equation (ref). The results for {\sc Vadore}.0\ and $\mathcal{U}$-rec, presented in Figure (ref) (see also Table (ref), taking the group receiving recommendations from the $\mathcal{U}$-rec \ as reference) are consistent with those observed in the previous experiment.\footnote{Preregistered outcomes for this beta-test are: ratings, clicks, applications and hirings on recommended vacancies.} {\sc Vadore}.0\ outperforms $\mathcal{U}$-rec\ in terms of hiring probability (\( \mathcal{P} \)), but, as expected by construction, underperforms in terms of the adequacy score (\( \mathcal{U} \)). It performs marginally better than $\mathcal{U}$-rec\ in terms of clicks, but not in terms of applications.

The most notable changes are observed in click-through rates and application rates. Relative to the benchmark algorithms $\mathcal{U}$-rec \ and {\sc Vadore}.0, the gains are considerable. Application rates for the {\sc Vadore}.2 and Application recommendations are approximately twice as high as those for $\mathcal{U}$-rec. The improvements are less dramatic for Mix-$\sfrac{1}{2}$\ and XGBoost, but still meaningful.

Figure (ref)-(f) (see also column (6) of Table (ref)) reports hiring outcomes on recommended vacancies. The baseline hiring rate in the reference group is very low: 0.42 \textpertenthousand. We detect a modest signal for the new Mix-$\sfrac{1}{2}$({\sc Vadore}.2) algorithm, significant at the 10% level: although the absolute hiring rate remains small, it is roughly three times higher than in the control group $\mathcal{U}$-rec. Except this group, no statistically significant differences are observed. This is not particularly surprising: despite their differences, all algorithms yield low application rates—below 1%—and, although the success rates of applications vary substantially across algorithms, they are capped around 7%.

To inform expectations about the effects of scaling these interventions, the final column of Table (ref) reports hiring rates conditional on application, that is, the ratio of hires to applications among recommended vacancies. Mix-$\sfrac{1}{2}$({\sc Vadore}.2) and XGBoost nearly double this efficiency relative to the other algorithms, including Application.

These low hiring rates on the recommendations are not unexpected and are consistent with the model. Since application probabilities are small (below 1% per recommendation), the joint hiring probability $p_h = p \times p_a$ is extremely low. Detecting welfare differences through hiring rates alone would require samples several orders of magnitude larger than a beta-test; the intermediate outcomes such as click-through rates, application rates, and the hiring rate conditional on application, are therefore the relevant margins for comparing recommendation principles at this scale. Accordingly, the primary objective of these experiments is not to estimate employment effects at scale, but to compare recommendation principles and generate the variation needed to estimate the structural model in Section (ref).

figure[figure omitted — 2,426 chars of source]

Estimation of the search model

This section uses the experimental variation generated by the random assignment of recommendation algorithms to serve two purposes. First, we estimate a structural model of application behavior and test whether the core behavioral predictions of the model, in particular that both utility and hiring probabilities shape application decisions, are consistent with the data. Second, we use the estimated model to construct the welfare metric $\Gamma(p,U)$ from Proposition (ref) and compare all tested recommendation rules against this common benchmark.

Preferences estimation using application behavior

The model in Section (ref) characterizes how job seekers’ application decisions depend on both the perceived utility of vacancies and their probability of success. In this section, we use observations on job seekers' applications to estimate these preferences. As in hitsch2010matching\footnote{See for example their Equation (9).} or le2021gender, given the threshold based decision rule in Equation (ref), these preferences can be estimated using a discrete choice model.

We primarily rely on the experimental data introduced in Section (ref), but also provide estimates based on observational data in the Appendix as a robustness check. In the experiments, for each job seeker in the sample, we observe:

itemize• The group corresponding to the algorithm used to generate the 10 recommendations: either one of the mixtures used in Experiment 1, or one of the six algorithms used in Experiment 2. We denote the associated group dummy variables by $T_i$; • The list of 10 vacancies selected by the corresponding algorithm; • The scores $\mathcal{U}_{i,j}$ and success probabilities $\mathcal{P}_{i,j}$; • The clicks and applications for each of the 10 vacancies.

We closely follow Equation (ref) in Proposition (ref). We estimate this structural model of application behavior by instrumenting $\mathcal{U}_{i,j}$ and $1/\mathcal{P}_{i,j}$ using the assignment variables $T_i$. Our binary choice model takes the following general form:

align[align omitted — 196 chars of source]

where $B_{2,i}\in\{0,1\}$ is an indicator for participation in Experiment 2, and the vector $W_{i,j}=(\mathcal{U}_{i,j},\, 1/\mathcal{P}_{i,j},\, B_{2,i})$ collects the utility score, the inverse hiring probability, and this experiment indicator. This specification is in the spirit of hitsch2010matching,chen2023reducing, who estimate similar equations in the context of marriage markets. However, we leverage here the randomization performed in our experiments. Note that the term $-\beta/\mathcal{P}_{i,j}$ enters with a negative sign in (ref), so $\beta>0$ implies a positive effect of $\mathcal{P}$ on the application probability, as expected. The individual effect $c_i$ captures heterogeneity across job seekers. As discussed in Section (ref), this effect captures systematic individual specific differences between the observed index $\mathcal{U}$ and the actual utility gain relative to a reservation value.

The functional form used in (ref) corresponds to a logit model. However, logit models with fixed effects are not easily compatible with the transparent identification structure provided by random assignment. We would also like to account for potential measurement error in $W_{i,j}$: following the standard model $W_{i,j} = W^*_{i,j} + e_{i,j}$, where $W^*_{i,j}$ is the true latent regressor and $e_{i,j}$ is error, the instruments $T_i$ allow consistent estimation under classical measurement error assumptions.

Given the low application probability (3\textperthousand), we adopt the approximation $\Lambda(x) \approx \exp(x)$ and estimate a Poisson IV model. Under this specification, the conditional expectation becomes:

equation[equation omitted — 181 chars of source]

where, from Equation (ref), $\theta = (\alpha, -\beta, \delta)$. The control function approach (see wooldridge2010econometric) can be used to address endogeneity.\footnote{In a nutshell, the control function method in this context works as follows. The potentially endogenous regressors $W_{i,j}$ are linked to instruments $T_i$ through a first-stage equation:

equation[equation omitted — 59 chars of source]

and the structural error term $\mu_{i,j}$ is modeled as depending on $v_{i,j}$ but not on $T_i$ (exclusion restriction):

equation[equation omitted — 73 chars of source]

with $\nu_{i,j}$ independent of $v_{i,j}$. Substituting (ref) into (ref) and integrating over the distribution of $\nu_{i,j}$ yields the control function moment condition:

equation[equation omitted — 139 chars of source]

In practice, the first-stage model (ref) is estimated, from which the residuals $\hat{v}_{i,j}$ are computed, and substituted into Equation (ref). Notably, this approach also provides an estimate of $\rho$, which captures the correlation between the structural error term $\mu_{i,j}$ and the endogenous variables $W_{i,j}$.}

An alternative approach is to use a first-order Taylor expansion of $\Lambda(x)$ around the sample mean $\overline{x}$: $\Lambda(x) \approx \Lambda(\overline{x}) + \Lambda'(\overline{x})(x - \overline{x}) = \tilde{\Lambda} + \tilde{\Lambda}' x$. This leads to a linearized model of the form:

equation[equation omitted — 239 chars of source]

This equation can be estimated either by instrumental variables or using the control function approach described above. In the linear context both methods yield exactly the same results. Note that, due to the linear approximation, coefficients are identified up to a scaling factor, which is not problematic for our purposes. We are primarily interested in the sign and statistical significance of the coefficients on $\mathcal{U}$ and $1/\mathcal{P}$, and the ratio of the two coefficients (see Proposition (ref)).

The results are presented in Table (ref). Columns (1) and (2) report estimates for the linear approximation in Equation (ref). Column (1) reports results using OLS, ignoring potential endogeneity and column (2) the estimates when instead using random RS-assignment as instruments. Column (3) reports the estimates for the exponential model of Equation (ref) using the control function method. Column (4) provides estimates of the Average Marginal Effects and can thus be compared more easily to columns (1) and (2).

All columns provide evidence consistent with the theoretical model. The two key variables of the model, $\mathcal{U}$ and $1/\mathcal{P}$, are both highly significant, with the coefficient on $1/\mathcal{P}$ being, as expected, negative, indicating a positive relationship between the probability of success and the decision to click or apply. Comparing columns (1) and (2) shows that the coefficient of $\mathcal{U}$ is not changed when using instrumental variables rather than OLS but that the coefficient of $1/\mathcal{P}$ is reduced by almost 40% in column (1) ignoring endogeneity compared to column (2). Indeed, when looking at the bottom panel of the table, reporting the $\rho$'s of the Control Function method (see footnote (ref)), we observe that the covariance between the residuals and the dependent variables is non-significant for $\mathcal{U}$, but significant and positive for $1/\mathcal{P}$.\footnote{If we use the benchmark error in variable model $y=a+bx^*+u$, a measurement model $x=x^*+e$ and an instrumental first stage regression $x=\alpha z+\varepsilon+e$ with $u$, $e$ and $\varepsilon$ uncorrelated and $u$ and $\varepsilon+e$ uncorrelated with $z$; if we assume errors in variable is the only source of endogeneity; then, in equation (ref), $\mu$ is equal to $u-be$ and $v$ to $\varepsilon+e$. Thus $\rho=-b\sigma^2_e/(\sigma^2_\varepsilon+\sigma^2_e)$. This would lead to a share of variance of the error $\sigma^2_e/(\sigma^2_\varepsilon+\sigma^2_e)$ of $0.008/0.014= 0.57$ for column (2) and $0.029/0.068=0.43 $ for column (3).} It is also worth highlighting the consistency of the results. As is standard in discrete choice models, the key quantity is the ratio of the coefficients. When considering the ratio between the coefficient on $1/\mathcal{P}$ and that on $\mathcal{U}$, we obtain very similar values across columns: 0.016 for column (2) and 0.024 for columns (3) and (4).

As stressed above, these results are especially important because they support the interpretation outlined at the beginning of Section (ref). They are consistent with viewing the score $\mathcal{U}_{i,j}$ as a signal of the utility gap $U-U^*$ and the probability $\mathcal{P}_{i,j}$ as a signal of the likelihood of success of an application. Taken together, these findings reinforce the idea that both dimensions are relevant inputs for the design of welfare-improving RSs.

table[table omitted — 2,260 chars of source]

As a robustness check, we also examine an alternative specification using observational data. We rely on data from the monitoring of job seekers’ search activity. All job postings on which a job seeker has clicked are identified and stored, along with subsequent actions—particularly whether an application was submitted. For each of these postings, we compute the indicators $\mathcal{U}$ and $\mathcal{P}$. We then estimate the model directly using this observational dataset. Table (ref) in Appendix (ref) presents these results, which are remarkably consistent with those in Table (ref). The main variables are all significant with the expected sign and, in addition, the coefficients from these estimations provide similar orders of magnitude. More precisely, the ratio of the two coefficients ranges from 0.018 to 0.025, very close to the previous values.

Comparison of different RSs

We use the data from our experiments, together with the model developed in Section (ref), to compare the performance of different recommendation systems (RSs).

Equation (ref) shows that the welfare-relevant score for a vacancy is given by the product of the probability of a successful application, $p$, and the function $\sigma \log\!\left(1+e^{\Delta(p,U)/\sigma}\right)$. The estimates from the previous section allow us to identify the surplus function $\Delta(p,U)$. In principle, one could therefore reconstruct the model-implied optimal score for each vacancy directly from the baseline utility and hiring scores.

We do not pursue this approach. Doing so would mechanically anchor the analysis on the two initial scores $\mathcal{U}$ and $\mathcal{P}$, and would prevent us from exploiting the additional information revealed by alternative RSs. Instead, we adopt the perspective that each RS provides a noisy signal about the two underlying components that matter for job seekers' welfare: the probability of being hired conditional on applying and the surplus from applying. Our objective is therefore to enrich the proxies for both dimensions by exploiting the full set of algorithms tested in the experiments.

To implement this strategy, we rely on data from Experiment 2. We briefly describe how we proceed and provide further details in Appendix (ref). For each enrolled job seeker, we observe the list of ten recommended vacancies and, for each of these vacancies, the scores produced by all algorithms considered in the second beta-test (with the exception of XGBoost). By combining these scores with observed application decisions and subsequent hiring outcomes on recommended vacancies, we estimate the components of the welfare-relevant score.

We proceed in four steps, detailed in Appendix (ref).

Step 1. Hiring conditional on applying. We identify the probability of a successful application, $p$. Restricting attention to vacancies to which job seekers actually applied, we estimate a logistic regression of the hire outcome on the scores produced by the different algorithms ({\sc Vadore}.0, {\sc Vadore}.2, Mix-$\sfrac{1}{2}$({\sc Vadore}.2), Application, $\mathcal{U}$-rec). This yields predicted hiring probabilities $p_{i,j}$ for each job seeker--vacancy pair in the recommendation lists.

Step 2. Applications on recommended vacancies. We identify the surplus component $\Delta$, or equivalently the application probability $p_a$. We estimate a logistic model for the probability that a recommended vacancy receives an application, again as a function of the same set of algorithmic scores. This yields predicted application probabilities $p_{a,i,j}$.\footnote{The results of these two first steps are reported in columns (1) and (2) of Table (ref). For the hiring probability (column (1)), only the coefficient for the {\sc Vadore}.2 score is significant. This supports the idea that it effectively incorporates the {\sc Vadore}.0 score in predicting hires. Moreover, the fact that the Application and $\mathcal{U}$-rec\ scores do not predict hires once the {\sc Vadore}.2 score is accounted for also strengthens the point that these scores contain information distinct from the hiring probability. For applications (column (2)), the Application score is as expected the most important predictor, even though the {\sc Vadore}.2 and $\mathcal{U}$-rec\ scores are also significant.}

Step 3. Reconstructing $\Gamma$ and related objects. Combining these two sets of predictions, we reconstruct for each job seeker $i$ and each recommended vacancy $j\in\{1,\ldots,10\}$ the composite score $\Gamma_{i,j}$ as well as its components $p_{i,j}$, $p_{a,i,j}$, and $p_{h,i,j}=p_{i,j}\times p_{a,i,j}$.

Step 4. Counterfactual optimal recommendations. We use the estimated $\Gamma$ function to construct a counterfactual benchmark. For each job seeker, we identify the set of vacancies that would have been recommended had it been possible to rank all available vacancies at the time of the experiment using the welfare-relevant score. This yields, for each job seeker, counterfactual scores $\Gamma_{i,j}^*$ for the (counterfactual) top ten vacancies.

We then evaluate each RS using the performance measures $\mu_{i,j}\in\{p_{i,j},p_{a,i,j},p_{h,i,j},\Gamma_{i,j},\Gamma_{i,j}^*-\Gamma_{i,j}\}$ and compute their averages within each experimental group assigned to algorithm $a\in \mathcal{A}_2$.

Since these estimates are subsequently used to construct predicted scores, we adopt a split-sample approach. One randomly selected half of the data, $S_1$, is used to estimate $p_a$, $p$, $p_h$, and $\Gamma$. The remaining half, $S_2$, is then used to compute average performance measures by experimental group: \[ \overline{\mu_{i,j}}^{\,i\in S_2,\; a_i=a}. \] To account for the uncertainty introduced by sample splitting, valid confidence intervals are constructed by taking the medians of the upper and lower bounds of the confidence intervals across multiple splits chernozhukov2018generic.

Figure (ref) summarizes the performance of the different RSs along four dimensions: the probability of being hired conditional on applying, $p$ (panel (a)); the probability of applying, $p_a$ (panel (b)); expected utility prior to application, $\Gamma$ (panel (c)); the joint probability of applying and being hired, $p_h=p \times p_a$ (panel (d)).\footnote{Table (ref) provides the associated estimates.} The figures also display the average value of the vacancies that would have been recommended by the optimal RS, as measured by the metric used in the figure (depending on the figure, either $p$, $p_a$, $p_h$ or $\Gamma$). This benchmark represents the performance that would be attained by the $\Gamma$-optimal recommendation set. It allows us to directly compare the performance of each recommendation rule to that of the optimal RS, using the metric implicit to each panel. Accordingly, the gap between the performance of a given RS and this reference captures how far the RS is from the optimal benchmark along each dimension ($p$ in panel (a), $p_a$ in panel (b), $\Gamma$ in panel (c), and $p_h$ in panel (d)). While this comparison is informative across all panels, it is particularly meaningful in panel (c), which relies on the $\Gamma$ metric: only for $\Gamma$ does the gap admit a direct welfare interpretation as a loss relative to the optimum. This panel therefore provides a direct assessment not only of how well each algorithm performs under the appropriate objective, but also of how close it comes to the optimal benchmark.

Panel (a) shows that all algorithms significantly increase the probability of a hire conditional on application $p$, relative to the baseline $\mathcal{U}$-rec\ system. Even the algorithm with the smallest gain, XGBoost, more than doubles this probability. The best-performing algorithm on this dimension is {\sc Vadore}.2, for which the probability of hire is 3.2 times higher than under $\mathcal{U}$-rec. Despite these relative improvements, it is important to note that absolute success rates remain very low. Even for the best algorithm ({\sc Vadore}.2), the conditional probability of hire on a recommended job posting remains below 1.5%.

Panel (b) presents similarly strong gains in the probability of application $p_a$. Again, relative to the $\mathcal{U}$-rec\ benchmark, improvements are substantial. The smallest gain is observed for {\sc Vadore}.0, which still increases the probability of application by 75%. The Application algorithm achieves the highest impact, increasing the probability by a factor of 1.7, with {\sc Vadore}.2 not far behind at 1.4. This can be seen as a consistency check given as the main predictor in our estimation of $p_a$ is the Application score (see Table (ref)). Yet, absolute levels remain modest: even the best-performing algorithm yields an application probability of only 0.4%.

Panel (d) shows the unconditional probability of hire $p_h$. The pattern mirrors that of the previous panels: the new RSs all substantially outperform $\mathcal{U}$-rec, with {\sc Vadore}.2 again providing the largest gain—tripling the likelihood of successful matches. The Application algorithm performs nearly as well. However, the absolute probability of a match remains extremely low, around 1 in 10,000.

Finally, panel (c) reports the optimal score $\Gamma$ which ranks the algorithms according to their expected value from the job seeker’s perspective. Two performance tiers emerge clearly: {\sc Vadore}.2 and Application form a top tier, generating large gains relative to the benchmark $\mathcal{U}$-rec, while {\sc Vadore}.0, Mix-$\sfrac{1}{2}$({\sc Vadore}.2), and XGBoost constitute a second tier. As shown in Section (ref), $\Gamma\approx p_h\times(1+p_a/2)$ for small $p_a$, so the two metrics are nearly identical in our empirically low-application setting. Overall, the strong performance of {\sc Vadore}.2 and Application highlights the importance of explicitly modeling application behavior when deriving optimal recommendations using the objective $\Gamma$. At the same time, these gains should be interpreted cautiously: in absolute terms, application and hiring rates for recommended vacancies remain low across all systems.

Across all panels, recommendations generated by the optimal algorithm strictly dominate those produced by any alternative algorithm. This ranking holds regardless of the metric considered. Among the feasible algorithms, {\sc Vadore}.2 consistently performs closest to the optimal benchmark, with one exception: for the application probability metric, the algorithm based solely on $p_a$ is the closest and in fact delivers nearly identical values. The most informative comparison is that based on the $\Gamma$ metric. Along this dimension, the performance achieved by {\sc Vadore}.2 is very close to that of the optimal algorithm, implying that the effective loss from using {\sc Vadore}.2 instead of the optimal recommendation rule is small.

figure[figure omitted — 1,580 chars of source]

Conclusion

In this paper, we study the design of job recommendation systems (RSs) by combining economic modeling, machine learning, and field experimentation. We develop a job-search framework in which vacancies are lotteries characterized by a hiring probability $p$ and payoff $U$. The model shows why recommendation rules based solely on proxies for $p$, proxies for $U$, or observed application behavior are incomplete, and that welfare-relevant rankings must combine both dimensions into an expected-surplus index. It also highlights an inversion problem: observed application choices reveal whether applying is privately profitable, but not the magnitude of the expected gains relevant for welfare.

We bring this framework to the data through collaboration with the French PES. Starting from two operational RSs, one reflecting stated preferences and one optimized to predict hiring outcomes, we conduct two randomized field experiments conceived as beta tests within a learning cycle. Guided by the model and experimental feedback, this process leads to an approximation of the welfare-optimal RS ({\sc Vadore}.2), whose performance in terms of clicks and applications substantially exceeds that of the initial systems.

Beyond reduced-form performance, the experiments generate exogenous variation in the characteristics of recommended vacancies, which we use to estimate a structural model of application behavior. The estimates support the key behavioral mechanisms emphasized in the theory and quantify the relative importance of hiring probabilities and utility in job seekers’ decisions. Combined with the model’s structure, the experimental data allow us to construct an empirically grounded benchmark for welfare-relevant rankings and compare all tested recommendation rules to it. We find that the approximation {\sc Vadore}.2 of the welfare-optimal algorithm delivers large gains relative to the initial systems and performs close to the model-implied optimum.

A broader lesson from our analysis concerns the performance of simple heuristic rankings. While the joint application-and-hiring probability $p_h$ is not welfare-optimal in theory, it emerges as a strong empirical benchmark in our setting. This result is structural rather than algorithmic: application probabilities are empirically small and remain so even under recommendation rules designed to stimulate applications. In this regime, the welfare-relevant index is well approximated by $p \times p_a$, explaining why hiring-based rankings dominate alternative heuristics. By contrast, rankings based solely on application behavior are theoretically fragile. Their reasonable performance in our setting may not generalize to environments where application behavior responds more strongly to recommendations. When recommendations substantially affect application decisions, the gap between behavior-based and welfare-based rankings may be much larger.

More broadly, our results suggest a general lesson for the design of algorithmic intermediation in labor markets. Machine-learning tools can substantially improve matching outcomes, but only when embedded in a framework that defines the economic objective and disciplines behavioral assumptions with experimental evidence. Without such a framework, RSs optimized for observable behaviors may perform well on predictive metrics yet remain misaligned with welfare-relevant outcomes.

Our findings suggest several directions for future research. First, improving RS performance requires better prediction of the primitives $p$ and $U$, especially job seekers’ utility for different jobs. Second, scaling recommendations may generate congestion and general-equilibrium effects gee2019more,altmann2022direct,bied2021congestion,SuBayoumiJoachims2022,roland2022,behaghel2024potential,lehmann2023, calling for market-level recommendation rules, for example based on optimal transport bied2021congestion. Third, fairness and inequality concerns remain central in labor-market RSs zhang2022understanding,bied2022fairness, and post-processing approaches with fairness constraints appear promising. Fourth, incorporating behavioral frictions may further improve recommendation design, for instance by combining RSs with elicited beliefs to study how recommendations shape expectations and search strategies almaas2023economics,anticipations. Finally, an important extension is to develop RSs that also benefit firms Horton2017,algan2020active, moving toward two-sided systems that jointly model applications and hiring decisions.

\FloatBarrier

\FloatBarrier

\setcounter{section}{0} \setcounter{table}{0} \setcounter{figure}{0} \setcounter{equation}{0}