Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
108,987 characters · 30 sections · 66 citation commands
Winner's Curse Drives False Promises in Data-Driven Decisions: A Case Study in Refugee Matching
\RUNAUTHOR{Bastani et al.} \RUNTITLE{False Promises in Data-Driven Decisions}
\TITLE{Winner's Curse Drives False Promises in Data-Driven Decisions: A Case Study in Refugee Matching}
\ARTICLEAUTHORS{ \AUTHOR{Hamsa Bastani} \AFF{Wharton School, \EMAIL{[email removed]}} \AUTHOR{Osbert Bastani} \AFF{University of Pennsylvania, \EMAIL{[email removed]}} \AUTHOR{Bryce McLaughlin} \AFF{Wharton School, \EMAIL{[email removed]}}
}
\ABSTRACT{A major challenge in data-driven decision-making is accurate policy evaluation---i.e., guaranteeing that a learned decision-making policy achieves the promised benefits. A popular strategy is model-based policy evaluation, which estimates a model from data to infer counterfactual outcomes. This strategy is known to produce unwarrantedly optimistic estimates of the true benefit due to the winner's curse. We searched the recent literature on data-driven decision-making, identifying a sample of 55 papers published in the Management Science in the past decade; all but two relied on this flawed methodology. Several common justifications are provided: (1) the estimated models are accurate, stable, and well-calibrated, (2) the historical data uses random treatment assignment, (3) the model family is well-specified, and (4) the evaluation methodology uses sample splitting. Unfortunately, we show that no combination of these justifications avoids the winner's curse. First, we provide a theoretical analysis demonstrating that the winner's curse can cause large, spurious reported benefits even when all these justifications hold. Second, we perform a simulation study based on the recent and consequential data-driven refugee matching problem. We construct a synthetic refugee matching environment (calibrated to closely match the real setting) but designed so that no assignment policy can improve expected employment compared to random assignment. Model-based methods report large, stable gains of around 60% even when the true effect is zero; these gains are on par with improvements of 22--75% reported in the literature. Our results provide strong evidence against model-based evaluation. } \KEYWORDS{winner's curse, data-driven decision-making, policy optimization and evaluation}
The past two decades have seen tremendous gains in the accuracy of statistical and predictive models, due to a confluence of data availability, computational resources, and algorithmic advances. A major application of such estimation is for data-driven decision-making---e.g., targeting interventions to patients, discounts to customers, or even matching refugees to cities---based on observed covariates. In practice, a decision-maker gathers a historical dataset, estimates a model of both outcomes and treatment effects, and then uses this model to learn a good decision-making policy.
Naturally, one would want to know if this learned data-driven policy improves the decision-maker's objective (e.g., revenue or social welfare) over the existing policy. This requires policy evaluation, which estimates the purported improvement of the data-driven policy over the existing policy. A variety of methods have been proposed for estimating policy improvement with varying assumptions and statistical guarantees. We posit that a key goal of policy evaluation is to guarantee that the true policy improvement lies within a reported confidence interval with high probability---i.e., a decision-maker can have confidence that the proposed policy has benefits consistent with what was promised. We say a policy evaluation methodology is valid if it provides this guarantee.
While this sounds straightforward, the key challenge is that we do not observe counterfactual outcomes for decisions that differ from the ones in the historical data. As a consequence, we cannot directly estimate the efficacy of a decision-making policy on a held-out test set the way we can estimate the predictive accuracy of a prediction model.
When the historical dataset uses random treatment assignment (e.g., the data was collected via a randomized controlled trial (RCT)), then the gold standard policy evaluation methodology is inverse probability weighting (IPW) horvitz1952generalization, rosenbaum_central_1983 on a held-out test set. Intuitively, IPW only evaluates decisions when the chosen decisions align with those in the historical data, altogether avoiding the need to estimate counterfactual outcomes. We call this a model-free approach. These estimates can be combined with standard statistical methodologies to establish valid confidence intervals for policy improvement. However, there are several reasons why IPW may not be feasible. When the action space is large or the covariates are high-dimensional, IPW can have very high variance; yet, these are often the settings where data-driven decisions are most desirable. Alternatively, the treatment assignment may not be random, in which case IPW is invalid.\footnote{We are also assuming the treatment probabilities are known; they can alternatively be estimated from data, but this strategy relies on the assumption that there are no unobserved confounders.}
In these cases, practitioners often adopt a popular alternative where they use a model to predict counterfactual outcomes, and then use these counterfactual outcomes to estimate performance improvement. We call this a model-based approach. The model can range from a simple linear model to a dynamic Markov decision process to a modern machine learning model such as a random forests. The key feature is that they estimate policy improvement by using their estimated model to impute unobserved outcomes under counterfactual treatment decisions.
It is well known that using the same estimated counterfactual model for both policy optimization and policy evaluation results in an issue known as the winner's curse harrison1984decision,smith2006optimizer,andrews_inference_2024,zrnic_flexible_2024, where the estimated policy improvements are optimistically biased. Intuitively, the optimization procedure exploits estimation errors, choosing decisions whose predicted quality exceeds their true quality. If we use the same model to evaluate these decisions, these estimation errors go uncorrected, often predicting substantial performance improvements where none exist. Existing model-based evaluation methodologies do not correct for this unwarranted optimistic bias, making them invalid.
While the winner's curse is well-known, existing work has only shown it to hold under specific conditions. As a consequence, many papers provide seemingly plausible justifications for using model-based evaluation. To understand the nature of these justifications, we performed a survey of recent Management Science papers on data-driven decision-making (detailed in Section (ref)). Of the 55 papers that report on data-driven policy improvement, all but two use the model-based (rather than model-free) method. Surprisingly, a large number of papers provided no justification. When justifications were provided, they typically fell into one of the following categories:
To see why these justifications are plausible, consider the following argument: (1) if either J1 and J2 hold or J3 holds, then the estimated counterfactual outcomes are unbiased or have small bias; (2) if, in addition, J4 holds, then the estimated policy improvement is independent from the learned policy; (3) therefore, the estimated policy improvement is accurate.
Unfortunately, we show both theoretically and through a simulation study that this argument is incorrect. First, we provide two stylized examples showing that the winner's curse can be arbitrarily large while simultaneously satisfying these justifications. Second, we perform a realistic simulation experiment in the context of the refugee matching problem bansak2018improving,ahani2021placement to show that model-based evaluation can report very large spurious effects even when the reported justifications are met. Our results show that the winner's curse is substantially more pervasive than previously understood. They provide strong evidence against the validity of model-based policy evaluation methods, suggesting that they should not be used.
We highlight the critical need for better policy evaluation methodologies that are valid while addressing the shortcomings of IPW. For instance, recent work has studied promising directions that integrate policy learning and evaluation to provide lower-variance policy improvement estimates while preserving validity---e.g., targeting statistically significant policy improvement instead of purely maximizing expected outcomes bastani2025beating,chernozhukov2025policy, reducing the dimensionality of the policy class to mitigate the winner's curse banerjee2025selecting, or rigorously integrating auxiliary datasets with real-world outcome data mandyam_perry:_2025.
\paragraph{Theoretical analysis.} Intuitively, the winner's curse does not arise if the estimated counterfactuals are unbiased and sample splitting is used. However, unbiased estimates are a very strong assumption. In practice, one of two possible shortcomings can arise:
The literature specifically on the winner's curse typically focuses on the well-specified setting without sample splitting harrison1984decision,smith2006optimizer; furthermore, they generally rely on the estimated model being noisy (i.e., achieves low accuracy on a held-out test set) and unstable (i.e., parameter estimates can change drastically when the model is trained on new or bootstrapped samples). Our analysis shows that due to regularization, the winner's curse can occur under much more general conditions. Separately, there has been work showing that misspecification can result in bias jiang2016doubly, kallus2019intrinsically; however, they do not show that this bias can be systematically positive, or that it can arise even when the estimated model is highly accurate and stable.
\paragraph{Simulation study.}
Empirically, we perform a simulation experiment in the context of data-driven refugee matching bansak2018improving,ahani2021placement, where the goal is to match refugees to capacity-constrained cities to optimize overall employment outcomes. While model-based evaluations are widely used in the literature, we focus on refugee matching both due to widespread recent interest in this problem as well as the fact that it has been deployed across several countries to make placement decisions for thousands of refugees. Both bansak2018improving and ahani2021placement use the model-based method to evaluate proposed data-driven refugee matching algorithms, estimating 22--75% gains in refugee employment outcomes from data-driven placement decisions. We design a simulated dataset that closely matches the setting described in bansak2018improving, but constructed so that no policy can outperform random assignment. Thus, any estimated performance gains are spurious by design. Applying the algorithm and model-based evaluation method from bansak2018improving to this dataset, we obtain (spurious) estimated gains in employment outcomes of around 60%, which is on par with the estimates reported in their paper. We additionally consider a more sophisticated variation of the model-based evaluation method using bootstrapping ahani2021placement, and find that it also continues to produce substantial, stable estimated improvements even when the ground truth is null.\footnote{To the best of our knowledge, all published estimates for this problem use some variation of the model-based method. Prior to posting this paper, we shared our critique of model-based evaluation methods in private correspondence with bansak2018improving and ahani2021placement; in response, bansak2018improving have started developing model-free estimates to rigorously establish performance gains, and have shared preliminary results that appear promising.}
Policy evaluation methods are broadly categorized into model-free and model-based approaches. Model-free evaluation, rooted in the foundational work of horvitz1952generalization and rosenbaum_central_1983, provides asymptotically unbiased estimates of policy performance. However, these methods exhibit substantial variance, particularly in settings with large action spaces saito2022off. The granularity required for data-driven decision-making undermines standard policy evaluation mechanisms: as the support of the historical data becomes sparse relative to the covariate/action space, inverse probability weighting (IPW) yields extreme weights and unstable estimates. While doubly robust estimation robins1994estimation, robins1995semiparametric improves efficiency, the variance remains problematic in finite samples.
Given these limitations, many studies instead rely on model-based evaluation, which is subject to the winner's curse. The winner's curse was originally studied in auction theory, where the winner typically overestimates an item's value capen1971competitive; the concept was formalized for decision analysis by harrison1984decision and smith2006optimizer. In this work, we rebut common justifications used in the literature to dismiss the winner's curse.
Recent literature addresses these challenges through three primary mechanisms. First, inference methods account for selection effects, either by constructing confidence intervals conditional on the selection event andrews_inference_2024 or by applying corrections zrnic_flexible_2024. Similarly, gupta2024debiasing and xu2025winner propose debiasing in-sample performance using gradient-based and bootstrap-based corrections respectively. Second, variance reduction can be achieved by pooling treatments banerjee2025selecting or leveraging auxiliary data mandyam_perry:_2025. Third, the optimization objective itself can be modified; for example, swaminathan2015batch penalize variance to constrain the proposed policy to be close to the support of the existing policy, so that it can be evaluated downstream. Similar principles underpin “pessimistic learning” in offline reinforcement learning kumar2020conservative. Most recently, chernozhukov2025policy and bastani2025beating propose learning policies that target statistical significance on a held-out test set, with the latter explicitly characterizing the Pareto frontier of policies that trade off expected performance and statistical significance.
To understand the extent to which model-based policy optimization is currently used in the literature---and consequently, the prevalence of the winner's curse---we surveyed recent papers published in Management Science. We focused on this journal as a leading outlet for quantitative research at the intersection of data analysis (estimation) and decision-making (optimization).
We constructed a sample of recent papers using the EBSCO Business Complete database. We searched for papers published in the last ten years (2015--2024) that contained keywords related to both data-driven estimation and policy optimization. Specifically, we used the query (data OR estim*) AND (optim* OR policy) and restricted the search to peer-reviewed articles in Management Science. This initial search yielded 875 results. To efficiently screen this large corpus, we employed a two-stage classification procedure involving a Large Language Model (LLM) followed by manual verification. In the first stage, we fed the titles and abstracts of the 875 papers to the GPT-5.2 API with high reasoning. We prompted the model to identify papers that likely met three specific criteria: (1) one of the primary contributions involves estimating a model (e.g., demand, welfare) from data; (2) the estimated model is used to optimize a decision (e.g., pricing, routing, allocation); and (3) the abstract explicitly states an estimated improvement resulting from this optimized decision, suggesting that this performance gain is a headline contribution. To calibrate the model, we provided 39 manually labeled examples (both positive and negative) and instructed it to follow the provided reasoning. This automated screening process identified 111 papers as potentially fitting our criteria (classified as “yes” or “unclear”).
In the second stage, we manually downloaded and reviewed the full text of these 111 papers. We retained only those fitting the “estimate-then-optimize” framework where evaluation was performed on historical data with unknown counterfactuals. We further excluded papers where the optimization was purely theoretical without empirical estimation, or where the estimation was not central to the policy construction (e.g., purely descriptive regression analysis followed by a simulation-based analysis).
As summarized in Table (ref), this process yielded 55 papers. Of these, 53 (96%) relied on model-based evaluation---specifically, they evaluated the performance of their proposed policy using the same data generating process or specific dataset used for estimation. As demonstrated in the rest of this paper, this methodology creates a direct channel for the winner's curse to inflate performance estimates. This survey suggests that the issue is not merely a theoretical curiosity but a pervasive feature of the current state of data-driven decision-making in the field.
Table (ref) summarizes the specific characteristics of these 53 papers. We start with the mechanics of the evaluation. Thirty-five of the 53 papers (66%) evaluate the performance of their optimized policy using the exact same estimated model that was used to generate the policy. This is the most severe form of the winner's curse, as the optimizer is free to exploit every residual error in the model. The remaining papers generally use the same model structure but re-estimate parameters, which, as we show in Section (ref), does not eliminate the bias.
How do these papers justify this approach? Most do not. Among those that do, the most commonly cited defense is predictive accuracy. Twenty-five papers (47%) argue that because the underlying estimated model achieves high accuracy (e.g., low MSE on a held-out test set), the resulting policy evaluation is valid. Nineteen papers (36%) appeal to stability (e.g., the estimated gains are stable across bootstrapped subsamples). Calibration of predictions is cited less frequently, appearing in only seven papers (13%). Sample splitting---which ensures that the evaluation and optimization procedures do not re-use the same data---is rare, appearing in only four papers (8%). As we demonstrate in Sections (ref) & (ref), none of these justifications---accuracy, stability, calibration, or sample splitting---are sufficient to rule out spurious gains from the winner's curse.
The vast majority of papers do not address the uncertainty of the estimated gains at all. Only 9 papers (17%) report confidence intervals for their policy improvements. None attempt to construct these intervals in a way that accounts for the optimization process itself. This lack of statistical inference is particularly troubling given the magnitude of the claimed improvements, which may rely on exploiting noisy predictions from tail events or rare subpopulations.
Two additional points are worth noting regarding the model classes and problem domains. First, the winner's curse is not limited to “black box” machine learning methods. Thirty-five of the papers (66%) use simple parametric models (such as linear or logit regressions, either via offline or online learning), 13 (25%) use structural models (such as MDPs or queueing), and only five (9%) use blackbox machine learning models (such as KNNs or DQNs). This suggests that the issue is fundamental to the estimate-then-optimize paradigm rather than the adoption of blackbox machine learning methods. Second, the practice is widespread across domains. Revenue management (45%) and healthcare (30%) are the most represented fields, likely due to the natural fit of optimization techniques in these areas. In short, our survey suggests that model-based evaluation is the default standard in the literature, despite its inherent susceptibility to the winner's curse.
In this section, we provide a theoretical construction where misspecification drives the winner's curse even when justifications J1 (accurate, stable, and calibrated model), J2 (historical data uses random treatment assignment), and J4 (sample splitting) all hold. While our example is stylized, as we discuss in Section (ref), the key aspects of our example that drive the winner's curse---namely, the exploitation of estimation errors in regions of sparse data---correspond to natural behaviors of simple estimators (e.g., ordinary linear regression) in high-dimensional settings.
We consider a simple setting with a single real-valued treatment $t\in[0,1]$ and an outcome $y\in\mathbb{R}$, without covariates. We consider a noiseless true model $y=f^*(t)$, where $f^*$ is a piecewise linear function with two hyperparameters $t_0\in(0,1)$ and $y_{\text{max}}\in\mathbb{R}$:
This function is illustrated as the solid black line in Figure (ref).
To estimate this model, we consider historical data collected using a uniform distribution over treatments $t\sim\text{Uniform}([0,1])$. Specifically, a single example in our historical dataset is a pair $(t,y)$, where $y=f^*(t)$. For simplicity, we consider the infinite data limit, where we observe the exact distribution over $t$. We assume we are fitting a linear function $\beta t$, a standard model family that achieves high out-of-sample accuracy in this setting.
\proof{Proof of Lemma (ref)} Note that the OLS estimator is $\hat\beta=\mathbb{E}[t^2]^{-1}\mathbb{E}[tf^*(t)]$. We have
and
The claim follows. \Halmos
Thus, our model is $\hat{f}(t)=\hat\beta t$; this function is visualized as the dashed line in Figure (ref). Our first result characterizes the winner's curse for this model.
\proof{Proof of Proposition (ref)} The best treatment according to our model is $\hat{t}=\operatorname*{\arg\max}\hat{f}(t)=1$. It has estimated outcome $\hat{y}=\hat{f}(\hat{t})=\frac{(1+t_0)y_{\text{max}}}{2}$ and true outcome $\tilde{y}=f^*(\hat{y})=0$, so the bias is $\hat{y}-\tilde{y}=\frac{(1+t_0)y_{\text{max}}}{2}$. \Halmos
For $t_0\ge\frac{1}{2}$, we have $\hat{y}-\tilde{y}\ge\frac{y_{\text{max}}}{2}$; thus, the winner's curse can be arbitrarily large. Note that the estimated performance is comparable to the true optimal performance (the optimal treatment is $t^*=\operatorname*{\arg\max}_{t\in[0,1]}f^*(t)=t_0$ and its true outcome is $y^*=f^*(t^*)=y_{\text{max}}$). Yet, the realized performance is zero—i.e., the estimated outcome is entirely illusory due to the winner's curse.
Crucially, this happens even though the model is highly accurate on the historical data distribution (i.e., achieving arbitrarily low MSE).
\proof{Proof of Proposition (ref)} Note that
as claimed. \Halmos
Intuitively, the discrepancy between high accuracy and large winner's curse occurs because the optimized treatment induces a shift in the data distribution---the estimated model performs significantly worse on this shifted distribution.
Our model illustrates that even when justifications J1, J2, and J4 in Section (ref) hold, the winner's curse can result in arbitrarily large, spurious policy improvement estimates. Here, we connect our example to these justifications:
While our example is highly stylized, it proves that the justifications described in Section (ref) (except the well-specified justification, J3) are insufficient to rule out the winner's curse from a theoretical perspective. To connect this stylized example to practical failure modes, consider the two key factors that drive the winner's curse:
Then, the optimizer can exploit the high-error region of the input space to maximize the predicted outcome, leading to the winner's curse.
One notable feature of our example is that the outcome goes from $y_{\text{max}}$ to zero with a tiny change in treatment, which might be unlikely to happen in practice. However, this construction was only necessary since we were using a simple, one-dimensional treatment and no covariates. In realistic examples, the input space is typically high-dimensional, providing significantly more room for this kind of bias to occur. For instance, in the refugee setting, the goal is to assign refugees to cities to maximize employment rate (e.g., after 90 days). In this problem, each city is a discrete treatment $t\in\mathcal{T}=\{1,...,k\}$. This problem additionally involves targeting the treatment based on individual covariates $x\in\mathcal{X}\subseteq\mathbb{R}^d$, such as country of origin, languages spoken, etc. The outcome for a single individual is an indicator for employment $y\in\{0,1\}$. In this case, a policy has form $\pi:\mathcal{X}\to\mathcal{T}$. A standard way to learn such a policy called the plug-in approach is to first estimate a model $f^{\beta}:\mathcal{X}\times\mathcal{T}\to[0,1]$ mapping covariate-treatment pairs to outcome probabilities. This function induces a policy
i.e., choose the best treatment according to our estimated model.
Consider using a low-capacity model family---e.g., linear models $f^{\beta}(x,t)=\beta^\top\phi(x,t)$ over a feature map $\phi(x,t)\in\mathbb{R}^m$. If the feature map is too simple, then even for the optimal linear model $\beta^*$, some covariate-treatment pairs will naturally be upwards biased---i.e., $f^{\beta^*}(x,t)>\mathbb{E}[y^*\mid x,t]$, where $y^*$ is the true employment outcome---and some will naturally be downwards biased. Even if this average bias is small (or zero), the optimizer can systematically exploit the upwards-biased covariate-treatment pairs to obtain spurious improvements. Furthermore, as the feature dimension $m$ becomes larger, we might expect the winner's curse to become larger since there are more opportunities for bias to arise.
In Section (ref), we showed that misspecification can result in the winner's curse. Thus, one may hope that the winner's curse is mitigated when using a high-dimensional (e.g., LASSO) or non-parametric model (e.g., random forest); then, we may argue that it is likely that the true data generating process lies in the model family, satisfying J3.
Unfortunately, we show that the winner's curse can still arise when the estimation problem is well-specified as long as the model is regularized. This case is substantially more delicate---indeed, it is true that in the well-specified setting (along with the other justifications), the winner's curse disappears asymptotically. However, in finite sample, these models must be regularized to avoid arbitrary overfitting, which creates bias in finite sample. Surprisingly, this bias can actually be stable---i.e., the estimated model is biased in the same way with high probability over data sub-samples. Intuitively, regularization is typically applied in a systematic way, forcing “high variability” components of the model family to zero until a sufficient amount of data is available to estimate them accurately.
We show via a simple well-specified linear model that in fact, regularization bias is sufficient to result in the winner's curse. Specifically, consider a linear model with a binary treatment $t\in\{0,1\}$, a single continuous covariate $x\in\mathbb{R}$, and an interaction term $tx$. We assume that $t\sim\text{Bernoulli}(p)$ (for some $p\in[0,1]$) and $x\sim\mathcal{N}(0,1)$ are independent random variables. Also, assume that the noise distribution is $\epsilon\sim\mathcal{N}(0,1)$. We can interpret this setting in the context of our refugee matching case study by considering two locations $A,B$, with $t$ indicating whether to assign a refugee to location $A$ (if $t=0$) or $B$ (if $t=1$). Each individual is represented by a single covariate $x$.
To estimate a useful targeting model, we need to have interaction terms between $t$ and $x$; otherwise, a linear model will conclude that a single treatment is best for all individuals. Thus, we consider the following feature map (which we assume to be well-specified):
which includes a single interaction term $tx$. In this section, we distinguish between covariates (i.e., the individual attribute $x$) and features (i.e., the components of the feature map $\phi(x,t)$ encoding covariate-treatment pairs into a real-valued vector). Next, we assume the true parameters are
Note that under these parameters, there is no value to targeting---the outcome only depends on $x$, and not on the interaction $tx$.
Now, suppose that $p\approx1$; that is, in our historical data, refugees are overwhelmingly assigned to location $B$ (e.g., due to stringent capacity constraints in $A$). In this case, there are very few refugees assigned to location $A$. In the data, we observe that the outcome $y$ is strongly correlated with both $x$ and $tx$ (since $x$ and $tx$ are also strongly correlated), but we cannot distinguish which of these two features is causal; thus, we cannot identify their coefficients. Mathematically, this is because the true covariance matrix
has $\lambda_{\text{min}}(\Sigma)\to0$ as $p\to1$, highlighting the strong correlation between $x$ and $tx$. As a consequence, any unbiased estimate of $\beta^*$ (i.e., via OLS) will be highly unstable. Specifically, while we can accurately estimate $\beta^*_1$ from just a few samples, estimating $\beta_2^*,\beta_3^*$ requires many more samples due to insufficient variation in the training data. This is not a problem for prediction---we can still identify the sum $\beta_2^*+\beta_3^*=b$. If our goal is only to achieve good out-of-sample accuracy on the distribution of the historical data, then identifying $\beta_2^*+\beta_3^*$ is sufficient, since $t=1$ with high probability, so
We can obtain a stable estimate of $\beta_2^*+\beta_3^*$ by using regularization, specifically, ridge regression. Intuitively, regularization forces the ridge regression estimator to evenly distribute $b$ between $\beta_2^*$ and $\beta_3^*$; indeed, consider the parameter vector
We can show that the ridge regression estimator $\hat\beta^{(\lambda)}$ concentrates tightly to $\bar\beta^{(\lambda)}$, implying that it is stable. Furthermore, $\bar\beta^{(\lambda)}$ satisfies the desired relation $\bar\beta^{(\lambda)}_2+\bar\beta^{(\lambda)}_3=b$; thus, this estimator has high accuracy. Specifically, define the excess mean-squared error (MSE) to be the generalization error relative to the true model (typically estimated on a held-out test set):
where $\mathbb{E}_{x,t}$ is the expectation with respect to the training covariates $X\sim\mathcal{N}(0,1)^n$ and training treatments $T\sim\text{Bernoulli}(p)^n$. Then, $\mathcal{E}(\bar\beta^{(\lambda)})$ can be small (e.g., on the order of $1/\sqrt{n}$) even when $\|\bar\beta^{(\lambda)}-\beta^*\|_2$ is large (e.g., on the order of $1$).
However, our goal is policy learning---thus, it is not enough to achieve good accuracy on the observed data distribution. Instead, we need to obtain good accuracy on the eventual data distribution induced by using an alternative policy $\pi:\mathcal{X}\to\mathcal{T}$ to assign treatments. Unfortunately, this requires that we identify $\beta_2^*$ and $\beta_3^*$ separately; the winner's curse arises from this disconnect. In particular, note that $\pi^{\bar\beta^{(\lambda)}}(x)=\mathbbm{1}(x\ge0)$. Because the interaction term $tx$ has estimated coefficient $\bar\beta^{(\lambda)}_3=b/2$, the policy believes it is beneficial for units with $x\ge0$ to receive treatment $t=1$; conversely, for $x\le0$, it believes it is beneficial for units to receive treatment $t=0$.
This optimistic bias for a unit with $x\ge0$ is $bx/2$, and the overall optimistic bias of $\pi^{\bar\beta^{(\lambda)}}$ is
Furthermore, note that stability of $\hat\beta^{(\lambda)}$ implies stability of the optimism bias (formalized in Proposition (ref)); thus, using sample splitting or a bootstrapped estimate would reliably produce the same optimism bias. We formalize these results in the following proposition.
We give a proof in Appendix (ref). If we take $\alpha=n^{1/4}$, then for any choice of $b$, as $n$ becomes sufficiently large, then the MSE goes to zero and the estimator is perfectly stable, whereas the optimism bias converges to $b/(2\sqrt{2\pi})$. Thus, accuracy and stability cannot definitively rule out the winner's curse.
Our model illustrates that even when all of the justifications in Section (ref) hold, the winner's curse can result in arbitrarily large, spurious policy improvement estimates. Here, we connect our example to these justifications:
Yet, by taking $b$ large, we can obtain arbitrarily large unwarranted optimism in our policy improvement estimates.
While our example is stylized, it proves that all of the justifications described in Section (ref) are insufficient to rule out the winner's curse from a theoretical perspective. In this case, the winner's curse is driven by a very similar mechanism as in Section (ref); the key difference is that the bias arises from regularization:
As before, the optimizer can exploit the high-error region of the input space to maximize the predicted outcome, leading to the winner's curse.
In realistic scenarios, low-probability regions of the input space can include not just rare treatments (as in our stylized example), but also individual covariates that rarely occur (e.g., in our refugee example, a rare country of origin). Our setting is very low-dimensional and used a single, very rare treatment. In practice, the feature dimension $m$ is usually much larger, so it is actually much more likely that there are features $\phi(x,t)_i$ along which variation in the historical data is low, thereby requiring regularization to avoid instability and ensure accuracy. For these features, the optimizer can exploit biases that arise from regularization to achieve spurious improvements.
Beyond linear models, model families such as random forests and gradient boosted machines (GBMs) also perform implicit regularization to avoid overfitting the data. For instance, random forests average over a large number of decision trees, thereby regularizing estimates to zero when uncertain. Alternatively, GBMs rely on fitting low-capacity base models such as shallow decision trees, thereby inheriting the biases of the base model family. Thus, these kinds of models may also exhibit the winner's curse; our simulation study in Section (ref) illustrates this possibility
We now provide empirical evidence that the justifications described in Section (ref) do not guarantee accurate policy evaluation using a synthetic environment modeled after the refugee matching problem. Resettlement agencies attempt to settle incoming refugees in locations that maximize their welfare, often measured in the short-run by employment outcomes. To improve their matching of refugees to locations, bansak2018improving and ahani2021placement have proposed optimizing future location assignments using employment prediction models trained on historical outcomes. Both papers claim this procedure leads to large gains in employment, using model-based evaluation and justifications J1, J2, and J3. We design a simulation environment calibrated to the U.S.-based setting studied by bansak2018improving, but generate employment outcomes in such a way that no assignment policy can impact the expected employment rate. We train multiple prediction models, use them to estimate optimal policies, then evaluate these policies using both model-based methods (biased by the winner's curse) and a model-free method (IPW, which is unbiased, but high variance). Despite being built off accurate, calibrated, and stable prediction models (J1), being trained on random historical assignments (J2), and (for some models) being well-specified (J3), all model-based methods falsely report improved employment rates due to the winner's curse.
The refugee matching problem provides a structure to analyze and improve assignments of incoming refugees to locations for resettlement. Resettlement agencies assign refugees to locations based on a small set of visible features and a variety of constraints. The authors argue that these constraints introduce random variation to the assignment process that enable cross-location comparisons of refugees with similar features. The goal of refugee matching is to identify feasible assignment policies that improve short-term refugee employment outcomes.
Upon entering the United States, incoming refugees are assigned to a resettlement agency that manages their arrival. Each resettlement agency sets up affiliate networks in various locations around the country to help integrate arriving refuges and provide them various services. Using a limited set of features for each refugee (such as age, nationality, and education), agencies must assign refugees to these locations.
Refugee assignments must satisfy various constraints. Some refugees have prior ties in the United States that dictate their placement. Other refugees require specific services, such as english language education, that only exist in some location's affiliate networks. Moreover, each location has limited resources, creating capacity constraints which may vary over time. These constraints complicate the assignment process, but provide the variation required to estimate the heterogeneous impact of locations on refugees.
Resettlement agencies want to assign refugees in a welfare-maximizing way, subject to their various constraints. In the short-run, few welfare-relevant outcomes are easily trackable, so many agencies record employment outcomes after 90-days as a proxy for a successful refugee-location match. The goal is thus to use past observations of refugee-location matches and their employment outcomes to design an assignment policy that leads to a higher refugee employment rate.
Following bansak2018improving and ahani2021placement, we formulate a simple offline version of the refugee matching problem and describe an estimate-then-optimize method for proposing refugee assignments. Each incoming refugee consists of (known) covariates and (unknown) counterfactual employment outcomes for each location. We use the covariates to estimate employment outcomes for each location, then use an integer program to assign refugees to locations based on these estimates. Additional details are provided in Appendix (ref).
The assignment problem consists of matching $N$ unrestricted refugees to $L$ locations in order to maximize the number of employed refugees. Each refugee $i \in [N]$ is associated with a set of potential outcomes $Y_i(t) \in \{0,1\} \ \forall \ t \in [L]$ which describe whether the refugee would find employment at location $t$.\footnote{To match the assumption by the original authors, we also assume the standard Stable Unit Value Treatment Assumption which says one refugees potential outcomes are not impacted by other refugees' location assignments.} A matching $\pi$ assigns refugees to locations where $\pi_{it}\in\{0,1\}$ records whether refugee $i$ is assigned to location $t$. Each location $t$ can support a maximum of $c_t$ refugees. The optimal matching can then be described by the following integer program.
However, the counterfactual outcomes under alternative locations for each refugee are unknown. Thus, the authors use a predictive model to estimate these counterfactuals based on refugee covariates $X_i$ (such as age, nationality, and education). Specifically we have access to observations of prior refugee-location matches where for refugee $i$ we observe their covariates $X_i$, the location they were assigned to $T_i \in [L]$, and their observed employment outcome $Y_{i}(T_i)$.\footnote{We similarly adopt the assumption of unconfoundedness: conditional on the covariates $X_i$ the assignment $T_i$ is independent of the potential outcomes. Later we will assume the exact policy used to generate these matches is known and also assigns positive treatment propensity to each location.} These prior observations may have had assignment restrictions, which should be captured by the covariates $X_i$. We can use these observations to build a prediction model $\hat{\beta}$ where $\hat{\beta}_t(X_i)$ estimates $\mathbb{E}[Y_i(t)|X_i]$ given refugee covariates $X_i$ and a location $t$. We then propose using the assignment which maximizes the expected number of employments according to the model $\hat{\beta}$, which can be described by the following integer program.
We refer to the assignment given by this program as $\pi^{\hat{\beta}}$.
bansak2018improving and ahani2021placement both evaluate their proposed policies using model-based methods---specifically, they use the same model for both policy estimation and policy optimization, which is prone to the winner's curse as discussed earlier. In this setting, under their assumptions of unconfoundedness and random assignment, a model-free approach (such as IPW), is guaranteed to be valid; however it is prone to high variance due to the size of the action space.
The model-based method of policy evaluation addresses the issue of missing counterfactual outcomes by imputing them. Specifically, it uses the model $\hat{\beta}$ (which optimized the proposed assignment), effectively reporting $\sum_{i = 1}^N \sum_{t=1}^L \hat{\beta}_{t}(X_i) \pi^{\hat{\beta}}_{it}$. ahani2021placement do the same, but also justify the evaluation using a set of $B$ alternative models $\{\hat{\beta}^{(b)}\} _{b=1}^B$ where model $b$ is built off an independently resampled bootstrap of the original training samples. While straightforward, these model-based approaches may produced biased evaluations if the prediction errors of the model used to propose the assignment $\pi^{\hat{\beta}}$ are correlated with the prediction errors of the model which evaluates the assignment. This is trivially true when $\hat{\beta}$ is reused for evaluation, but may occur even (i) if the evaluation models are built off bootstrapped samples of the training data or (ii) if the evaluation models are built off a completely independent sample. This can occur due the specification or regularization the prediction models must apply to provide stable insights when training samples are limited.
IPW is a model-free method of policy evaluation which addresses the issue of missing counterfactual outcomes by only evaluating the proposed policy $\pi^{\hat{\beta}}$ when the employment outcome is available. That is we assume each refugee $i \in [N]$ was actually assigned to location $T_i$ resulting in employment outcome $Y_i(T_i)$ where $T_i$ was generated independently from a known distribution where $\Pr(T_i = t) = p_{it}>0$ for all locations $t\in[L]$. IPW then rescales the outcome data based on how likely each assignment was to occur in the data according to $p$ and estimates the number of employed refugees using the cases the proposed assignment aligns with the observed assignment $\pi^{\hat{\beta}}_{iT_i} = 1$. Thus the expected number of employments under $\pi^{\hat{\beta}}$ according to IPW is $$\sum_{i:\pi^{\hat{\beta}}_{iT_i}=1} \frac{Y_i(T_i)}{p_{iT_i}}.$$ This estimate of the proposed assignment's impact is unbiased, but can suffer from excessive variance if $p_{iT_i}$ is small for some refugees. In our environment there are 43 locations, leading to a volatile IPW estimator.
We develop a synthetic environment to explore the potential impact of the winner's curse bias on model-based evaluation methods in the refugee matching problem. We tune our setting to mimic the US-based setting of bansak2018improving and J2: we generate refugee covariates according the marginal distributions reported by their supplemental material, then assign the same number of historical locations randomly in their observed proportions. We generate counterfactual employment outcomes in a way that depends on both refugee covariates and locations, but not their interaction, preventing any assignment policy from improving the expected employment rate.
The synthetic environment generates observations of prior refugee-location matches suitable for the policy optimization and evaluation methods described above. That is for each refugee $i$, our environment generates covariates $X_i$, a location $T_i$, and an employment outcome $Y_i(T_i)$. First, the environment assigns the covariates used by bansak2018improving (age, gender, education, english-speaking, case restriction, country of origin, arrival year, and arrival month) independently at random according to the marginal distributions reported in their Supplemental Material. Next the location $T_i$ is drawn independently at random from a set of 43 possible locations according to the empirical distribution $p$ that bansak2018improving reports. We report the covariate and location distributions in Appendix (ref). Finally the employment outcome is drawn independently according to $Y_i(T_i)\sim \mathrm{Ber}(f(X_i,T_i))$ for our causal model $f$, which dictates the employment probability for a refugee with covariates $X_i$ placed in location $T_i$.
We generate employment outcomes using a causal model that induces the same expected employment rate for all feasible refugee-location matchings when all locations are at capacity. Specifically we determine the probability a refugee attains employment as $$f(X_i,T_i) =\frac{1}{2}f_X(X_i) +\frac{1}{2}f_L(T_i),$$ where $f_X\in[0,1]$ represents the impact of the refugee's covariates on employment and $f_L \in[0,1]$ represents the impact of the location on employment. Under a causal model of this form no feasible matching can impact the expected employment rate if all locations are at capacity. We may rearrange the probability of employment across refugees through different assignments, but the lack of covariate-refugee interaction effects prevents us from improving the employment rate via better matchings. The refugee effect, $f_X$, is determined by first using a logit regression with random coefficients\footnote{We manually assign the covariate indicating case restriction=free to have a large positive impact. We do this as the test set of bansak2018improving only includes free case placements and achieves $34\%$ employment while their training set includes both types of cases and achieves $23\%$ employment. To replicate this difference we gave free case placement a positive impact on employment probability.} over refugee covariates and locations to assign an employment outcome to each refugee, then training a random forest regression using the refugee covariates to determine the underlying refugee effect. The location effect, $f_L$, for each of the 43 locations is drawn independently from a Beta distribution with $\alpha = 1$ and $\beta = 2$.\footnote{We chose these parameters to replicate the employment differences bansak2018improving observes across locations as faithfully as possible. Employment rates by location are reported in Appendix (ref).}
We test the estimate-then-optimize refugee matching method using three prediction model classes, then apply the evaluation methods described in Section (ref) to the proposed policies. In each case we train the estimation model on 33,000 refugees and evaluate the performance on 1,000 refugees with no prior placement restrictions (as done in bansak2018improving). We perform each simulation 250 times to obtain histograms of performance estimates for each approach.
We compare the model-based and model-free methods described above for refugee assignments built using a variety of prediction model classes. First, we examine a LASSO-constrained logit regression over refugee covariates, locations, and their interactions (as done by ahani2021placement). This model class is very stable and generally performs well out of sample, but is improperly specified based on our causal model $f$. Second, we examine a family of Honest Random Forests (i.e. causal machine learning wager2018estimation); one for each location. The models are known from their ability to produce non-parametric unbiased estimates, but reduce the prediction power of the data in the process. Finally we examine a family of gradient-boosted classification trees (referred to as GBM); one for each location (as done by bansak2018improving). These models are very powerful, but not very stable and are prone to overfitting.
For each model class our simulation:
In addition to histograms which report estimates of the policy's performance over the 250 runs, we report ROC and calibration curves on a testing dataset for each prediction model in Appendix (ref) to address J1. We generate these curves using SKLearn methods as we do with the logit models and gradient boosted classifiers. We generate Honest Random Forests using the EconML library. We implement the integer optimization problem in CVXPY using the SCIP solver. All simulations are performed in Python.
Our simulations reveal that the model-based methods exhibit large biases from the winner's curse for all three prediction model classes, and that those biases are not eliminated when instead evaluated by models trained on bootstrapped versions of the training dataset. These simulations do not prove that the policies learned by bansak2018improving, ahani2021placement are ineffective; rather, they show that this form of policy evaluation has a very high FPR---consistently promising non-existent gains even when none are present---underscoring the need for valid policy evaluation methodologies in practice.
Figure (ref) presents histograms comparing evaluations of the optimized policy's impact using both the prediction model informing the policy's construction and IPW estimation. Both LASSO and Honest Random Forests present a moderate winner's curse bias, while GBM exhibits an extreme bias (likely due to its lack of stability). On the other hand, while IPW estimation remains unbiased regardless of the prediction model used to optimize the policy, its variance, driven by both small sample size and large action space, precludes effective evaluation of the policy's impact. We caution readers from using these simulations to conclude that LASSO and Honest Random Forest provided less biased evaluation, as alternative parameterizations of our simulation lead to larger biases from LASSO and Honest Random Forest (compared to GBM). Instead, these simulations highlight the need to use provably valid policy optimization and evaluation methodologies that can limit variance in the presence of large action spaces and small sample sizes bastani2025beating.
Figure (ref) presents similar histograms, only now we compare evaluation of the optimized policy using models trained on bootstrapped versions of the training dataset, as performed by ahani2021placement. For LASSO, we find that the bootstrapped models perpetuate as much bias from the winner's curse as evaluating using the original model used for policy optimization. We believe this follows from misspecification of the stable LASSO model: it is consistently reproducing the predictions of the original model in its bootstraps due to the complicated impact of the refugees' covariates on employment outcomes. Meanwhile Honest Random Forest and GBM have no obvious misspecification and consequently are able to reduce, but not eliminate, the bias from their evaluations. Instead these techniques reproduce some of the winner's curse bias that the assignment model suffers due to correlation between the training dataset and its bootstraps. As such we would expect this bias to decrease as the power of the training set increases. However, it is unclear how the size of this bias relates to the size of the training dataset, making it very difficult to use the bootstrapping approach for evaluation.\footnote{We once again caution readers from using these simulations to conclude that Honest Random Forest is the “best” prediction model to use for bootstrapped policy evaluation as other simulation setups using Honest Random Forest models have retained large biases.}
Overall our simulations highlight that in order to confidently assess complex data-driven assignment policies, we need better policy evaluation methods. Without improved methods, the only way to confidently evaluate these policies is to retain large testing datasets to counteract the variance of IPW estimation, which is often economically infeasible.
Our results demonstrate that the winner’s curse is not merely a theoretical curiosity, nor is it confined to obviously misspecified models. Rather, it is a systematic failure mode of model-based policy evaluation in estimate–then–optimize pipelines. In our refugee-matching case study, state-of-the-art model-based methods produced large and stable---but entirely spurious---estimated gains, despite standard checks such as cross-validation and calibration. This is because standard diagnostics validate average predictive performance, whereas policy optimization is designed to exploit precisely those parts of the input space where residual errors remain. To clarify why these spurious gains can survive familiar best practices, it is useful to contrast model-based and model-free evaluation across three data regimes.
The fundamental dilemma in evaluating data-driven policies is a bias–variance trade-off. Model-free methods such as inverse probability weighting (IPW) are asymptotically unbiased but suffer from prohibitively high variance when propensity scores are small or the action space is large. Model-based methods instead impute counterfactual outcomes, reducing variance by smoothing over noise. The problem, however, is that model-based estimates are biased in a systematically optimistic direction. The optimizer preferentially selects actions and subpopulations that appear best under the model---i.e., “wins” are disproportionately driven by favorable estimation error.
In “small data” settings, estimated models are noisy and unstable, discouraging any kind of policy evaluation. At the other extreme, in “large data” regimes with sufficient coverage, model-free estimators become precise, making it possible to accurately evaluate policy improvements.
The most problematic regime---and the one where the winner’s curse is most acute---is “medium data.” Here, the data are sufficient to estimate models that appear accurate, stable, and well-calibrated on average, yet insufficient to support low-variance model-free policy evaluation. In this regime, standard diagnostics create a false sense of security: a practitioner may observe high AUC, stability, and calibration, and conclude the model is a reliable proxy for reality. But policy optimization is effectively an adversarial search process that identifies and exploits residual model errors, especially in regions that are rare under the observed data distribution. Even a model with strong aggregate performance may overestimate outcomes for a small subset of cases; the optimizer will systematically select precisely those cases.
Crucially, because model-free estimates remain high-variance in medium data settings---particularly with large or combinatorial action spaces---they often cannot deliver a decisive rebuttal to a precise (but biased) model-based estimate. The result is a predictable pattern: the optimistic model-based gain looks stable and persuasive, while unbiased checks are too noisy to serve as an effective counterweight. The refugee matching case study exemplifies this regime: roughly 30,000 observations are sufficient to support seemingly reasonable predictive models, yet the sparsity of the assignment space ensures that model-free evaluation is too noisy to robustly validate the optimizer’s purported gains.
Because the medium-data regime is pervasive, the winner’s curse is a systemic threat. Many high-stakes decision problems possess enough data to estimate reasonable statistical models, but not enough to rigorously validate the tails of the distribution where optimized policies concentrate. Thus, we argue against using model-based imputation for policy evaluation, even when the estimated models appear strong by conventional metrics. Effectively doing so requires rigorously integrating model-based data with real-world outcomes mandyam_perry:_2025.
Rather, emerging research addresses the challenge of the high variance of model-free evaluation by rethinking the learning process itself; for instance, chernozhukov2025policy,bastani2025beating propose algorithms that seek statistically guaranteed improvements rather than simple expected value maximization, while banerjee2025selecting demonstrate that constraining the complexity of the policy class can effectively curb the winner's curse.
The authors are grateful to Gad Allon, Jackie Baek, Kirk Bansak, Mohsen Bayati, Ron Berman, Rob Bray, Gerard Cachon, Vishal Gupta, Dean Knox, Elisabeth Paulson, Neha Sharma, Yannis Stamotopoulos, Alexander Teytelboym, Stefan Wager, and others for helpful feedback. This research was supported by generous funding from the Wharton AI & Analytics Initiative.
\ECSwitch \ECHead{Appendix}
We begin by proving a number of general results on accuracy and stability of ridge regression; our proof of Proposition (ref) relies on the general theory to establish accuracy and stability. Our proof that the winner's curse happens is specific to the problem construction established in Section (ref).
We begin by formalizing the problem of policy evaluation where the counterfactual outcomes are estimated using a ridge regression model. We consider a covariate space $\mathcal{X}\subseteq\mathbb{R}^d$, a compact space of treatments $\mathcal{T}$, a parameter space $\beta\in\mathcal{B}\subseteq\mathbb{R}^m$, a feature space $\mathcal{Z}=\mathbb{R}^m$, and a feature map $\phi:\mathcal{X}\times\mathcal{T}\to\mathcal{Z}$. The decision objective is $f^{\beta}(x,t)=\beta^\top\phi(x,t)$, and the $\beta$-optimal policy is $\pi^{\beta}(x)=\operatorname*{\arg\max}_{t\in\mathcal{T}}f^{\beta}(x,t)$. We consider a probability measure $\mathbb{P}_x$ (capturing the covariate distribution) and $\mathbb{P}_t$ (capturing the treatment distribution in the historical data), defining the product measure $\mathbb{P}_{x,t}=\mathbb{P}_x\times\mathbb{P}_t$. We assume $\phi$ is measurable and bounded in expectation: $\mathbb{E}_x\left[\max_{t\in\mathcal{T}}\|\phi(x,t)\|_2\right]\le\phi_{\text{max}}$, and we assume that $\phi(x,t)$ is $\eta$-subgaussian. We let $\beta^*\in\mathcal{B}$ denote the parameters for the true decision objective. We evaluate parameter estimates $\beta\in\mathcal{B}$ using two metrics: (1) global prediction accuracy, and (2) the optimistic bias of the resulting policy $\pi^{\beta}$.
To estimate $\beta^*$, we consider a historical dataset $W=\{(z_i,y_i)\}_{i=1}^n$, where $y_i=f^{\beta^*}(x_i,t_i)+\epsilon_i$ with i.i.d. noise $\epsilon_i \sim \mathcal{N}(0,\sigma^2)$; we let $\mathbb{P}_{x,t,\epsilon}=\mathbb{P}_{x,t}\times\mathbb{P}_{\epsilon}$ denote the distribution of a single historical example. We employ the classical ridge regression estimator, which serves as a proxy for any stable, regularized machine learning method:
where $\lambda\in\mathbb{R}_{>0}$ is the regularization hyperparameter, and
Let $\Sigma=\mathbb{E}_{x,t}[zz^\top]$ be the true covariance matrix and $\hat\Sigma=n^{-1}Z^\top Z$ be the empirical covariance matrix, and let $\Sigma_{\lambda}=\Sigma+\lambda I$ and $\hat\Sigma_{\lambda}=\hat\Sigma+\lambda I$, where $I\in\mathbb{R}^{m\times m}$ is the identity matrix. Then,
where $E=
^\top$ is the vector of noise terms. Given $\Delta\in\mathbb{R}_{>0}$, define the event
In addition, given $\zeta\in\mathbb{R}_{>0}$ satisfying $\zeta\in\mathbb{R}_{>0}$, define the event
We let $\mathbb{P}_E=(\mathbb{P}_{\epsilon})^n$ denote the measure of $E$, and $\mathbb{P}_Z=(\phi_*\mathbb{P}_{x,t})^n$ of $Z$ (where $f_*\beta$ denotes the pushforward measure).
First, we prove that $E_{\Delta}$ holds with high probability as $n$ grows large.
Next, we prove a high-probability bound for $E_{\zeta}'$.
Next, we prove a standard formula for the excess MSE of any parameter vector $\beta$.
Our next result establishes a straightforward bound on the sensitivity of the estimated policy improvement $\mathcal{P}(\beta;\beta')$ to $\beta'$ (note, however that $\mathcal{P}$ can be arbitrarily sensitive to $\beta$).
Another simple result bounds the maximum eigenvalue of $\hat\Sigma_{\lambda}^{-1}$.
Next, we turn to analyzing the quantity $\hat\Sigma_{\lambda}^{-1}\Sigma$; this quantity is a key part of our analysis of the excess MSE of $\hat\beta^{(\lambda)}$.
Now, we state and prove our first main result, which bounds the accuracy of $\hat\beta^{(\lambda)}$ on events $E_{\Delta}$ and $E_{\zeta}'$. Intuitively, the first term in the bound is the bias term (which becomes small as $\lambda\to0$), and the second term is the variance term (which becomes small as $n\to\infty$). We generally think of $\lambda$ as scaling as $1/\sqrt{n}$.
Our second major result provides a stability guarantee for $\hat\beta^{(\lambda)}$; in particular, it says $\hat\beta^{(\lambda)}$ concentrates around a deterministic quantity $\bar\beta^{(\lambda)}$ on events $E_{\Delta}$ and $E_{\zeta}'$ (which, by Lemmas (ref) & (ref), holds with high probability over $W$). This result straightforwardly implies that all $\hat\beta^{(\lambda)}$ are pairwise close together. One thing to note in this result is that both $\Delta$ and $\lambda$ scale as $1/\sqrt{n}$; thus, to obtain a bound that goes to zero, we need $\lambda/\Delta$ to be sufficiently large.
Our next result is a straightforward consequence of our previous result---it says that since $\hat\beta^{(\lambda)}$ concentrates to $\bar\beta^{(\lambda)}$, then the optimism bias when evaluating under $\hat\beta^{(\lambda)}$ correspondingly concentrates to the optimism bias under $\bar\beta^{(\lambda)}$.
Let $\Delta=\sqrt{18\log(8/\delta)/n}$ (so $\lambda=\alpha\Delta$) and $\zeta=\log(8/\delta)$. Note that $\|\beta^*\|_2=b$, and
Next, by Lemma (ref), we have $\mathbb{P}_Z[E_{\Delta}]\ge1-\delta/8$, and by Lemma (ref), we have $\mathbb{P}_W[E_{\zeta}']\ge1-\delta/8$. By a union bound, these events hold for all $\beta\in\{\hat\beta^{(\lambda)},\hat\beta^{(\lambda)\prime},\hat\beta^{(\lambda)\prime\prime},\hat\beta^{(\lambda)\prime\prime\prime}\}$. Now, we prove the three results on the events $E_{\Delta}$ and $E_{\zeta}'$ for all of these $\beta$.
\paragraph{Accuracy.}
By Proposition (ref), on event $E_{\Delta}$ and $E_{\zeta}'$, we have
\paragraph{Stability.}
Let
so $\bar\Sigma=\Sigma+\Gamma$. Further define $\bar\Sigma_{\lambda}=\bar\Sigma+\lambda I$. Also, note that $\gamma=\|\Gamma\|_2=2(1-p)=\Delta$. Furthermore, let $\bar\beta^{(\lambda)}=(I-\lambda\bar\Sigma_{\lambda}^{-1})\beta^*$. Then, on events $E_{\Delta}$ and $E_{\zeta}'$, by Proposition (ref), we have
so the bound on $\|\hat\beta^{(\lambda)}-\hat\beta^{(\lambda)\prime}\|_2$ follows by the triangle inequality. Next, by Proposition (ref), we have
Now, we bound $|\mathcal{P}(\hat\beta^{(\lambda)};\bar\beta^{(\lambda)})-\mathcal{P}(\bar\beta^{(\lambda)};\bar\beta^{(\lambda)})|$. To this end, let $\hat\xi=\hat\beta^{(\lambda)}-\bar\beta^{(\lambda)}$, and note that
On events $E_{\Delta}$ and $E_{\zeta}'$, we have $\|\hat\xi\|_2\le\xi_{\text{max}}$, where
By our assumptions on $n$ and $\alpha$, we have $\xi_{\text{max}}\le b/4$. Thus, defining $I=[-4\xi_{\text{max}}/b,4\xi_{\text{max}}/b]$, then $\pi^{\hat\beta^{(\lambda)}}(x)=\pi^{\bar\beta^{(\lambda)}}(x)$ for $x\not\in I$. As a consequence, we have
By the inequality $(a+b)^2\le2a^2+2b^2$, we have
Finally, by repeated application of the triangle inequality, we have
\paragraph{Winner's curse.}
Note that
so
Finally, we have $\pi^{\bar\beta^{(\lambda)}}(x)=\mathbbm{1}(x\ge0)$, so
The claim follows by the stability of $\mathcal{P}(\hat\beta^{(\lambda)};\hat\beta^{(\lambda)\prime})$ established above together with the fact that by our choice of $\beta^*$, for all $\beta\in\mathbb{R}^m$, we have $\mathcal{P}(\beta;\beta^*)=0$, so $\mathcal{R}(\beta;\beta')=\mathcal{P}(\beta;\beta')$ . \qed
This section provides the implementation details for our refugee matching simulation environment. Wherever possible, we calibrated parameters, sample sizes, and marginal distributions using those reported for the US site in the Supplemental Material of bansak2018improving.
Recent work proposes to use historical data to design algorithmic assignment rules that improve short-run employment relative to existing practice.
Using U.S. and Swiss data, bansak2018improving train machine learning models to predict employment for each refugee--location pair using features such as country of origin, language, age, gender, education, and placement restrictions. In their main specification, they fit separate prediction models by location and select gradient-boosted trees based on out-of-sample classification accuracy and calibration. They then:
To evaluate the algorithm, they run backtests: train on earlier arrivals, treat later arrivals with no placement restrictions as a test set, and compare predicted employment under the algorithmic assignment to the observed employment rate under the historical assignment. Reported gains are large, between 40--75% improvements in predicted employment relative to existing procedures.
Similarly, ahani2021placement develop Annie MOORE, an integrated machine learning and integer-optimization tool for a U.S.\ resettlement agency. They focus on “free cases”---families without U.S. ties who can be assigned flexibly across affiliates. Using administrative data on demographics, household structure, origin, language, health indicators, and macroeconomic conditions, they train predictive models of employment for refugee-affiliate pairs.
They compare pooled logistic regression, affiliate-specific logistic regression, LASSO-logit with hand-specified interactions, and gradient-boosted trees. In their test set, LASSO and gradient-boosted trees achieve substantially lower misclassification error and higher AUC than simpler models; LASSO is chosen as the primary model because it combines good discrimination with reasonable calibration.
Given these predictions, they:
These backtests suggest improvements in predicted employment on the order of 20--40%. In an e-companion, they also conduct a bootstrap analysis: they draw many bootstrap samples of the training data, re-estimate the LASSO model on each, and evaluate the same optimized allocation under each bootstrapped model. The resulting distribution of estimated gains is tight and positive, which they interpret as evidence of robustness.
The calibration and ROC curves reflect that each prediction model is well-calibrated and makes fairly accurate predictions.