Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
84,682 characters · 21 sections · 100 citation commands
Compound Selection Decisions: An Almost SURE Approach
\pagestyle{plain}
In settings with parallel noisy estimates for many parameters, researchers are often interested in selecting units with high values of the parameter. For these selection decisions, it is popular to screen based on shrunken estimates. Yet the objectives for shrinkage and for selection are misaligned: it is not obvious that good selection decisions must rely on good estimates, nor that good estimates lead to good selections manski2021econometrics. This paper directly studies these compound selection problems\footnote{In this paper, we refer to compound decision---often used interchangeably with “empirical Bayes” and “shrinkage” in the literature---specifically as decision-theoretic frameworks in which (i) the true parameters are fixed or conditioned upon and (ii) the loss function averages over units. Empirical Bayes additionally posits a random effect model for the distribution of the true parameters, and evaluates decisions by integrating over the distribution of the parameters. } and proposes methods that optimize selection decisions while still “borrowing strength” across different noisy estimates robbins1985asymptotically.
To introduce this procedure, we consider a standard setup for empirical Bayes and compound decision jiang2009general,efron2012large,jiang2020general,walters2024empirical,soloff2024multivariate,chen2022empirical. Researchers observe noisy estimates $Y_ {1:n} = (Y_1,\ldots, Y_n)'$ and standard errors $\sigma_{1:n}$ for unknown parameters $\mu_{1:n}$. Motivated by central limit theorems applied to the procedure generating $Y_i$, we model the signals as Gaussian with known variance: $ Y_i \mid \mu_i, \sigma_i, X_i \sim \Norm (\mu_i, \sigma_i^2).$ Here, $X_i$ are covariates that do not predict the noise $Y_i - \mu_i$. We seek binary decisions $a_i \in \br{0,1}$ so as to maximize the compound utility
The objective (ref) is the average payoff of selection decisions over units $1,\ldots,n$, where the payoff of selecting unit $i$ is $\mu_i - K_i$, for some known cost $K_i$. It can be viewed as a compound version of the treatment choice problem manski2004statistical, where normality is motivated by local asymptotic approximations hirano2009asymptotics,https://doi.org/10.3982/ECTA20364.
For concreteness,\footnote{This type of problem also appears in other economic applications such as the meta-analysis of experiments azevedo2020b, teacher value-added chetty2014measuring,kwon2023optimal,cheng2025optimal, identifying discrimination kline2022systemic, and treatment choice kitagawa2018should, athey2021policy,moon2025optimalpolicychoicesuncertainty. This problem also appears in the statistical literature under the name empirical Bayes testing with linear loss liang1988convergence, liang2000empirical, liang2004optimal, karunamuni1996optimal.} in bergman2024creating, $Y_i$ is the estimated economic mobility of a Census tract, estimated on Census microdata. $\mu_i$ is true economic mobility, defined as the population economic outcome of poor children in adulthood in the tract. We may imagine that the costs $K_i$ are equal to some value $K$ that represents the economic mobility level for which a social planner is indifferent between incentivizing a low-income family to move to a tract with mobility $K$ and not doing so. The objective (ref) then rewards selecting an above-breakeven tract, punishes selecting a below-breakeven tract, and weighs the rewards and penalties according to the distance to $K$. Since $Y_i$ are fixed effect estimates, their normality---similarly invoked in armstrong2022robust,mogstad2024inference,andrews2024inference, chen2022empirical---is motivated by the central limit theorem applied to regression specifications that generate them chetty2018opportunity.
Because the Gaussian family has monotone likelihood ratio karlin1956theory, admissible selection decisions are increasing in $Y_i$.\footnote{See (ref) for a formal result.} We thus identify selection decisions with threshold rules $\delta_i(Y_{-i})$ such that $a_i = \one(Y_i > \delta_i)$.\footnote{The threshold rules depend on $i$ through the contextual information $Z_i = (\sigma_i, K_i, X_i)$, which we suppress in the notation.} For a given threshold rule $\delta_ {1:n} (Y_ {1:n})$, define the compound welfare \[ W(\delta_{1:n}, \mu_{1:n}) = \frac{1}{n} \sum_{i=1}^n \E_{Y_i \sim \Norm(\mu_i, \sigma_i^2)} \bk{ \one(Y_i > \delta_i(Y_{-i})) \left (\mu_i - K_i \right) } \label{eq:welfare_expected} \numberthis \] as the expected utility, integrating solely over $Y_{1:n}$. This is the compound decision analogue of frequentist risk wald1950statistical. Optimal decisions should maximize (ref), but doing so is infeasible since welfare depends on the unknown parameters $\mu_{1:n}$. For estimation problems interested in minimizing MSE, one feasible approach is Stein's unbiased risk estimate stein1981estimation. SURE estimates $\frac{1} {n}\sum_{i=1}^n \E_ {Y_i \sim \Norm(\mu_i, \sigma_i^2)}[(\delta(Y_i) - \mu_i)^2]$ and produces decisions by optimizing the estimated objective over a parametrized class of decisions xie2012sure,kwon2023optimal.
Inspired by SURE, we construct a novel estimator $\hat W(\delta_{1:n})$ of $W(\delta_ {1:n})$. Unlike with MSE, it turns out that an exactly unbiased estimator for (ref) does not exist, but an almost unbiased one---whose bias decays exponentially in a tuning parameter---does. We thus term our approach { ASSURE}, for Almost SURE. This almost unbiasedness is an appealing and nontrivial property. One could instead estimate (ref) by (synthetically) sample splitting the estimates oliveira2024unbiased,chen2022empirical,ignatiadis2025empirical.\footnote{That is, $Y_{i+} = Y_i + \epsilon W_i$ and $Y_{i-} = Y_i - \frac{1}{\epsilon} W_i$ for independent $W_i \sim \Norm(0,\sigma_i^2)$ satisfy $Y_{i+} \indep Y_{i-} \mid \mu_i, \sigma_i^2$. Proposition 1 in chen2022empirical uses this coupled bootstrap idea to estimate (ref) for decisions that depend on $Y_{i+}$.} Doing so incurs bias that does not vanish as quickly as { \textsc{ASSURE}}, and thus we can view { \textsc{ASSURE}} as a further debiased version of this approach.
Like SURE, we then optimize this estimated welfare over a user-chosen class of decision rules $\delta \in \mathcal D$. For instance, $\mathcal D$ may be the class of decisions that screens on linear shrinkage rules \[\mathcal D_{\textsc{close-gauss}} =\br{\delta_i : \one (y > \delta_i) = \one\br{a(\sigma_i, X_i, K_i; \beta) y + b (\sigma_i, X_i, K_i ; \beta) > K_i}}\] for functions $a(\cdot), b (\cdot)$, parametrized by a finite-dimensional $\beta \in B \subset \R^d$, that depends on contextual information $ (\sigma_i, X_i, K_i)$ for some covariates $X_i$ weinstein2018group,chen2022empirical. Such restrictions are necessitated by familiar bias-variance concerns---optimizing over too large a class of decisions risks overfitting to the noise $\hat W - W$. In practice, these restrictions can come from external preference for simple decisions, similar to kitagawa2018should,sudijono2024optimizing,crippa2025regret, or from a benchmark (empirical) Bayesian model on the parameters $\mu_{1:n}$, similar to kwon2023optimal and cheng2025optimal. $\mathcal D_{\textsc{close-gauss}}$, for instance, can be motivated either by simplicity preferences or by a correlated random effects model in which $\mu_i \mid \sigma_i, K_i, X_i \sim \Norm (m(\sigma_i, K_i, X_i; \beta), s^2(\sigma_i, K_i, X_i; \beta))$, under which posterior means $\mu_i \mid Y_i, \sigma_i, K_i, X_i$ are linear in $Y_i$.\footnote{chen2022empirical refers to this random effects model as close-gauss.}
Decisions selected by { ASSURE} enjoy strong optimality guarantees. As the number of parameters diverges ($n \to\infty$), we show that the performance gap between the estimated optimal decision and the optimizer of (ref) within $\mathcal D$---i.e., regret---converges to zero at minimax optimal rates, up to log factors. Both the upper and lower bounds for regret are novel analyses. Applied to $\mathcal D_{\textsc{close-gauss}}$, since many popular shrinkage methods are nested within $\mathcal D_{\textsc{close-gauss}}$, decisions tuned by { ASSURE} would asymptotically improve over screening on these status-quo shrinkage methods. The bulk of the paper concerns Gaussian estimates, but similar arguments pertain to Poisson $Y_i$ as well; we present these related results as an extension.
Importantly, our analysis differs from and complements empirical Bayes analyses jiang2020general,gu2023invidious,walters2024empirical,kline2024discrimination,chen2022empirical by focusing on the compound objective (ref) rather than its empirical Bayesian counterpart \[W_{\text{EB}}(\delta_{1:n}) = \E_{\mu_{1:n} \sim P}[W(\delta_ {1:n}, \mu_{1:n})] \text{ for } \mu_{1:n} \mid \sigma_{1:n}, K_{1:n}, X_{1:n} \sim P.\] Empirical Bayes analyses treat $\mu_ {1:n}$ as random effects and integrates over their distribution; in contrast, our compound perspective is analogous to fixed effects in panel data dano2025binary. The benefit of treating $\mu_{1:n} \sim P$ as random is that the optimal selection decision screens on the posterior mean $\E_{P} [\mu_i \mid Y_i]$: Under $P$ and $W_{\text{EB}}$, there is no misalignment between estimation and selection decisions.
That benefit comes with two costs that our framework avoids. First, empirical Bayesians have to carefully model $P$. Failing to model this distribution well---either misspecifying how $\mu_i$ is predicted by the contextual information chen2022empirical or how $\mu_i$ correlates with $\mu_j$ bonhomme2024estimating---could harm the selection decision. The compound decision framework avoids these problems entirely by conditioning on $\mu_{1:n}$. Second, optimality for $W_ {\text{EB}}$ is with respect to a new draw of units $\mu_i \sim P$. In bergman2024creating, for instance, empirical Bayes evaluates performance over a randomly drawn new Census tract, whereas optimality with respect to (ref) keeps fixed the sample of Census tracts, and only imagines alternative draws of the estimation error in economic mobility.
To be sure, a SURE-based compound decision is limited to optimality within a class of decision rules $\mathcal D$. Empirical Bayes methods can match the performance of the optimal decision under $P$, without constraining to a class $\mathcal D$. That shortcoming of compound decisions can be overcome by selecting $\mathcal D$ via some working empirical Bayes model. When a model of $P$ implies a parsimonious class of decisions rules through $a(Y_i) = \one (\E_P [\mu_i \mid Y_i] > 0)$,\footnote{Posterior means with Gaussian $Y$ are monotone in $Y$, and thus this threshold rule can be equivalently written as $Y_i > \delta_i$.} optimizing among this class using { ASSURE} weakly improves over the empirical Bayes decisions under an estimated prior $\hat P$. Thus, our method can be combined with (parametric) empirical Bayesian modeling: it fine-tunes such a model without harming performance.
We illustrate \assure in three empirical applications: selecting census tracts to maximize economic mobility chetty2018opportunity,bergman2024creating, picking innovations in A/B testing azevedo2020b, and identifying discrimination in large firms kline2022systemic, kline2024discrimination. These applications show that \assure offers improvements and robustness over common plug-in empirical Bayes methods. The \assure-estimated welfare $\hat W$ additionally assesses whether a status-quo decision rule has reasonable welfare performance under assumptions on decision costs $K_i$. It also enables researchers to infer $K_i$ under an assumption that a compound decision-maker performs optimally.
This paper is structured as follows. (ref) defines the { ASSURE} estimator and discusses its bias and variance. (ref) presents several theoretical guarantees on selecting a decision using { ASSURE}, showing upper and lower bounds on regret. We next highlight extensions to other observation distributions like the Poisson and to complicated decision procedures in (ref). Simulation results and empirical applications are discussed in (ref) and (ref).
Let $Y_i \sim \Norm(\mu_i,\sigma_i^2), i \in [n] = \br{1,\ldots,n}$ with known $\sigma_i$, known costs $K_i$, with potentially related known covariates $ X_i$. The contextual information $Z_i := \left(X_i, K_i, \sigma_i \right)$ and the parameters $\mu_i$ are treated as fixed, and all probability and expectation statements are over the joint distribution $Y_i \sim \Norm (\mu_i,\sigma_i^2)$. For simplicity, we assume that $Y_i$ mutually independent across $i$.\footnote{If $ (Y_1,\ldots, Y_n) \sim \Norm((\theta_1,\ldots,\theta_n)', \Sigma)$ for some known $\Sigma$, then our arguments generalizable by studying the conditional distributions $Y_i \mid Y_{-i}$.} We also restrict to separable thresholds of the form \[ \delta_i(Y_{-i}) =\delta(Z_i; \beta). \] in all classes $\mathcal D$ that we consider. That is, the threshold for unit $i$ does not depend on information for unit $j \neq i$ (except through estimation of $\beta$).
Having parametrized thresholds with $\beta$, we can rewrite (ref) as
Optimizing (ref) directly is infeasible since it depends on the unknown $\mu_{1:n}$. However, if we had access to an estimator $\hat W_n(\beta)$, we can optimize $\hat W_n$ over a class of threshold rules $\mathcal D = \br{\delta (\cdot;\beta): \beta \in \mathcal B \subset \R^d}$ (e.g., $\mathcal D = \mathcal D_{\textsc{close-gauss}}$), yielding: \[ \hat\beta \in \argmax_{\beta \in \mathcal B} \hat W_n(\beta).\numberthis \label{eq:sure_type} \] Our subsequent theoretical results control the difference, or regret, between the welfare of the best decision rule in $\mathcal D$ and the expected utility achieved by $\hat\beta$: \[ \regret_n = \sup_{\beta \in \mathcal B} W(\beta) - \E_{\mu_{1:n}}[u(\hat\beta)]. \numberthis \label{eq:regret} \] Intuitively, for $\regret_n$ to be small, it is critical that our estimator $\hat W_n$ for welfare is accurate. A good welfare estimator for the compound selection problem is the focus of this paper.
We propose the following estimator for (ref) with good regret properties, building on the deconvolution literature kolmogorov1950unbiased,tate1959unbiased,liang2000empirical,pensky2017minimax,zhou2019fourier. Define the sinc, sine integral, and cumulative sinc functions as follows: \[ \sinc(x) = \frac{\sin x}{\pi x} \qquad \Si(x) = \int_0^x \frac{\sin t}{t} dt \qquad \Csinc(x) = \int_{-\infty}^x \sinc(t)\,dt = \frac{1}{2} + \frac{1}{\pi} \Si(x). \]
(ref) visualizes the function $w_h$ in (ref) for $K=0$ and $\sigma = 1$. The estimator $w_h$ is chosen due to its excellent bias properties: The following proposition shows that its bias decays exponentially in $1/h$.
Unlike SURE, { ASSURE} is biased. The bandwidth term $h_n$ is chosen to ensure the bias has negligible contribution towards mean squared error for estimating $W(\beta)$. For $W(\beta)$, it turns out that unbiased estimators (with reasonable growth behavior) do not exist: we formalize in (ref) an argument in stefanski1989unbiased.
The next subsection provides heuristic motivation for the functional form (ref) in { ASSURE} and connects it to SURE, sample splitting, and to a literature on estimation using Fourier transforms. The subsection after compares our SURE-type analysis (ref) to other approaches to compound selection problems, such as empirical Bayes.
Inspecting (ref), estimating $W(\beta)$ can be reduced to a Gaussian estimation problem: Given $Y \sim \Norm(\mu, \sigma^2)$ and $K, \delta \in \R$, we would like to estimate the parameter $(\mu-K) \Phi\pr{\frac{\mu-\delta}{\sigma}}$. Inspired by SURE, a natural starting point is Stein's identity: For differentiable $F(Y)$ with derivative $f(Y)$, \[ \E_{Y\sim \Norm(\mu, \sigma^2)}[(Y-\mu) F(Y)] = \sigma^2 \E_{Y\sim \Norm(\mu, \sigma^2)}[f(Y)]. \] We can rearrange to obtain \[ \E[(Y-K) F(Y) - \sigma^2 f(Y)] = (\mu-K) \E[F(Y)]. \numberthis \label{eq:stein_identity} \] Thus, $(Y-K) F(Y) - \sigma^2 f(Y)$ is an unbiased estimator for $(\mu-K) \E[F(Y)]$. To obtain a low-bias estimator for $(\mu-K) \Phi\pr{\frac{\mu-\delta}{\sigma}}$, we could find some differentiable $F$ whose expectation $\E[F(Y)]$ is approximately $\Phi \pr{\frac{\mu-\delta}{\sigma}} = \E[\one(Y > \delta)].$
Having reduced the problem in this way, by rescaling if necessary, we may assume $\sigma = 1, \delta = 0$ without loss of generality. A natural idea is to smooth the Heaviside function $H(y) := \one(y > 0)$ with some kernel $k_h(y) = \frac{1}{h} k\pr{\frac{y} {h}}$---for some kernel $k(\cdot)$ and bandwidth $h$---so that the resulting function is differentiable. That is, we may consider convolving $H$ with the kernel: \[ F_h(y) := (H \star k_h)(y) = \int_{-\infty}^{y/h} k(t) \,dt. \]
Which kernel should we choose? A natural but suboptimal idea is to use the Gaussian kernel $k (\cdot) = \varphi (\cdot)$. Doing so yields an estimator that has $O(h^2)$-bias \[ \E\bk{(Y-K) \Phi\pr{\frac{Y}{h}} - \varphi\pr{\frac{Y}{h}}} = (\mu-K) \Phi(\mu) + O (h^2). \numberthis \label{eq:coupled_bootstrap_form} \] Interestingly, this choice of kernel has a natural interpretation as coupled bootstrap oliveira2024unbiased,leiner2023data,ignatiadis2025empirical,chen2022empirical. For $Y \sim \Norm (\mu, 1)$, we can construct two independent samples by adding and subtracting an independent Gaussian noise $Q$ \[ Y_1 = Y + h Q \quad Y_2 = Y - \frac{1}{h} Q \quad Q \sim \Norm(0,1) \implies Y_1 \indep Y_2. \numberthis \label{eq:cb} \] This coupled bootstrap procedure is a synthetic version of sample-splitting.\footnote{If $Y$ is a sample mean of Normally distributed micro-data, $Y = \frac {1}{m} \sum_{j=1}^m Y_ {(j)}$ for $Y_ {(j)} \sim \Norm (\mu, m)$, then sample-split means over $Y_{(1)},\ldots, Y_{(m)}$ can be represented in the form (ref).} With (ref), the welfare of selection decisions based on $Y_1 > 0$ can be unbiasedly estimated by $(Y_2 - K)\one(Y_1 > 0)$, since $Y_2$ acts as fresh “testing data” that is independent of the “training data” $Y_1$. The estimator (ref) is exactly the Rao--Blackwellization of coupled bootstrap: \[ \E[(Y_2 - K)\one(Y_1 > 0) \mid Y] = (Y-K) \Phi\pr{\frac{Y}{h}} - \varphi\pr{\frac{Y}{h}}. \] The $O(h^2)$ bias, however, means that the regret rate for coupled bootstrap is $\tilde O(n^ {-4/5})$ ((ref)).\footnote{Throughout, we use $\tilde O(\cdot)$ to denote the analogue of big-O notation but ignoring logarithmic terms in $n$. }
Instead of a Gaussian kernel, the \assure estimator uses the sinc kernel for improved bias davis1975mean, Tsybakov2009IntroNonparametricEstimation. Observe that the expectation $\E[F_h]$ convolves $F_h$ with a Gaussian kernel: \[ \E_{Y \sim \Norm(\mu,1)}[F_h(Y)] = (H \star k_h \star \varphi) (\mu), \] whereas the target parameter is a convolution without $k_h$, $ \Phi(\mu) = (H\star \varphi) (\mu). $ Thus, a low-bias kernel is one that barely smoothes the Gaussian density: $k_h \star \varphi \approx \varphi$. The sinc kernel is motivated by this approximation in frequency space:\footnote{A large literature on Gaussian estimation and deconvolution follows similar ideas kolmogorov1950unbiased,zhou2019fourier,pensky2017minimax,tate1959unbiased,stefanski1989unbiased.} By taking the Fourier transform of both sides, we would like a kernel for which \[ \hat k_h e^{-\omega^2/2} \approx e^{-\omega^2/2} \text{, for } \hat k_h(\omega) := \int_\R k_h (t) e^{-it\omega}\,dt. \] The sinc kernel has Fourier transform $\one(|\omega| < 1/h)$, which only truncates the high frequency signals in $\varphi$. Since $\hat\varphi(\cdot)$ has Gaussian tails, this truncation leaves it essentially unchanged---allowing the bias to decay exponentially in $1/h$ ((ref)).
The identity (ref) is also suggestive of \assure's robustness to viewing $Y$ as approximately Gaussian---though we leave a formal analysis to future work. Let $Z \sim \Norm (\mu, \sigma^2)$ and suppose $Y$ has mean $\mu$ and variance $\sigma^2$, but is not necessarily Gaussian. Then
The term $A$ measures the discrepancy between the distributions of $Y$ and $Z$ in terms of $\E[F(\cdot)]$. This term is controlled with various weak convergence arguments, such as controlling the Wasserstein-1 distance between $Y$ and $Z$. The term $B$ measures the extent to which $Y$ does not obey Stein's identity. It often controlled directly by central limit theorems using Stein's method. As long as $F_h$ is sufficiently well-behaved so that $A$ and $B$ are small by the approximate Gaussianity of $Y$, we would like $\E[F_h(Z)] \approx \Phi(\mu)$, for which the sinc kernel is optimizing. Simulation studies in Sections (ref) and (ref) show robustness to the Gaussian assumption.
We now contextualize our approach in the SURE literature and compare it against a few alternatives. We also disambiguate our regret notion (ref) from the regret notions in these literatures.
\noindentStein's unbiased risk estimate. Our approach exactly mimics SURE: For $Y \sim \Norm (\mu, \sigma^2)$ and a differentiable estimator $\delta(Y)$, the squared error risk of $\delta$ is unbiasedly estimable by $T (Y; \delta, \sigma) = (\delta (Y) - Y)^2 + 2 \sigma^2 \delta'(Y) - \sigma^2$ stein1981estimation, so that \[ \E_{\mu}\bk{ T(Y; \delta, \sigma) } = \E_{\mu}[(\delta(Y) - \mu)^2]. \tag{SURE} \] Thus, for compound estimation problems---where one is interested in predicting $\mu_i$ using $\delta(Y_i; Z_i, \beta)$---one could form a risk estimator $\hat R_ {\mathrm{SURE}}(\beta) := \frac{1} {n} \sum_{i=1}^n T(Y; \delta(\cdot, \beta) , \sigma_i)$. This is the strategy pursued in xie2012sure,kwon2023optimal,cheng2025optimal. This research kwon2023optimal,cheng2025optimal often takes selection problems as motivation and illustrates that resultant shrinkage estimators $\hat\delta(Y)$ for $\mu$ lead to different selection decisions from naive selection rules. Yet, these shrinkage estimators are motivated by estimation---it is not obvious that they are simultaneously good for selection manski2021econometrics. Our contribution is to directly target the selection problem.
\noindentEmpirical Bayes. A popular framework for compound decision problems is empirical Bayes efron2012large, which models $\mu_i$ as random effects. This additional distributional structure characterizes optimal decision rules as posterior quantities involving the estimable distribution of $\mu_i$ liang1988convergence, liang2000empirical, liang2004optimal, karunamuni1996optimal,gupta2005empirical,gu2023invidious,kline2024discrimination. However, the additional structure hinges on modeling and estimating the distribution of $\mu_i$ well, which can make subsequent procedures and guarantees less robust.\footnote{A similar concern---about the robustness of correlated random effect models---motivates the literature of fixed effect nonlinear panels dano2025binary.} In particular, our procedure can be viewed as a compound analogue of liang2000empirical, which robustifies empirical Bayes selection.
Our procedure is also complementary to empirical Bayes. Empirical Bayes-style modeling for $\mu_i$ is powerful at generating classes of decision rules $\mathcal D$ that are reasonable---indeed optimal if the model on $\mu_i$ happens to be correct. Choosing a member of $\mathcal D$ with { ASSURE} provides guarantees for the compound selection problem directly---thus not requiring researchers to take the random effects model fully seriously. At the same time, if the random effect model is correct, tuning via { ASSURE} does little harm.\footnote{For concreteness, a model for $\mu_i$ that performs well in the empirical exercise in chen2022empirical is \[ \mu_i \mid Z_i \sim \Norm(m_0(Z_i; \beta), s_0^2(Z_i; \beta)), \] where $m_0, s_0$ are flexibly parametrized by $\beta$. Under this prior,
This motivates using { ASSURE} to choose among the following threshold decision rules, derived from inverting $\E[\mu_i \mid Y_i, Z_i] > K_i$ in $Y_i$ :
}
From a formal perspective, the empirical Bayes literature jiang2020general,soloff2024multivariate,chen2022empirical often considers the difference in $W_{\mathrm{EB}}$ between the oracle Bayes decision rule and the empirical Bayes decision rule:\[ \mathsf{EBRegret}_n = \E_{\substack{\mu_{1:n} \sim P \\ Y_i \mid \mu_i, Z_i \sim \Norm (\mu_i, \sigma_i^2)}} \bk{\frac{1} {n}\sum_ {i=1}^n \max(0, \E_P[\mu_i \mid Y_i, Z_i] - K_i) - u(\hat P) }, \] where $u(\hat P) = \frac{1}{n} \sum_{i=1}^n \one(\E_{\hat P}[\mu_i \mid Y_i, Z_i] > K_i) (\mu_i - K_i)$. Relative to the regret criterion (ref) that we subsequently study, $ \mathsf{EBRegret}_n$ (i) takes another expectation under $\mu_{1:n} \mid Z_{1:n} \sim P$, (ii) chooses the benchmark as the oracle decision rule $\one(\E_P[\mu_i \mid Y_i, Z_i] > K_i)$, and (iii) typically considers the class of posterior decision rules indexed by an estimate of $P$. When the class of decision rules $\mathcal D$ is exactly the class of posterior thresholding decisions, controlling (ref) automatically controls $\mathsf{EBRegret}_n$: \[ \mathsf{EBRegret}_n = \E_{\substack{\mu_{1:n} \sim P \\ Y_i \mid \mu_i, Z_i \sim \Norm (\mu_i, \sigma_i^2)}}\bk{W(P) - u(\hat P)} \le \E_{\mu_{1:n}\sim P}\bk{\regret_n} = \E_{\mu_{1:n}\sim P}\bk{\eqref{eq:regret}}. \]
\noindentTreatment choice and policy learning. The decision problem (ref) is a compound version of the “testing an innovation” decision problem in manski2009identification, where $\mu_i$ represents the average treatment effect of a new treatment relative to a known status quo. To connect to the treatment choice literature manski2004statistical,stoye2012minimax,kitagawa2018should,athey2021policy, the compound aspect may represent $n$ parallel innovations that the decision-maker is simultaneously entertaining, or it may represent discrete covariates taking on $n$ values, wherein $\mu_i$ is the conditional average treatment effect on covariate value $i$.\footnote{Though, from this perspective, the objective function (ref) assumes that each covariate cell is equally probable. } The normality assumption on $Y_i$ corresponds to the limit experiment in hirano2009asymptotics, though rigorously justifying this approximation through Le Cam--Hajek-style arguments remains important future work https://doi.org/10.3982/ECTA20364.
Our statistical perspective is distinct from the treatment-choice literature. First, we aim for decision rules whose regret (ref) converges to zero as $n\to\infty$, uniformly over sequences $(\mu_i, Z_i)_{i=1}^n$. These decision rules are not necessarily minimax-$\mathsf{FreqRegret}$ for any finite $n$, for $\mathsf{FreqRegret}$ defined as follows ishihara2022shrinkage: \[ \mathsf{FreqRegret}_n = \E_{\mu_{1:n}}\bk{ \frac{1}{n} \sum_{i=1}^n \max(0, \mu_i - K_i) - u(\hat\beta) }. \] This means that \assure may suffer from higher worst-case $ \mathsf{FreqRegret}$ for any finite $n$. Conversely, minimax-$\mathsf{FreqRegret}$ decision rules, like the empirical success rule $a(Y_i) = \one(Y_i > 0)$ stoye2012minimax, have non-vanishing worst-case regret $\sup_ {\mu_ {1:n}} \regret_n = \Omega (1)$---meaning that these decisions can be substantially improved from the perspective of (ref). See (ref) for further comparisons and intuition.
Second, our asymptotic regime is different from the policy learning literature kitagawa2018should,athey2021policy,mbakop2021model. If a unit $i$ is taken to be a covariate cell, we consider a regime where the number of covariate cells grow large, perhaps due to finer discretization of continuous covariates. In contrast, the standard asymptotic regime in the policy learning literature considers increasing sample size within each covariate cell, which can be approximately written as $\sigma_i^2 \rateeq c_i/m$ for $m \to \infty$. Thus the regret rates we obtain are not directly comparable to the regret rates in kitagawa2018should.
This section presents statistical guarantees on { ASSURE}. The main results are upper and matching lower bounds for expected {regret} (ref). We first present general upper and lower bounds for $\tilde O(1/\sqrt{n})$ regret. Additional assumptions on $\mu_{1:n}$ allow for a faster rate by exploiting an analogue of the margin condition audibert2005fast.
Our main result is the following $(1/\sqrt{n})$-regret bound.
Thus, assuming that $s_2,m_2,\nu_4$ are bounded and the envelope function $D$ is controlled, the procedure achieves $\frac{\log n}{\sqrt{n}}$ regret in general. Constants in the regret rate depend on various norms of the parameters. If all parameters are bounded, we have the following corollary, showing a simple $\log n /\sqrt{n}$ rate.
(ref) builds on a standard argument in empirical risk minimization:\footnote{ See, e.g., 8.4.3. in vershynin2009high. Regret guarantees (ref) are related to, but are distinct from, excess risk control in empirical risk minimization. Standard empirical risk minimization (e.g., Theorem 8.4.4 in vershynin2009high) controls---in our notation---the welfare gap $W (\beta^*) - \E[W(\hat\beta)]$. This is a different quantity than (ref) since $\E[W (\hat\beta)]\neq \E[u(\hat\beta)]$. $\E[W (\hat\beta)]$ imagines evaluating $\hat\beta$ on a new draw of $Y_i$, whereas $\E[u (\hat\beta)]$ evaluates $\hat\beta$ on the same sample of draws $Y_{1:n}$.} from the optimality of $\beta^*$ and $\hat\beta$,
The $1/\sqrt{n}$-rate follows from analyzing the empirical processes on the right-hand side---in particular, $\beta \mapsto \hat W (\beta) - W(\beta)$. (ref) shows that $|\hat W (\beta) - W(\beta)| = O_P(\sqrt{\log n} / n)$ pointwise. Empirical process arguments control this uniformly at the price of an additional $\sqrt{\log n}$ factor.\footnote{The VC subgraph assumption is standard but not crucial, as long as an appropriate covering number of the decision class is controlled. We state our results in terms of the VC subgraph dimension for simplicity.} We modify standard empirical process theory to accommodate independent but not identically distributed data in our setting.
We now give several examples of function classes with finite VC dimension.
(ref) is unimprovable, up to log factors, in the worst case.
Here, the regret notion in (ref) compares the welfare of an arbitrary decision $a_i(Y_1,\dots,Y_n)$ to the expected welfare of the best decision in the thresholding class restricted to $|\beta| \leq M$. Lower bounds for this quantity thus imply lower bounds when we allow the oracle to be more flexible than using $\delta = \beta$.\footnote{Since the setting in (ref) does not have contextual information, the optimal separable decision rule does take the threshold form. An analogous regret notion was considered by polyanskiy2021sharp in their Eq. (12) for the mean square case.}
(ref) is derived by considering cases in which, for some $h > 0$, either $\mu_ {i} = h/ \sqrt{n}$ for all $i$ or $\mu_i = -h/\sqrt{n}$ for all $i$. In this case, the optimal $\beta^*$ is either $+M$ or $-M$, meaning that $W(\cdot)$ is maximized on the boundary instead of at a local maximum. In fact, when $\beta^*$ is a local maximum, the upper and lower bounds can be improved to $\tilde O(1/n)$.
In particular settings, the performance for { ASSURE} can be better than the $\tilde{O}(n^{-1/2})$ rate predicts. Technically, this is due to the fact that the derivatives of { ASSURE} also estimate the derivatives of welfare (ref) as shown in (ref). This analysis shares similarities with the kernel estimator introduced by liang2000empirical for the random effect case.
The assumptions needed for the improved rate are reminiscent of the margin condition used to achieve fast rates in classification and treatment choice audibert2005fast,audibert2007fast,kitagawa2018should,ponomarev2024lower,crippa2025regret. There, the margin condition bounds the density of difficult examples near the optimal classification boundary. Intuitively, with many units near the classification boundary, it is difficult to distinguish between different candidate parameter values, meaning that the objective function lacks a cleanly separated maximizer. This is exactly the failure mode that (ref) exploits for the lower bound, which we rule out in (ref).
For simplicity in the proof, we will restrict to the $d=1$ case which covers a range of examples. We leave the multidimensional extension to future work. Suppose the following conditions hold.
For convenience in the proof, we additional conditions on the decision rules and search space.
(ref) in (ref) illustrates a class of decision rules and lower-level assumptions on $\mu_{1:n}$ for which (ref) applies. A heuristic argument for (ref) is as follows. If $\hat\beta$ is a local maximum of $\hat W$, then from a first-order Taylor expansion,
It can be shown that $\hat W'$ estimates $W'$ uniformly at rate $\tilde O(1/\sqrt{n})$, from which it follows that $\hat\beta - \beta^* = \tilde O_P(1/\sqrt{n})$. Another Taylor expansion $W (\hat\beta)\approx W (\beta^*) + W''(\beta^*)(\hat\beta - \beta^*)^2$ shows that the welfare gap is $\tilde O (1/n)$: $W (\hat\beta) - W(\beta^*)= \tilde O_P(1/n)$. Finally, we consider a leave-one-out stability argument to control $\E[W(\hat\beta) - u(\hat\beta)]$.\footnote{The log factors are likely not tight and removable with a more refined application of empirical process theory. }
Finally, we present a matching lower bound for the fast rate.
The intuition for (ref) is to lower bound the compound regret by the Bayes regret for a well-chosen prior and then further lower bound the Bayes regret by Le Cam's two point argument.\footnote{Section 3.1 of liang2004optimal states a $\Omega(n^{-1})$ lower bound for a related but distinct regret quantity, where the decision rule is learned from data points $Y_1,\dots,Y_n$ and evaluated on a newly drawn decision problem $(\mu_{n+1},Y_{n+1})$. We are concerned with an “in-sample” version of liang2004optimal's regret, motivating a distinct proof strategy.} We will show that the Bayes regret for this problem can be interpreted as a type of weighted classification loss, building on an insight by polyanskiy2021sharp. As a result, (ref) may be of independent interest for empirical Bayes testing with linear loss.
The approach---finding an estimator for $W$ and optimizing this estimator over decisions---can be extended to settings where the observation distribution is non-Gaussian. For certain exponential families like the Poisson distribution and exponential distribution, one can find an exactly unbiased estimate for the welfare $W$. The intuition for these cases is to exploit a formula akin to Tweedie's formula. We focus on the Poisson setting as it frequently appears in empirical Bayes and compound decisions.\footnote{See Chapter 6 of efron2021computer for an overview of classic applications of the Poisson model to insurance claims, the missing species problems, and medical applications. See jana2022optimal, jana2023empirical for modern methods to this problem and also montiel2021empirical for econometric applications.}
In this setting, let $\mu_i \geq 0$ and $Y_i \sim \Poi(\mu_i)$. Again, associated with each decision problem is a cost $K_i$ and auxiliary information $Z_i$.\footnote{For example, analogous to $\sigma_i$ in the Gaussian case, $Z_i$ might include the known length $t_i$ of the observation window for the Poisson outcome $Y_i$. If the unknown parameter $\mu_i$ denotes the Poisson rate per unit time, then $Y_i \sim \Poi(\mu_i t_i)$. The decision rule might incorporate the observation window $t_i$. } Consider any integer-valued decision rule $\delta(Z_i;\beta)$ parametrized by $\beta$. Then the estimator
is the analog for { ASSURE}.
Under mild conditions, { ASSURE} achieves $1/\sqrt{n}$ regret, which is optimal. In contrast to the Gaussian case, the discreteness of the Poisson distribution suggests that there is no fast-rate regime.
We prove the upper bound by appealing to the similar empirical process theory arguments as in the Gaussian case. We suspect some of the assumptions of (ref) may be further relaxed by. Mirroring the argument of the Gaussian case, we have a matching lower bound. Proofs of these results are in (ref).
We discuss how to extend the { ASSURE} framework to accommodate complex, non-separable decision rules $\delta(\beta,Z_{1:n},Y_{-i})$ which may depend on the outcomes and auxiliary information of other units. For example, an analyst might consider making decisions using empirical Bayes posterior mean estimates $\E_{\hat{G}}[\mu_i \mid \sigma_i, Y_i]$ and flexible machine-learning models $\hat{f}(X_i)$ that estimate the regression $\E[Y_i\mid X_i]$. In both of these cases, it would be desirable to estimate $\hat{G}$ and model $\hat{f}$ using the data itself.
By a similar argument to (ref), one can create a near unbiased estimator of the welfare of this decision rule using a leave-one-out cross fitting construction. Redefine the { ASSURE} summand (ref) with the decision rule $\delta(\beta,Z_{1:n},Y_{-i})$, so that \[ w_h(Y_i; \beta,Z_i,Y_{-i}) := (Y_i - K_i) \Csinc\pr{ \frac{Y_i - \delta(\beta,Z_{1:n},Y_{-i})}{\sigma_i h} } - \frac{\sigma_i}{h} \sinc\pr{ \frac{Y_i - \delta(\beta,Z_{1:n},Y_{-i})}{\sigma_i h} }. \] Then,
which can be interpreted as the average welfare of using the non-separable decision rule. (ref) shows that the bias in the above approximation is small.
Given this observation, a simple construction is the following. Let $m_{\hat G_1}^{(-i)}(y,\sigma)$ be a posterior mean function obtained by applying an empirical Bayes method on the data $(Y_{-i},\sigma_{-i})$. Let $m_{\hat G_2}^{(-i)}(y,\sigma)$ denote the same for a different empirical Bayes method. Finally, let $\widehat{f}^{(-i)}$ denote a machine learning model trained using the data $(Y_{-i},X_{-i})$. Consider the class of decision rules which selects a unit whenever the ensemble prediction of these models is greater than the implementation cost:
where the parameters $\beta = (b_1,b_2,b_3)$ are restricted to be in the unit simplex. Extensions to multiple empirical Bayes posterior mean models or machine learning models is straightforward. By the monotonicity of the posterior mean function, the left hand side of (ref) is increasing in $Y_i$: the decision is equivalent to $Y_i \geq \delta(\beta,Z_{1:n},Y_{-i})$ for some function $\delta$. As a result, { ASSURE} can be used to tune the parameter $\beta$. Under mild stability assumptions and for large enough samples, { ASSURE} is expected to perform no worse than the constituent models since the decision class nests each model individually. This ensemble class thus immediately enables { ASSURE} to improve on any fixed empirical Bayes decision method and illustrates the utility of the framework. We defer a detailed theoretical analysis of the ensembling method for future work.
This section conducts two calibrated simulations where the $\mu_i$'s are constructed by sampling from some estimated empirical Bayes model. In each draw of the simulation, we then sample $Y_i \sim \Norm(\mu_i ,\sigma_i^2)$. Code to replicate these simulations and the empirical applications in the next section may be found \href{https://github.com/tsudijon/CompoundWelfareMaximization/tree/main}{on Github}.
We first consider a calibrated simulation based on the Opportunity Atlas (OA) dataset from chetty2018opportunity,bergman2024creating. See (ref) for further background on the underlying dataset. The data for this simulation comes from a simulation exercise of chen2022empirical and is generated by sampling $\mu_{1:n}, n \approx 10^4,$ from a Monte Carlo sample of an empirical Bayes prior fitted on the OA dataset. We compare several classes of methods by evaluating the in-sample welfare obtained by the chosen decision using the known $\mu_i$. In particular, we compare (a) the linear shrinkage class in Eq. (ref), (b) the Fay--Herriot class of Eq. (ref), and (c) the { CLOSE-GAUSS} decision class of Eq. (ref). In each of these three classes, we compare (i) empirical Bayes with plug-in estimates of prior parameters against the decisions selected by (ii) { ASSURE} and (iii) coupled bootstrap (ref). In addition we include two { NPMLE} methods, and finally an ensemble tuned using { ASSURE} as in (ref). Complete details on the simulation setting are given in (ref), alongside further numerical results.
(ref) presents our results. For the linear shrinkage and Fay-Herriot classes, { ASSURE} consistently produces decisions that improve over the respective empirical Bayes baselines. These improvements highlight that \assure robustifies empirical Bayes procedures, improving performance when empirical Bayes models are misspecified. Indeed, these respective empirical Bayes methods assume that the $\mu_i$ are drawn from Gaussian priors with a fixed constant variance. This assumption is questionable in the present dataset, as can be seen in (ref). When the model is misspecified, { ASSURE} selects a better-performing set of parameters within the class. This gain can be significant, as seen in the linear shrinkage class.
For empirical Bayes models which are fairly well-specified, { ASSURE} does not do harm and performs comparably to the respective empirical Bayes method, as shown by the comparison between methods {\footnotesizeCLOSE-GAUSS} and the { ASSURE}-chosen decision in the class. As expected, ensembling produces decisions which outperform {\footnotesizeCLOSE-GAUSS} and {\footnotesizeCLOSE-NPMLE}. Since the covariates in this setting are predictive of the true effects, we expect the ensemble method to strictly outperform the empirical Bayes counterpart. Across these simulations, { ASSURE} and coupled bootstrap have similar performance. Though the theory for { \textsc{ASSURE}} suggests a better rate of $\tilde O(n^{-0.5})$ as opposed to $\tilde O(n^{-0.4})$, the difference appears negligible in practice unless $n$ is very large. However, a small difference might be economically significant, especially for large $n$. The performance of our methods appear robust to the normality assumption, as shown in (ref). (ref) details additional specifications where both normality is misspecified and the covariates are uninformative.
Next, we present results on a semisynthetic dataset derived from the experimentation program dataset of Section (ref). The specification is somewhat challenging due to smaller size of the dataset ($n \approx 330$) and the large heterogeneity in the $\sigma_i$, as seen in (ref). On this dataset, we evaluate three classes of decision rules: (a) the $t$-statistic/$p$-value threshold rules given in (ref), (b) linear shrinkage rules, and (c) the Fay-Herriot decision rules. We also consider the performance of standard nonparametric maximum likelihood jiang2020general.
(ref) shows the results. Importantly, both { ASSURE} and the coupled bootstrap method consistently identify a better choice of $t$-statistic threshold than the fixed decision rule corresponding to $p < 0.05$. The performance of both { ASSURE} and coupled bootstrap are comparable to the corresponding empirical Bayes methods in the linear shrinkage class and Fay--Herriot class, but somewhat more variable due to the smaller size of this dataset. In contrast to the data from the Opportunity Atlas in (ref), the empirical Bayes assumptions underlying the plug-in methods are not unreasonable for this dataset, and thus { ASSURE} does not provide improvements. Nevertheless, even with a good empirical Bayes model, applying { ASSURE} does not significantly harm performance.
The Opportunity Atlas chetty2018opportunity provides census-tract level estimates of economic mobility as measured by a suite of children's outcomes in adulthood. For each tract, multiple economic indicators such as earnings and incarceration rates are estimated by strata such as parental income, race, and sex. Building on the Opportunity Atlas dataset, bergman2024creating conducted the randomized trial Creating Moves to Opportunity (CMTO), offering low-income families housing vouchers and assistance to move to neighborhoods with historically high upward mobility. Implicit in defining highly upward-mobile tracts is a compound selection problem where the goal is to select the top 1/3 upwardly-mobile tracts. bergman2024creating and chen2022empirical approach this selection problem using empirical Bayes.
We consider the related selection task as defined in (ref) on a particular economic mobility outcome pictured in (ref), where the mobility measure is the average adult income rank of children whose parents are at the 25th percentile of income. For each Census tract $i$, we decide whether or not to select it as high-mobility.
We consider the class of linear shrinkage decision rules given in (ref) with specifications of $K_i$ to be discussed below. Recall that this class of decisions can be motivated by assuming that $\mu_i \sim \mathcal{N}(m_0, s_0^2)$, popular in many empirical economic contexts walters2024empirical, which motivates the decision class:\footnote{With constant costs, the linear shrinkage decision class simplifies to this form.}
We compare fitting the plug-in empirical Bayes estimator discussed in (ref) to the { ASSURE} chosen decision in the class (ref). We consider three settings for costs $K_i$, which will be assumed constant for simplicity, given by $\set{0.2, 0.361, 0.369}$. The first specification corresponds roughly to the treatment cost in bergman2024creating. The latter two choices are different specifications which emulate the top 1/3 selection exercise in Bergman. Details are given in (ref).
(ref) shows the results. For $K = 0.2$, the majority of census tracts likely have a parameter $\mu_i$ which is greater than the cost $K = 0.2$, so the optimal decision is to select more tracts. In this case, the empirical Bayes plug-in and { ASSURE} chosen decision are comparable in performance. In the two other specifications $K \in \set{0.361, 0.369},$ the cost tradeoffs are more meaningful. The { ASSURE} welfare estimate suggests that the empirical Bayes plug-in is meaningfully suboptimal compared to the optimal decision, which takes $\beta$ slightly less than zero. (ref) visualizes the difference, where { ASSURE} selects more aggressively.
This example shows that the { ASSURE} estimator has utility beyond just selecting a near-optimal decision through a black box. Consider again the left-hand side of (ref). The estimated curve suggests that decisions where $\beta \leq 0$ are all roughly equivalent in terms of expected welfare. Policy-makers can then choose a decision amongst this near-optimal subset on the basis of other criteria, such as to maximize another metric, or for interpretability. { ASSURE} therefore enables decision-makers to certify a status-quo decision if the estimated welfare is near-optimal. In other settings, it may also reveal that a status quo decision might be far from optimal, as discussed in the next empirical application.
In the technology industry, experimentation programs are large collections of related A/B tests where the goal is to move a metric of interest azevedo2020b, sudijono2024optimizing, chou2025evaluating. By explicit randomization, the treatment effect estimates $Y_i$ are plausibly normally distributed around their true unknown average treatment effects $\mu_i$ goldstein2007l1,li2017general. $\mu_i$ are measured in some metric of practical business interest. We take a dataset of anonymized treatment effects from Netflix sudijono2024optimizing, consisting of treatment effects and standard deviations for $331$ feature experiments. (ref) shows some descriptive statistics for the data.
Our goal is to select a subset $S$ of innovations $i$ to implement in the platform. azevedo2020b explicitly introduce the problem of choosing a subset of indices $i$ in order to maximize the welfare (ref) in the Bayes setting where $\mu_{1:n} \iid G$. Past work typically takes the empirical Bayes approach where $G$ is estimated from large repositories of past A/B tests. In some cases, the Bayes assumption may be too strong. For example, innovations are often created as variants of the same idea, suggesting correlation between the returns to innovations. In new experimentation programs, no past data may be available to estimate $G.$ In the absence of a trustworthy prior, the compound decision formulation is a compelling alternative.
The status quo decision procedures in the industry is the $t$-statistic decision class of (ref). Traditionally, the chosen threshold corresponds to a one-sided $p$-value of $0.05$. We use { ASSURE} to select the threshold $\beta$ that maximizes the expected welfare of the decision $W(\beta) := \sum_{i=1}^n (\mu_i - K_i) \Prob(Y_i \geq \beta\sigma_i)$ with $K_i = 0.$ (ref) shows the { ASSURE} estimate $\widehat{W}(\beta)$ as a function of the curve $C$. By maximizing the curve $\widehat{W}(\beta)$, { ASSURE} suggests to use a decision corresponding to $C \approx 0$.
The data in this application are anonymized: both the treatment effect estimates and standard deviations are obfuscated by random multiplicative factors. Thus, the optimal decision cannot directly be interpreted, though we hope this section serves as a template for the reader to conduct similar exercises on their own datasets. Despite this, the shape of the estimated welfare curve suggests that a lenient threshold is more warranted than a conservative one, questioning the optimality of the industry-standard $p < 0.05$ decision rule. Other papers in the experimentation literature find similar conclusions under different frameworks azevedo2020b, berman2022false, sudijono2024optimizing.
As reflected in the literature, the suboptimality of $p < 0.05$ in the compound setting is not surprising when $K=0$. In general, measuring costs for A/B tests is a difficult firm-specific exercise. Often costs are implicit, reflecting a host of factors including upfront costs of implementing and increasing future complexity of the platform. A useful exercise is to back out the implicit costs that justify the $p < 0.05$ decision rule and to use this as a heuristic benchmark when choosing a $p$-value threshold in practice. As a start, we assume that costs associated with each decision are given by $\kappa$ times the median treatment effect size across the dataset, for various levels of $\kappa$. The left panel shows the estimates $\widehat{W}(\beta)$ for various levels of $\kappa$. As $\kappa$ grows, the optimal decision generally becomes more conservative. The right panel of (ref) shows the { ASSURE}-optimal $t$-statistic threshold as a function of the cost multiplier. The true optimal threshold is generally non-decreasing.\footnote{The estimated threshold in (ref) is decreasing due to finite-sample effects and the form of the { ASSURE} estimator. The discontinuous jumps are however also present for the true optimal thresholds.} Crucially, (ref) suggests that the $p < 0.05$ can be rationalized by costs which are roughly twice the median treatment effect.
In this empirical illustration, we demonstrate how selection decisions at the individual level can differ when using ASSURE versus thresholding on nonparametric EB estimates. kline2022systemic conducted a large correspondence experiment that sent up to 1,000 job applications to each of 108 firms. Here $\mu_i$ is the firm contact gap and $Y_i$ is the average difference in callback across pairs of resumes within a firm. Specifically, each firm in the sample had approximately 125 job listings throughout the experiment. For each job, applications were sent in pairs, one with a distinctly white-sounding name and the other with a distinctly black-sounding name. Other attributes in the resume were randomized. The difference in callback rates between races is then averaged over jobs to create $Y_i$. Due to the explicit paired randomization, the assumption of normality---similarly made in kline2022systemic---is fairly justified.
Based on an auditor with an “extensive margin” preference, kline2022systemic found 23 firms to discriminate against black applicants, controlling false discovery rates to the 5% level. To adapt their setting for our illustration, we instead consider an auditor with an “intensive margin” preference, whose objective is to maximize the expected welfare gain from an investigation selection rule $S$ given by $\sum_{i=1}^n (\mu_i -K) \Prob\set{i \in S}$ where $K$ is the cost of an investigation. Using the nonparametric EB estimates of kline2022systemic, adapted from efron2016empirical, a Bayesian interpretation implies setting $K=0.025$ to match with the selection of 23 firms,
To demonstrate how ASSURE can make different selections in practice, we consider the linear shrinkage class with constant cost $K=0.025$ as specified in Eq. (ref).
As shown in Figure (ref)(b), comparing with the nonparametric EB estimates of kline2022systemic, the selections by { ASSURE} are largely concordant, including a few more but excluding a firm with noisy estimate. As argued before, the linear shrinkage class can also be motivated with a parametric normal prior. Figure (ref)(a) shows that directly plugging in estimates for $(m_0, s_0)$ would give a higher threshold and select fewer firms than { ASSURE}.
In this work, we introduce { ASSURE}, a data-driven approach for compound selection decisions. By using a near-unbiased estimate of the welfare, { ASSURE} targets an instance-specific optimal decision while making only mild assumptions on the unknown parameters $\mu_{1:n}$. This robustifies empirical Bayes procedures where it may be difficult to justify distributional assumptions on $\mu_i$. Section (ref) discusses the pertinence of the Bayes assumption in our empirical applications. Another strength of { ASSURE} is its flexibility and ability to validate the performance of a wide range of decision procedures which may involve covariates and other auxiliary information. The construction of Section (ref) allows decision-makers to improve on any black-box empirical Bayes decision by ensembling. In applications with large $n$, the { ASSURE}-optimal decision procedure consistently improves over plug-in empirical Bayes estimates. At the same time, the { ASSURE} estimator is also a useful tool for heuristic decision making: the nearly-unbiased welfare estimates produced by { ASSURE} allows a policy-maker to screen out clearly poorly-performing decision procedures or validate a fixed decision rule. For these theoretical and empirical reasons, we envision { \textsc{ASSURE}} to be a useful tool for selection decisions in practice.