EconBase
← Back to paper

Estimating the Value of Evidence-Based Decision Making

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

54,457 characters · 0 sections · 20 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

\xdef\@thefnmark\@footnotetext{Alberto Abadie, Department of Economics, MIT, [email removed]. Anish Agarwal, Department of Industrial Engineering and Operations Research, Columbia University, [email removed]. Guido Imbens, Graduate School of Business, Stanford, [email removed]. Siwei Jia, Amazon.com, [email removed]. James McQueen, Amazon.com, [email removed]. Serguei Stepaniants, Amazon.com, [email removed]. Santiago Torres, MIT, [email removed]. We are grateful to Haruki Kono, Soonwoo Kwon, Vira Semenova and seminar participants at Amazon.com, Brandeis, Princeton, and Toronto for helpful comments. The views expressed in this article are solely those of the authors and do not necessarily reflect those of Amazon.com. This research is partially supported by ONR grants N00014-24-1-2687 (Abadie) and N00014-19-1-2468 (Imbens).} \vskip 20pt \centerline{\bf Estimating the Value of Evidence-Based Decision Making}

center[center omitted — 466 chars of source]
center[center omitted — 43 chars of source]
quote{In an era of data abundance, statistical evidence is increasingly critical for business and policy decisions. Yet, organizations lack empirical tools to assess the value of evidence-based decision making (EBDM), optimize statistical precision, and balance the costs of evidence-gathering strategies against their benefits. To tackle these challenges, this article introduces an empirical framework to estimate the value of EBDM and evaluate the return on investment in statistical precision and project ideation. The framework leverages parametric and nonparametric empirical Bayes methods to account for parameter heterogeneity and measure how statistical precision changes the value of evidence. The value extracted from statistical evidence depends critically on how organizations translate evidence into policy decisions. Commonly used decision rules based on statistical significance can leave substantial value unrealized and, in some cases, generate negative expected value.}

{0.5\baselineskip}

\@startsection{section}{1}{0mm}{-\baselineskip}{0.25\baselineskip}{\normalfont\bf}{Introduction}

Many organizations use randomized experiments and observational studies to improve their decision making. For example, gupta2019top write “Together [Airbnb, Amazon, Booking.com, Facebook, Google, LinkedIn, Lyft, Microsoft, Netflix, Twitter, Uber, Yandex, and Stanford University] have tested more than one hundred thousand experiment treatments last year.” The fact that so many organizations conduct such a large number of experiments suggests that these organizations believe that data evidence provides significant value in guiding business and policy decisions. However, we are unaware of empirical tools that organizations can use to assess the actual value of their EBDM practices. In the absence of such tools, it is difficult to determine whether too much experimentation is being conducted or too little, whether experiments are too large or too small, and whether the right experiments are being undertaken. Part of the challenge in evaluating the value of EBDM lies in the need to describe the role of evidence in the business and policy decision-making process. In other words, estimating the value of EBDM requires assumptions about what organizations will do with and without various amounts of evidence, which they can choose to generate at some cost.

In this article, we propose an empirical Bayes estimator for the value of EBDM. We study the problem of a decision maker choosing whether to adopt a particular policy intervention. We use the term “agent” to refer to the decision maker, and “policy,” “intervention,” and “treatment” interchangeably to refer to the policy intervention under scrutiny. The agent can implement the intervention based on prior information or gather additional information at some cost—for example, by running an experimental or observational evaluation of the intervention's effect. At the stage where the agent decides whether to implement the intervention, they aim to maximize utility based on the available information. We derive expressions for the value of additional information and demonstrate how to estimate this value using metadata on estimates of the effects of business and policy interventions, along with their standard errors.

Our framework allows decision makers to assess in a principled way the value of experimental and non-experimental studies, and how design choices affect that value. Currently, many organizations decide on the precision of their studies based on power calculations. These do not take into account the costs and benefits of EBDM and instead rely on statistical conventions (for example, requiring 80% power for tests at the 5% level). Additionally, our framework enables decision makers to {\it ex ante} assess whether an experiment is worth conducting based on its cost and expected benefits (with benefits increasing with the decision maker's initial uncertainty about the intervention’s effect).

{\em Related literature.}---This article builds on foundational work by blackwell1951experiments and howard1966voi on the value of information. It extends that framework to the applied domain of EBDM, combining costs, precision, and empirical Bayes estimation into a flexible tool for real-world decisions and counterfactual analysis. The resulting methods are particularly suited to data-rich environments, where organizations must balance the benefits of additional information against the costs of generating it.

The empirical setting is also closely related to that of meta-analytic studies hedges1985statistical, higgins2019cochrane in that it leverages information from many individual studies. However, unlike meta-analysis, the goal of this article is not to assess the effectiveness of a set of policies but to quantify the value brought by empirical evidence in guiding better policy decisions.

A key component of the EBDM estimand is the expected value of the positive part of the predicted policy payoff, that is, the expectation of the maximum of the predicted policy payoff and zero. A related but distinct object is considered in semenova2023generalized in the context of estimating the size of a latent population whose outcomes are observed regardless of treatment exposure. The estimand in semenova2023generalized targets the distribution of an expectation given observed covariates, and does so in a single study setting. In contrast, our estimand pertains to the distribution of predicted policy payoffs across many studies.

This article also contributes to the growing literature on empirical Bayes methods morris1983parametric,Efron_2010. Empirical Bayes and related shrinkage techniques play a central role in applied economics, informing research on teacher and school value-added chetty2014measuring, angrist2017leveraging, neighborhood effects Chetty2018impact, income dynamics gu2017unobserved, and racial discrimination kline2024discrimination, among other topics. As empirical Bayes methods gain traction in applied economics, a parallel methodological literature has emerged in econometrics koenker2014convex, abadie2019choosing, fessler2019how, armstrong2022robust, kwon2023optimal, koenker2024empirical, chen2024empirical. WALTERS2024183 provides a comprehensive account of empirical Bayes methods and their applications in economics. This article applies both parametric and nonparametric empirical Bayes techniques to estimate the distribution of policy payoffs in settings where the data contain information about the effects of many policies.

\@startsection{section}{1}{0mm}{-\baselineskip}{0.25\baselineskip}{\normalfont\bf}{The value of EBDM} The notation $X \sim (\theta, \sigma^2)$ indicates that the random variable $X$ has mean $\theta$ and variance $\sigma^2$. When $X$ follows a Gaussian distribution with mean $\theta$ and variance $\sigma^2$, it is denoted as $X \sim N(\theta, \sigma^2)$. $f_X(\cdot)$ represents a probability density function of the random variable $X$, while $f_{X|W}(\cdot|w)$ denotes a conditional probability density function of $X$ given $W = w$. $F_X(\cdot)$ denotes the cumulative distribution function of $X$, and $F_{X\mid W}(\cdot \mid w)$ is the cumulative distribution function of $X$ given $W = w$. The functions $\phi(\cdot)$ and $\Phi(\cdot)$ refer to the probability density function and cumulative distribution function of the standard Gaussian distribution, respectively.

\@startsection{subsection}{2}{0mm}{-\baselineskip}{0.25\baselineskip}{\normalfont\bf}{Setup} Consider the problem of a risk-neutral decision maker tasked with choosing whether to adopt a particular policy for a population of units. The ex ante unknown per-unit payoff of the policy, $\tau$, follows a distribution with known probability density function $f_\tau(\cdot)$ and mean $\mu=E[\tau]$. In an organization where teams generate new ideas for policies or interventions, $f_\tau(\cdot)$ can be thought of as the distribution of the quality of those ideas. For retrospective estimation tasks, we take this distribution as fixed. However, organizations can shift it, for example, by prioritizing high-risk projects with substantial upside, protecting agents from failure, or allocating resources to exploratory projects.

In the absence of additional information, the agent launches the policy if the expected payoff from launching is positive, \[ \mu-c_L> 0, \] where $c_L$ is the cost of launching per-unit. The expected value of this decision is $\max\{\mu-c_L,0\}$.

Now, suppose the agent has the option to obtain additional information about the policy payoff at some cost. Specifically, the agent can acquire a signal, $\widehat{\tau}$, distributed as

equation[equation omitted — 98 chars of source]

at a cost of $c_F + c(\sigma^2)$, where $c_F \geq 0$, $c(\cdot) \geq 0$, and $c'(\cdot) \leq 0$. This setup captures the information obtained from studies estimating policy effects using experimental or observational data. The constant $c_F$ reflects the fixed cost of conducting a data-driven policy evaluation, while the function $c(\cdot)$ captures the cost of precision, which partly depends on the study's sample size. The restriction on the derivative $c'(\cdot)$ indicates that obtaining more precise information is weakly more expensive. The assumption of Gaussianity for the distribution of $\widehat\tau\,|\,\tau, \sigma^2$ is motivated by the approximate Gaussian nature of the large-sample distributions of many commonly used estimators of treatment effects.

After observing the signal $\widehat\tau$, the expected payoff of the policy is

align*[align* omitted — 251 chars of source]

If a signal is observed, the agent launches the policy if \[ E[\tau|\widehat\tau] - c_L> 0. \]

For any set $\mathcal A$, let $I_{\mathcal A}(x)$ be the function that takes value one if $x\in \mathcal A$, and value zero otherwise. The expected payoff with EBDM for a fixed value of the variance of the signal $\sigma^2$ is

align[align omitted — 266 chars of source]

Define $V(\infty)$ as the expected payoff with no information beyond the distribution of $\tau$, that is, \[ V(\infty) = \max\{\mu-c_L,0\}. \] Because $\max\{x,0\}$ is a convex function of $x$, Jensen's inequality implies, \[ \text{V}(\sigma^2)\geq \max\{\mu-c_L,0\}=V(\infty). \] The value of evidence (VoE) is the difference in expected payoffs $V(\sigma^2)$ and $V(\infty)$, which is nonnegative, minus the cost of acquiring the information, which is generally positive:

align[align omitted — 96 chars of source]

A second version of $\text{VoE}$, which we term $\text{VoID}$ (for Value of Information under Default adoption) is obtained when, in the absence of additional information about the effect of the intervention, the intervention is always deployed and so $V(\infty)=\mu -c_L$:

align[align omitted — 99 chars of source]

$\text{VoID}(\sigma^2)$ is motivated by settings with ex-ante (pre-evaluation) ambiguity on the distribution of $\tau$, and agents who have a bias for action in the presence of such ambiguity.

\@startsection{subsection}{2}{0mm}{-\baselineskip}{0.25\baselineskip}{\normalfont\bf}{A motivating example}

A common instance of the setting described above is one where the decision maker obtains experimental evidence on the effect of the policy. Consider an experiment with $N$ units: $i=1, \ldots, N$. The experimenter assigns $N_1$ units at random to treatment and the remaining $N_0=N-N_1$ to control. If unit $i$ is treated, an outcome is drawn

align*[align* omitted — 49 chars of source]

If unit $i$ is untreated, the outcome is drawn

align*[align* omitted — 49 chars of source]

Let $W_i$ be an indicator of treatment for unit $i$. We observe $Y_i=Y_i(1)W_i+Y_i(0)(1-W_i)$. The average effect of the treatment is

align*[align* omitted — 41 chars of source]

A simple estimator of $\tau$ is the difference in mean outcomes between treated and nontreated,

align*[align* omitted — 112 chars of source]

Then, for large $N_0$ and $N_1$, equation (ref) holds approximately, with

align*[align* omitted — 98 chars of source]

When a fraction $p=N_1/N$ of units are assigned to treatment, and assuming that the only variable cost of the experiment comes from recruiting subjects at a cost $\kappa$ per subject, the total cost of the experiment is \[ c_F + \frac{\kappa}{\sigma^2} \left(\frac{\sigma_1^2}{p} + \frac{\sigma_0^2}{1-p}\right), \] where $c_F$ is the fixed cost of the experiment.

Neyman's allocation rule, $p/(1-p) = \sigma_1/\sigma_0$, minimizes the variance $\sigma^2$ for a fixed total number of experimental subjects, $N$. When this rule is applied to allocate subjects between a treatment and a control group, the cost of the experiment becomes \[ c_F + \kappa \frac{(\sigma_1 + \sigma_0)^2}{\sigma^2}. \]

In some settings, researchers favor treatment effect parameters free of units of measurement, such as lift $\tau = (\theta_1-\theta_0)/\theta_0$. Let \[ \widehat\tau = \frac{\displaystyle\frac{1}{N_1}\displaystyle\sum^{N}_{i = 1} W_i Y_i - \displaystyle\frac{1}{N_0}\displaystyle\sum^{N}_{i = 1} (1-W_i)Y_i}{\displaystyle\frac{1}{N_0}\displaystyle\sum^{N}_{i = 1} (1-W_i)Y_i}. \] Then, for large $N_0$ and $N_1$, equation (ref) holds with \[ \sigma^2 = \frac{1}{\theta_0^2}\left(\frac{\sigma_1^2}{N_1}+(1+\tau)^2\frac{\sigma_0^2}{N_0}\right). \] For values of the lift parameter close to zero, as is common in many online experimentation settings, we can approximate \[ \sigma^2 \approx \frac{1}{\theta_0^2}\left(\frac{\sigma_1^2}{N_1}+\frac{\sigma_0^2}{N_0}\right). \] In this case, Neyman allocation yields, \[ c_F + \kappa\frac{(\sigma_1 + \sigma_0)^2}{\theta_0^2\sigma^2}, \]

\@startsection{subsection}{2}{0mm}{-\baselineskip}{0.25\baselineskip}{\normalfont\bf}{A Gaussian distribution for $\tau$}

This section derives a simple closed-form expression for $V(\sigma^2)$ under the assumption that the distribution of $\tau$ is Gaussian,

align[align omitted — 66 chars of source]

While equation (ref) is supported by the Central Limit Theorem in studies with large samples, equation (ref) imposes two important restrictions. First, $\tau$ is a Gaussian random variable. Second, implicit in the notation is the assumption that the distribution of $\tau$ is independent of $\sigma^2$. The first is a strong parametric restriction. The Gaussian approximation for $\tau$ could be valid in some settings but questionable in others. The second restriction could be violated, for example, if researchers adapt the power of individual studies to take into account prior information about the effect on the treatment. We dispose of these two restrictions later in the article. We adopt them in this section, however, to obtain closed-form formulas for the value of EBDM in a simple setting.

If equations (ref) and (ref) hold, the marginal distribution of $\widehat{\tau}$ is $\widehat\tau\sim N(\mu, \gamma^2 + \sigma^2)$. The posterior for $\tau$ is given by

align*[align* omitted — 180 chars of source]

The expected payoff with EBDM is

align*[align* omitted — 145 chars of source]

Let \[ Z= \frac{\mu / \gamma^2 + \widehat\tau / \sigma^2}{1 / \gamma^2 + 1 / \sigma^2}-c_L. \] Recall that the marginal distribution of $\widehat\tau$ is Gaussian with mean $\mu$ and variance $\gamma^2+\sigma^2$. As a result,

equation[equation omitted — 112 chars of source]

Now, $V(\sigma^2)$ is the the first moment of the Gaussian distribution in (ref) censored from below at zero,

align[align omitted — 245 chars of source]

The derivatives of $V(\sigma^2)$ with respect to $\sigma^2$, $\gamma^2$, and $\mu$ are

align[align omitted — 499 chars of source]

Higher precision of the signal $\widehat\tau|\tau$ increases the expected payoff from EBDM. In addition, the value of experimentation increases when an organization increases the variance of the distribution of true effects---that is, the variance of idea quality. The derivative of $V$ with respect to $\mu$ implies

align[align omitted — 286 chars of source]

$\text{VoE}$ peaks at $\mu = c_L$ and decreases monotonically with $|\mu - c_L|$. When $|\mu - c_L|$ is large, a simple rule that selects policies based solely on the sign of $\mu - c_L$ gets most decisions right, leaving little room for additional evidence to add value. $\text{VoID}$, the value of EBDM when policies are adopted by default in the absence of additional information, decreases monotonically with $\mu$. $\text{VoID}$ is particularly large when the distribution of $\tau$ is concentrated on negative values, as additional information winnows out many ineffective or counterproductive policies.

Notice that

equation*[equation* omitted — 92 chars of source]

and \[ \lim_{\sigma^2\rightarrow 0} V(\sigma^2) = (\mu -c_L)\Phi\left(\frac{\mu-c_L}{\gamma}\right)+\gamma\phi\left(\frac{\mu-c_L}{\gamma}\right).\vspace*{.2cm} \] As $\sigma^2\rightarrow\infty$, we lose any additional information about the value of $\tau$ beyond its distribution, and $V(\sigma^2)$ converges to $V(\infty)$. As $\sigma^2\rightarrow 0$, the information gathering process reveals the value of $\tau$. In this case, $V(\sigma^2)$ converges to $E[\max\{\tau - c_L, 0\}]$, the mean of the distribution of $\tau - c_L$ censored at zero.

So far, we have treated $\sigma^2$ as a constant. We now allow $\sigma^2$ to have a non-degenerate distribution, independent of $\tau$. In this case, the average payoff of EBDM is

align[align omitted — 409 chars of source]

with the expectation taken over the distribution of $\sigma^2$.

\@startsection{subsection}{2}{0mm}{-\baselineskip}{0.25\baselineskip}{\normalfont\bf}{A Gaussian mixture distribution for $\tau$}

In Section (ref) we use a Gaussian mixture distribution to evaluate the effects of misspecification of the distribution of $\tau$ on EBDM value estimates. Suppose $\tau$ follows a mixture of $k$ Gaussian distributions with parameters $(\mu_1,\gamma^2_1), \ldots, (\mu_k,\gamma^2_k)$, and mixture probabilities $p_1, \ldots, p_k$. Let $\widehat\tau=\tau + \varepsilon$, where $\varepsilon$ is independent Gaussian noise with variance $\sigma^2$. Conditional on $\tau\sim N(\mu_j,\gamma^2_j)$, we have \[ E[\tau|\widehat\tau, \tau\sim N(\mu_j,\gamma^2_j)] = \frac{\mu_j / \gamma_j^2 + \widehat{\tau} / \sigma^2}{1 / \gamma_j^2 + 1 / \sigma^2}. \] As a result,

align[align omitted — 607 chars of source]

The simulations in Section (ref) use the Gaussian mixture model to capture deviations from normality in the distribution of $\tau$.

\@startsection{section}{1}{0mm}{-\baselineskip}{0.25\baselineskip}{\normalfont\bf}{Empirical Bayes estimation} In this section, we analyze a setting with $n$ realizations from the distribution of $(\tau,\sigma,\widehat\tau,\widehat\sigma)$, where only the estimates $\widehat\tau$ and $\widehat\sigma$ are observed. We use the observations on $\widehat\tau$ and $\widehat\sigma$ to estimate $\mu$ and $\gamma^2$ using an Empirical Bayes strategy. $(\widehat\tau,\widehat\sigma)$ represent point estimates and their corresponding standard errors for a set of policy evaluations in the dataset. We examine both the homoskedastic case, where $\sigma^2$ is constant, and the heteroskedastic case, where $\mbox{var}(\sigma^2)>0$. Throughout our analysis, we approximate the per-unit launch cost as $c_L\approx 0$. Alternatively, we can interpret $\tau$ as representing the net benefits of the treatment after accounting for the launch cost.

We employ a database of thousands of online experiments run by Upworthy to illustrate the applicability of our methods. Upworthy is a U.S. online news and media publisher that built a large following in the 2010s by pairing positive, uplifting stories with optimized headline-and-image packages designed to drive clicks and shares. Upworthy pioneered large-scale A/B testing of these packages, routinely randomizing visitors across alternative headlines and images for the same article preview and using click-through performance to guide editorial and distribution choices.

The Upworthy Research Archive records thousands of randomized experiments run by Upworthy between January 24, 2013 and April 30, 2015. We restrict our analysis to the Upworthy Exploratory Dataset, which contains outcomes for $4{,}873$ online experiments. Each experiment compared alternative headline-and-image packages for the same article preview. We analyze outcomes for the first two packages deployed in each experiment, labeling the first as the control arm and the second as the treatment arm. For each package, the archive reports impressions and clicks. We remove from the sample all experiments with fewer than 100 impressions in one of the experimental arms, which yields a sample of $n=4{,}857$ online experiments. We then compute each experimental arm’s click rate per thousand impressions and define $\widehat{\tau}$ as the treatment–control difference in those rates. We also compute the standard error of $\widehat{\tau}$.

Figure (ref) shows the distribution of $\widehat{\tau}$. The average value of $\widehat{\tau}$ is $-0.7621$ (clicks per thousand impressions), with range $[-54.17,,45.13]$. The standard errors have mean $2.9727$ and range $[0.2202,,7.942]$. For the remainder of this section, we use the Upworthy data to illustrate empirical Bayes estimation of the value of EBDM. Sections (ref) through (ref) delve deeper into the EBDM value estimates for the Upworthy dataset.

figure[figure omitted — 232 chars of source]

\@startsection{subsection}{2}{0mm}{-\baselineskip}{0.25\baselineskip}{\normalfont\bf}{Parametric empirical Bayes} In this section, we adopt a Gaussian specification for the distribution of $\tau$. Because $\widehat\tau$ is unbiased, we can estimate $\mu$, the mean of the distribution of $\tau$, as the mean of $\widehat\tau$ across evaluations. To estimate $\gamma^2$, the variance of the distribution of $\tau$, we deconvolute the distribution of $\widehat\tau$ as follows. By the Total Law of Variance, \[ \mbox{var}(\widehat\tau)=E[\mbox{var}(\widehat\tau|\tau)]+\mbox{var}(E[\widehat\tau|\tau]). \] Unbiasedness of $\widehat\tau$ conditional on $\tau$ implies

align*[align* omitted — 140 chars of source]

As a result, we define $\widehat\gamma^2$ as the difference between the variance $\widehat\tau$ across experiments in the data minus the mean of the squares of the standard errors. This estimator is not guaranteed to be non-negative morris1983parametric. In the Upworthy dataset, $\widehat\gamma^2=36.4608-10.2437=26.2171$.

\@startsection{subsubsection}{3}{0mm}{-\baselineskip}{0.25\baselineskip}{Homoskedastic case}

For the homoskedastic case, we estimate $\sigma^2$ as the average of the squares of the standard deviations of $\widehat\tau$ across studies. In the Upworthy data, this estimate is 10.4919. Plugging in this value in (ref) along with estimates of $\mu$ and $\gamma^2$, we obtain $V(10.2437)=1.3691$, with the value of information measured in clicks per thousand impressions.

\@startsection{subsubsection}{3}{0mm}{-\baselineskip}{0.25\baselineskip}{Heteroskedastic case}

We now relax the assumption that $\sigma^2$ is constant. It can be shown (see appendix) that $V(\sigma^2)$ is convex. Then, by Jensen's inequality, $V(E[\sigma^2])\leq E[V(\sigma^2)]$. This result implies that the assumption of homoskedasticity may lead to underestimation of the average payoff when $\sigma^2$ is not constant. Under heteroskedasticity, we evaluate the expression in (ref) plugging in study-specific estimates of $\sigma^2$. Relative to the calculations in the previous section, now the value of experimentation is computed for each value of $\sigma^2$ and then integrated over the distribution of $\sigma^2$. An estimator of $V$ based on a set of policy estimates can be calculated in two steps: {\it (i)} use the square of the standard error of $\widehat\tau$ to approximate $\sigma^2$, and estimate the value of each study separately, and {\it (ii)} take the average over all studies in the sample. For the Upworthy data, this procedure yields $V=1.4057$.

\@startsection{subsection}{2}{0mm}{-\baselineskip}{0.25\baselineskip}{\normalfont\bf}{Nonparametric empirical Bayes} We next relax the parametric restriction $\tau\sim N(\mu,\gamma^2)$ of Section (ref) and consider a nonparametric distribution for $\tau$.

\@startsection{subsubsection}{3}{0mm}{-\baselineskip}{0.25\baselineskip}{NPMLE under precision independence}

Suppose $\tau \mid \sigma \sim G_\sigma$, where $G_\sigma$ is an unspecified distribution, and \[ \widehat\tau \mid (\tau,\sigma) \sim N(\tau,\sigma^2). \] Under a precision-independence assumption (namely, that $\tau$ is independent of $\sigma$) we have $\tau \mid \sigma \sim G_\sigma = G$, so the distribution of $\tau$ does not depend on $\sigma$. It follows that for any $s>0$, the conditional distribution of $\tau/\sigma$ given $\sigma=s$ coincides with the distribution of $\tau/s$.

Given $n$ studies with observed pairs $\{(\widehat\tau_i,\sigma_i)\}_{i=1}^n$, we estimate the mixing distribution $G$ using nonparametric empirical Bayes methods. A nonparametric maximum likelihood estimator (NPMLE) of $G$ solves the problem

equation[equation omitted — 145 chars of source]

where $\mathcal{G}$ denotes a class of discrete distributions supported on a fixed grid $u_1,\ldots,u_m$, with probabilities $g_1,\ldots,g_m$ (see appendix for details). Modern implementations of NPMLE (e.g., koenker2017rebayes) solve (ref) efficiently and come with theoretical guarantees jiang2020general,soloff2025multivariate. The solution is computed over a grid $u_1, \ldots, u_m$ representing support points of $G$, with corresponding probabilities, $\widehat g_1, \ldots, \widehat g_m$.

Moreover, as shown in the appendix, the posterior mean satisfies

equation[equation omitted — 205 chars of source]

from which we derive a sample analog of $E[\tau|\widehat\tau=z,\sigma=s]$ as \[ \frac{\displaystyle\sum_{j=1}^m u_j\phi \left(\frac{z-u_j}{s} \right)\widehat g_j}{\displaystyle\sum_{j=1}^m \phi \left(\frac{z-u_j}{s} \right)\widehat g_j}. \]

For the Upworthy dataset, we approximate $\sigma_1,\ldots,\sigma_N$ using the reported standard errors, estimate $G$ via the algorithm of koenker2014convex as implemented in koenker2017rebayes, and obtain $\widehat V=0.7621$.

\@startsection{subsubsection}{3}{0mm}{-\baselineskip}{0.25\baselineskip}{Relaxing the precision independence assumption} We relax the precision-independence assumption by partitioning the range of $\widehat\sigma_1, \ldots, \widehat\sigma_n$ into five intervals and performing the NPEB calculations from the previous section within each interval. We refer to this approach as binning. Applied to the Upworthy data, binning yields $\widehat V=0.9285$.

As an alternative, we use the CLOSE-NPMLE framework of chen2024empirical, which models the conditional distribution of the estimates given their standard errors as a flexibly parameterized location–scale family and estimates the mixing distribution nonparametrically via NPMLE. For the Upworthy data, CLOSE-NPMLE yields $\widehat V=0.9560$.

\@startsection{section}{1}{0mm}{-\baselineskip}{0.25\baselineskip}{\normalfont\bf}{Simulations}

We consider two data generating processes (DGP). In DGP1, the parameters $\tau$ have a standard Gaussian distribution $\tau\sim N(0,1)$, so the parametric empirical Bayes model of Section (ref) applies. In DGP2, the parameters $\tau$ follow the mixture distribution as in Section (ref). In particular, in DGP2, \[ \tau \sim \left\{

array[array omitted — 123 chars of source]

\right. \] DGP1 and DGP2 both produce a distribution of $\tau$ with mean zero and variance one. We generate $\widehat\tau$ as $\widehat\tau=\tau + \sigma u$, where $u$ is independent standard Gaussian and $\sigma=0.1$. To calculate the expected payoff of EBDM, we consider the case of $c_L=0$.

We run 1000 simulations for DGP1 and DGP2 with $n=500$. Equation (ref) with $\mu=0$, $\gamma^2=1$, $\sigma=0.1$, and $c_L=0$ gives the true expected payoff of EBDM under DGP1. To calculate the true expected payoff of EBDM under DGP2, we first use equation (ref) to compute $E[\tau|\widehat\tau]$ over the $n\times 1000=500{,}000$ realizations of $\widehat\tau$ in the simulations, and report the average of $\max\{E[\tau|\widehat\tau],0\}$. In each of the simulations, we calculate parametric and nonparametric empirical Bayes estimates of the average payoff of EBDM. The parametric empirical Bayes estimator is the sample analog of equation (ref). This estimator is valid under the assumption that the true distribution of $\tau$ is Gaussian. The nonparametric empirical Bayes estimator is as in Section (ref). For the simulations in this section, we treat $\sigma^2$ as known.

table[table omitted — 886 chars of source]

Table (ref) reports the true values of the expected payoff of EBDM along with means and 95 percent intervals for the distribution of the estimates across simulations. When $\tau$ is Gaussian, the distributions of the parametric and nonparametric empirical Bayes estimates across simulations are both centered near the true value of the expected EBDM payoff. Moreover, there is no evidence of substantial gains from knowledge of the parametric form of the distribution of $\tau$. The 95 percent interval for the nonparametric estimator is only 1.5 percent wider than the interval for the parametric estimator.

For the case when the distribution of $\tau$ is a mixture, the results for the parametric estimator reveal a clear bias, while the distribution of the nonparametric estimator remains centered at the true value of the expected payoff. Moreover, the 95 percent interval for the nonparametric estimator is 28.3 percent narrower than the interval for the parametric estimator.

\@startsection{section}{1}{0mm}{-\baselineskip}{0.25\baselineskip}{\normalfont\bf}{Application to the Upworthy dataset}

Table (ref) reports parametric and nonparametric empirical Bayes estimates of the value of EBDM in the Upworthy data. In the parametric case, the table reports estimates computed under heteroskedasticity (Section (ref)). In the nonparametric case, it reports three estimates: the precision-independence NPMLE (Section (ref)), and binning and CLOSE-NPMLE estimates (Section (ref)) that relax the assumption of precision independence. Below each of the estimates of the value of EBDM, Table (ref) reports 95 percent intervals computed over $1{,}000$ bootstrap draws from the distribution of $(\widehat\tau,\widehat\sigma^2)$ in the data. In our calculations, we impose $c_L=c_F=c(\sigma^2)=0$.

Because $\widehat\tau$ has a negative mean and large dispersion relative to its mean, this is a setting where we expect to have substantial gains from EBDM. Indeed, the parametric model suggests a $\text{VoE}$ of about $1.8$ times the magnitude of $\widehat\mu$ (but with a positive sign), while nonparametric empirical Bayes under the most restrictive specification yields a value only slightly larger than $\widehat{\mu}$ in magnitude. The most flexible specifications (binning and CLOSE) deliver similar results, implying a $\text{VoE}$ that is 25.4 percent larger than $\widehat{\mu}$ in magnitude and with a positive sign and a $\text{VoID}$ that is 125.4 percent larger than $\widehat{\mu}$ in magnitude, again with a positive sign.

table[table omitted — 1,732 chars of source]

\@startsection{section}{1}{0mm}{-\baselineskip}{0.25\baselineskip}{\normalfont\bf}{Estimation of counterfactual EBDM values}

This section provides estimates of the value of EBDM under alternative levels of statistical precision and under alternative levels of dispersion in the distribution of the effects of the policies.

First, we estimate how the value of EBDM would change as a result of a change in $\sigma^2$. In the resulting counterfactuals, the variance of the estimators is equal to the variance of $\widehat\tau_1\,|\,\tau_1, \ldots, \widehat\tau_n\,|\,\tau_n$ in the original sample multiplied by $\lambda$. That is, $\lambda=0.5$ represents a counterfactual scenario where the variances of the estimators are 50 percent smaller than the variance estimates in the original sample, while for $\lambda=1.5$ the variances of the estimators are 50 percent larger than in the original sample.

For simplicity, we consider only counterfactual scenarios such that $\tau$ is independent of estimation variance, $\sigma^2$, and estimate the value of EBDM using the parametric empirical Bayes estimator of Section (ref). It is conceptually straightforward to extend this procedure to more general settings (e.g., by modeling the dependence between $\tau$ and $\sigma^2$ and/or using nonparametric empirical Bayes estimators).

For each estimate $i=1, \ldots, n$ in our sample, we draw a value from the empirical Bayes estimate of the distribution of $\tau$. Let $\tau_1^*, \ldots , \tau_n^*$ be the resulting values for the draws. Next, for $i=1, \ldots, n$, we obtain $\widehat\tau_i^*=\tau_i^*+\sigma_i^* U_i$, where $U_1, \ldots, U_n$ are independent draws from the standard Gaussian distribution, and $\sigma_i^*=\sqrt{\lambda}\widehat\sigma_i$. We use the new sample $(\widehat\tau_1^*,\widehat\sigma_1^*), \ldots, (\widehat\tau_n^*,\widehat\sigma_n^*)$ to compute an estimate of the value of EBDM. We repeat this procedure multiple times to obtain the distribution of EBDM-value estimates for a particular value of $\lambda$. The average of this distribution is our estimate of the value of EBDM under variance modification factor $\lambda$. For the parametric empirical Bayes case, this average can also be computed directly using an empirical counterpart of equation (ref) that applies the variance modification factor $\lambda$ to $\sigma^2$.

figure[figure omitted — 203 chars of source]

Panel A of Figure (ref) reports the results obtained from applying the procedure described above to the Upworthy data. The solid line represents the value of EBDM as a function of the variance modification factor, $\lambda$. The shaded area represents 95 percent intervals from the distribution of EBDM estimates. An investment that reduces estimation variance by half (about a 29.3 percent decrease in standard errors) leads to an increase in the value of EBDM by $8.33$ percent, from $1.4057$ to $1.5228$. Conversely, an increase in estimation variance by half (about a 22.5 percent increase in standard errors) decreases the value of EBDM by 6.56 percent, from $1.4057$ to $1.3133$.

We next compute counterfactual EBDM values for different levels of $\gamma^2$, the variance of the distribution of $\tau$. For each observation $i = 1, \dots, N$ in our sample, we draw a value from a distribution with the same mean as the empirical Bayes estimate of the distribution of $\tau$ but with variance adjusted by a factor $\lambda$. Let $\tau_1^*, \dots, \tau_n^*$ denote these draws. Next, for each $i$, we compute $\widehat\tau_i^* = \tau_i^* + \widehat\sigma_i U_i$, where $U_1, \ldots, U_n$ are independent standard Gaussian draws. Using the sample $(\widehat\tau_1^*, \widehat\sigma_1), \dots, (\widehat\tau_n^*, \widehat\sigma_n)$, we estimate the value of EBDM. Repeating this procedure multiple times yields a distribution of EBDM estimates for a given $\lambda$. The mean of this distribution is our estimate of EBDM under variance modification factor $\lambda$ for $\gamma^2$.

Panel B of Figure (ref) shows how changes in the heterogeneity of true policy effects influence the value of EBDM. When the variance modification factor $\lambda$ is greater than one, meaning that the variance of the true effects increases, the value of EBDM rises. For instance, when $\gamma^2$ is increased by 50 percent ($\lambda = 1.5$), the estimated value of EBDM increases from $1.4057$ to $1.8882$, reflecting a 34.32 percent gain. Conversely, when $\gamma^2$ is reduced by half ($\lambda = 0.5$), the value of EBDM declines to $0.7841$, representing a 44.22 percent decrease from the baseline.

\@startsection{section}{1}{0mm}{-\baselineskip}{0.25\baselineskip}{\normalfont\bf}{Significance testing reduces the value of information}

Organizations commonly implement interventions only when they {\em (i)} deliver positive estimated treatment effects and {\em (ii)} attain statistical significance at a prespecified level. This section shows that significance-based decision rules fail to exploit the full informational content of the data and are therefore suboptimal from an EBDM perspective.

Intuitively, statistical significance decision rules prioritize Type I error control rather than maximizing expected payoff. Therefore, they may discard interventions with large but imprecisely estimated effects that could generate substantial rewards, or systematically favor interventions with small but precisely estimated effects, thereby biasing decisions toward low-variance interventions rather than those with high expected payoffs. Moreover, conditioning implementation decisions on statistical significance induces a winner's curse effect: selected interventions are disproportionately likely to be those whose effects were overestimated in the sample due to sampling variation andrews2023inference. Empirical Bayes methods address these challenges by shrinking extreme estimates toward the prior distribution of treatment effects, thereby removing substantial noise from signals. This enables decision-makers to better distinguish signal from noise and rank interventions according to their posterior expected value, thus favoring interventions with genuinely positive expected payoffs.

Define the value of EBDM under a statistical significance rule with one-sided significance level $\alpha$ as

align*[align* omitted — 203 chars of source]

where $z_{1-\alpha}$ is the $(1-\alpha)$-th quantile of a standard normal distribution.

Consider first the case of a fixed value of $\sigma$. For the parametric model $\tau\sim N(\mu,\gamma^2)$, the marginal distribution of $\widehat\tau$ is $\widehat\tau \,|\, \sigma \sim N(\mu, \gamma^2 + \sigma^2)$. As a result, conditional on $\sigma^2$, the value of EBDM under a statistical significance rule at significance level $\alpha$ is

align*[align* omitted — 308 chars of source]

Calculations in the appendix show

equation[equation omitted — 292 chars of source]

Compare this expression to $V(\sigma^2)$ in (ref). It holds that

equation[equation omitted — 81 chars of source]

provided $\gamma^2>0$, with strict inequality except for the case in which the arguments of the functions $\Phi(\cdot)$ and $\phi(\cdot)$ coincide in the expression of $V(\sigma)$ and $V_{\text{sig}}(\sigma)$. The appendix contains a detailed comparison of the payoffs $V(\sigma)$ and $V_{\text{sig}}(\sigma)$.

A plug-in procedure yields the estimator of $V_{\text{sig}}$ \[\dfrac{1}{n} \sum \limits_{i=1}^n\left[(\widehat{\mu} - c_L) \Phi\left(\frac{\widehat{\mu} - c_L-z_{1-\alpha}\sigma_i}{\sqrt{\widehat{\gamma}^2+\sigma_i^2}} \right) + \frac{\widehat{\gamma}^2}{\sqrt{\widehat{\gamma}^2+\sigma_i^2}}\phi\left(\frac{\widehat{\mu} - c_L-z_{1-\alpha}\sigma_i}{\sqrt{\widehat{\gamma}^2+\sigma_i^2}}\right)\right].\]

As an alternative, for any estimator of the posterior mean $\widehat{E}[\tau \,|\, \widehat{\tau}_i,\sigma_i]$ (parametric or nonparametric) an estimator of the value of EBDM under a statistical significance decision rule is given by

align*[align* omitted — 172 chars of source]
table[table omitted — 1,105 chars of source]

Table (ref) compares the value delivered by significance-based EBDM to the $\text{VoE}$ for the Upworthy data. Across specifications, the empirical Bayes procedures in Section (ref) deliver substantially higher value than significance-based decision making. Under the parametric heteroskedastic model, EBDM yields $1.4057$, compared with $1.0287$ under the 5 percent significance rule—a reduction of about 27 percent. The gap is larger under nonparametric methods: binning and CLOSE yield values between $0.93$ and $0.95$, whereas the corresponding significance-rule values cluster around $0.66$$0.67$, implying that significance screening discards roughly 30 percent of attainable value. Overall, these results show that significance-oriented decision rules systematically underperform value-based policies, especially in settings with substantial heterogeneity and estimation noise, where empirical Bayes methods can extract value from interventions that significance tests would discard.

\@startsection{section}{1}{0mm}{-\baselineskip}{0.25\baselineskip}{\normalfont\bf}{Conclusions}

This article develops an empirical framework to quantify the value of evidence-based decision making and to examine how statistical precision and heterogeneity in policy effects moderate that value. Using both parametric and nonparametric empirical Bayes methods, we estimate the benefit of incorporating data-driven evidence into decision-making processes by balancing the trade-off between the costs of acquiring information and the expected improvements in outcomes. Higher statistical precision (i.e., lower $\sigma^2$) and greater heterogeneity in policy effects (i.e., higher $\gamma^2$) increase the value of evidence-based decision making.

Our framework provides a principled approach for organizations to evaluate whether investing in additional data collection or policy exploration is worthwhile given the expected gains in decision quality. The proposed methods are particularly relevant for organizations that frequently conduct experiments to optimize policies and business strategies. The availability of many experimental evaluations---that is, the availability of many instances of $(\widehat\tau, \widehat\sigma^2)$---makes it possible to estimate the distribution of policy effects.

Future research could extend our framework by incorporating more flexible empirical Bayes estimators and exploring settings where estimation variance and treatment effect heterogeneity are determined endogenously by prior information. Additionally, applying our approach to firm-level, governmental, and healthcare decision-making contexts could further validate its usefulness in diverse policy environments.

Currently, many organizations rely on power calculations to guide their study designs, without taking into account the cost-benefit trade-offs associated with evidence-based decision making. As organizations continue to expand their reliance on data for policy decisions, our proposed methods offer a practical procedure for optimizing information acquisition strategies.

\centerline{\bf Appendix}