Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
92,842 characters · 27 sections · 40 citation commands
Statistical protocols are fundamental to various regulatory approval processes, serving as the backbone for validating the efficacy and reliability of new products or interventions. For example, hypothesis testing is widely employed to analyze data from clinical trials during drug approval processes. Scientific journals rely on $p$-values to assess the significance of study results. Technology and financial companies use A/B testing to evaluate the performance of new features. Meanwhile, the regulatory landscape is inherently complex, characterized by the presence of multiple stakeholders, each with distinct goals and incentives.
In this environment, agents such as pharmaceutical companies, researchers, and engineers typically bear the costs of collecting evidence to support their proposals. As a consequence, agents are likely to make strategic decisions, such as determining what evidence to gather and which proposals to advance, with the aim of maximizing their returns. These returns may take the form of financial gain, career advancement, or professional prestige. On the other hand, regulators seek to design decision-making protocols that achieve controlled statistical error and high social utility. A key challenge lies in the interaction—and possible misalignment—between these goals. The statistical protocols used by regulators inform decision-making and reward distribution, thereby impacting the interests of the involved agents. In many cases, the chosen protocol can incentivize agents to act in ways that are not aligned with the regulator's goals. These challenges are often exacerbated by information asymmetry between agents and regulators.
In this paper, we study a game-theoretic formalization of statistical decision-making with strategic agents, known as principal-agent hypothesis testing. In this framework, the statistician—also referred to as the principal or regulator—seeks to design a hypothesis testing protocol with controlled error. To achieve this, she interacts with a population of agents capable of generating data and possessing prior information (known only to the agents) about its quality. The agents' preferences are represented by a utility function. For a given testing protocol, each agent decides whether or not to opt in based on maximizing their expected utility. The statistician observes data only from agents who opt in; consequently, the agents' strategic behavior influences the distribution of the data observed by the statistician and, therefore, the associated testing error.
As a concrete illustration, consider the hypothesis testing procedures employed by the Food and Drug Administration (FDA) for approving medical treatments. In its regulatory role, the FDA acts as the statistician, or principal: it issues guidelines outlining how evidence from clinical trials will be evaluated, specifying a $p$-value threshold for testing. On the other side, pharmaceutical companies are the agents: they develop potential drug candidates and face the decision of whether to conduct clinical trials, the cost of which can range from millions to billions of dollars. If the FDA approves a drug based on trial evidence, the company can market the drug and generate substantial profits.
Given their internal research divisions, companies possess considerable prior information about the possible effectiveness of proposed drugs; this information is not available to the FDA, creating a form of information asymmetry. In addition to this prior information, companies make decisions based on the potential profits associated with an approved drug. The possibility of high profits provides an incentive to conduct trials even for drugs with uncertain effectiveness, thereby increasing the false positive rate for the FDA's testing protocol. At the same time, most companies exhibit various forms of risk aversion giambona2018theory, including avoidance of reputational damage or legal liability, and accounting for high uncertainty or delays in future profits. Such risk-averse behavior also influences the data observed by the FDA. Overall, the FDA's testing error is determined by a delicate interaction between the incentives created by any protocol and companies' prior information and risk preferences.
Thus, we are led to the central question tackled in this paper: for a given hypothesis test proposed by the principal, how can the probability of false positives be controlled? This question is both conceptually interesting, as it quantifies the delicate interplay between incentives and statistical performance described above, and practically important, as it provides guidance for designing statistical tests with guaranteed error control. We analyze this question in a Bayesian setting, where the null and alternative spaces are endowed with priors known only to the agents. In this context, it is most natural to control the Bayesian false discovery rate (FDR), which corresponds to the posterior probability of a false positive; see equation (ref) for the precise definition.
A notable feature of our framework, compared to past work on principal-agent testing tetenov2016economic,bates2023incentivetheoretic,bates2022principalagent, is its incorporation of both risk aversion and uncertainty in future rewards. More specifically, we allow for agents who make decisions based on an arbitrary concave and increasing utility function. This family includes risk-averse agents, who are influenced by the stochasticity of future rewards. By contrast, the risk-neutral agents analyzed in prior work—whose underlying utility is linear—are indifferent to such uncertainty. Within this general framework, our main result, stated as (ref), provides an upper bound on the Bayesian FDR for any choice of the principal's hypothesis test. When specialized to risk-neutral agents (linear utility) and constant rewards, this general bound offers a tighter guarantee than prior work by a subset of the current authors bates2023incentivetheoretic. This improvement is quantified in (ref) and illustrated in (ref).
Remarkably, the upper bound provided in this paper is unimprovable in a rather strong sense. First, in the simpler setting where the principal interacts with a single agent, we prove that the upper bound on the Bayes FDR is achieved with equality at some point; see (ref)(b) for the precise statement. When the principal interacts with a population of agents, we establish an even stronger result—namely, there exists an agent population for which the bound can be made $\epsilon$-tight at any countable number of points. See (ref) and the accompanying (ref) for an illustration of this “staircase sharpness" guarantee. Furthermore, these guarantees have significant implications for the principal's social utility. In particular, we demonstrate how to design a test that is maximin optimal for the principal (see (ref)).
Finally, an interesting by-product of our analysis is a connection to the notion of prior elicitation (see (ref) for related work). In particular, as an intermediate step in our proof, we show how, for a given testing protocol, an agent's decision to opt in implies a certain upper bound on their prior probability of being null. Thus, the principal can be viewed as eliciting information about the agent's prior through her design of the testing protocol and the associated payoff structure. See (ref) for the details of this result.
The principal-agent form of interaction is a classical game-theoretic model in contract theory laffont2001,bolton2004contract,salanie2005economics. It can be understood as a Stackelberg game in which the principal moves first by declaring a protocol, and the agents produce the best responses according to their personal utility. Our work falls within the broader literature that uses game-theoretic approaches to characterize how strategic behavior (including incentive misalignment) can influence statistical procedures. As one instance of this type of interplay, there is a rich literature on the phenomenon of “$p$-hacking” within the publication process: researchers selectively report statistically significant results, thereby increasing the likelihood of false positives. Classical work on this problem (e.g., sterling1959publication,tullock1959publication,leamer1974false) studied how reported estimates reflect not only data but also researcher preferences; various corrections for such bias have been proposed in the statistical literature on selective inference (e.g., taylor2015statistical,berk2013valid). Most related to this paper is the econometrics literature, in which researchers tackle $p$-hacking from a game-theoretic perspective, modeling incentive structures explicitly. For instance, mccloskey2024critical provided critical values that remain robust even when researchers engage in $p$-hacking. Other studies also focus on scenarios in which agents could influence data collection or specify part of the statistical protocol. chassang2012selective,ditillio2017 examined randomized controlled trials where agents' hidden actions affect outcomes. Working within the principal-agent formalism, spiess2018optimal studied a problem in which the principal delegates statistical inference tasks to the agent, and recommended restricting the agent to fixed-bias estimators for accurate estimates. mcclellan2022experimentation suggested adjusting the approval threshold sequentially to incentivize continued experimentation.
The version of principal-agent testing studied here was introduced by tetenov2016economic, who considered risk-neutral agents, bounded the type I error in terms of a cost-profit ratio, and provided a certain type of maximin guarantee. We discuss connections to this latter guarantee in (ref). Working in a Bayesian setting, bates2023incentivetheoretic proved an upper bound on the false discovery rate for risk-neutral agents, whereas bates2022principalagent demonstrated how the principal can utilize $e$-values in statistical protocols to protect against null strategic agents. Moving from binary to multiple hypothesis testing, viviano2024modelmultiplehypothesistesting also considered incentive misalignment, and characterized the optimal critical value based on the agent's cost and the number of hypotheses.
As discussed previously, underlying our analysis is a characterization of the principal's ability to elicit the agent's private information. For this reason, our work also has connections to the literature on information elicitation mechanisms. These studies focus on designing payment structures to ensure truthful reporting from the agent; for example, see the paper kong2019information on information-theoretic approaches to designing systems for information elicitation. An important class of elicitation techniques includes proper scoring rules (e.g., brier1950verification,good1952rational,savage1971elicitation,gneiting2007strictly), which are used to elicit probabilistic estimates from risk-neutral agents. Unlike proper scoring rules, our model does not assume that the principal can adjust the agent's rewards and costs, nor does it require the agent to report their private information directly. Instead, this information is inferred through an analysis of the agent's actions. From another perspective, the agent's decision to opt in can be viewed as a signal to the principal, making research relevant to the growing literature (e.g., carpenter2007regulatory,andrews2021model,banerjee2020theory,henry2019research,williams2021preregistration) on persuasion and signaling in scientific communication.
\paragraph{Paper organization:} The remainder of this paper is organized as follows. We begin in (ref) with a precise formulation of principal-agent testing. (ref) is devoted to our main results. To clarify connections with past work, we start in (ref) by discussing the consequences of our general results for the special case of risk-neutral agents and constant rewards. (ref) is devoted to the statement of our main result ((ref)), which provides a general upper bound on the Bayes FDR of the principal's test. In (ref), we prove the staircase-sharpness property of this upper bound, whereas (ref) is devoted to maximin properties. In (ref), we illustrate some consequences of our theory, both for synthetic problems and for the testing problem faced by the FDA. We conclude with a discussion in (ref).
In this section, we provide a precise formulation of the problem under study, including the structure of the hypothesis spaces, the principal-agent interaction, and the utility functions that drive the agents' decision-making.
We consider a game-theoretic form of interaction involving two parties: the principal, who acts as a statistical regulator, and the agent. The principal's role is to decide whether to approve or deny any proposal submitted by the agent. Accordingly, we represent the principal's action space as $\{\approves,\denies\}$. The interactive process begins with the principal, who declares the form of the statistical hypothesis test she will use to make her decision. Based on shared knowledge of this test, as well as their private information and utility (discussed below), the agent decides whether to submit a proposal or opt out. Accordingly, we represent the agent's action space as $\{\optin,\optout\}$.
To formalize the binary testing problem, we assume that there exists a hidden random variable $\theta$, which takes values in some space $\Theta$ and represents the latent quality of any proposal. Neither the agent nor the principal knows the true value of $\theta$, but the agent has knowledge of a prior distribution $\priorDist$ over $\Theta$. The principal has no access to $\priorDist$ and cannot ask the agent directly about $\priorDist$. (Indeed, a strategic agent has no incentive to respond truthfully and might choose to report incorrect information for their own benefit.) Given the parameter space $\Theta$, the hypothesis test is defined by a disjoint partition of the form
The set $\Theta_0$ corresponds to ineffective proposals, whereas the alternative set $\Theta_1$ corresponds to effective proposals. An important quantity in our analysis is the prior null probability $\priorNull \defn \priorDist(\theta \in \Theta_0)$.
The principal's goal is to design testing rules that, as much as possible, lead to the \approves decision for effective proposals and the \denies decision for ineffective ones. We now specify the form of the tests and the variables to which they apply. Let $\{ \Pdist_\theta \mid \theta \in \Theta \}$ be a family of (conditional) probability distributions indexed by $\theta$. We use this family to model both (a) the evidence $\evidence$ used by the principal for decision-making; and (b) the random reward $\ensuremath{R}$ received by an agent for approved proposals. When an agent with parameter $\theta$ chooses to submit a proposal, they conduct an experiment that yields a random variable $\evidence$, drawn according to the distribution $\Pdist_\theta$. In typical cases, the variable $\evidence \in [0,1]$ is a $p$-value, meaning that its distribution under the null is uniform. Our analysis allows for a slight relaxation, applying to variables $\evidence$ that are super-uniform under the null, meaning that for any $\theta \in \Theta_0$, we have
In order to evaluate proposals, the principal declares in advance that she will perform a binary hypothesis test with a threshold $\threshold \in [0,1]$---that is, she \approves the proposal if $\evidence \leq \threshold$, and \denies it otherwise. The principal is not permitted to change the threshold $\threshold$ once it is published. Given knowledge of the threshold $\threshold$, the agent then decides whether to spend $\cost$ dollars to run a trial (\optin) and collect evidence $\evidence \sim \Pdist_\theta$, or to \optout. When an agent chooses to \optin and the principal \approves the proposal, the agent receives a randomly drawn reward $\reward \sim \Pdist_\theta$ in dollars. Otherwise, when the principal \denies the proposal, the agent loses $\cost$ dollars (deterministically) for running the trials.
The sequence of interaction between the principal and the agent can be summarized as follows:
The formulation given here allows the reward $\reward$ to be both random and dependent on $\theta$. Formally, for any value $\theta \in \Theta$, we assume that $\reward$ is drawn from the distribution $\Pdist_\theta$. In addition, we impose a few regularity conditions. First, we assume throughout that
Condition (ref) implies that for any threshold $\tau \in [0,1]$, the principal's decision $\Ind(X \leq \threshold)$ is also conditionally independent of the reward $\rfunc$.
Having set up the hypothesis test, rewards, and the form of principal-agent interaction, the final ingredient is the specific mechanism by which the agent makes decisions. He does so to serve his own interests, particularly by maximizing his expected utility.
More precisely, suppose that each agent has some initial wealth $\Wealth_0 > \cost$, and let $\wealth \mapsto \util(\wealth)$ be a concave and non-decreasing utility function. It measures the value that the agent ascribes to a particular wealth level $\wealth$. For instance, a risk-neutral agent is defined by the linear utility function $\util(\wealth) = \wealth$. A more general family of utility functions consists of those with constant relative risk aversion arrow1965riskaversion,PrattJohn1964, given by
By convention, we set $\util_1(\wealth) = \log(\wealth)$ for $\ensuremath{\gamma} = 1$; this logarithmic utility underlies betting strategies that seek to maximize the growth rate under repeated trials (i.e., the Kelly criterion). Otherwise, note that the risk-neutral choice is the special case with $\ensuremath{\gamma} = 0$, whereas this specification captures risk-seeking behavior for $\ensuremath{\gamma} < 0$ and risk-aversion for $\ensuremath{\gamma} \in (0, 1]$. There is empirical evidence giambona2018theory showing that most companies exhibit some degree of risk-aversion, meaning that choices $\ensuremath{\gamma} \in (0,1]$ are the most practically relevant within the family (ref).
Given an arbitrary concave and non-decreasing utility function $\util$, the agent chooses the action in the set $\{\optin, \optout \}$ that maximizes expected utility. Based on the previously specified principal-agent interaction, we have the following three possibilities:
These outcomes define the expected utility for both actions, and we assume that the agents are utility-maximizing, meaning that they choose to \optin if and only if
where $\Ind(\evidence \leq \threshold)$ is a binary indicator function for the event $\evidence \leq \threshold$. To be clear, the left side of this inequality (ref) involves a three-part expectation: conditioned on the unknown value of $\theta$, we take an expectation over the evidence $\evidence$ and reward $\reward$ drawn from $\Pdist_\theta$; and second, we take an expectation over $\theta$ distributed according to the prior $\priorDist$ of the agent. Thus, as shown by our analysis, the agent's decision to \optin implicitly imposes a constraint on the prior distribution, which can be inferred by the principal.
We now present our main results and discuss their consequences. Recall that the principal moves first by declaring a threshold $\threshold$ to be used in the hypothesis test. We provide guarantees on the posterior null probability, also referred to as the Bayes false discovery rate (FDR), associated with any such test:
We have followed the terminology of efron2012large by using Bayes FDR. This term is appropriate because the quantity $\posNull$ is closely related to the usual definition of false discovery, which is the expected ratio of false positives to the total number of positives when testing multiple hypotheses. The key difference is that our setup also includes a Bayesian prior over the parameter space $\Theta$.
The remainder of this section is organized as follows. In (ref), we begin by presenting an upper bound on the Bayes FDR in a simplified setting to provide intuition for our general result and establish connections to past work. Building on these insights, we extend our analysis in (ref) to derive a bound on the Bayes FDR in a more general setting (see (ref)). This extension encompasses the fully general framework described in (ref), including stochastic rewards and general utility. In (ref), we establish a sharpness guarantee for our upper bound, whereas (ref) provides maximin guarantees for the principal.
In this section, we consider a risk-neutral agent who makes decisions using the linear utility function $\util(\wealth) = \wealth$, and constant deterministic rewards, meaning that there is a fixed scalar $\constR$ such that $\rfunc \stackrel{(a.s.)}{=} \constR$ for all $\theta \in \Theta$. Moreover, we consider a simple-versus-simple hypothesis test, meaning that both the null set $\Theta_0 \equiv \{\theta_0\}$ and the non-null set $\Theta_1 \equiv \{\theta_1 \}$ are singletons. This scenario allows us to demonstrate how the Bayes FDR depends on key factors such as the cost $\cost$, the reward $\constR$, the threshold $\threshold$, and the power of the hypothesis test. It also facilitates concrete comparisons with past work bates2023incentivetheoretic.
In this simple setting, our bound depends on four quantities: the cost $\cost$ paid by the agent for choosing to \optin; the reward $\constR > \cost$ received when the principal \approves; and the two functions
When $\evidence$ is actually uniform under the null, we have $\beta_0(\threshold) = \threshold$, but our analysis only requires the super-uniformity condition (ref), equivalently stated as $\beta_0(\threshold) \leq \threshold$. Note that $\threshold \mapsto \beta_{1}(\threshold)$ is the power function of the test, and we assume the test has non-trivial power, which means
We prove this claim using a more general result to be stated in the sequel; see the discussion following (ref).\\
Equation (ref) provides two upper bounds on the Bayes FDR. Inequality (I) is a sharper result, but computing it requires knowledge of the quantities $\beta_0(\threshold)$ and $\beta_1(\threshold)$. Depending on the application, these quantities may or may not be known to the principal. Inequality (II) is a weaker result that is established using the facts that $\beta_0(\threshold) \leq \threshold$ by assumption, and the upper bound $\beta_1(\threshold) \leq 1$ by definition (ref). The advantage of the bound (II) is that it can be calculated without knowledge of $\beta_0(\threshold)$ and $\beta_1(\threshold)$.
When $\beta_0(\threshold) = \threshold$, inequalities (I) and (II) provide non-trivial information---that is, an upper bound on the posterior null less than $1$---only when $\threshold < \cost/\constR$. As noted in prior work by tetenov2016economic, the fraction $\cost/\constR \in (0, 1]$ corresponds to the threshold at which profit-maximizing agent who knows with certainty that he is null would still \optin. As such, there exist (worst-case) testing problems for which the Bayes FDR is equal to $1$ when $\threshold = \cost/\constR$, so that non-trivial bounds are impossible above this threshold.
In past work on risk-neutral agents, bates2023incentivetheoretic proved that the Bayes FDR is upper bounded by $\threshold \frac{\constR}{\cost}$. The bound (ref) provides a refinement of this claim: in particular, for any $\tau \in [0, \frac{\cost}{\constR}]$, we have
where inequality (a) is bound (II) from equation (ref); and inequality (b) holds since $\frac{\constR - \cost}{\constR - \threshold \constR} \leq 1$ whenever $\threshold \leq \frac{\cost}{\constR}$. In fact, the bound (ref)(I)\xspace is unimprovable in a very strong sense; see part (b) of (ref) for a sharpness guarantee for a general utility function.
Underlying our bounds on the Bayes FDR---both in (ref) and the more general (ref) in the sequel---is an interesting connection to prior elicitation. Here we provide an informal description for the case of linear utility; see (ref) to follow for a precise statement in the general setting.
Consider an agent who chooses to \optin for a trial with a threshold $\threshold \in (0, \cost/\constR]$. This agent has their private prior $\priorDist$ over the parameter space $\Theta = \Theta_0 \cup \Theta_1$; of particular interest is the prior null probability $\priorNull = \priorDist(\theta \in \Theta_0)$ associated with this distribution. Intuitively, the fact that this agent is utility-maximizing (ref) and has chosen to \optin implies that their prior null probability cannot be too large. In particular, our analysis shows that for a risk-neutral agent with constant reward, any agent who decides to \optin must have the prior null probability upper bounded
The bound (II)\xspace is obtained from the first inequality via the upper bounds $\beta_0(\threshold) \leq \threshold$ and $\beta_1(\threshold) \leq 1$, and certain monotonicity properties. See the proof of (ref) in (ref) for details.
The bounds (ref) reveal two phenomena that are intuitively reasonable. First, as the reward $\constR$ increases (with other problem parameters held fixed), these upper bounds also increase. Larger potential rewards for the \approves decision mean that an agent needs less a priori confidence in the quality of their proposal to \optin, suggesting that their prior null $\priorNull$ could potentially be larger. Second, for hypothesis tests with larger power (i.e., larger values of $\beta_1(\threshold)$), the upper bound (I)\xspace increases. Larger power implies that the agent has more certainty that any $\theta \in \Theta_1$ they submit will lead to the \approves decision, and hence a higher expected profit.
In this section, we examine Gaussian mean testing as an example to compare various bounds on the Bayes FDR.
Given this class of hypothesis tests, we consider agents that operate with the linear utility function, cost $\cost = 10$ and constant reward $\constR = 25$. For a given choice of $\theta_1$, we compute the bounds (I) and (II) from equation (ref); the bates2023incentivetheoretic bound; and the exact Bayes FDR as a function of the threshold $\threshold$. We plot the results for different choices of the alternative mean $\theta_1$ in the Gaussian testing problem: $\theta_1 = 1$ and $\theta_1 = 2$ in panels (a) and (b), respectively, of (ref). For computing the exact Bayes FDR, we consider a mixture ensemble of agent types, in which 10% of the agents are “good” with $\priorNull^g = 0.3$, and the remaining $90\%$ are “bad” with $\priorNull^b = 0.8$.
Focusing first on the exact Bayes FDR (yellow line with shading underneath for emphasis) in panel (a), note that it exhibits two step-like transitions. The first occurs at the threshold $\threshold \approx 0.16$; below this threshold, no agents \optin so that the FDR is zero by definition, whereas above this threshold, agents start to \optin. At and above this threshold, the “good” agents in our mixture ensemble \optin, and as $\threshold$ increases above this transition, the exact Bayes FDR also increases, since the same population of agents are provided with a progressively less stringent hypothesis test. The second step transition occurs at $\threshold \approx 0.32$, at which point the “bad” agents also choose to \optin. Turning to panel (b), observe that it also exhibits the same two-step phenomenon; the difference here is that the transitions occur for smaller values of the threshold, since the larger mean in the alternative ($\theta_1 = 2$ in panel (b) versus $\theta_1 = 1$ in panel (a)) means that the power function $\beta_1(\threshold)$ grows more quickly, which leads to greater expected utility (and hence greater incentive) for agents to \optin.
Now let us discuss the three upper bounds on the exact Bayes FDR shown in panel (a). By definition, the Bates et al. upper bound is linear in the threshold $\threshold$ with slope $\cost/\constR$. Both bounds (I) and (II) from equation (ref) are non-linear and monotonic functions of $\threshold$. Notice that bound (I) is zero near $\threshold \approx 0.1$, which represents the minimum threshold required to ensure an opt-in agent. Below this threshold, even an ideal agent with $\priorNull = 0$ will not \optin, causing the bound to evaluate to negative values. Moreover, bound (I) equals the exact Bayes FDA at the first transition, highlighting the sharpness of our bound. Consistent with our earlier analysis (ref), both bounds improve upon the original Bates et al. result. The bounds in panel (b) are qualitatively similar in nature; the main difference to note here is that the gap between Bounds (I) and (II)---for which the functions $\{\beta_j \}_{j=0}^1$ are known and unknown, respectively---is smaller than in panel (a). This difference can be predicted by returning to the statement of (ref), where we see that bound (II) is obtained by replacing the unknown $\beta_1(\threshold)$ with the upper bound $1$. Panel (b) corresponds to a higher signal-to-noise ratio, or SNR for short, since the size of mean shift $\theta_1$ is doubled in moving from panel (a) and (b), so that this approximation is more accurate for the high SNR problem.
In this section, we extend the results developed above to the more general setting described in (ref), allowing for a concave and non-decreasing utility function $\wealth \mapsto \util(\wealth)$, stochastic rewards $\reward \sim \Pdist_\theta$, and composite hypothesis testing (so that the sets $\Theta_0$ and $\Theta_1$ need not be singletons). Recall from (ref) that the rewards are assumed to satisfy three conditions. The non-dominating cost condition (ref) is necessary for the principal to observe some \optin agent; the stochastic monotonicity condition (ref) ensures that non-null hypotheses are more rewarding than null; the conditional independence condition (ref) guarantees that the evidence $\evidence$ and reward $\reward$ are conditionally independent given $\theta$.
There are two key differences between this more general setting and the simpler setting from the previous section. Whereas linear utility allows for a straightforward comparison of costs and rewards on the wealth scale, risk-averse agents consider additional factors beyond wealth when deciding whether to \optin. Consequently, the agent's cost and reward must be evaluated on the utility scale. More risk-averse agents derive less utility from monetary rewards and therefore are less likely to \optin. Second, since the reward depends on $\theta$, the reward when the principal \approves a non-null proposal may differ significantly from that when the principal \approves a null proposal. If the reward for a non-null proposal is substantially higher, agents with greater confidence in their proposals are more likely to \optin.
From our discussion of the simple-versus-simple hypothesis test, recall the two quantities $\beta_{j}(\threshold)$ for $j \in \{0,1\}$, as defined in equation (ref). Moving to the composite-versus-composite hypothesis test, our bound depends on the related functions
For a general utility function, our results depend on three forms of difference in expected utility. The first two correspond to differences between the agent's expected utility, comparing the \approves to the \denies decision. More precisely, we define
Note that $\util(\Wealth_0 - \cost)$ is the utility of an agent who decides to \optin but receives a \denies decision, whereas the quantity $\E[\util(\WealthGain) \mid \theta \in \Theta_j]$ is the agent's expected utility for an \approves decision, taken over $\Theta_j$ for $j \in \{0,1\}$. Finally, the loss incurred by the agent when the principal \denies, for $\theta$ in either the null or the alternative, is given by
The following guarantee applies to a principal-agent testing problem with rewards satisfying conditions (ref) through (ref), as well as the power condition (ref).
We prove this claim in (ref). \\
{\bf{Remark:}} There is no loss of generality in requiring that $\bar \beta_0(\threshold) \in \big(0, \; \utilLoss/\nullUtilGain \big)$. For a threshold such that $\bar \beta_0(\threshold) > \utilLoss/\nullUtilGain$, even the “worst agent” with $\priorNull = 1$ would choose to \optin, leading to a Bayes FDR equal to one.
\paragraph{Implications for risk-neutral agents:} We can gain helpful intuition for the general bound (ref) by specializing it to a simple-versus-simple test, a risk-neutral agent (with linear utility function $\util(\wealth) = \wealth$) and constant reward $\constR$. With these choices, we have $\nullApp = \beta_0(\threshold)$, $\altApp = \beta_1(\threshold)$, $\UtilGain{j} = \constR$ for $j \in \{0,1 \}$, and $\utilLoss = \cost$, and the bound (ref) becomes
Thus, we have recovered inequality (ref)(I)\xspace in (ref) as a special case of the general result (ref).
\paragraph{Role of assumptions:} Returning to the general setting, let us clarify how our assumptions are related to the appearance of various terms in the bound (ref). The bound includes three differences---all of which should be positive to yield a valid upper bound on a probability. First, from the non-trivial power condition (ref), the difference $\bar \beta_1(\threshold) - \bar \beta_0(\threshold)$ is strictly positive; moreover, the non-decreasing property of the utility and the assumption $\cost \geq 0$ imply that $\utilLoss \geq 0$. Second, it follows from the stochastic monotonicity condition (ref) that $\altUtilGain - \nullUtilGain \geq 0$. Lastly, we need to verify that
The validity of this inequality depends on the opt-in assumption implicit in (ref)---namely, that there is some agent that chose to \optin at threshold $\threshold$---along with the utility-maximizing nature of agents. These conditions ensure that the “ideal agent”---meaning an agent with prior null probability $\priorNull = 0$---would certainly \optin. Combining these two conditions allows us to establish inequality (ref); see the argument surrounding equation (ref) in (ref) for the details.
If the principal lacks precise knowledge of the quantity $\nullApp$, $\altApp$, $\nullUtilGain$, and $\altUtilGain$, she can still obtain a conservative estimate of the guarantee (ref) by upper bounding these unknowns. Recall that $\nullApp \leq \threshold$ by the super-uniformity condition (ref); moreover, suppose that there is a known envelope function $\threshold \mapsto \kappa(\threshold)$ such that
where $\Exs_\theta$ denotes expectation under the distribution $\Pdist_\theta$. Finally, recall our previous definition $\utilLoss \defn \util(\Wealth_0) - \util(\Wealth_0 - \cost)$.
See (ref) for the proof of this claim. Here we discuss its consequences for risk-neutral agents and compare it to (ref).
\paragraph{Implications for risk-neutral agents:} Let us show how (ref), when specialized to the linear utility function $\util(\wealth) = \wealth$, constant reward $\constR$ and envelope function $\kappa(\threshold) = 1$, recovers the bound (II)\xspace from equation (ref) in (ref). With these choices, we have $\DelBar_0 = \DelBar_1 = \constR$ and $\utilLoss = \cost$. Setting these values and $\kappa(\threshold) = 1$ into inequality (ref), we find that
as claimed in inequality (ref)(II)\xspace.
\paragraph{Comparing with\texorpdfstring{ (ref)}{ Theorem 1}:} It is worthwhile comparing the upper bound (ref) with our earlier upper bound (ref) from (ref). For discussion, let us refer to the latter as the known-$\beta$ bound, and the former as the unknown-$\beta$ bound. Suppose that the rewards are deterministic, equal to separate values $\constR_0$ and $\constR_1$ in the null and alternative cases, respectively. For such rewards, the two guarantees only differ in that the potentially unknown values $(\beta_0(\threshold), \beta_1(\threshold))$ are replaced by the corresponding upper bounds $\big( \threshold, \kappa(\threshold) \big)$. When the $p$-value is actually uniform under the null, we have $\beta_0(\threshold) = \threshold$, so nothing is lost in the upper bound.
The unknown-$\beta$ upper bound (ref) always holds with $\kappa(\threshold) = 1$. With this choice, the gap between the known and unknown-$\beta$ cases should depend on the “easiness” of the testing problem. For instance, consider the case of a simple null $\Theta_0 = \{0 \}$ versus the simple alternative $\Theta_1 = \{\theta_1\}$ for Gaussian mean testing (see (ref) for the model set-up). As $|\theta_1|$ increases, the power function $\beta_{\theta_1}(\threshold)$ approaches $1$, so that the unknown-$\beta$ bound should become a better approximation to the known $\beta$-bound. (ref) provides an illustration of this phenomenon in the special case of linear utility.
Otherwise, returning to general rewards, in order to compute the upper bound (ref), the principal requires uniform upper bounds $\constR_0$ and $\constR_1$ of the reward means $\Exs_\theta[\ensuremath{R}]$ over the null and alternative spaces (cf. equation (ref)). In the proof, these upper bounds are combined with Jensen's equality, and the assumed concavity and non-decreasing property of the utility function $\util$, to argue that
In the case of a simple-versus-simple hypothesis test---say with $\Theta_j = \{\theta_j \}$ for $j \in \{0,1\}$---we have $\constR_j = \Exs_{\theta_j}[\rfunc]$, so that the bounds (ref) are sharp for a risk-neutral agent ($\util(\wealth) = \wealth)$. For risk-averse agents, the bounds (ref) become progressively weaker as the degree of risk aversion increases. This gap can be understood via the lens of certainty equivalence (or risk premia); for a risk-averse agent, the utility associated with stochastic rewards is always less than the corresponding utility with mean-matched deterministic rewards.
The proof of (ref) involves an auxiliary result that is of independent interest. In particular, a key analytical step is to show that an agent's decision to \optin implies an upper bound on the prior probability $\priorNull = \priorDist(\theta \in \Theta_0)$ of the proposal being null. More precisely, we establish the following fact:
See (ref) for the proof of this claim. \\
\paragraph{Implications for risk-neutral agents:} It is useful to study the implications of this claim for a risk-neutral agent (with $\util(\wealth) = \wealth)$) and constant reward $\constR$. For the linear utility, we have $\UtilGain{0} = \UtilGain{1} = \constR$ and $\utilLoss = \cost$. For a simple-versus-simple test, we have $\nullApp = \nullAppSimple$ and $\altApp = \altAppSimple$. Substituting these choices into equation (ref), we recover our previously stated claim (ref) for risk-neutral agents. \\
(ref) lies at the heart of the upper bounds on Bayes FDR from (ref). In particular, we prove that for any fixed threshold, the Bayes FDR is an increasing function of $\priorNull$, so that any upper bound on $\priorNull$ induces an upper bound on the Bayes FDR. Thus, it is also intimately connected to the sharpness guarantee stated in (ref). In particular, our proof shows that the FDR bound in (ref) is met with equality for any agent who chooses to \optin with a prior null probability $\priorNull$ that makes inequality (ref)(I)\xspace hold with equality. Any such agent is “worst-case” in the sense that its prior null probability $\priorNull$ is as large as possible while still being consistent with his decision to \optin at the threshold $\threshold$.
In this section, we present a stronger sharpness result that is naturally formulated in the setting of multiple agents. Recall that (ref)(b) provides a sharpness guarantee for the single-agent case: there exists a prior null probability $\priorNull = \priorDist(\theta \in \Theta_0)$ under which the upper bound (ref) on the Bayes FDR is exactly attained at a single point. The bound also applies to a mixture of different agent types; (ref) illustrates such a mixture with $K = 2$ types of agent---say “good” versus “bad”. As shown in these plots, there is some $\threshold$ at which only good agents choose to \optin, and the upper bound matches the upper bound at this point (consistent with (ref)(b)). Moreover, the Bayes FDR is actually very close to the bound at the threshold at which the bad agents also choose to \optin. This phenomenon shown in (ref) is, in fact, quite general: with a population of agents, the exact Bayes FDR can be made arbitrarily close to the known-$\beta$ bound at a countable number of points. \\
To understand this fact, we begin by formalizing the meaning of Bayes FDR for a mixture of agents. Doing so requires two functions that depend on a prior null probability $\priorNull$ and a threshold $\threshold$. In particular, the function
corresponds to the Bayes FDR induced by agents with prior null $\priorNull$ who choose to \optin under threshold $\threshold$. We also define the indicator function
which tracks whether or not an agent with prior null $\priorNull$ chooses to \optin.
Now consider a mixture of $\numAgent$ types of agent, where each type is characterized by a prior null probability $\priorNull^i \in [0,1]$ and a mixture proportion $\agentWeight{i} \in [0,1]$, normalized so that $\sum_{i=1}^\numAgent \agentWeight{i} = 1$. The exact Bayes FDR under for this $\numAgent$-mixture, which we refer to as $\kfdr$, is given by
given the definition (ref) of $\kfdr(\threshold)$ as a convex combination.
In the following theorem, we establish that there are $\numAgent$-mixtures of agents, as specified by some collection of prior null probabilities $\{\priorNull^i\}_{i=1}^\numAgent$ and weights $\{\agentWeight{i}\}_{i=1}^K$, for which the Bayes FDR $\kfdr(\threshold)$ is arbitrarily close to $\Psi(\tau)$ at any $\numAgent$ points.
See (ref) for the proof.\\
The sharpness guarantee (ref)---exact at a point and $\epsilon$-close at $(K-1)$ points---is substantially stronger than the result in (ref). It shows that it is impossible to improve (ref) in any substantive way without further constraints on the agent mixtures. \\
To gain intuition for (ref), it is helpful to revisit the Gaussian mean testing example from (ref). Following the same setup, we consider agents with a linear utility function, a cost $\cost = 10$, and a constant reward $\constR = 25$. For each choice of $K \in \{20, 40 \}$, we construct a $K$-mixture ensemble in which: (i) the agents have prior null probabilities $\{\priorNull^i\}_{i=1}^K$ evenly spaced over the interval $[0.02, 0.97]$; and (ii) the mixture proportions $\{\agentWeight{i}\}_{i=1}^K$ are constructed to satisfy the condition
This condition ensures that at each threshold at which a new type of agents with a higher prior null probability \optin, they account for 99% of the opt-in population.
In (ref), we plot the resulting $\kfdr(\threshold)$ as the threshold $\threshold$ is varied, along with bounds (I) and (II) from equation (ref); and the bates2023incentivetheoretic bound. Bound (I) is equivalent to the bound defined by the function $\Psi$ from equation (ref), and panels (a) and (b) correspond to $K = 20$ and $K = 40$, respectively.
In panel (a), the exact Bayes FDR exhibits 20 step-like transitions, each corresponding to the \optin of agents with progressively higher prior null probabilities. At these transitions, the exact Bayes FDR closely aligns with the known-$\beta$ bound, with the maximum gap given by $\epsilon = 0.002$. In panel (b), we increase the number of mixtures and observe a diminishing gap between the known-$\beta$ bound and the exact Bayes FDR, with the maximum gap in this case decreasing to $\epsilon = 0.0016$.
Our construction of the mixture ensemble in this example also offers insight into (ref) and its connection to the sharpness result in (ref)(b). Consider an increasing sequence of $K$ thresholds $\{\threshold_i\}_{i=1}^K$ and $K$ types of agents, where agents of type $i$ have a prior null probability $\priorNull^i$ that makes inequality (ref)(I)\xspace hold with equality under $\threshold_i$. These agents represent the “worst-case" agents that would \optin under $\threshold_i$. According to (ref)(b), if only agents of type $i$ \optin under $\threshold_i$, the known-$\beta$ bound is exact. With a mixture ensemble, agents with prior null probabilities smaller than $\priorNull^i$ will also \optin, thereby reducing the exact Bayes FDR. (ref) implies that if agents of type $i$ constitute the vast majority of opt-in agents under $\threshold_i$, the known-$\beta$ bound remains nearly sharp, providing a highly accurate estimate of the exact Bayes FDR. This highlights the “worst-case" perspective of (ref): in the presence of information asymmetry, the principal can use agents' incentives to guard against the worst-case scenario. If these worst-case agents dominate the set of proposal submissions, then the upper bound on the Bayes FDR established in (ref)(a) is saturated.
In our discussion thus far, we have focused purely on the agent's utility. The principal might also wish to optimize some form of social welfare, and in this section, we show how our theory leads to a maximin guarantee for the principal's choice of threshold.
In particular, when dealing with an agent with a prior $\priorDist$---call it a $\priorDist$-agent for short---the principal might wish to maximize the societal benefits of approving effective treatments ($\theta \in \Theta_1$) while minimizing the societal harms of approving ineffective treatments ($\theta \in \Theta_0$). We view the principal as taking an action $a \in \{\approves, \denies\}$, and we now explicitly define the principal's utility as
where $\omega_0$ and $\omega_1$ are positive weights. Recall that the principal's policy is to play action $\approves$ when $X \le \tau$. The principal's average utility with threshold $\tau$ when confronted with a $\priorDist$-agent is then given by
Now of course, the principal does not know the prior distribution $\priorDist$; consequently, it is natural to consider the worst-case $\underbar{\ensuremath{\mathcal{V}}}(\threshold) \defn \inf_{\priorDist} \ensuremath{\mathcal{V}}(\threshold; \priorDist)$. It follows from the definition of $\ensuremath{\mathcal{V}}$ that for any choice of threshold $\threshold$, we have $\underbar{\ensuremath{\mathcal{V}}}(\threshold) \leq 0$. Accordingly, we say that a threshold choice $\threshold$ is maximin optimal if
The following result characterizes the structure of maximin rules, in terms of the upper bound (ref)(I)\xspace from (ref).
To state the result, we need additional notation. For any $\tau \in (0,1)$, let $\widetilde \beta_0(\tau) = \sup_{\theta \in \Theta_0} \Pdist_\theta(X \le \tau)$ and $\widetilde \beta_1(\tau) = \sup_{\theta \in \Theta_1} \Pdist_\theta(X \le \tau)$. These are the largest probabilities that the principal \approves when the threshold is set to $\tau$ under the null and alternative, respectively. We further assume that the distribution of $R$ depends only on whether or not $\theta$ is in the set of nulls (i.e., on $\Ind(\theta \in \Theta_0)$) so that $\Delta_0$ and $\Delta_1$ do not depend on $\priorDist$. Define the function
which (ref) guarantees to be a tight bound on the posterior probability of null. The following result states that it characterizes the set of all maximin thresholds.
We provide the proof here, since it provides useful insight into the connection with (ref)(b).
In short, the upper bound from (ref) implies that $\threshold$ such that $\widetilde \Psi(\tau) \le \omega_1 / (\omega_0 + \omega_1)$ is maximin optimal. In the other direction, the sharpness result shows that for any looser $\tau$, there exists some distribution with Bayes FDR larger than $\omega_1 / (\omega_0 + \omega_1)$, so such $\tau$ is not maximin optimal.
We introduce this maximin optimality result in part to facilitate comparison with the findings of tetenov2016economic. That work shows that if agents have perfect information about $\theta$, then taking $\tau$ as the agent cost divided by the agent reward is maximin optimal. By contrast, our result shows that if agents have a prior distribution rather than full information about $\theta$, a stricter value of $\tau$ is typically required to achieve maximin optimality. In addition, that work also derives a maximin optimality result for agents with prior distributions, but with a different assumption about the principal's utility function. Roughly, under certain assumptions about the underlying distributions and that the principal's utility increases in effect size, the threshold of the agent cost divided by reward is again shown to be maximin optimal. Our result shows that with the utility function described above, maximin optimality instead corresponds to controlling the Bayes FDR, and so the sharp bound on the Bayes FDR yields an exact characterization of maximin optimal rules.
In previous sections, we present numerical results with two objectives: comparing our upper bound on the Bayes FDR to prior work and illustrating its sharpness. In (ref), we shift focus to examine how the upper bound is influenced by the agents' risk sensitivity and the stochasticity of the reward. Revisiting the binary hypothesis test introduced in (ref), where the proposal quality $\theta$ represents the mean of two normal distributions, we analyze the effects of agents' risk preferences and reward variability. This analysis allows us to explicitly demonstrate how agents' incentives can affect the inference on the Bayes FDR under various type I error thresholds.
In (ref), we analyze the FDA's current clinical trial approval policies within our principal-agent framework. Using estimates of trial costs, profitability of drug development, and the wealth and risk sensitivity of pharmaceutical companies, we assess the false positive rate of the FDA's testing procedures. Here, we show how our framework leads to an improved understanding of how current regulatory policies, which shape the financial incentives of pharmaceutical companies, ultimately impact the quality of approved drugs.
Recall that we consider the binary hypothesis test defined by the parameter space $\Theta = \{0, \theta_1\}$ with the null set $\Theta_0 = \{0\}$ and the non-null set $\Theta_1 = \{\theta_1\}$ where $\theta_1 > 0$, and the test statistic $Z \sim \mathcal{N}(\theta, 1)$. Throughout our experiments, we set $\theta_1 = 1$. Furthermore, we consider a mixture ensemble of two types of agents: $10\%$ of the agents are “good” with $\priorNull^g = 0.3$, and the remaining $90\%$ are “bad” with $\priorNull^b = 0.8$. The cost of \optin is $\cost = 10$, and both types of agents have an initial wealth of $\Wealth_0 = 20$.
In order to explore risk-sensitive behaviors, we perform experiments using the class (ref) of utility functions with constant relative risk aversion arrow1965riskaversion,PrattJohn1964, indexed by some parameter $\ensuremath{\gamma} < 1$. Recall that $\ensuremath{\gamma} = 0$ corresponds to risk neutrality, whereas $\ensuremath{\gamma} \in (0, 1]$ yields a risk-averse loss. Based on the analysis of holt2002risk, we pick $\ensuremath{\gamma} = 0$ for risk-neutral agents, $\ensuremath{\gamma} = 0.35$ for slightly risk-averse agents, and $\ensuremath{\gamma} = 0.7$ for highly risk-averse agents. For each choice of $\ensuremath{\gamma} \in \{0, 0.35, 0.7\}$, we compute the known-$\beta$ bound as a function of the threshold $\threshold$. Assuming both types of agents are highly risk-averse, we also compute the exact Bayes FDR. These results are shown in (ref), where panel (a) uses a constant reward $\constR = 25$ and panel (b) uses a constant reward $\constR = 100$.
In both panels, we observe that the known-$\beta$ bounds for more risk-averse utility functions reach a value of $1$ at significantly larger threshold values, compared to the upper bound assuming risk-neutral utility. As a direct consequence, the principal can afford to set a looser $\threshold$ without incurring a high Bayes FDR when the agents are risk-sensitive. However, if the principal underestimates the agent's risk aversion---for instance, mistaking a highly risk-averse agent for a risk-neutral one---then the principal might set a threshold $\threshold$ that is too low for any agent to benefit from opting in. Comparing panel (a) to panel (b), we observe that for a fixed threshold $\threshold$, risk aversion introduces a greater reduction in the known-$\beta$ bounds as the reward increases from $25$ to $100$. This aligns with the concept of diminishing marginal utility of wealth in economics. Since a more risk-averse agent has a more concave utility function, their marginal utility decreases more rapidly as wealth increases. As a result, a more risk-averse agent derives a lower utility gain per unit of additional wealth, which explains the widening gaps between the known-$\beta$ bounds. This implies that for highly risk-averse agents, the Bayes FDR would not increase significantly, even if the agent is offered a much larger reward.
In this section, we investigate how random rewards influence the known-$\beta$ bound in two key ways: (1) by increasing the gap between the null and the non-null rewards, and (2) by introducing stochasticity into both rewards. Based on our previous discussion (see (ref) and associated comments), stochastic rewards exert a greater influence on agents with higher risk aversion. Thus, we assume both types of agents are highly risk-averse in the following simulation.
First, we examine the influence of the ratio between null and non-null rewards. For illustrative purposes, we assume both rewards are deterministic constants, with the non-null reward set to 100 and the null reward selected from the set $\{100, 75, 50, 25\}$. In (ref)(a), we plot the known-$\beta$ bounds for various reward ratios and the exact Bayes FDR assuming the null reward is $25$. For a fixed threshold $\threshold$, the known-$\beta$ bound decreases as the null reward decreases. In other words, for a fixed level of the Bayes FDR, the principal can set larger thresholds as the null reward decreases. Notably, this change in threshold $\threshold$ is not linear: the increase in $\threshold$ is much more pronounced when the null reward drops from 50 to 25 than when it decreases from 100 to 75. This suggests that even if the non-null reward is 10 times the trial cost, agents who doubt the effectiveness of their proposals will likely \optout, as their profits largely depend on the null rewards. If the principal knows that the null reward is substantially lower than the non-null reward, she can leverage this information to achieve tighter control over the Bayes FDR.
Next, we investigate the effects of reward stochasticity. We consider three levels of stochasticity: deterministic, slightly stochastic, and highly stochastic. For the deterministic reward, we set the null reward to 50 and the non-null reward to 150. To model stochasticity, we use a truncated normal distribution. Let $\tnormal(\mu, \sigma, [a, b])$ denote a normal distribution with mean $\mu$ and variance $\sigma^2$ that lies within the interval $[a, b]$. For slightly stochastic rewards, we set:
For highly stochastic rewards, we set:
In (ref)(b), we plot the known-$\beta$ bounds for varying levels of reward stochasticity and the exact Bayes FDR assuming highly stochastic rewards. As the reward becomes more stochastic, the known-$\beta$ bound decreases, with the reduction being more pronounced at larger thresholds. Risk-averse agents prefer certain outcomes over uncertain ones with the same expected value; see equation (ref) and the associated discussion. Consequently, if the principal knows that the agent is highly risk-averse and the reward is highly volatile, the agent's decision to \optin signals strong confidence in the proposal.
Next, we return to the motivating problem described in the introduction---namely, the testing problem faced by the Food and Drug Administration (FDA). In particular, we study some implications of our results using the range of costs, profits, and existing type I error levels associated with clinical trials in the United States.
Let us begin with some relevant background. The FDA mandates that pharmaceutical companies conduct clinical trials to provide evidence of the safety and efficacy of new drugs. The financial burden of conducting these trials falls on the pharmaceutical companies. Based on the trial results, the FDA then decides whether to grant approval. If approved, the company can market the drug to the public and generate profits.
The FDA retains some flexibility in the type and amount of evidence that suffices for approval; see the papers bates2022principalagent,bates2023incentivetheoretic for a more detailed discussion of the FDA's approval guidelines. Following bates2022principalagent, we consider three simplified statistical protocols that align with current FDA practice. We evaluate (1) a {\it standard} protocol that requires a $p$-value below 0.05 in two independent trials; (2) a {\it modernized} protocol that approves a drug if the $p$-value is less than 0.01 in a single trial; (3) an {\it accelerated} protocol where two trials are conducted and the drug is approved if either trial results in a $p$-value below 0.05. All $p$-values referenced above are two-sided.
The cost of conducting a Phase III clinical trial in the U.S. can vary significantly based on the therapeutic area, the number of patients, and the trial's complexity. A 2020 study by the Institute for Safe Medication Practices estimated the median cost of a pivotal clinical trial to be around \$48 million, with an interquartile range of \$20 million to \$102 million moore2020variation. dimasi2016innovation estimated an average out-of-pocket cost of \$255 million for a Phase III trial. Additionally, wouters2020estimated provided a higher estimate of \$291 million. Note that these estimates pertain only to pivotal trials; the overall cost of drug development also includes expenses related to preclinical research, regulatory and legal costs, and investments in manufacturing and production. Given these estimates from the literature, we use an estimate of $\cost = \$200$ million as the cost of a trial in our analysis.
Regarding the values of a company's reward, we note that the profitability of a drug follows a long right-skewed distribution, with substantial profits possible for the most commercially successful drugs. For instance, Keytruda, a leading cancer immunotherapy, generated annual sales of over \$25 billion in 2023. In a more typical case, a successful new drug would generate annual sales of around \$500 million to \$1 billion during its patent-protected lifespan---see ledley2020profitability for further analysis of the profitability of major pharmaceutical companies. Given substantial differences in drug profitability, we examine a case where the average profit of a null drug is capped at \$800 million, {\it i.e.}, $R_0 = \$800$ million, and the average profit of an effective drug $R_1$ ranges from \$1 billion to \$50 billion.
To model the risk aversion of pharmaceutical companies, we again utilize the utility function specified in equation (ref) with parameters $\ensuremath{\gamma} \in \{ 0, 0.35, 0.7 \}$ representing risk-neutral, slightly risk-averse, and highly risk-averse companies, respectively. We assume all companies have the same level of risk sensitivity. Turning to the initial wealth, major pharmaceutical firms, such as Pfizer and Johnson & Johnson, typically generate revenues ranging from \$50 billion to \$100 billion, while small to medium-sized companies have estimated annual revenues between \$1 billion and \$10 billion. Given that large companies often provide a broad range of medical services, the portion of the budget allocated specifically for clinical trials is likely to be significantly lower. Accordingly, we adopt an estimate of \$5 billion for the initial wealth $\Wealth_0$.
Based on these estimates, we report the FDR bounds derived in bates2023incentivetheoretic and implied by our (ref) in (ref). To apply (ref), we suppose $\kappa = 1$. We also assume that the drug will receive a reward of $R_1$ even if it might be ineffective when applying the result in \citereset bates2023incentivetheoretic. A value of “n/a" indicates that the FDR bound is larger than 1, in which case the theory does not indicate reasonable control of FDR. Compared to \citeresetbates2023incentivetheoretic, our result suggests a much lower fraction of false positives for risk-neutral agents across the three protocols. In regimes where \citereset bates2023incentivetheoretic suggests a false discovery rate above 1, our bound still provides a nontrivial bound on the FDR. When the companies are risk-sensitive, we see a significant reduction in FDR especially for highly profitable drugs that can earn \$25 billion or more. This analysis suggests that the FDA might consider loosening the significance level for less lucrative drugs.
It is crucial to note that (ref) is based on our simplified theoretical model of how pharmaceutical companies interact with the FDA. In reality, the utility functions of these companies could be more complex and not fully captured by (ref). Our analysis seeks to clarify how false discovery rates in clinical trials might be affected by the incentives of various stakeholders. It provides guidance on how factors such as risk sensitivity and variations in rewards together with the nominal type I error rate impact the average quality of the approved drugs.
In this paper, we have studied the problem of principal-agent hypothesis testing, a specific instance of statistical decision-making involving strategic agents. This problem---and others of its type---establishes connections between statistical inference and incentive alignment in economic theory. A fundamental challenge in principal-agent testing lies in the fact that the quality of the data seen by the principal---and hence the false discovery rate (FDR) associated with her ultimate decision-making---is impacted by the agents' incentives. These incentives are determined by the interplay of their utility functions, private prior information, and the probabilistic structure of the experiment and rewards.
Among the main results of this paper are a general upper bound on the Bayes FDR of the principal's test ((ref)) and the “staircase sharpness” property it satisfies ((ref)). Both results hold for general utility functions and stochastic rewards, making them relevant to a broad class of decision models and environments. An important implication of these two findings is that, without making further assumptions on agent behavior and priors, it is impossible to guarantee any tighter control on the Bayes FDR. Given this insight, one promising future direction is to explore different forms of information asymmetry, and/or mechanisms by which the principal might enhance test performance by leveraging additional knowledge about the distribution of agents.
Finally, aside from the specific technical contributions of this paper, at a higher level, one interesting takeaway message is the connection between information asymmetry in statistical inference with strategic agents and mechanisms for prior elicitation. In the current work, since the principal is unaware of the agents' priors, inferential conclusions about the FDR can only be drawn if the agents' actions are information-revealing. Thus, as formalized in our analysis, the performance of statistical protocols hinges on their ability to extract this information and on how this information influences statistical inference. In future work, it would be interesting to broaden the range of possible mechanisms for information elicitation---for instance through the use of multiple contracts (as opposed to the single-contract model studied here) or via sequential forms of interaction, in which the principal and agent engage in multiple rounds of communication.
The authors thank Alberto Abadie, Isaiah Andrews, Lihua Lei, Whitney Newey, Ashesh Rambachan, and Davide Viviano for their comments on this work. In addition, this work was partially supported by NSF-DMS-2413875 to SB and MJW, and the Cecil H. Green Chair to MJW.
\AtNextBibliography \printbibliography