Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
83,613 characters · 21 sections · 18 citation commands
Hypothesis testing involves making a discrete choice and constitutes the most basic form of decision-making under uncertainty. It occupies a central role in practice, serving as an assessment of evidence in most scientific research. Testing, however, rarely occurs in isolation. In a broader system-level context, outcomes of a test have downstream implications, such as deploying a new drug, adding a new feature to a commercial product, or publishing a scientific finding. While all parties involved are affected, stakeholders may be differentially impacted by the results. In the other direction, upstream of the phase of data collection, these same parties may have additional, potentially private information about the object of study and make choices that affect the data generation. For example, a pharmaceutical company sponsoring a phase III clinical trial for a candidate drug implicitly has a belief about the drug's efficacy and chooses to run a trial accordingly. In this way, the distribution of the data observed, and thus the performance of any method for statistical decision-making, is altered by strategic behavior of the agents involved.
In this paper, we study hypothesis testing based on data collected from a heterogeneous population of agents, and analyze how their strategic behavior impacts the design of optimal statistical tests. More specifically, we consider a principal that seeks to carry out a collection of hypothesis tests. Each test is controlled by an agent that possesses private information---not known to the principal---and this private information can differ across agents. Our main contribution is to show how to design a protocol for hypothesis testing that yields statistically optimal results even when the private information of the agents is not known in advance.
As a simple illustration, suppose there are two agent types: a “good” type with a high prior probability of being non-null and a “bad” type with a low prior probability of being non-null. The decision-maker evaluates agents using $p$-value thresholds, which determine the stringency of the test: smaller thresholds reduce false positives but also reduce true discoveries. For a heterogeneous collection of tests, one natural measure of aggregate performance is based on the false discovery rate and true discovery rate; cf. equations (ref) and (ref). The false discovery rate (FDR) corresponds to the fraction of rejections that are false, while the true discovery rate (TDR) corresponds to the rate of true discoveries across the mixture of types, and captures a notion of overall power.
(ref) illustrates how heterogeneity undermines efficiency. The light-blue region shows all achievable (FDR, TDR) pairs when the decision-maker is permitted to apply different tests to each agent type; the upper boundary of this region (marked in a solid red line) corresponds to the oracle Pareto frontier that could be achieved if the principal were given access to the hidden agent type. A single threshold applied uniformly across both types (orange curve) lies substantially below the Pareto frontier. We can also compare to two other single threshold protocols that do require knowledge of agent type. Applying a test to only the “good” agents yields the purple curve; it has a more favorable trade-off, but still lies below the Pareto frontier. Testing only the “bad” type (green curve) defines the lower boundary of the achievable region.
When the agent types are known, it is possible to tailor thresholds to different agent types so as to balance error trade-offs more effectively: moderate thresholds for good type to increase power, and stricter thresholds for bad types to limit false discoveries. It is exactly this tailored choice of thresholds that yields the oracle Pareto frontier (red curve). However, in practice, agent types are not directly observable, so we are led to the following question:
{
}
The main contribution of this paper is to answer this question in a constructive way: we both characterize when it is possible to match the oracle and provide a concrete procedure for constructing testing protocols that do so. The key insight in our work is that, because agents are assumed to behave strategically, their actions can reveal private information. We leverage this by offering a menu of possible statistical tests, each paired with an associated payoff structure---together forming what we call a contract. We show that it is possible to design a menu of contracts such that an agent’s choice of contract reveals their private information, and moreover, such that each agent selects the statistical threshold that is optimal for their type. This design yields a testing protocol whose statistical error trade-offs match the oracle Pareto frontier, even though the principal has no a priori knowledge of agent types.
We study a game-theoretic model of hypothesis testing with a heterogeneous population of strategic agents. The decision-maker (principal) designs a hypothesis testing protocol, while agents decide whether to participate based on private information and expected payoffs. Participation requires a fixed cost, and agents receive a reward if they pass the test. We formalize this interaction through a menu of statistical contracts, as defined precisely in equation (ref), where each contract specifies the testing protocol, the upfront cost, and the reward.
There is an evolving line of work on principal-agent testing tetenov2016economic, bates2022principalagent, bates2023incentivetheoretic, shi2024sharp, hossain2025, viviano2024modelmultiplehypothesistesting, and within this context, our analysis makes the following contributions:
{1em}
Taken together, our results illustrate a key principle: heterogeneity and strategic behavior, which are often viewed as obstacles, can instead be leveraged as tools. By carefully designing testing protocols and incentives, the principal can transform private information and self-interested participation into a mechanism for statistically optimal decision-making---and remarkably, this can be achieved with minimal financial cost.
Our work connects to several strands of literature at the intersection of contract theory, statistical decision-making, and strategic behavior. In addition, by comparing statistical performance to an oracle, our work shares the spirit of statistical literature on adaptive estimation and testing.
Let us begin by discussing the connections to screening and incentive design in contract theory laffont2001, bolton2004contract, salanie2005economics. Here the focus of the study is how a principal can elicit private information from heterogeneous agents by offering a menu of contracts. Classical applications include nonlinear pricing mussa1978monopoly, maskin1984monopoly, auctions myerson1981optimal, maskin1984optimal, and regulation baron1982regulating, laffont1993theory. Our work shares this idea of elicitation, but shows that hypothesis tests themselves, paired with suitable payoff structures, can act as screening devices. Related recent work has used experimental designs as screening mechanisms wittbrodt2025delegating, yoder2022designing, wang2023contractingheterogeneousresearchers, jagadeesan2025publication, underscoring the broader potential of statistical objects for information elicitation. Elicitation is naturally connected to the notion of proper scoring rules brier1950verification, good1952rational, mccarthy1956measures, savage1971elicitation, gneiting2007strictly, which correspond to utility functions designed to elicit truthful reports. Our mechanism also encourages truthful elicitation, but with the important distinction that we have only indirect and partial control over the utility via the specification of the hypothesis test.
More broadly, our work relates to a growing literature on strategic behavior in decision-making, where agents may manipulate data, analysis, or participation to influence outcomes. A prominent example is $p$-hacking, where researchers adjust analyses to obtain significant results, inflating false positives. The statistics literature has responded with selective inference methods taylor2015statistical, berk2013valid, while economic analyses take a game-theoretic perspective. For instance, mccloskey2024critical designed critical values robust to strategic manipulation, and jagadeesan2025publication studied how differences in private costs and incentives shape optimal publication rules. spiess2018optimal showed that restricting researchers to fixed-bias estimators can improve inference when planner's and researcher's objectives diverge. Beyond $p$-hacking, the machine learning community has explored mechanism design under strategic manipulation of data, such as altering features or labels to influence classifiers hardt2016strategic, dong2018strategic or regressions dekel2010incentive, perote2004strategy, chen2018strategyproof. In contrast, the strategic behavior in our setting takes the form of participation decisions: agents cannot alter test outcomes directly, but they can choose whether to engage. This captures environments where experiments are preregistered or data collection is externally verified, so participation is the primary lever for influence.
As noted above, our work contributes directly to the evolving literature on principal–agent hypothesis testing. tetenov2016economic provided an early analysis linking type I error control to a cost–profit ratio. Subsequent work bates2022principalagent, bates2023incentivetheoretic, shi2024sharp, hossain2025 expanded the framework to incorporate risk aversion, stochastic rewards, and simultaneous control of type I and II errors. Our contribution departs from this line by focusing on heterogeneous populations and showing how menus of contracts can achieve type-optimal thresholds and optimal error trade-offs. In doing so, we integrate heterogeneity, screening, and strategic participation into a unified framework that ties statistical objectives to incentive-compatible design.
Finally, we define optimality in this paper with reference to an oracle that is given direct access to the agents' private information. Our work thus shares the spirit of work on adaptive estimation and testing (e.g., CaiLow2006, Lep90,Spo96). This classical work considers tests or estimators that adapt to unknown problem structure such as smoothness or sparsity, whereas we study adaptation to the unknown private information of a collection of agents.
\paragraph{Paper organization:}
(ref) formalizes the menu-based principal–agent testing framework, specifying agents’ strategic behavior and the principal’s statistical objectives. (ref) is devoted to the statement of main results on separating menus ((ref)), along with results on menu constructions that satisfy pre-specified criteria ((ref)). Our central theorem ((ref)), stated in (ref), characterizes all separating menus that elicit private information and implement type-optimal thresholds for arbitrary heterogeneous populations. In (ref), we show how this theorem allows the principal to match the statistical performance of the oracle ((ref)). We highlight a connection to proper scoring rules in (ref), and analyze the financial costs of elicitation in (ref). In (ref), we study menu constructions under two practical design criteria: one that minimizes the principal’s financial cost, showing that information elicitation can be achieved “for free," and another that addresses settings where the principal can only partially specify contracts. We illustrate the implications of these results through synthetic examples. We conclude in (ref) with a discussion. We examine the robustness of menu-based testing to model misspecification in (ref).
In this section, we formalize the problem by specifying the structure of the hypothesis tests, the principal–agent interaction, the agent’s utility maximization problem, and the principal’s statistical decision problem.
We consider a game-theoretic framework involving two parties: a principal, who serves as a statistical regulator, and an agent. The principal’s role is to approve or deny proposals submitted by agents. To make these decisions, the principal offers the agent a menu of statistical contracts at different costs (described below). Agents observe the entire menu and choose whether to select one contract from the menu or to opt out, based on their private information and utility. Upon selecting a contract, data is collected and reported to the principal, who conducts a hypothesis test according to the terms specified in the chosen contract. We make all of these steps precise below.
We assume that there exists a hidden random parameter $\theta$, which takes two possible values $\theta \in \Theta = \{\theta_0, \theta_1\}$ and represents the latent quality of any proposal. The null value $\theta_0$ corresponds to an ineffective proposal and the non-null value $\theta_1$ corresponds to an effective proposal. Neither the principal nor the agent knows the true value of $\theta$. However, the agent has private knowledge of a prior distribution $\priorDist$ over $\Theta$, which is inaccessible to the principal. While the principal could attempt to ask the agent directly about $\priorDist$, there is no guarantee the agent would report it truthfully. As we later show, the principal can instead offer a menu of contracts designed to elicit information about $\priorDist$.
We consider a setting with a population of agents, each characterized by a potentially distinct prior distribution $\priorDist$. Since the parameter space is binary, this distribution is fully specified by the prior null probability $\priorNull \defn \priorDist(\theta = \theta_0)$. We can classify agents according to the value of their prior null: in particular, we say that an agent is of type $\priorNull$ if he has prior null probability equal to $\priorNull$. The overall population of agents consists of different types occurring with a certain frequency, and we let $\typeDist$ denote a distribution over possible types $\priorNull$. We do not assume that the principal knows $\typeDist$.
We now formalize the notion of a statistical contract and its role in the principal-agent interaction. The principal aims to approve effective proposals and deny ineffective ones by conducting hypothesis tests. To this end, she designs a contract menu, meaning a collection of contracts indexed by reported type $\reportNull \in [0,1]$, where each contract specifies how the hypothesis test will be conducted. Any contract menu takes the form
where $\ensuremath{\mathcal{P}} \subseteq [0,1]$ is a subset of values associated with contracts, and any contract is defined by a triplet with the following components:
The agent observes the full contract menu (ref) before making any selection. Based on his private information $\priorNull$, the agent decides whether to opt in or opt out. If the agent opts in, he selects a type $\reportNull \in \ensuremath{\mathcal{P}}$ to the principal. The contract corresponding to this reported type $\reportNull$ is then executed.
We now describe how each possible contract is executed. Let $\{\Pdist_\theta \mid \theta \in \Theta \}$ be a family of probability distributions indexed by $\theta$. These are the possible distributions generating the data $\evidence$ that the principal uses for decision-making. The family of distributions is known to both the principal and agent. Without loss of generality, we assume the variable $\evidence$ is a $p$-value, meaning that distribution under the null hypothesis is uniform:
When an agent with parameter $\theta$ chooses a statistical contract $(\fthreshold{\reportNull}, \reward_\reportNull, \cost_\reportNull)$, he spends $\cost_\reportNull$ dollars to conduct an experiment, which yields a random variable $\evidence$ drawn from distribution $\Pdist_\theta$. The principal then approves the proposal if $\evidence \leq \fthreshold{\reportNull}$ and denies it otherwise.
If the principal approves, the agent receives $\reward_\reportNull$ dollars in return, resulting in a net payoff of $\reward_\reportNull - \cost_\reportNull$ dollars after accounting for the up-front cost. Otherwise, when the principal denies the proposal, the agent receives no reward, thus losing a net of $\cost_\reportNull$ dollars. The principal-agent interaction is summarized as follows:
We assume that agents are strategic decision-makers who act to maximize expected utility. Each agent is risk-neutral, meaning that their utility is equivalent to their expected gain in wealth. The agent’s utility, determined jointly by their own actions and the principal’s decision, can fall into one of three possible outcomes:
For an agent who opts into the contract indexed by $\reportNull$, his change in wealth is given by the random variable $\WealthAfter{\reportNull}$. Since the prior distribution $\priorDist$ of any agent is specified by the null probability $\priorNull = \priorDist(\theta = \theta_0)$, we define the agent’s expected utility as
Given a contract menu $\ensuremath{\mathscr{M}}$ with support $\ensuremath{\mathcal{P}}$, we assume that agents are wealth-maximizing, meaning that the behavior of an agent of type $\priorNull$ is characterized by the selection function
See (ref) for the proof of this claim.
Notice the utility (ref) is linear in the agent's prior null probability $\priorNull$ with slope $\reward_\reportNull [\nullAppSimple{\threshold_\reportNull} - \altAppSimple{\threshold_\reportNull}]$. We are interested in the case with non-trivial power, i.e., $\altAppSimple{\threshold_\reportNull} > \nullAppSimple{\threshold_\reportNull}$, in which case the slope is negative. Thus, agents with smaller $\priorNull$ (more optimistic priors) derive greater expected utility from any given contract. Moreover, the utility is nondecreasing in $\tau_p$, since both the type I error $\nullAppSimple{\threshold_\reportNull}$ and the power $\altAppSimple{\threshold_\reportNull}$ are nondecreasing.
At a high level, the principal’s objective is to design a menu that controls the two types of error in a binary hypothesis test. The menu should achieve a desired balance between type I errors (where $\theta = \theta_0$ but the principal approves) and type II errors (where $\theta = \theta_1$ but the principal denies). We consider two specifications of this trade-off below: (a) she may wish to minimize a weighted sum of the errors, or (b) to maximize total discovery rate subject to a constraint on the false discovery rate. In either case, the key feature is that the principal's optimal decision rule depends on the agent’s prior null probability $\priorNull$. When the null is more likely (large $\priorNull$), a lenient test with a large threshold risks too many false positives, so the optimal test should be more stringent. Conversely, when the null is unlikely (small $\priorNull$), the principal can afford to be more permissive to avoid unnecessary false negatives.
We use $\tau_q$ to denote the principal's ideal threshold for an agent with prior $q$, and refer to it as the type-optimal threshold. We describe two concrete cases next.
\paragraph{Weighted sum of type I and type II error.}
In many settings, the principal's cost of making a type I error is not the same as that of type II error. Suppose type I errors incur cost $\FPcost$ and type II errors incur cost $\FNcost$. For an agent of type $\priorNull$, if the principal deployed a test at level $\threshold$, the associated Bayes risk (BR) is given by
Classical results show that the optimal decision rule for the Bayes risk $\ensuremath{\operatorname{BR}}_\omega$ is to threshold the likelihood ratio statistic at level $\priorNull \FPcost / (1 - \priorNull) \FNcost$. Thus, the optimal $p$-values $\evidence$ for the principal are those obtained from the likelihood ratio statistic. Concretely, if we let $\evidence \mapsto \mathcal{L}(\evidence)$ be the function mapping $\evidence$ to the likelihood ratio, then the type-optimal threshold on the $p$-value scale is
\paragraph{Maximizing TDR subject to FDR control.} Alternatively, the principal might wish to maximize the expected true discovery rate (TDR) while controlling the false discovery rate (FDR) at a target level $\fdrLevel$. For an agent of type $\priorNull$, the FDR associated with the hypothesis test with threshold $\threshold$ is given by
which represents the probability of correctly approving non-null proposals from agent $\priorNull$ when using threshold $\threshold$. Since the FDR is increasing in $\threshold$, the constraint $\fdr(\priorNull, \threshold) \leq \fdrLevel$ places an upper bound on the threshold $\threshold$ that agent type $\priorNull$ can be allowed to choose. The TDR is also increasing in $\threshold$, and thus maximized by choosing the largest such $\threshold$. We conclude that the type-optimal threshold for an agent of type $\priorNull$ is given by
\paragraph{Defining the oracle performance:} In either case, we are left with the following problem: if the agent types were known, then given an agent of type $\priorNull$, the principal would implement the test with type-optimal threshold $\fthreshold{\priorNull}$. Doing so over the full population of agents (as $\priorNull$ varies) would achieve a statistical guarantee that we refer to as the heterogeneous agent oracle performance. More specifically, given any distribution $\typeDist$ over the set of possible agent types $\priorNull \in [0,1]$, the oracle value of the $\omega$-weighted Bayes risk is given by
based on the type-optimal threshold $\threshold_\priorNull(\fdrLevel)$ from equation (ref). In words, this is the maximum of the $\ensuremath{\operatorname{TDR}}$ subject to the $\fdr$ being $\fdrLevel$-bounded over all possible tests when agent types are known.
Of course, the key challenge is that types are unobserved. Agents, motivated by their own payoffs, may misreport their type to secure more favorable terms. The principal’s problem is therefore to design a contract menu that simultaneously uses the type-optimal thresholds while also making truthful reporting optimal for the agent. We refer to this object as a separating menu, since it acts to separate agents according to their type. If a separating menu can be constructed, it allows the principal to match the oracle statistical performance, as defined in equations (ref) and (ref), without knowing the types or their proportion $\typeDist$ a priori. Designing such separating menus is the focus of the next section.
In this section, we develop the main results: how the principal can construct menus $\ensuremath{\mathscr{M}}$ that implement type-optimal thresholds $\fthreshold{\priorNull}$. In (ref), we begin with a formal definition of a separating menu, identifying the incentive compatibility and participation constraints for truthful reporting. We then show how to construct such menus for arbitrary agent distributions, establish their link to proper scoring rules, and examine implementation costs. Finally, (ref) considers optimal designs under two practical criteria: minimizing financial cost and handling constrained contract parameters.
Throughout, we assume the principal knows the statistical properties of the test, namely the type I error rate $\nullAppSimple{\threshold} = \threshold$ and the power function $\threshold \mapsto \altAppSimple{\threshold}$ (cf. equation (ref)). We restrict attention to cases where the power is non-trivial:
We assume that the principal's type-optimal threshold assignment $q \mapsto \tau_q$ is non-increasing but make no further restrictions. As such, the menu construction that follows applies when the principal wishes to control the weighted combination of type I and type II errors or to maximize power given an FDR constraint.
As discussed in (ref), if the principal could observe each agent's true type, she would simply assign the type-optimal threshold $\fthreshold{\priorNull}$. However, since types are private, type-optimal thresholds cannot be assigned directly. Instead, the principal must design a menu of contracts---each specifying a $p$-value threshold, a reward, and a cost---that incentivizes agents to truthfully reveal their types.
We now formalize the requirements for such menus. Given a subset $\mathcal P \subset [0,1]$ of agent types, a contract menu $\ensuremath{\mathscr{M}}$ is said to be separating if each agent opts into the contract designed for their true type. In terms of the selection function (ref), a menu $\ensuremath{\mathscr{M}}$ is separating if and only if
We refer to condition (ref)(a) as an incentive compatibility constraint, which implies that it is optimal for any agent type to report their type $\priorNull$ truthfully, since misreporting will not yield higher utility. Condition (ref)(b) is a participation constraint: since agents can always opt out, their selected contract should guarantee nonnegative expected utility.
Thus, if the principal can construct a separating menu $\ensuremath{\mathscr{M}}$ with type-optimal threshold $\fthreshold{\priorNull}$, then for each type of agent that opts in, she can conduct statistically optimal tests. We now turn to the construction of separating menus. In what follows, we show that every separating menu is characterized by a real-valued convex function $\ensuremath{\mathcal{G}}$ and establish a connection between separating menus and proper scoring rules.
We first set up the notation needed to provide our general characterization. Consider a function $\ensuremath{\mathcal{G}}: \supp(\typeDist) \to \R$ defined on the support of the type distribution $\typeDist$ that satisfies the following two properties. First, for each \( \reportNull \in \supp(\typeDist) \), there exists a scalar $\ensuremath{g^*_{\reportNull}} < 0$ such that
When $\supp(\typeDist)$ is an interval in $[0,1]$ and $\ensuremath{\mathcal{G}}$ is differentiable, these conditions are equivalent to $\ensuremath{\mathcal{G}}$ being strictly convex, nonnegative, and decreasing. In the non-differentiable case, we can understand $\ensuremath{g^*_{\reportNull}}$ as playing the role of a subgradient of $\ensuremath{\mathcal{G}}$ at $\reportNull$, and when $\supp(\typeDist)$ is a discrete set, a function $\ensuremath{\mathcal{G}}$ satisfying (ref) can be interpreted as satisfying a form of discrete convexity. \\
Our main result is that the class of all such functions $\ensuremath{\mathcal{G}}$ characterizes the set of separating menus:
See (ref) for the proof of this result. We also provide an illustration of this construction when the agent types are finite in (ref). \\
(ref) provides a systematic way to construct separating menus for any distribution over agent types. Moreover, it shows that there is a one-to-one correspondence between convex functions $\ensuremath{\mathcal{G}}$ and separating menus.
Given this one-to-one correspondence, one might suspect that $\ensuremath{\mathcal{G}}$ has a fundamental meaning. As we show in (ref), the function value $\ensuremath{\mathcal{G}}(\priorNull)$ represents the expected utility attained by an agent of type $\priorNull$ when reporting truthfully---that is, we have the equivalence $\ensuremath{\mathcal{G}}(\priorNull) = \utilfunc(\priorNull;\priorNull)$. The reward function (ref) is constructed so that when an agent of type $\priorNull$ reports type $\reportNull$, his utility (ref) induced by the contract $(\threshold_\reportNull, \reward_\reportNull, \cost_\reportNull)$ has a slope given by $\ensuremath{g^*_{\reportNull}} = \reward_\reportNull[\nullAppSimple{\threshold_\reportNull} - \altAppSimple{\threshold_\reportNull}]$. In parallel, the cost function is chosen so that his utility (ref) under this contract coincides with the supporting hyperplane of $\ensuremath{\mathcal{G}}$ at $\reportNull$, given by
As a result, property (ref) guarantees incentive compatibility (ref)(a): each agent type achieves maximal expected utility by reporting truthfully.
The condition $\ensuremath{g^*_{\priorNull}} < 0$ is imposed to maintain consistency with the non-trivial power assumption (ref); it implies that utility declines with type $\priorNull$, so more optimistic agents (those with lower $\priorNull$) obtain higher expected utility under the same contract. Lastly, property (ref) ensures that participation constraints (ref)(b) are met: once the highest type secures nonnegative utility, all agents with lower types will strictly prefer to participate.
The statistical motivation of our work was to determine when it is possible for the principal to match the statistical performance of the oracle given access to agent types a priori. In particular, recall the definitions of the oracle risk for the $\omega$-weighted Bayes risk (ref), and the TDR risk (ref). A straightforward but important consequence of (ref) is to provide a prescriptive means for the principal to match the oracle. In stating this result, we say that a function $\ensuremath{\mathcal{G}}$ is valid if it satisfies the conditions of (ref). We summarize as follows:
Thus, we have shown that it is always possible for the principal to match the oracle performance. Note that there are no assumptions on the structure of the underlying testing problem, apart from having a $p$-value uniform under the null, as well as non-trivial power (ref). The caveat here is that (ref), and hence (ref), assumes that the principal has full control over all three components (cost, reward, and threshold) of each statistical contract. In (ref), we visit a constrained setting in which the principal has only partial control, and see that the answer then depends on more fine-grained features of the test.
A separating menu ensures that any utility-maximizing agent type opts into the contract designed for their true type. In this way, the principal can elicit an agent’s private belief by observing their choice. This resembles the logic of proper scoring rules: reward functions that yield truthful elicitation under expected utility maximization. In what follows, we demonstrate that an incentive-compatible menu can be interpreted as inducing a strictly proper scoring rule.
In order to make this connection precise, consider the binary random variable $Y \defn \indicator(\theta = \theta_0)$, where $Y=1$ indicates that the agent's proposal is ineffective ($\theta = \theta_0$) and $Y=0$ indicates effectiveness. The agent’s belief is summarized by the prior null probability $\priorNull = \priorDist(\theta = \theta_0)$, while the reported type is denoted by $\reportNull$. For each report $\reportNull$, the principal offers a contract consisting of a reward $\reward_\reportNull$, a $p$-value threshold $\threshold_\reportNull$, and a cost $\cost_\reportNull$. The agent’s expected payoff depends on the chosen contract and the state $Y$ as
This function can be interpreted as a scoring rule: the agent reports a probability $\reportNull$ and receives a payoff that depends on the realized outcome $Y$. If the agent’s true belief is $\priorNull$, his expected payoff from reporting $\reportNull$ is $\ensuremath{\mathcal{S}}(\reportNull, \priorNull) = \priorNull \big[\reward_\reportNull \nullAppSimple{\threshold_\reportNull} - \cost_\reportNull \big] + (1-\priorNull) \big[\reward_\reportNull \altAppSimple{\threshold_\reportNull} - \cost_\reportNull \big]$, which coincides with our previously defined utility function (ref).
A scoring rule is said to be proper if truthful reporting maximizes expected payoff, meaning that
and strictly proper if equality holds if and only if $\reportNull = \priorNull$. By construction, an incentive-compatible menu guarantees that each type $\priorNull$ strictly prefers the contract intended for them. Equivalently, it induces a scoring rule that is strictly proper. Thus, the design of an incentive-compatible menu in the hypothesis testing context can be viewed as the design of a strictly proper scoring rule tailored to the principal's statistical performance criterion. This connection also clarifies why separating menus can be constructed using convex real functions: every proper scoring rule admits a convex potential representation mccarthy1956measures, savage1971elicitation, gneiting2007strictly.
Our setting, however, differs from the abstract scoring-rule framework in the following way: the principal has only partial control of the agent's utility via the menu design. As evident from definition (ref), the induced scoring rule depends not only on contract parameters $(\threshold_\reportNull, \reward_\reportNull, \cost_\reportNull)$ but also on the type I error and power function determined by the hypothesis test itself. The principal can adjust contract parameters so that the resulting utility behaves like a proper scoring rule, but she cannot fully dictate payoffs independently of the test. This distinction foreshadows an important limitation: if additional constraints are imposed on the contract parameters, it may no longer be possible to induce a strictly proper scoring rule. We illustrate this in (ref), where requiring the reward to be constant restricts the class of hypothesis tests and the range of agent types for which the principal can implement separating menus.
Constructing separating menus enables the principal to assign type-optimal thresholds and elicit agents’ private beliefs. However, this design is not costless: in order to satisfy incentive compatibility, the principal must leave some surplus to agents. To understand what is lost by eliciting information through menus, we introduce two complementary measures of the trade-off, each defined with respect to a different benchmark.
\paragraph{Information rent.} The first benchmark is the full-information setting, in which the principal directly observes each agent’s type. With full information, she could implement the first-best contract by assigning the type-optimal threshold and adjusting the reward and cost so that each agent earns zero utility. Agents would still participate, but they would not retain any surplus.
However, under asymmetric information, the principal cannot condition contracts directly on types and must instead design a menu that induces truthful reporting. To do so, she must offer higher utility to agents with lower prior null to prevent misreporting. This additional utility is the information rent---the surplus that agents receive in order to implement incentive compatibility (see laffont2001,bolton2004contract). In our setting, the utility that an agent of type $\priorNull$ gains from participating is exactly $\ensuremath{\mathcal{G}}(\priorNull)$, so the total information rent for a separating menu constructed by $\ensuremath{\mathcal{G}}$ is
When the type distribution $\typeDist$ is uniform on $[0,1]$, the information rent corresponds to the area under the curve $\ensuremath{\mathcal{G}}$.
\paragraph{Screening cost.} Information rent highlights the gap to the unattainable first-best. Since it is the unavoidable price of asymmetric information, a more practical benchmark is the pooling contract, where the principal offers the same contract to all agents. The natural candidate is a baseline contract designed for the worst agent type $\worstPriorNull$ (the highest prior null), since this ensures participation of all agent types. To ensure incentive compatibility, however, the menu must give each agent with better type (that is, with lower $\priorNull$) strictly higher utility than he would obtain from pretending to be type $\worstPriorNull$. In this situation, the principal must sacrifice some surplus compared to simply offering a single contract. Following the contract theory literature, we refer to this surplus loss as the screening cost, since it is the cost of designing a menu that separates (or “screens") types rather than pooling them. Formally, we define the screening cost as
where $\utilfunc(\priorNull; \worstPriorNull)$ is the utility an agent of type $\priorNull$ would receive from the contract intended for $\worstPriorNull$. \\
In classical contract theory, there is typically a trade-off: a menu enables more precise targeting of types, but at the expense of screening costs; a single contract avoids such costs but results in inefficiencies from misallocation. One might expect the same tension here. Surprisingly, as we show later in (ref), there exist separating menus that achieve incentive compatibility with arbitrarily small financial cost to the principal. In some situations, the principal even realizes financial gains by offering separating menus rather than a single contract.
The key difference in our setting lies in how utility is transferred across types. In standard contracting problems, transfers are monetary rewards or cost adjustments, which reduce the principal’s payoff. In our setting, however, much of the utility increase for better-type agents comes from being assigned looser statistical thresholds, which raise the probability that their proposals are accepted.
This is not zero-sum; assigning more permissive thresholds to better types increases the likelihood of true discoveries while maintaining control of type I errors, which simultaneously benefits both the principal and the agent. Menu-based testing can therefore be Pareto-improving: agents receive strictly higher utility when presented with more options, while the principal is able to achieve better statistical performance.
(ref) characterizes the full class of separating menus that can elicit agents’ private beliefs and implement type-optimal thresholds. Yet this characterization leaves open a practical question: which separating menu should the principal actually implement? Because infinitely many functions $\ensuremath{\mathcal{G}}$ induce menus that satisfy participation and incentive compatibility, additional design criteria are needed to guide selection.
We focus on two such criteria, each linked to the discussion in earlier sections. First, as noted in (ref), compared to offering a single contract, a separating menu generally entails a screening cost to the principal. A natural question is therefore whether menus can be constructed to minimize this cost. In (ref), we show that there exist families of $\ensuremath{\mathcal{G}}$ that make the screening cost arbitrarily small, effectively rendering elicitation “free." Second, as emphasized in (ref), separating menus must induce strictly proper scoring rules. This requires flexibility in contract parameters, and when some parameters are fixed exogenously, the set of separating menus can shrink dramatically. In (ref), we examine the case where rewards are held constant and show that under this restriction, there is a unique separating menu.
Throughout this section, we restrict attention to differentiable functions $\ensuremath{\mathcal{G}}$ defined on $[0,\worstPriorNull]$, where $\worstPriorNull$ denotes the least favorable agent type that the principal is willing to accommodate. Any $\ensuremath{\mathcal{G}}$ satisfying conditions (ref) and (ref) is then differentiable, strictly convex, and nonnegative on this interval. Since separating menus guarantee statistical efficiency by construction, our analysis focuses on their financial cost relative to a benchmark. As in the definition (ref) of the screening cost, we compare each separating menu to a base contract $(\threshold_{\worstPriorNull}, \reward_{\worstPriorNull}, \cost_{\worstPriorNull})$ designed for the worst type $\worstPriorNull$. To ensure participation of all types up to $\worstPriorNull$ and non-participation of types exceeding $\worstPriorNull$, we impose a worst-case participation condition on the base contract
so that $\ensuremath{\mathcal{G}}(\worstPriorNull) = 0$.
Implementing a separating menu to elicit agents’ private beliefs generally entails a screening cost to the principal. By definition (ref), this cost is the extra utility agents receive relative to the base contract. From the principal’s perspective, screening cost is the expected financial loss that arises when operationalizing a menu rather than offering a single base contract. To see this, recall the benchmark in which the principal offers a fixed base contract $(\threshold_{\worstPriorNull}, \reward_{\worstPriorNull}, \cost_{\worstPriorNull})$ to all agents. Under this contract, the principal pays a reward $\reward_{\worstPriorNull}$ upon approval and collects a fixed payment $\cost_{\worstPriorNull}$ from each agent to cover trial costs. By contrast, under a separating menu, the principal tailors the contract terms to each type. For an agent of type $\priorNull$, the principal may adjust both the statistical threshold and the financial terms: additional rewards $\reward_\priorNull - \reward_{\worstPriorNull}$ upon approval, and extra payments $\cost_\priorNull - \cost_{\worstPriorNull}$ from the agent. These adjustments differentiate the menu from the base contract and implement type-specific targeting.
The principal’s expected financial return from offering the tailored contract $(\threshold_\priorNull,\reward_\priorNull,\cost_\priorNull)$ to type $\priorNull$, relative to the base contract, is
which simplifies to $- \big[\ensuremath{\mathcal{G}}(\priorNull) - \utilfunc(\priorNull;\worstPriorNull)\big]$. Integrating over the agent population $\typeDist$, the principal’s expected return is therefore the negative of the screening cost (ref). Since incentive compatibility requires the screening cost to be strictly positive, the principal necessarily incurs a financial loss from implementing a separating menu. In practice, resource constraints may make such losses problematic. Thus, an important design goal is to construct menus that minimize screening cost, which is equivalent to minimizing the principal’s financial loss while still guaranteeing incentive compatibility.
Using (ref), we can deduce an entire family of cost-reward pairs that yield separating menus. By careful choice, these menus can be designed to incur an arbitrarily small screening cost. In particular, let $z \mapsto \ensuremath{\epsilon}(z)$ be any differentiable function such that
See (ref) for the proof. \\
To elaborate upon the connection to (ref), designing a separating menu with a small screening cost is equivalent to designing a function $\ensuremath{\mathcal{G}}$---one which satisfies the properties of the theorem---such that it lies just above the utility curve $\ensuremath{\smallsub{\Psi}{\mbox{base}}}(\priorNull) \defn \utilfunc(\priorNull;\worstPriorNull)$ under the base contract. So that the theorem can be applied, the function $\ensuremath{\mathcal{G}}$ should be strictly convex. There are many possible choices of $\ensuremath{\mathcal{G}}$ that satisfy this requirement. Given a function $\ensuremath{\epsilon}$ satisfying the conditions (ref), one choice of $\ensuremath{\mathcal{G}}$---the one used in our proof---is
By choosing $\epsilon$ to be small, the principal achieves statistical efficiency while incurring negligible financial loss relative to the single-contract benchmark. Thus, the principal effectively elicits agents' private types for free.
So far, our construction of separating menus has assumed that the principal retains full control over all components of contracts---namely, the $p$-value threshold, reward, and cost. However, when one or more parameters are constrained, it may no longer be possible to induce a strictly proper scoring rule for all agent types. Such constraints are not merely theoretical: in practice, rewards or costs may be determined by forces outside the principal’s control. For example, in regulatory settings, the reward may be determined by market forces rather than by the regulator. In our running example of drug testing, the market generates the revenue for an approved drug, which is typically similar across firms producing treatments for the same condition. While the Food and Drug Administration (FDA) cannot set these rewards, it could, in principle, adjust the cost by imposing additional trial requirements or requiring a financial down payment.
With this motivation, we now study what separating menus can be constructed when contract parameters are constrained. Since the $p$-value thresholds are dictated by the principal’s statistical objective, the natural candidates are rewards or costs. In this section, we focus on the case in which all contracts share the same reward. The case of fixed costs is less compelling: it requires the power function $\altAppSimple{\cdot}$ to satisfy a form of convexity (see discussion in (ref)), which is unnatural since power functions are usually concave in $\threshold$.
\paragraph{Assumptions.} Fixing rewards makes the design of separating menus more challenging. In particular, we require additional structure on the power function $\altAppSimple{\threshold}$. Our result applies when $\altAppSimple{\threshold}$ satisfies two conditions:
This is a mild requirement. For Bayes-optimal decision rules, the mapping (ref) is strictly decreasing whenever the likelihood ratio is continuous and strictly monotone. For the FDR-control rule (ref), strict monotonicity follows under condition (ref) (see proof in (ref)).
See (ref) for the proof. \\
In this case, the connection to (ref) is provided by the function
Since we assume $\priorNull \mapsto \threshold_{\priorNull}$ is strictly decreasing, condition (ref) ensures $\ensuremath{\mathcal{G}}''(\priorNull) \defn \reward\big[1 - \beta_1'(\threshold_{\priorNull})\big] \threshold'_{\priorNull}$ is strictly positive, so $\ensuremath{\mathcal{G}}$ is strictly convex. Since $\ensuremath{\mathcal{G}}$ is uniquely determined within the class of differentiable functions under these conditions, the resulting separating menu is unique among all constructions derived from differentiable $\ensuremath{\mathcal{G}}$.
When the principal controls all contract parameters, she can design menus for any test satisfying the non-trivial power assumption (ref) and for all agent types on $[0,1]$. This flexibility is powerful---it allows full elicitation of private beliefs---but it also obscures how elicitation depends on the structure of the hypothesis test itself, since rewards and costs can always be adjusted to make things work. By contrast, when the reward is fixed, the dependence on the test becomes much clearer: the power function directly determines both the level of information rent and the extent of elicitable types.
\paragraph{Elicitation depends on test structure:} Recall that $\ensuremath{\mathcal{G}}(\priorNull)$ represents the information rent allocated to type $\priorNull$. Under the fixed-reward construction (ref), $\ensuremath{\mathcal{G}}$ is pinned down directly by the power function $\altAppSimple{\cdot}$, so that any change in the test translates immediately into a change in information rent. For example, if the principal employs a more powerful test $\tilde \beta_1(\threshold)$ with $\tilde \beta_1(\threshold) > \altAppSimple{\threshold}$ for all $\threshold \in [0,1]$, the information rent increases strictly: greater statistical power raises agents’ probabilities of receiving approvals and hence their utilities under the base contract, which in turn requires higher rents to maintain incentive compatibility. While information rent is always driven by incentive compatibility, the fixed-reward setting makes its dependence on the structure of the statistical test fully transparent. The structural condition (ref) also highlights the limits of elicitation under fixed rewards. Given a concave, differentiable power function, the inequality $\beta_1'(\threshold) > 1$ holds only on an interval $[0,\bar \threshold]$ (see (ref)(a) for an illustration). This places an upper bound $\bar \threshold$ on type-optimal thresholds in any separating menu. Since $\fthreshold{\priorNull}$ is decreasing in $\priorNull$, this implies that only agent types in $[\bestPriorNull,\worstPriorNull]$ can be elicited, where $\bestPriorNull$ is the type assigned threshold $\bar \threshold$. Hence, although the principal intends to elicit all types $\priorNull \in [0, \worstPriorNull]$, fixing the contract reward to be constant imposes structural limits on which types can be elicited, and these limits depend directly on the hypothesis test.
\paragraph{Financial cost of fixed-reward menus:} Compared to the menus constructed from $\ensuremath{\mathcal{G}}$ in (ref), the fixed-reward setting entails a non-trivial screening cost. With reward $\reward$ held constant, the curvature of $\ensuremath{\mathcal{G}}$ is large, since $\ensuremath{\mathcal{G}}''(\priorNull)$ cannot be tuned close to zero. As a result, the utility curve $\ensuremath{\mathcal{G}}$ lies strictly above the linear baseline $\ensuremath{\smallsub{\Psi}{\mbox{base}}}(\priorNull)$ associated with the base contract, and the screening cost cannot be made arbitrarily small. In some regulatory environments, however, the constant reward $\reward$ is not borne by the principal---for example, when it reflects the average market profits accruing to successful firms. In such cases, the principal can implement the menu by charging $\cost_{\worstPriorNull}$ to the worst type and requiring better types to pay a surcharge $\cost_{\priorNull} - \cost_{\worstPriorNull}$ in exchange for looser thresholds $\threshold_{\priorNull}$. Under this arrangement, the principal earns strictly positive revenue from better types, while maintaining incentive compatibility. This design is Pareto-improving: the principal can screen agents by type and implement type-optimal thresholds, while agents receive contracts tailored to their beliefs and strictly higher utilities than under the base contract. \\
Next, we present numerical results illustrating the convex functions $\ensuremath{\mathcal{G}}$ that underlie (ref), defined in equations (ref) and (ref), respectively. We do so in the context of Gaussian mean testing: consider a testing problem defined by the parameter space $\Theta = \{0, \theta_1\}$, with a null value $\theta_0 = 0$ and a non-null value $\theta_1 > 0$. A given observation $Z \sim \mathcal{N}(\theta, 1)$ can be converted into a $p$-value via the transformation $\evidence = 1 - \Phi(Z)$, where $\Phi$ is the standard normal cumulative distribution function. With this set-up, the null rejection probability is $\beta_0(\threshold) = \threshold$ by construction, while the power function under the alternative is given by:
where $\Phi^{-1}$ denotes the quantile function of the standard normal distribution.
First, consider a type-dependent reward with the function $\ensuremath{\mathcal{G}}$ from equation (ref), chosen to maximize the principal's expected financial return. For illustration, we focus on an alternative with $\theta_1 = 1$. We consider a worst agent type with prior null probability $\worstPriorNull = 0.8$ and use type-optimal thresholds (ref) for FDR control. The false discovery rate is constrained at level $0.25$, which implies an optimal $p$-value threshold of $\threshold_{\worstPriorNull} = 0.004$ for this agent. The base contract fixes the reward at $\reward_{\worstPriorNull} = 100$, with the cost $\cost_{\worstPriorNull}$ calibrated so that the worst type obtains zero utility under this contract.
As an illustration of the behavior of $\ensuremath{\mathcal{G}}$, we consider the perturbation function $\ensuremath{\ensuremath{\epsilon}{(z)}} \defn \eta\cdot(1-z)^2$ for some $\eta > 0$; in (ref), we plot the principal's expected financial return (ref), relative to the base contract, when using a separating menu constructed from $\ensuremath{\mathcal{G}}$ for the four choices $\eta \in \{0.01, 0.1, 0.5, 1\}$. Here, the area under each curve represents the principal's total financial loss for a population of agents with uniformly distributed types. Observe that setting $\eta$ close to zero results in smaller and smaller screening cost.
We next turn to the case of a constant reward across all contracts. We fix the reward at $100$ and again use type-optimal thresholds (ref) with an FDR constraint at level $0.25$. As analyzed in (ref), the function $\ensuremath{\mathcal{G}}$ now depends on the power function $\altAppSimple{\threshold}$. We consider three different alternatives for the nonnull hypothesis: $\theta_1 \in \{0.5, 1, 2\}$, with the null still given by $\theta_0 = 0$.
In (ref)(a), we plot the power functions corresponding to these values of $\theta_1$, and mark the maximum $p$-value threshold for which the condition $\beta_1^\prime(\threshold) > 1$ from equation (ref) holds. As $\theta_1$ increases, the power function becomes steeper, and the maximum threshold satisfying this condition decreases. This restricts the range of agent types for whom contracts with constant reward are incentive-compatible. As discussed in (ref), this maximum threshold imposes a lower bound on the agent types. The effect is exhibited in panels (b) through (d) of (ref), where we plot the corresponding functions $\ensuremath{\mathcal{G}}$. For example, when $\theta_1 = 0.5$, incentive-compatible contracts can be offered to agents with prior null beliefs in the interval $[0.33, 1]$. As $\theta_1$ increases to 2, this range shifts toward agents with higher $\priorNull$, i.e., agents who are more likely to face the null hypothesis. This shift occurs because when $\theta_1 = 2$, the maximum separating threshold is approximately 0.16, which is too stringent to be optimal for agents with lower $\priorNull$. The shaded regions in panels (b)--(d) represent the total information rent incurred when offering the menu to a uniformly distributed population of agent types. As previously discussed, we see that a more powerful test increases the information rent required to screen better types.
In this paper, we studied hypothesis testing over a heterogeneous population of strategic agents with private information. In this setting, any single uniform test yields sub-optimal performance, due to the underlying heterogeneity in agent types. At the other extreme, an oracle given a priori access to agent types can construct an optimal test for each type. Our main result was to show that it is possible for the principal to design a separating menu of type-tailored tests, coupled with appropriately designed payoffs, that induce agents to self-select according to their private information, thereby implementing type-optimal thresholds and achieving statistical efficiency. Strikingly, this improvement comes at negligible additional cost relative to single-test designs, demonstrating that information elicitation through incentive design can lead to better statistical performance with little cost.
Beyond the technical results specific to this paper, our analysis also offers some broader conceptual insights. First, it underscores the importance of treating agents as strategic and heterogeneous: conventional approaches that assume passive participants miss key opportunities for improving statistical outcomes. Second, it demonstrates how hypothesis tests themselves can serve as instruments of information elicitation. By embedding contracts in statistical design, the principal can shape agents’ utilities into a proper scoring rule, ensuring incentive compatibility. More broadly, the results further develop the connection between mechanism design and statistical decision-making and highlight that design in strategic settings requires jointly considering statistical objectives and agent behavior.
There are a number of natural extensions to the current analysis. First, we assumed that each agent’s private information is one-dimensional, summarizing only their belief about type. In practice, agents may hold richer, multi-dimensional information; for example, a researcher’s prior knowledge across multiple outcomes or a drug developer’s beliefs about several treatment effects. Second, the current analysis is predicated upon the principal knowing the test's power function; it would be interesting to explore extensions involving partial knowledge. A final opportunity lies in the scope of design and behavior considered. We focused on $p$-value thresholds and binary participation, abstracting away other levers. In reality, a regulator may be able to influence sample sizes, stopping rules, or test statistics, while agents may strategically respond to these choices. Extending the analysis to such richer action spaces could capture a wider range of interactions.
This work was partially funded by NSF-DMS-2413875 to SB and MJW, and the Cecil H. Green Chair and Ford Professorship to MJW.
\AtNextBibliography \printbibliography