Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
115,478 characters · 22 sections · 63 citation commands
A model of multiple hypothesis testing
\thispagestyle{empty}
\setcounter{page}{1}
Hypothesis testing plays a prominent role in evidence-based decision-making. Typically, researchers report results from more than one test, and there has recently been increasing interest in and debate over whether the testing procedures they employ should reflect this in some way---that is, whether to apply some form of multiple hypothesis testing (MHT) adjustment. As a concrete example, consider pharmaceutical companies reporting the results of clinical trials to regulators when seeking approval to market new drugs: the U.S. regulator (the Food and Drug Administration, FDA) recently released guidelines calling for MHT adjustments on the grounds that omitting them could “increase the chance of false conclusions regarding the effects of the drug” fda2022guidance. Analogous concerns arise in other settings, including experimental program evaluation in economics. As a result, a number of procedures for MHT adjustment have been proposed, and their statistical properties are well-understood romano2010hypothesis.
What is less clear is whether and when these procedures are economically desirable. That is, under what conditions does MHT adjustment lead to better decision-making from the point of view of the actor designing the process? The answer is far from obvious. It is certainly true, for example, that without MHT adjustments, the chance of making at least one type I error increases with the number of tests. But this is analogous to the truism that the more decisions one makes, the more likely one is to make at least one mistake. It is indisputable, but sheds no light on the pertinent questions, which are whether and how the rule for making individual decisions should change with the total number being made.
This paper provides a framework for analyzing such questions. We focus, in particular, on whether and when MHT adjustments arise as a solution to incentive misalignment between a researcher and a mechanism designer. Our interest in this case reflects two primary considerations. The first is substantive: incentives are clearly an issue in real-world cases of interest (e.g., in the drug approval process, which we will use as a running example). The instinctive concern many seem to have is that without MHT adjustments, the researcher would have an undue incentive to test many hypotheses in the hopes of getting lucky. We would like to formalize and scrutinize that intuition. And the second is pragmatic: to have a theory of MHT adjustments, we must have a theory that rationalizes hypothesis testing at standard levels in the first place, which (as we discuss below) is hard to do convincingly in a non-strategic setting tetenov2012statistical,tetenov2016economic.
Specifically, we study a model in which a benevolent social planner chooses norms with respect to MHT adjustments, taking into account the way this shapes researchers' incentives. The model embeds two core ideas. First, social welfare is affected by the summary recommendations (in particular, hypothesis tests) contained in research studies, and the planner also cares about the more generic benefits to society and to the researcher of conducting research per se.\footnote{We describe the case where hypothesis rejections affect welfare; under a straightforward reinterpretation the framework can also accommodate situations in which “precise null” results do so.} Second, while this makes the research a public good, the costs of producing it are borne privately by the researcher. She decides whether or not to incur these costs and conduct a (pre-specified) experiment based on the private returns to doing so. The planner must, therefore, balance the goals of (i) limiting the possibility of harm due to mistaken conclusions and (ii) motivating the production of research. We represent these preferences with a utility function that includes both ambiguity-averse and expected-utility components gilboa1989maxmin,banerjee2020theory, which (we show) turn out to have intuitive connections with the statistical concepts of size control and power. We focus on cases where multiplicity takes the form of testing multiple treatments or estimating effects within multiple sub-populations;\footnote{These forms of multiplicity are common in practice. For example, the majority of the clinical trials reviewed in pocock2002subgroup tested for effects in more than one subgroup. In economics, 27 of 124 field experiments published in “top-5” journals between 2007 and 2017 feature factorial designs with more than one treatment muralidharan2025factorial.} multiple outcomes are an economically distinct case covered in an earlier version of the paper VWN2025MHTv8.
We start by characterizing optimal hypothesis testing protocols. We show that the class of optimal protocols is the class of maximin optimal and unbiased protocols, where maximin optimality is closely connected to size control while unbiasedness requires the power of the protocols to exceed their size. We then prove that separate $t$-tests, which are ubiquitous in applied work, are maximin optimal and unbiased, and we provide an explicit characterization of the optimal critical value in terms of the researcher's costs.\footnote{We focus on one-sided $t$-tests in the main text and consider two-sided tests in Appendix (ref).}
We next characterize the role of multiplicity, drawing two broad conclusions. First, it is generically optimal to adjust testing thresholds (i.e., critical values) for the number of hypotheses. A loose intuition is as follows. The worst states of the world are those in which the status quo of no treatment is best; in these states, a research study has only a downside, and it is desirable to keep the benefits from experimentation low enough that the researcher chooses not to experiment. If the hypothesis testing protocol were invariant to the number of hypotheses being tested, then for sufficiently many hypotheses, this condition would be violated: the researcher's expected payoff from false positives alone would be high enough to warrant experimentation. Some adjustment for hypothesis count may thus be needed. This logic aligns fairly well with the lay intuition that researchers should not be allowed to test many hypotheses and then “get credit” for false discoveries. Interestingly, the same logic immediately implies that critical values should adjust for other factors that influence cost such as the sample size, though these have not attracted the same degree of attention.
Second (and as this suggests), economic fundamentals---in particular, the research cost function---determine exactly how much adjustment is required. When hypotheses are equally important, for example, the optimal critical values for the separate $t$-tests are given by
where $J$ is the set of hypotheses tested (with $|J|$ denoting its cardinality), $\Sigma$ captures features of the experimental design such as the sample size, $C(J, \Sigma)$ the cost of the experiment, and $b$ the benefit to the researcher of rejecting a null. When research costs are fixed, so that $C$ is invariant to $J$, this implies a Bonferroni correction.\footnote{Including, for example, in subgroup analysis contexts where experimentation costs are sunk.} When costs scale in proportion to the number of hypotheses, on the other hand, no MHT adjustment is required. Intuitively, the researcher has no undue incentive to test many hypotheses in this scenario because doing so is costly. This cost-based perspective also helps to clarify confusion about the boundaries of MHT adjustment and whether researchers should adjust for multiple testing across different studies. It suggests that MHT adjustments may be appropriate when there are cost complementarities across studies but not otherwise.\footnote{The broader principle is that optimal MHT adjustments depend on how exactly hypotheses interact. Our base model emphasizes interactions in the research cost; Appendix (ref) considers interactions of other kinds, through non-linearities in the researcher's payoff and through interactions in the planner's objective.}
To illustrate the quantitative implications of the model, we apply it to our running example, regulatory approval by the FDA. Applying the formulae implied by the model to published data on the cost structure of clinical trials, we calculate adjusted critical values that are neither as liberal as unadjusted testing, nor as conservative as those implied by some of the procedures in current use. If the appropriate level in the single-hypothesis case is 5%, for example, then the optimal level according to our formulae is 3.2% with two tests, 2.6% with three tests, and tends to 1.4% as $|J|\rightarrow\infty$. By comparison, the level implied by Sidak's correction vsidak1968multivariate for controlling the Family-Wise Error Rate (FWER) under independence (and, up to rounding, also Bonferroni), is 2.5% for two tests, 1.7% for three tests, and tends to zero as $|J|\rightarrow\infty$.\footnote{The Sidak correction is a natural benchmark because it is exact for controlling the FWER with independent tests, and FWER control is common in practice (see Figure (ref)).} These results suggest both that some adjustments are warranted but also that standard practices may be overly-conservative. Moreover, because costs also scale with the sample size, optimal adjustments must be less conservative for larger samples in order to induce researchers to incur the correspondingly larger costs.
It is also natural to wonder about applicability to economic research. The share of experimental papers published in “top 5” journals that used some form of MHT adjustment grew rapidly, from 0% in 2010 to 39% in 2020, so that there is now wide variability in whether (and how) these papers adjust (see Figure (ref)) and little consensus on what the norms should be. In Appendix (ref), we develop an additional empirical application to experimental program evaluation, using a unique dataset on the costs of projects submitted to the Abdul Latif Jameel Poverty Action Lab (J-PAL) from 2009 to 2021 which we assembled for this purpose. In this application, we find that the estimated adjustments implied by our model are less conservative than FWER control using Sidak's correction, but only slightly so. This is because the relationship between costs and number of treatment arms, while significant, is relatively weak in this setting, with a cross-sectional elasticity of approximately 15%.
Our paper draws inspiration from other work using economic models to select statistical procedures, in which researchers' preferences and incentives drive the analysis. This includes work on scientific communication andrews2021model,frankel2022findings, several aspects of which have been studied in more recent papers including bates2022principal,bates2023incentive and kasy2023optimal.\footnote{See also chassang2012selective,banerjee2017decision,spiess2018optimal,henry2019research,banerjee2020theory,williams2021preregistration,mccloskey2022incentive,yoder2022designing.} None of these papers analyze multiple hypothesis testing, however. The most closely related paper is the insightful work by tetenov2016economic, who shows that $t$-tests are maximin optimal and uniformly most powerful in the single-hypothesis case. Our extension to a multiple-hypothesis setting requires us to deal with two major challenges. First, the notions of maximin optimality and the corresponding theoretical results are more complex because the effects of different treatments may have opposite signs. Second, within the (large) class of maximin optimal protocols, none uniformly dominates all others, requiring us to develop new notions of optimality suitable to the context.
Our paper also relates to an extensive literature at the intersection between decision theory and hypothesis testing, dating back to wald1950statistical and robbins1951asymptotically. Previous work has motivated notions of compound error control in single-agent non-strategic environments; see in particular kline2022systemic,kline2024discrimination for recent examples in economics based on a Bayesian interpretation of the false discovery rate (FDR), as well as storey2003positive, lehmann2005testing, efron2008simultaneous, and hirano2020handbook for further examples.\footnote{The literature on statistical treatment choice has similarly focused for the most part on non-strategic planner problems. See manski2004 and tetenov2012statistical as well as hirano2009asymptotics,KitagawaTetenov_EMCA2018,athey2017efficient for recent contributions.} We complement this literature (as well as the statistical literature discussed below) by developing a model that explicitly incorporates the incentives and constraints of the researchers. Relative to the decision-theoretic approach, this has two main advantages. First, it lets us characterize when MHT adjustments are appropriate---and also when they are not---as a function of measurable features of the research process. Second, it allows us to justify and discriminate between different notions of compound error (e.g., average error rate or the FWER) in the same framework based on these same economic fundamentals.
Finally, we aim to provide guidance for navigating the extensive statistical literature on MHT. This literature provides procedures for controlling particular notions of compound error,\footnote{See efron2008microarrays and romano2010hypothesis for overviews. } but few statistical optimality results spjotvoll1972optimality, lehmann2005optimality,romano2011consonance, and none in which MHT procedures address an incentive problem. We maximize a different (social planner's) objective, subject to incentive compatibility constraints. We also draw on list2019multiple's helpful distinction between different types of multiplicity, and show how these lead to different optimal testing procedures.
We study MHT in a game between a social planner who chooses statistical procedures and a representative experimental researcher with private incentives. In our running example, we can think of the planner as a regulator (e.g., the FDA) who defines testing protocols for studies submitted in support of applications for the approval of new drugs, and the researcher as a pharmaceutical company running a pre-specified clinical trial of a new drug hoping to obtain such approval. Multiple testing issues arise whenever research informs multiple decisions. We focus on settings with multiple treatments (e.g., multiple drugs) or different subpopulations (e.g., multiple demographic groups for which a drug may be approved); for brevity we will refer throughout to treatments, taking this to refer to multiplicity of both types.
To say something coherent about MHT, a framework must be able to rationalize conventional (single) hypothesis testing in the first place. This is known to be a challenging problem, requiring non-trivial restrictions on the research process tetenov2016economic---in particular, strong asymmetries to match the inherently asymmetric nature of null hypothesis testing. For example, tetenov2012statistical shows that justifying testing at conventional levels in a single agent model with minimax regret requires extreme degrees of asymmetry: statistical tests at the 5% (1%) level correspond to decision-makers placing 102 (970) times more weight on type I than type II regret. Here the asymmetry necessary for rationalizing hypothesis testing will arise naturally from the planner's desire to prevent the implementation of harmful treatments.
We consider a two-stage game between the planner and the researcher. In the first stage, the planner prescribes and commits to a hypothesis testing protocol, restricting how the researcher can report findings. In the second stage, given this protocol, the researcher decides whether or not to run one of several possible experiments by comparing the private benefits of experimentation to the private costs. Importantly, these private benefits may differ from the planner's objective. Unless noted otherwise, we will assume that the researcher's preferences are common knowledge and that she is not allowed to mis-characterize them.
Hypothesis testing protocols take as input the data from the experiment and output multiple binary findings indicating whether the treatments were found to be effective. These findings, in turn, affect social welfare: we will interpret a finding as equivalent to the planner's decision to implement the corresponding treatment. Since the planner selects the hypothesis testing protocol, this is equivalent to the planner pre-committing to a decision rule and the researcher truthfully reporting the decisions it implies, given the observed data.
The researcher takes the hypothesis testing protocol as given and decides, before observing data, whether and how to experiment. We first describe the experiment and hypothesis testing protocol and then introduce the researcher's optimization problem.
Experiment. Let $\mathcal{J}$ denote the finite set of all combinations of non-exclusive treatments that can be tested in an experiment (the “power set”), with $\emptyset \in \mathcal{J}$ denoting no experimentation. The parameter vector $\theta \in \Theta$, where $\Theta$ is a compact parameter space, captures the effects corresponding to all possible combinations of treatments in $\mathcal{J}$.
An experiment consists of a set of treatments $J \in \mathcal{J}$ and a design $\Sigma \in \mathcal{S}(J)$, which are chosen by the researcher. Here $\mathcal{S}(J)$ is the set of all possible designs given $J$. If the researcher experiments, she draws a vector of statistics $X$ from a distribution $F_{\theta, J, \Sigma}$, indexed by $J$, $\theta$, and $\Sigma$. The design $\Sigma$ summarizes all the relevant features of the distribution of $X$ the researcher can choose ex-ante, such as the sample size of the experiment. The researcher pre-specifies and reports $J$ and $\Sigma$ before running the experiment, so that they become common knowledge. $J$ and $\Sigma$ may depend on the researcher's prior knowledge and private incentives, but not on the realized statistics $X$. We thus abstract from issues of $p$-hacking and selective reporting. This case is relevant for considering decision-making at the FDA, for example, which requires pre-registration.\footnote{Specifically, the summary of the Final Rule for Clinical Trials Registration and Results Information Submission (42 CFR \S 11) on ClinicalTrials.gov states that “[r]egistration is required for studies that meet the definition of an `applicable clinical trial' (ACT) and either were initiated after September 27, 2007, or initiated on or before that date and were still ongoing as of December 26, 2007” final_rule. Registration must specify, among other things, the intervention(s), primary outcomes, and intended enrollment and study design 42cfr1128.}
Hypothesis testing protocols. As described above, the researcher first chooses and pre-specifies an experiment $(J,\Sigma)$. She then runs the experiment, as a result of which the vector of statistics $X$ is realized and becomes publicly available. The results from the experiment are reported in the form of a vector of non-exclusive findings or recommendations,
where $r_{j}(X;J, \Sigma) = 1$ if and only if the treatment corresponding to $J_{j}$ is found to be effective, with $J_{j}$ denoting the $j^{th}$ entry of $J$. If $J = \emptyset$, no findings are reported, $r(X,\emptyset,\Sigma)=0$. If there are no findings ($r(X,J,\Sigma)=0$), the status quo prevails. We will refer to $r$ as a hypothesis testing protocol.
To simplify notation when describing the researcher's payoff and welfare below, it is useful to introduce the selector function $\delta(r(X;J, \Sigma); J) \in \{0,1\}^{2^{|J|} - 1}$. Each entry of $\delta(r(X;J, \Sigma); J)= \left(\delta_1(r(X;J, \Sigma);J),\dots,\delta_{2^{|J|} - 1}(r(X;J, \Sigma);J)\right)$, corresponds to one of the $2^{|J|} - 1$ possible combinations of the $J$ treatments. Specifically, for $k=1,\dots,2^{|J|} - 1$, $\delta_k(r(X;J, \Sigma);J)=1$ if the treatment combination $k$ is found to be effective and $\delta_k(r(X;J, \Sigma);J)=0$ otherwise. We let $\delta(r;\emptyset) = 0$ for all $r$. By definition of $\delta$, we have that
Here $\delta(\cdot;J)$ only takes into account combinations of the treatments in the set $J$, ignoring combinations not studied in the experiment. Example (ref) provides an illustration of $\delta$.
The researcher's objective. For each $(J,\Sigma)$, the researcher takes as given the corresponding hypothesis testing protocol $r(\cdot;J,\Sigma)$, which is chosen by the planner in the first stage. For simplicity, we assume that the researcher knows $\theta$, but our main results continue to hold when the researcher is imperfectly informed and has a prior about $\theta$ (see Section (ref)).
We consider settings where the researcher's and the planner's objective are misaligned. We model misalignment using researcher's utility of the form $p(J)^\top \delta(r(X;J,\Sigma)) - c_\theta(J,\Sigma)$ where $p_k(J)\geq 0$, $k=1,\dots, 2^{|J|} - 1$, is the benefit from getting approval for the $k$th combination of treatments, conditional on the set of treatments $J$ being tested, and $c_{\theta}(J, \Sigma)$ captures the costs of research, which can be a function of $(J,\Sigma,\theta)$.\footnote{For instance, we could write $c_{\theta}(J,\Sigma) = C(J,\Sigma) - b(J,\theta)$ for some function $b(\theta, J)$ of $(J,\theta)$ that is the part of benefits that depend on $\theta$ ($b(\theta,J)$ can be an implicit function of $p$) and the costs $C(J,\Sigma)$ as a function of the design. Practically speaking however this requires being able to measure $b(J, \theta)$ in addition to the costs.} Taking expectations over $X$ yields the following class of expected researcher payoff functions,
Thus, $\beta_r(\theta, J,\Sigma)$ captures the net benefits of experimenting. We let $c_\theta(J,\Sigma) \ge 0$ for all $(J,\Sigma,\theta)$ and $c_\theta(\emptyset, \Sigma) = 0$ for all $(\theta,\Sigma)$, so that the researcher's net benefits are normalized to zero if no experiment is conducted. Misalignment arises because the researcher's payoff is different from the planner's objective. In the following, whenever we consider settings where the costs do not depend on $\theta$, we will write $c_{\theta}(J,\Sigma) \equiv C(J,\Sigma)$ for a function $C(J,\Sigma)$ that does not depend on $\theta$.
The researcher chooses which treatments to study and how to design the experiment so as to maximize her net benefits. Formally, the researcher's problem is
where $J_{r,\theta}^* = \emptyset$ corresponds to no experimentation. To state the theoretical results, we also define the experiment the researcher would choose when forced to run an experiment,
We impose a standard tie-breaking rule: whenever the researcher is indifferent regarding whether to experiment (i.e., when $\beta_r(\theta, J_{r,\theta}^*, \Sigma_{r,\theta}^*) = 0$), she experiments if the planner's utility (defined below) is weakly positive. Let $e_r^\ast(\theta)$ indicate whether the researcher experiments, $e_r^\ast(\theta)=1\left\{J_{r,\theta}^* \neq \emptyset\right\}$.
Example (ref) continued. Suppose that $\mathcal{J}=\left\{\emptyset, \{1\}, \{2\}, \{1,2\}\right\}$. Let $p(\{1\}) \in \mathbb{R}_+$, $p(\{2\}) \in \mathbb{R}_+$, and $p(\{1,2\}) \in \mathbb{R}_+^3$ denote the researcher's (vector of) benefits from getting approval, conditional on the set of treatments $J$ being tested. In this example, the researcher's payoff for $J = \{1,2\}$ equals $$
$$ \qed
The social planner chooses a hypothesis testing protocol $r \in \mathcal{R}$ to maximize her utility, where $\mathcal{R}$ is the class of all (pointwise measurable) protocols. That is, $\mathcal{R}$ contains all protocols typically found in practice, including (and not limited to) standard $t$-tests. The planner's utility will depend on the welfare effects of implementing the recommended treatments as well as a measure of the more generic benefits to society and to the researcher of conducting research per se.
Welfare. Welfare depends on whether the researcher experiments and on her findings if she does. To define welfare, for $k=1,\dots,2^{|J|}-1$, let $u_k(\theta; J)$ denote the effect on welfare that would result from implementing the combination of treatments $k$.
Given $r(X;J, \Sigma)$, the overall welfare is $u(\theta; J)^\top \delta(r(X;J, \Sigma);J)$. We normalize $u(\theta;\emptyset) = 0$ for all $\theta \in \Theta$. That is, welfare is equal to zero under the status quo when no experimentation occurs. For a given experiment $(J,\Sigma)$, the expected welfare is
If the researcher does not experiment, $v_r(\theta, \emptyset, \Sigma) = 0$ for all $\theta \in \Theta$.
Example (ref) continued. Suppose that $\mathcal{J}=\left\{\emptyset, \{1\}, \{2\}, \{1,2\}\right\}$ and denote by $\theta_1$ and $\theta_2$, the welfare from implementing treatment 1 and 2, respectively. In this case, $u(\theta;\emptyset)=0$, $u(\theta;\{1\})=\theta_1$, and $u(\theta;\{2\})=\theta_2$. If there are no interaction effects, so that the welfare effect from implementing both treatments is equal to the sum of effects from implementing each of them separately, then
\qed
The planner's objective. We consider a planner who wishes to increase welfare while also limiting the possibility of harm due to mistaken conclusions, and to encourage research. Specifically, the planner chooses $r$ to maximize
where $\lambda\ge 0$ and $\pi(\theta) \ge 0$ for all $\theta \in \Theta$. The first component, which depends on which treatments are actually implemented, captures the desire to raise welfare while limiting harm using a standard ambiguity-averse (maximin) formulation. The second component depends on whether or not the researcher experiments. (If $\pi$ is a probability density, that second component is equal to the probability of experimentation since $\int e_r^\ast(\theta) \pi(\theta) d\theta=\int 1\{J_{r,\theta}^\ast\ne \emptyset\} \pi(\theta) d\theta$). It can be interpreted as capturing the benefits of scientific research per se, and as internalizing some aspects of the researcher's utility. In Appendix (ref), we formalize the latter interpretation, showing that protocols that are optimal under $U$ remain approximately optimal if the second component is replaced by $\int\beta_r^\ast(\theta)\pi(\theta)d\theta$, the expected researcher utility (if $\pi$ is a density). The parameter $\lambda$ allows us to trade-off each of these components. We show below that under suitable assumptions on $\pi$, the two components of $U$ have intuitive connections to the statistical concepts of size control and power. Moreover, $U$ admits optimal protocols that do not depend on $\lambda$ and $\pi$. This is important because the relative importance of the two components may be difficult to determine and choosing high-dimensional weights $\pi$ is difficult and often somewhat arbitrary. Working with $U$ thus provides a cogent rationale for testing protocols that control size, have non-trivial power, and do not depend on $\lambda$ and $\pi$.
In the regulatory approval example, the structure of the planner objective $U$ is motivated by regulators such as the FDA being tasked by legislators with several distinct objectives fda_mission. Each component of Equation (ref) relates to a distinct objective. The first captures the desire to avoid implementing harmful treatments, as for example under the “do no harm” principle (since, we show, our framework naturally rationalizes one-sided hypothesis testing). The second captures the broader value of scientific research, which need not be directly related to the immediate regulatory decision being made.\footnote{As the international guidelines for clinical trials state, for example, “the rationale and design of confirmatory trials nearly always rests on earlier clinical work carried out in a series of exploratory studies” lewis1999statistical. More broadly, the results of one study may lead to new conceptual insights or scientific hypotheses which are valuable independent of any immediate clinical application.} The relative importance of these two objectives is generally not specified, however, which motivates focusing on protocols that are optimal for all $\lambda\ge 0$.
The weighting of ambiguity-averse and expected-utility components in the planner objective $U$ echoes a long tradition in economic theory gilboa1989maxmin,banerjee2020theory. The planner objective $U$ differs from (but, as we discuss below, approximates) the objectives in gilboa1989maxmin and banerjee2020theory, which, in our notation, correspond to
for weights $w(\theta)$. The second component of $U'$ captures the welfare from implementing the treatments, whereas the second component of $U$ captures a preference for experimentation. The objective $U'$ has a decision-theoretic interpretation gilboa1989maxmin and is related to Huber's $\varepsilon$-contamination model banerjee2020theory. However, Appendix (ref) shows that exact solutions under $U'$ do not necessarily guarantee size control (as the solution depends on $\lambda$ and $w$). By contrast, working with the planner objective $U$ allows us to justify notions of size control and power and to obtain optimal protocols that do not depend on $w$ and $\lambda$, while retaining an approximate decision-theoretic justification.
In this section, we characterize optimal hypothesis testing protocols without imposing additional functional form restrictions on the researcher's payoff or the planner's utility.
Null space and alternative space. For a given set of treatments $J$, define the (global) null space, the set of parameters for which the welfare effect of implementing any combination of treatments is negative, as
Similarly, define the null space given the treatments chosen by the researcher ex-ante after excluding the option not to experiment, $J_{r,\theta}^+$, as $$ \Theta_0^*(r) = \left\{\theta: u_k(\theta, J_{r,\theta}^+) < 0 \text{ for all } k \in \{1, \dots, 2^{|J_{r,\theta}^+|}\} \right\}. $$
Moreover, define the (global) alternative space, the set of parameters for which welfare is always positive, as
A graphical illustration of $\Theta_0(J)$ and $\Theta_1(J)$ is provided in Figure (ref). We impose the following assumption.
Assumption (ref) states that for all combinations of treatments $J$, welfare is strictly negative for some values of $\theta$ and weakly positive for some other values of $\theta$.\footnote{This precludes interventions that everyone believes are sure to do good or those sure to cause harm, which are not of interest and would be precluded by, for example, rules regarding research ethics.}
Finally, denote by $\bar{\Theta}_1$ the set of parameters for which welfare is weakly positive for each choice of treatments $J$, $ \bar{\Theta}_1 = \bigcap_{J \in \mathcal{J} \setminus \emptyset} \Theta_1(J). $
Notions of optimality. The main notion of optimality we consider is uniform global optimality. We say that a protocol $r^*$ is uniformly globally optimal if
for a given set $\Pi$. Uniformly globally optimal protocols do not depend on $\lambda$ and $\pi$, which is important in practice, as argued above. In Appendix (ref), we discuss the relationship between uniform global optimality and alternative notions of optimality.
We also introduce two additional definitions that will be helpful for characterizing uniformly globally optimal protocols: maximin optimality and unbiasedness. We say that $r^*$ is maximin optimal if it maximizes the planner's objective (ref) for $\lambda = 0$, that is, $$ r^* \in \arg \max_{r \in \mathcal{R}} \min_{\theta \in \Theta} v_r(\theta, J_{r,\theta}^*, \Sigma_{r,\theta}^*). $$ We say that $r^*$ is unbiased if
Here $\beta_r^*(\theta)$ is the largest net benefit the researcher can achieve when conducting an experiment. We use the term “unbiased” because, as we show below, this definition has a close connection to the definition of unbiased tests in the hypothesis testing literature, where a test is called unbiased if its power exceeds its size lehmann2005testing.
Here, we characterize uniformly globally optimal and maximin optimal protocols. We impose the following assumption on the planner's weights $\pi$.
Assumption (ref) restricts the support of the planner's weights to the alternative space $\bar{\Theta}_1$. In other words, the planner desires to promote experimentation if she expects that treatments will generate a positive welfare effect and are therefore worth exploring. She derives no benefit (but also no harm), on the other hand, from exploration of parts of the parameter space in which some treatment effects are negative.
Using classical hypothesis testing terminology, Assumption (ref) allows for arbitrary alternative hypotheses over the positive orthant, including those in Section 4 of romano2011consonance and Chapter 9.2 of lehmann2005testing. Economically speaking, one can think of this assumption as ensuring that the components of the planner's utility function (ref) cleanly separate the two motives we wish to capture: avoiding harm and pursuing benefit. Doing so has the benefit that the components of utility will then map directly into the classical statistical concepts of size and power. As we discuss in more detail in Remark (ref), Assumption (ref) is necessary to justify testing protocols that control size, and we thus see Assumption (ref) as natural when size control is a desideratum.
The following proposition characterizes uniformly globally optimal protocols.
Proposition (ref) shows that uniformly globally optimal protocols are maximin optimal and unbiased. It assumes that a maximin and unbiased protocol exists. While the existence of such protocols is not guaranteed in general, we will show in Section (ref) that such protocols exist in leading cases. In the remainder of this section, we study maximin optimality and unbiasedness in more detail and show that these properties are related to notions of size control and power of testing protocols.
The next proposition provides necessary and sufficient conditions for maximin optimality. It generalizes Proposition 1 in tetenov2016economic (discussed in Appendix (ref)) to a setting in which $|J| > 1$ and where researchers can choose the experimental design and hypotheses to test.
Proposition (ref) shows that maximin optimality is equivalent to two conditions. First, as in the case with a single hypothesis (i.e., $|J|=1$), maximin protocols deter experimentation over the null space $\Theta_0^*(r^*)$, where each configuration of treatments reduces welfare. This captures a notion of size control, as we discuss below. Second, welfare must be non-negative for $\theta\in\Theta\setminus \Theta_0^*(r^*)$. This condition requires that if some treatments reduce welfare there must be others that compensate for them; it always holds in the single-hypothesis case but is non-trivial in the MHT case.
Building on Proposition (ref) and the sufficient conditions in Proposition (ref), the following corollary provides sufficient conditions for uniform global optimality stated in terms of the global null and alternative spaces.
Conditions (ref) and (ref) in Corollary (ref) mimic those in Proposition (ref), but are required to hold for all potential choices of $J$ and $\Sigma$. These conditions may be easier to check than those in Proposition (ref) because they are stated in terms of $\Theta_0(J)$ rather than $\Theta_0^\ast(r^*)$, which is itself a function of the set of treatments pre-specified by the researcher in response to $r^\ast$. Note that in Condition (ref), we need weakly positive welfare only when the researcher finds it beneficial to experiment, since otherwise the researcher will not experiment and hence welfare will be zero. Condition (ref) captures a stronger notion of unbiasedness that implies unbiasedness as defined in Equation (ref).
The theoretical results in this section establish close connections between uniform global optimality on the one hand, and size control and the power of testing protocols on the other. Specifically, maximin optimality captures a notion of size control, whereas unbiasedness captures a notion of power.
As discussed in Section (ref), rationalizing hypothesis testing (let alone multiple testing) is difficult in practice. The results and examples in this section show that one can write down a coherent economic objective function that rationalizes the standard statistical practice of choosing protocols that both control size and have non-trivial power. Specifically, optimal protocols must guarantee size control (encoded in the maximin optimality requirement) and also guarantee sufficient power against alternatives (encoded in the unbiasedness). Expressing these requirements in the form of an optimization problem has the benefit that it will allow us to then link the notions of size control and compound error rates directly to economic fundamentals, depending on the researcher's private costs and benefits and on how these scale with the number of hypotheses.
So far, we have assumed that after observing the protocol $r$, the researcher can choose any experiment $(J,\Sigma)$ with $J\in\mathcal{J}$ and $\Sigma\in S(J)$. In some applications, however, researchers may face constraints that restrict the menu of treatments and designs they can implement. Because the planner may not know ex ante which experiments are feasible, we consider a refined notion of global optimality, referred to as design-robust global optimality, that builds in robustness to design constraints.
Definition (ref) requires protocols to be optimal for every potential set of experiments $(J,\Sigma)$ available to the researcher. Without the refinement of global optimality in Definition (ref), the planner can afford to select protocols that are not unbiased (and hence may lead to low power) for designs that she expects the researcher not to choose. With it, she must consider the possibility that the researcher may be forced to choose any design, and thus must ensure that protocols are always unbiased.
Definition (ref) is appealing in settings where the planner has limited ex-ante knowledge of the researcher’s feasibility constraints, and thus wishes to mandate a protocol that performs well uniformly across all designs.
The following proposition provides a characterization of design-robust globally optimal protocols.
Proposition (ref) shows that the sufficient conditions for global optimality in Corollary (ref) become necessary once we strengthen the notion of global optimality to design-robust global optimality.
So far, we have assumed that the planner knows the researcher's payoff. The following proposition shows that maximin optimality for protocols satisfying Equations (ref) and (ref) is preserved if the planner knows only an upper bound on the researcher's payoff. Uniform global optimality is preserved under additional restrictions on $\Pi$.
Proposition (ref)((ref)) demonstrates an important robustness property of our maximin optimality results in settings where the researcher's payoff function is unknown. Proposition (ref)((ref)) states that protocols that are uniformly globally optimal with respect to an upper bound $\beta_r(\theta,J,\Sigma)$ are also uniformly globally optimal for weights $\pi \in \tilde\Pi$. That is, uniform global optimality is preserved when considering a (weakly) smaller class of alternatives. For example, for the optimal separate $t$-tests in Section (ref), the set of weights $\tilde{\Pi}$ is a subset of the set of weights on strictly positive treatment effects. Note that $\tilde{\Theta}_1(r^*)$ and $\tilde{\Pi}$ do not need to be known for the planner to implement the optimal protocol.
Proposition (ref) is particularly useful when applied to settings in which the planner's uncertainty about the researcher's payoff hinges on the researcher's costs.
Corollary (ref) states that in settings with uncertainty over the true cost function $C'(J,\Sigma)$, the planner may use sensible lower bounds $C(J,\Sigma) \le C'(J,\Sigma)$. This result is important in empirical applications such as the ones we consider in Section (ref).
Which (if any) specific hypothesis testing protocols are uniformly globally optimal? The answer depends on the functional form of the researcher's payoff, the functional form of welfare, and the distribution of $X$. Here, we show that separate $t$-tests are optimal in settings with a linear researcher payoff function (Assumption (ref)), a linear welfare function (Assumption (ref)), and a normally distributed vector of statistics $X$ (Assumption (ref)).
Linearity assumptions. Let $\omega$ denote a vector of weights and define $\bar{\omega}(J) = \sum_{j= 1}^{|J|} \omega_{J_j}$ for $J\in \mathcal{J}$. These weights will let us capture factors that affect the importance of the different treatments symmetrically from the point of view of both the researcher and the planner; they also nest the case in which all treatments are equally important ($\omega_{J_j} = 1$ $\forall j$). We consider the following assumption on the researcher's payoff.
The payoff function (ref) in Assumption (ref) is a special case of the general payoff function (ref), and the condition $b \bar{\omega}(J) > C(J,\Sigma)$ guarantees that the experiment $(J, \Sigma)$ is a relevant option; otherwise the researcher would never conduct this experiment, regardless of $r$. Note that Assumption (ref) rules out interactions between treatments in the researcher's utility, which we discuss in Appendix (ref). In addition, we assume that the researcher costs do not depend on $\theta$ and write them as $ C(J,\Sigma)$.
We consider the following linearity assumption on welfare.
Assumptions (ref) and (ref) capture a setting in which the researcher's and planner's objectives differ: the expected researcher's payoff depends on the (weighted) expected number of findings, whereas the welfare component of the planner's objective depends on the welfare effects that such findings generate. Before stating results, we provide an interpretation of these assumptions in the context of our leading example, the drug approval process.
With multiple subgroups, we interpret $r_j(X;J,\Sigma)$ as indicating whether the drug was found to be effective for subgroup $J_j$, which is of size $\omega_{J_j}$. The component $b \omega_{J_j}$ denotes the expected profits from selling the drug to subgroup $J_j$, where $b$ denotes the average per-sale profit. Assumption (ref) then states that researchers care about the sum of the expected profits they can earn by selling the drug to each of the subpopulations for which its use is approved. Our specification of welfare in Assumption (ref), meanwhile, corresponds to the utilitarian welfare from approving the drug, where $u(\theta;J_j)$ denotes the per unit treatment effect on members of subgroup $J_j$. The economic import of the assumption is that there are no spillovers between different subgroups.
With multiple treatments the interpretation is similar, but here each treatment denotes a different drug in the same market. $\omega_{J_j}$ denotes the expected number of individuals that would purchase and use drug $J_j$ if approved, and $u(\theta;J_j)$ is the effect of drug $J_j$ on those individuals. As before, $b$ denotes the average per-sale profit. The economic import of Assumption (ref) is that the sets of people who would use the different drugs are disjoint (or that the drugs do not exhibit interaction effects). Symmetry between the planner's and researcher's weights here implies that per-user profit and per-user consumer surplus are proportional across drugs. Without this property it is no longer necessarily the case that the separate $t$-tests defined in Equation (ref) below guarantee uniformly non-negative welfare (and therefore are maximin optimal); we do not characterize optimal testing protocols in that scenario, but note that this could be a fruitful direction for future work.
\noindentNormality assumption. We focus on the leading case where $X$ is normally distributed, which is motivated by standard normal approximations (e.g., Berry-Esseen bounds). Let $\theta_J$ denote the subvector of $\theta$ corresponding to treatments $J$.
Assumption (ref) imposes that the vector of statistics $X$ is normally distributed, centered around the vector of welfare effects $\theta_J$ and with finite sample covariance matrix $\Sigma$.\footnote{Because $\Sigma$ is the finite sample covariance matrix, it is proportional to the inverse of the square root of the sample size in the experiment.} The researcher can choose the covariance matrix, but we restrict the class of designs she can choose from to designs in which the variances (the diagonal entries of $\Sigma$) are positive and equal to each other. This implies that the researcher can choose the overall sample size of the experiment, for example, but is constrained in allocating sample across experimental arms. See Section (ref) for an extension to designs with heterogeneous variances.
The following proposition shows that separate $t$-tests are maximin optimal under Assumptions (ref), (ref), and (ref) and uniformly and design-robust globally optimal if, in addition, Assumption (ref) holds.
Proposition (ref) shows that separate one-sided $t$-tests with critical values $t\ge \Phi^{-1}\left(1 - \frac{C(J, \Sigma)}{b\bar{\omega}(J)} \right)$ are maximin optimal, where we write $t$ in lieu of $t(J,\Sigma)$ to simplify notation. The key technical step in the proof is to show that under this protocol welfare is non-negative even when parameters have different signs (Equation (ref)). Proposition (ref) further shows that standard separate one-sided $t$-tests are maximin optimal and unbiased and thus uniformly globally optimal by Proposition (ref) and also design-robust globally optimal.\footnote{It may seem surprising that the result in Proposition (ref) holds for any $\lambda$. Intuitively, once we focus on $\Pi$ defined in Assumption (ref), we can always find a maximin protocol that guarantees experimentation for each value of $\theta$ in the positive orthant. As a result, no matter how much weight the planner puts on her subjective utility, at the optimum, any optimal protocol will maximize separately (and therefore jointly) maximin welfare and subjective utility from experimentation.} While $t$-tests with all thresholds larger than $\Phi^{-1}(1 - \frac{C(J, \Sigma)}{b\bar{\omega}(J)})$ are maximin optimal, such tests are not uniformly globally optimal. Only $t$-tests with threshold $t=\Phi^{-1}(1 - \frac{C(J, \Sigma)}{b\bar{\omega}(J)})$ for at least some $(J,\Sigma)$ are uniformly globally optimal, and only $t$-tests with threshold $t=\Phi^{-1}(1 - \frac{C(J, \Sigma)}{b\bar{\omega}(J)})$ for all experiments $(J,\Sigma)$ are design-robust globally optimal. This demonstrates that uniform global optimality is a refinement of maximin optimality, restricting attention to protocols with sufficient power for at least some experiment, and that design-robust global optimality is in turn a refinement of uniform global optimality, restricting attention to protocols with sufficient power for all experiments.
Proposition (ref) also shows that whether and to what extent the level of these separate tests should depend on the number of hypotheses being tested depends on the structure of the research production function $C(J, \Sigma)$ and on $\bar{\omega}(J)$, and in particular on how they vary with $J$. For example, suppose that $\omega_j = 1$ for all $j$ (all treatments are equally important) so that $\bar{\omega}(J)=|J|$. If $C(J, \Sigma) = \alpha$ for some constant $\alpha$ then a Bonferroni correction is optimal. This corresponds to a stylized case in which the costs of experimentation are fixed regardless of the number of treatments tested ($|J|$) or the precision of the estimates ($\Sigma$). If, on the other hand, $C(J, \Sigma) = \alpha |J|$ then the optimal level of the test is $\alpha$, irrespective of $|J|$. This might correspond, for example, to a case in which there are no fixed costs and testing each additional treatment requires the same increment to the sample size. The former case arguably captures the lay intuition that if the researcher can test many hypotheses in the hopes of securing some private benefit, then the planner should require a hypothesis testing protocol that discourages this. In the latter case it is still true that the researcher obtains a higher expected reward from taking on projects that test more hypotheses, ceteris paribus, but the appropriate correction for this is already “built in” to the costs of conducting research, so that no further correction is required.
While our original motivation for obtaining the result in Proposition (ref) was to study the consequences of multiplicity ($|J|$), it follows immediately that, because the costs $C(J, \Sigma)$ may depend on the design $\Sigma$ in addition to $J$, the optimal critical values may as well. For example, if the researcher can choose the number of treatments to test and also the sample size $\bar{n}$ to use per treatment, then her cost structure might take the form $C(J, \Sigma) = c_f + c_{|J|} |J| + c_{\bar{n}} |J| \bar{n}$, where $c_f$ is a fixed cost (e.g., the costs of staff scientists), $c_{|J|}$ a cost that varies with the number of treatments tested (e.g., the cost of training clinical staff on various treatment protocols), and $c_{\bar{n}}$ a cost per experimental subject (e.g., recruitment costs). In this case (normalizing $b = 1$ and again assuming for simplicity that $\bar{\omega}(J)=|J|$) we obtain optimal thresholds $t=\Phi^{-1}(1 - c_f/|J| - c_{|J|} - c_{\bar{n}} \bar{n})$ which are decreasing in the (per-treatment) sample size $\bar{n}$ as well as in the cost per treatment. Intuitively, to the extent that large-sample experiments are more costly to the researcher to run, the planner need worry less about discouraging the researcher from running such experiments when doing so would not be socially optimal. This point follows from exactly the same economic logic that rationalizes adjusting testing thresholds with respect to $J$, but has not (to our knowledge) come up in past discussions of MHT adjustment. In that sense, it illustrates the value of working out the economic logic underlying MHT adjustment carefully, to make sure we have fully grasped the consequences of any implicit assumptions.
The practical value of Proposition (ref) lies in the fact that it connects optimal testing protocols to measurable properties of the cost function $C(J,\Sigma)$. We illustrate this in Section (ref) where we develop the application to clinical trials, using publicly available data on moments of their cost structure to derive specific testing thresholds. Appendix (ref) provides a second illustration, applying the framework to experimental program evaluation research in economics using unique data from the Jameel Poverty Action Lab (J-PAL). Readers primarily interested in implications for practice may wish to skip ahead to these exercises.
Here we present two extensions of our main results: a variant of our model where the researcher knows $\theta$ only imperfectly (Section (ref)) and a relaxation of the variance homogeneity requirement in Assumption (ref) (Section (ref)). Appendix (ref) presents additional extensions.
So far, we have assumed that the researcher is perfectly informed and knows $\theta$. Here, we show that our main results continue to hold in settings where the researcher has imperfect information in the form of a prior about $\theta$.\footnote{In the single-hypothesis testing case, tetenov2016economic gives results under imperfect information. However, these results rely on the Neyman-Pearson lemma, which is not applicable to multiple tests.} Denote this prior by $\pi' \in \Pi'$, where $\Pi'$ is the class of all distributions over $\Theta$.\footnote{The assumption that $\Pi'$ is unrestricted is made for simplicity. For our theoretical results, we only need that the class of priors $\Pi'$ contains at least one element that is supported on the null space $\bigcap_{J \in \mathcal{J} \setminus \emptyset} \Theta_0(J)$.} The prior $\pi'$ captures knowledge about $\theta$ that is available to the researcher but not to the planner.
We assume that the vector of statistics $X$ is drawn from a normal distribution conditional on $\theta$, where $\theta$ itself is drawn from the prior $\pi'$ with $\int_\Theta \pi'(\theta) d\theta = 1$: $$ X \mid \theta \sim \mathcal{N}(\theta, \Sigma), \quad \theta \sim \pi', \quad \pi' \in \Pi', $$ where $\Sigma$ is positive definite and assumed to be known after being chosen by the researcher.
The researcher acts as a Bayesian decision-maker and chooses
The researcher's prior is correctly specified, and welfare is given by
Under imperfect information, we define maximin protocols with respect to the prior $\pi'$.
Definition (ref) generalizes the notion of maximin optimality in Section (ref), which is stated in terms of the parameter $\theta$. When $\Pi'$ contains only point mass distributions, the two notions of maximin optimality are equivalent.
The next proposition shows that one-sided $t$-tests with appropriately chosen critical values are maximin optimal under imperfect information.
Proposition (ref) shows that the maximin optimality of separate $t$-tests continues to hold under imperfect information. Intuitively, maximin optimality is preserved here because the worst-case prior is a point mass prior over $\theta$, so that the same reasoning as under perfect information applies. This follows from classical results on linear programs (see Appendix (ref)). Uniform global optimality of $r^{t}$ follows as a corollary of Proposition (ref).
Assumption (ref) restricts the class of designs the researcher can choose from to designs where all $X_j$ have the same variance, $\Sigma_{i,i}=\Sigma_{j,j}$. This assumption may seem strong since one might expect the researcher to choose a design with unequal variances, especially if she believes that the outcomes are more variable under some treatments than under others. Here, we consider settings where the researcher can choose designs with heterogeneous variances. Specifically, we study a version of our model in which the researcher can choose the sample size for each treatment arm, and thus the variances of the test statistics, but has only imperfect knowledge of the underlying heterogeneous outcome variances in the treatment and control groups. See Remark (ref) for a discussion of settings with heterogeneous variances where the researcher has full information and can thus choose the sample size as a function of such variances.
We assume that for a given $(J,\Sigma)$, $X\sim \mathcal{N}(\theta_J,\Sigma)$, where for each $j$, we interpret $\theta_j$ as a ratio between the treatment effect $\tau_j$ and $\sigma_j := \sqrt{\sigma_{1j}^2/2 + \sigma_{0j}^2/2}$, where $\sigma_{1j}^2$ and $\sigma_{0j}^2$ are the variances of the treated and the control outcomes in the experiment, respectively. Define the vector of sample sizes in an experiment with treatments $J$ as $n(J) = \{n_{1j}, n_{0j}\}_{j \in J}$. Here, $n_{1j}$ and $n_{0j}$ are the number of treated and control units in arm $j$ (with $n_{0j}$ potentially constant across $j$ if all arms share the same control group). In this model, $\Sigma_{j,j} = \frac{\sigma_{1j}^2}{\sigma_j^2 n_{1j}} + \frac{\sigma_{0j}^2}{\sigma_j^2 n_{0j}}$.
Suppose that the researcher only knows $\theta$, the effect measured in standard deviations $\tau_j/\sigma_{j}$ for each $j$, but not $(\tau_j,\sigma_{1j}^2,\sigma_{0j}^2)$ separately. To encode such uncertainty, we assume that the researcher has a prior $(\tau_j,\sigma_{1j}^2,\sigma_{0j}^2) \sim \mathcal{P}_{\theta,j}$, which depends on $\theta$. Here $u(\theta, \{j\}) = \mathbb{E}[\tau_j | \theta]$ is the expected treatment effect $\tau_j$ given $\theta$ under the prior $\mathcal{P}_{\theta,j}$. We assume that (with a slight abuse of notation) $C(J,\Sigma)=C(J,n(J))$ (where $c_\theta(J,\Sigma) = C(J,\Sigma)$ is constant in $\theta$), so that the costs are known to the researcher. That is, the costs can depend on the sample sizes in the experiment, $n(J)$, but not on the unknown to the researcher variances $\{\sigma_{1j}^2,\sigma_{0j}^2\}_{j\in J}$.
The following assumption summarizes the model with unknown heterogeneous variances.
We consider a setting in which the researcher has limited knowledge, captured by the assumption that $(\sigma_{1j}, \sigma_{0j})$ have a common expectation across $j$, conditional on $\theta$.
Under Assumption (ref), the researcher expects the standard deviations to be the same before running the experiment. Importantly, however, Assumption (ref) allows the realized variances to be heterogeneous.
Under Assumption (ref), $\{n_{1j},n_{0j}\}_{j\in J}$ are chosen by the researcher before running the experiment and observed by the planner. As a result, the planner can de-facto mandate any choice of sample sizes by only rewarding designs that maximize her utility.
We say that the design $\Sigma$ has a sample-equalizing allocation if $n_{1j} = n_{0j} = \bar{m}$ for all $j \in \{1, \dots, |J|\}$ and for some constant $\bar{m} \in (0,n]$ (that can be chosen by the researcher). The next proposition shows that designs with sample-equalizing allocations are optimal.
Proposition (ref) shows that separate one-sided $t$-tests with critical values $t = \Phi^{-1}\left(1 - \frac{C(J, n(J))}{b \bar{\omega}(J)} \right)$ based on sample-equalizing allocations are maximin and unbiased. The planner “forces” the researcher to choose designs with sample-equalizing allocations by not rewarding any results from experiments with other designs. Intuitively, since the researcher expects $\sigma_j$ to be the same for all $j$, the optimal protocol equalizes the expected standard errors by requiring a sample-equalizing allocation. Proposition (ref) provides guidance both on which testing protocol to implement and which design to incentivize.
This section discusses the scope for applying and implementing the framework's implications. Section (ref) considers our running example, the regulatory approval process, while Section (ref) comments on potential applications to program evaluation in economics.
Proposition (ref) showed that the policymaker's preferred MHT adjustments hinge on how research costs $C(J,\Sigma)$ vary with the set of chosen treatments $J$ and the experimental design $\Sigma$. In particular, denote the level of the separate $t$-tests in Proposition (ref) as $$ \alpha(J, \Sigma) := \frac{C(J, \Sigma)}{b \bar{\omega}(J)}. $$ The planner will therefore want to obtain information about $C(J,\Sigma)$, $b$, and $\bar{\omega}(J)$ to compute this level. In the FDA approval context with multiple subgroups, we can think of $b$ as the expected profit per customer, and $\bar{\omega}(J)$ as the total number of customers that would buy the drug if it were approved for every subgroup. To build intuition, it will be helpful to impose the simplifying assumption $\bar{\omega}(J) = |J|$, which holds for example if each subgroup $j$ receives equal weight $\omega_j=1$.
Adjustment factor in general form. Let $\bar{C}$ denote the cost of a benchmark experiment with a single treatment.\footnote{This benchmark experiment could, for instance, be an experiment with the minimum sample size for a Phase III trial according to FDA guidance (see fda_sample_size).} Without loss of generality we can write
where $\bar{\alpha} = \bar{C} / b$ denotes the size of the hypothesis test in the benchmark experiment. This formulation shows that the appropriate size for tests in a study with $|J|$ hypotheses can be calculated as the product of two quantities. The first is the size of the optimal test in the benchmark, single-hypothesis case. The second is the MHT correction factor $[ C(J,\Sigma)/\bar{C} \times 1 / |J| ]$, which captures how the cost per test varies as the number of hypotheses tested grows (keeping in mind that this may affect the design $\Sigma$ as well as $J$). Unless all costs are fixed ($C(J, \Sigma) = \bar{C}$) this correction factor will differ from the standard Bonferroni correction factor $1/|J|$. Notice also that if costs are strictly proportional to the number of hypotheses ($C(J, \Sigma) = \bar{C} \times |J|$) then standard inference without adjustment for MHT is optimal.
Choice of $\bar{\alpha}$. There are competing benchmarks one might consider for $\bar{\alpha}$, the test size for a benchmark study with a single treatment. FDA guidelines currently recommend a size of 2.5% for one-sided single hypothesis tests fda2022guidance, but tetenov2016economic, using data on the costs and expected profits from Phase III trials, proposes a value of 15%. Given this, and the dispersion in the cost of trials for different drugs grabowski2002returns, we provide results for a range of values between 2.5% and 15%.
Modeling costs. In principle the regulator could evaluate the MHT adjustment term in ((ref)) separately for different categories (e.g., therapeutic classes) or even using data on each study individually. They might require pharmaceutical companies to declare the fixed and variable costs of a study (information about which is often contained in contracts with the hospital or contract research organization organizing a trial) when pre-registering it. Here we wish to illustrate the potential consequences of doing so using existing, published estimates of moments of the cost structure of clinical trials. This requires that we model the cost function. We consider a simple formulation with both fixed and variable costs:
where $c_f$ is a fixed cost invariant to $|J|$ and $c_v$ is a variable cost. This is a special case of the specification in Section (ref), where we assume here for simplicity that variable costs vary in proportion to the number of subjects $n_j$. Additional costs that vary with $|J|$, independent of $n_j$, could be accommodated with the appropriate data, and would imply adjustments less conservative than those reported below.
It will be convenient to rewrite this expression (without loss of generality) as $$ C(J,\Sigma) = c_f + c_v |J| \bar{m} \frac{\bar{n}}{\bar{m}}, \quad \bar{n} = \frac{1}{|J|} \sum_{j \in J} n_j, $$ where $\bar{n}$ is the average sample size across subgroups in the trial in question and $\bar{m}$ is the average overall sample size of the single arm benchmark experiment (with size $\bar{\alpha}$). With an abuse of notation, we can then write the optimal level as a function $\alpha\left(\bar{\alpha}, |J|, \frac{\bar{n}}{\bar{m}}\right)$ of the size for the benchmark single-hypothesis experiment $\bar{\alpha}$, the number of treatment arms $|J|$, and the ratio of the average subgroup sample size to the benchmark experiment size $\bar{n}/\bar{m}$.
Cost calibration. SertkayaEtAl2016drivers, using data on the costs of 31,000 pharmaceutical clinical trials conducted in the United States between 2004 and 2012, estimate that the average fixed costs of a Phase 3 trial were 46% of the average total cost, with the rest varying either directly with the number of subjects enrolled or with the number of sites at which they were enrolled.\footnote{According to SertkayaEtAl2016drivers, average variable costs (i.e., the per-patient and per-site costs) were USD 10,826,880, and average total costs were USD 19,890,000, so that the fraction of fixed costs is $(19,890,000-10,826,880)/19,890,000\approx 0.46$.} As an approximation, we set $\bar{m} \bar{J}$ equal to the average historical overall sample size across clinical trials. It follows that $c_f / (c_f + c_v \bar{m} \bar{J}) = 0.46$. Using $\bar{J} = 3$ based on the tabulations in pocock2002subgroup yields an MHT correction factor of $(\frac{\bar{n}}{\bar{m}} + 2.56 / |J|) / 3.56$.\footnote{We use the median estimate multiplied by the probability of reporting more than one subgroup. The critical values are not particularly sensitive to $\bar{J}$; if for example we fix $\bar{\alpha}(1) = 0.025$ and double $\bar{J}$ from 3 to 6 this decreases $\alpha(2)$ from $0.016$ to $0.015$, $\alpha(3)$ from $0.013$ to $0.011$, and $\alpha(\infty)$ from $0.007$ to $0.004$.}
Inserting this into ((ref)), we thus arrive at
The correction factor here has two terms. The first is a “pure” correction for multiple hypothesis testing, accounting for the influence of the number of treatment arms $|J|$. The second corrects for the effects of sample size on study cost: studies with sample sizes larger than $\bar{m}$ ($\bar{n} > \bar{m}$) are more expensive to run and thus require less strict testing thresholds.
Table (ref) illustrates the implications quantitatively. It tabulates the test level implied by ((ref)) for a range of values of $|J|$ (rows), $\bar{\alpha}$ (columns), and $\bar{n} / \bar{m}$ (column groups). For example, for studies with a benchmark sample size ($\bar{n} = \bar{m}$) and assuming $\bar{\alpha}=0.15$ as in tetenov2016economic, those with $|J|=2$ would use size $0.096$, those with $|J|=3$ would use $0.078$, and so on, asymptoting to $0.042$ at $|J| = \infty$. If instead we set $\bar{\alpha} = 0.025$, consistent with FDA guidance, then studies with $|J| = 2$ would use $0.016$, those with $|J| = 3$ would use $0.013$, and so on, asymptoting to $0.007$ at $|J| = \infty$. These thresholds are more conservative than unadjusted ones, but less conservative than those implied by Bonferroni corrections ($\bar\alpha/|J|$).
For further comparison, the last two columns of Table (ref) report the adjustment corresponding to FWER control based on the Sidak correction vsidak1968multivariate. The Sidak correction, which sets the level of each test to $1 - (1 - \bar{\alpha})^{1/|J|}$, is a useful benchmark because it is exact with independent tests (as for the case of subgroup analyses). It implies more conservative inferences than our tabulated values. For instance, with $\bar{\alpha} = 0.025$ and $|J|=9$, the optimal level is $0.009$ while that under the Sidak correction is $0.003$.
The table also illustrates how the adjustment factor depends on the (relative) average per-treatment sample size $\bar{n}/\bar{m}$. The fifth, sixth and seventh columns vary this while holding $\bar{\alpha}$ fixed at 0.025. When $\bar{n}$ is smaller than $\bar{m}$ we require more stringent size control, while for larger $\bar{n}$ we require less stringent size control. For example, with $|J|=2$, studies with half the benchmark sample size would use a $0.012$ threshold, while studies with double the benchmark sample size would use $0.023$.
The results in Table (ref) should be read as illustrative of what might happen if the FDA were to require cost disclosure and base testing thresholds on the disclosed trial-specific costs, which could be obtained, for example, from privately-owned contract data SertkayaEtAl2016drivers. If instead it were to base testing thresholds directly on the number of treatments $|J|$ and samples sizes $\bar{n}$, using the values indicated in the table, this would create incentives for gaming---e.g., reporting results obtained from a single experiment (in the sense that $c_f$ was incurred only once) as if they came from multiple distinct experiments (implying that $c_f$ was incurred more than once). Detecting and deterring such gaming might be easier in some cases than in others. Tests of the same drug in different populations, for example, could be matched based on the chemical formula of the compound in question; tests of different compounds on the same population might be harder to match.
Hypothesis testing norms are also a salient issue within economics, given the publication trends noted in Figure (ref). Our framework's assumptions arguably correspond most closely to economics papers that report the results of experimental program evaluations, as in that case there are clear policy decisions that the research is explicitly designed to inform (i.e., whether or not to implement or scale the programs being evaluated). Indeed, researchers often conduct studies like these in collaboration with implementation partners, such as governments or NGOs, precisely in order to evaluate the impact of treatments the partners are considering. The paper’s findings may thus affect social welfare because, in addition to potentially being published in an academic journal, they can influence those decisions. Pre-specification of the analysis to be conducted in a pre-analysis plan (corresponding to our assumption that researchers pre-specify their tests) is now common in this genre of work miguel2021evidence. And it is also common for such “policy experiments” to test more than one treatment as part of the same study. muralidharan2025factorial document at least 27 such experiments published in top-5 journals alone between 2007 and 2017.\footnote{Their list includes only studies with interaction arms, so provides a lower bound on the total number of multi-armed evaluations.}
With these points in mind, we also conducted a second quantitative application to program evaluation experiments in development economics, using unique data on their costs, sample sizes, and treatment arm counts which we obtained from the universe of funding proposals submitted to the Abdul Latif Jameel Poverty Action Lab (J-PAL) from 2009 to 2021. In the interests of brevity, we describe the data and analysis in depth in Appendix (ref) and briefly restate the main findings here. We estimate that research costs are significantly and substantially less than proportional to the number of treatments tested, with elasticities ranging from 0.13 to 0.22.\footnote{We obtain these estimates from descriptive regressions; they need not be causal to characterize the cost function $C(J,\Sigma)$ in our model, provided that function is invariant to $r$. } But they are also not invariant to scale: projects with more arms cost significantly more ($p < 0.05$). As a result the appropriate testing thresholds vary with the number of treatment arms. They are similar to but slightly less conservative than those that result from a Bonferroni correction and those implied by Sidak's correction (which is exact for controlling the FWER for independent tests). Finally, the testing thresholds also vary moderately with the sample size, with larger samples implying (ceteris paribus) less conservative procedures.
This analysis focuses on a specific type of multiplicity, namely multiplicity of treatments. Economists also often deal with multiple outcomes. These do not necessitate multiple tests; indeed, researchers often aggregate the outcomes into summary statistics instead. An earlier version of this paper VWN2025MHTv8 studied this problem within our framework, showing that optimal rules $r^*$ test for effects on an index formed using statistical weights anderson2008multiple when the outcomes are noisy proxies for some common underlying measure, but using economic weights bhatt2024predicting when they capture distinct components of the planner's utility. The fact that multiple outcomes justify different techniques than do multiple treatments (or subgroups) is noteworthy in the context of historical narratives about MHT practices: the multiple-treatment case---genetic association testing in particular---has often been cited to motivate new MHT procedures dudoit2003multiple, efron2008microarrays, while the multiple-outcomes case seems to have been bundled with it subsequently, and less intentionally.
One could also move away from the frequentist paradigm (which we have presumed) entirely, towards a Bayesian alternative. Proposals to control the FDR are interesting in this regard. Several papers have pointed out a Bayesian rationale for doing so: controlling the (positive) FDR can be interpreted as rejecting hypotheses with a sufficiently low posterior probability storey2003positive,gukoenker2020invidious,kline2022systemic. In fact, these arguments apply even in the case of a single hypothesis. The essential idea is to balance the costs of false positives and false negatives, rather than prioritize size control at any (power) cost. We thus interpret these arguments less as support for a particular solution to the MHT problem per se, and more as a reminder of the merits of Bayesian approaches generally.
We are grateful to the Editor and anonymous referees for constructive comments that helped improve the paper. We are also grateful to Nageeb Ali, Isaiah Andrews, Tim Armstrong, Oriana Bandiera, Sylvain Chassang, Kevin Chen, Tim Christensen, Graham Elliott, Stefan Faridani, Will Fithian, Paul Glewwe, Peter Hull, Guido Imbens, Lawrence Katz, Toru Kitagawa, Pat Kline, Michael Kremer, Michal Kolesar, Ivana Komunjer, Damian Kozbur, Lihua Lei, Adam McCloskey, Konrad Menzel, Francesca Molinari, Jose Montiel-Olea, Ulrich Mueller, Mikkel Plagborg-M\oller, David Ritzwoller, Joe Romano, Adam Rosen, Jonathan Roth, Andres Santos, Azeem Shaikh, Jesse Shapiro, Joel Sobel, Sandip Sukhtankar, Yixiao Sun, Elie Tamer, Aleksey Tetenov, Winnie van Dijk, Tom Vogl, Quang Vuong, Michael Wolf, and seminar participants for valuable comments, and to staff at J-PAL, Sarah Kopper and Sabhya Gupta in particular, for their help accessing data. Aakash Bhalothia and Muhammad Karim provided excellent research assistance. Viviano is also affiliated with Y-RISE, and W\"uthrich is also affiliated with CESifo. All errors are our own.
Niehaus and W\"uthrich gratefully acknowledge funding from the UC San Diego Academic Senate. Viviano gratefully acknowledges support from the Griffin Fund at Harvard University and from NSF Grant SES 2447088.
Non-financial support was provided by J-PAL in the form of the data used in Appendix (ref), which we used under the provisions of a Data Use Agreement (DUA) between J-PAL and UC San Diego. Key provisions are that we are allowed to publish data sufficient to meet the requirements of scholarly journals, subject to various provisions, that J-PAL may curate these data using its established data publication processes to remove any Confidential Data, and that J-PAL may review drafts of the paper prior to publication to ensure that no Confidential Data are disclosed.
Niehaus is co-chair of the Science for Progress Initiative at J-PAL, which provided the data used in Appendix (ref). The role is uncompensated. He is not an officer, director or board member of J-PAL.
\spacingset{1.25}
\spacingset{1.5}