Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
53,872 characters · 28 sections · 70 citation commands
Critical Values Robust to P-hacking
\paragraph{Definition of p-hacking} P-hacking occurs when scientists engage in various behaviors that increase their chances of reporting statistically significant results SNS14,WL16. Typical p-hacking practices include running many small-sample studies rather than one large-sample study; reporting studies with significant results but suppressing studies with insignificant results; collecting data until a significant result is obtained; dropping inconvenient observations or outcomes from a study; and searching for statistical specifications that produce significant results NSM12,L15,CFM19,SS23.
\paragraph{Prevalence of p-hacking} P-hacking is prevalent in science (appendix (ref)). Scientists readily admit to it. It is visible in meta-analyses: the distributions of test statistics in entire literatures show that scientists tinker with their analyses to obtain significant results. And it appears by tracking cohorts of scientific studies: studies finding significant results are almost certain to be reported, whereas studies finding insignificant results are likely to remain unreported.
\paragraph{Reasons for p-hacking} That p-hacking is so prevalent is unsurprising because scientists face strong incentives to p-hack. First, significant results are more rewarded than insignificant ones (appendix (ref)). This is because scientific journals prefer publishing significant results. Publications, in turn, determine a scientist's career path, including promotions, salary, and honorific rewards. Second, scientists enjoy a lot of flexibility in data collection and analysis (appendix (ref)). Hence, even when the null hypothesis is true, they have ample opportunity to obtain significant results without violating scientific norms.
\paragraph{Problems caused by p-hacking} Despite its prevalence, p-hacking is not accounted for in classical hypothesis testing theory. Therefore, classical critical values set a standard for significance that is too lax: a true null hypothesis is rejected more often than purported by the test's significance level. This is problematic because hypothesis tests are informative only insofar as a true null hypothesis is not rejected more often than the significance level. For instance, hypothesis tests are used to evaluate scientific theories and paradigms K57,AM18. They allow scientists to identify instances when theory does not accord well with empirical observations. Unbridled p-hacking threatens scientific progress. It leads to excessive rejection of established paradigms and to the unwarranted adoption of new paradigms. As such, it threatens the credibility of science. One manifestation of uncontrolled p-hacking is the replication crisis in science IGG14,CM18.
\paragraph{Existing corrections for p-hacking} A few corrections for p-hacking in hypothesis testing have been discussed A54,Lo83,G06. But these corrections take the scientist's p-hacking behavior as fixed, whereas in reality the scientist would change her p-hacking behavior as soon as the correction is implemented. Consider for instance a hypothesis test with a significance level of 5%. Classical critical values are constructed such that if the scientist conducted one experiment, a true null hypothesis would be rejected no more than 5% of the time. But if a scientist conducted more than one experiment, performed hypothesis tests in each experiment separately, and reported the best result, a true null hypothesis would be rejected more often than 5% of the time. Existing corrections take the number of experiments as given and compute a more stringent critical value based on this number. But this is insufficient to resolve the problem. Just as scientists may conduct more than one experiment under the classical critical value, they may conduct more experiments than anticipated under the new critical value, overwhelming the proposed correction.
\paragraph{This paper's correction for p-hacking} In this paper, we start by developing a model of hypothesis testing with p-hacking. We then use the model to construct critical values such that, if these values are used to determine significance, and if scientists optimally p-hack in response to the new significance standards, then significant results occur with the desired frequency. Unlike classical critical values, these robust critical values deliver the promised rate of type 1 error. Once the robust critical values are in place, scientists continue to p-hack, but readers can be confident that true null hypotheses are not rejected more often than the advertised significance level.
\paragraph{Model of hypothesis testing with p-hacking} We consider a scientist who tests a hypothesis by conducting an experiment. If she obtains a significant result from the experimental data, she obtains a high payoff. By contrast, if she obtains an insignificant result, she obtains a lower payoff. The difference in payoff between significant and insignificant results reflects the facts that significant results are more likely to be published, and publications yield rewards to scientists. Therefore, if the scientist obtains an insignificant result, and if she still has resources to devote to the project, she has the incentive to conduct another experiment to try to obtain a significant result using the second experiment's data. Conducting a second experiment without revealing the existence of the first experiment constitutes p-hacking.\footnote{Because the number of experiments is not observable, multiple-testing corrections cannot be used to correct for p-hacking.}
\paragraph{Optimal p-hacking strategy} Using optimal stopping theory, we find that the scientist's optimal strategy is to conduct experiments until finding a significant result F07. Not all projects yield significant results, however, because the resources that a scientist can devote to any project are finite C21. If the scientist runs out of resources before reaching significance, she reports an insignificant result.
\paragraph{Probability of type 1 error} We begin by computing the expected number of experiments run by a scientist when the null hypothesis is true, as a function of the prevailing critical value. From this we compute the probability of type 1 error as a function of the critical value. The critical value influences the rate of type 1 error in two ways. First, it determines the probability that a true null hypothesis is rejected in each experiment---as in classical statistics. Second, it influences the number of experiments that the scientist collects---a feature unique to our model.
\paragraph{Computation of robust critical value} From these results we compute the critical value such that type 1 errors occur at the desired rate---given by the significance level. This critical value is robust to p-hacking, and it is given by a nonstandard form of Bonferroni correction: for any test statistic and any significance level, the robust critical value is the classical critical value for the same test statistic with the significance level divided by the expected number of experiments when the robust critical value is in place. Accordingly, the robust critical value is larger than the classical critical value for the same test statistic and significance level. An advantage of the model is that the expected number of experiments when the robust critical value is in place, and the robust critical value itself, are solely determined by two parameters: significance level and probability of completing an experiment before running out of resources.
\paragraph{Numerical illustration} To illustrate the amount of correction that p-hacking might require, we calibrate the completion probability using evidence from the medical sciences DAA08. We obtain the rule of thumb that the robust critical value for any test statistic is the classical critical value for the same test statistic with one fifth of the significance level. Hence, the robust critical value for a significance level of 5% is the classical critical value for a significance level of $5\%/5 = 1\%$. For a $z$-test with a significance level of 5%, and similarly for a large-sample $t$-test with a significance level of 5%, this means that the robust critical value is $2.33$ instead of $1.64$ if the test is one-sided, and $2.58$ instead of $1.96$ if the test is two-sided.
\paragraph{Extensions of the model} Our model of hypothesis testing is quite stylized, but it can be extended in various ways. In appendix (ref), we add a cost of doing research, incurred by the scientist at each new experiment. In appendix (ref), we add time discounting, which reduces the value of significant results obtained far into the future. And in appendix (ref), we assume that consecutive experiments become more and more difficult to run, and thus less and less likely to be completed. In all these extensions, the robust critical value computed in the basic model continues to be operational: it maintains the rate of type 1 error below the significance level.
\paragraph{Other p-hacking strategies} In the model, scientists p-hack by repeatedly running experiments until they reach significant results. This p-hacking strategy appears to be quite common BVW12. However, the model can be adapted to describe a wider range of p-hacking strategies. In appendix (ref), we consider scientists who pool data across experiments. In appendix (ref), we consider scientists who remove more and more outliers until they reach significant results. In appendix (ref), we consider scientists who successively examine different regression specifications so as to obtain significant results. Finally, in appendix (ref), we consider scientists who successively examine different instruments to reach significant results. We find that the robust critical value computed under the repeated-experiment strategy remains useful under these other p-hacking strategies because it maintains the type 1 error rate below the significance level.
\paragraph{Control of type 1 error rate for generic p-hacking strategies} More generally, the robust critical value derived in the basic model controls the type 1 error rate for any p-hacking strategy that induces positive dependence across test statistics (appendix (ref)). While the basic model assumes independent test statistics---each obtained from a separate experiment---real-world p-hacking often yields dependent test statistics. Nonetheless, our robust critical value remains valuable by maintaining the type 1 error rate below the significance level even when p-hacking induces positive dependence across test statistics. Positive dependence results from various p-hacking strategies encountered in practice: when scientists pool data across experiments, when they remove outliers, or when they search across various statistical specifications. Our robust critical value can therefore be used even if the particular p-hacking strategies used by scientists are unknown, as long as these strategies can be expected to generate positive dependence across test statistics, and the completion probability is calibrated to the upper bound of plausible completion probabilities across strategies.
This section develops a simple model of hypothesis testing with p-hacking. A scientist runs experiments with the aim of reaching a significant result. Running experiments takes time, stamina, and money, which are all in finite supply. Because scientists must report results before running out of resources, not all projects yield significant results.
The scientist tests a null hypothesis $H_0$ against an alternative hypothesis $H_1$. The data are governed by a different probability distribution under each hypothesis. The scientist sets the test's significance level to $\a\in(0,1)$. The significance level gives the desired probability of type 1 error---the error that occurs when a true null hypothesis is rejected. Common significance levels are 10%, 5%, and 1%.
To conduct the hypothesis test, the scientist collects a dataset from an experiment. From this dataset she constructs a test statistic $T$, whose realization is $t$. Under $H_0$, the cumulative distribution function of the test statistic is $F$, its survival function is $S = 1-F$, and its inverse survival function is $Z = S^{-1}$.\footnote{For simplicity we focus on simple null hypotheses. For composite null hypotheses, we would use the distribution under the null hypothesis's configuration that is the easiest to reject. For example, when testing $H_0: \E(X) \leq \m_0$ versus $H_1: \E(X)>\m_0$, we would use the distribution of the test statistic at the point $\E(X)=\m_0$.}
The null hypothesis is rejected when the test statistic exceeds the critical value $z$. If the scientist obtains a test statistic $t > z$, the null hypothesis is rejected: the result is significant. But if she obtains a test statistic $t \leq z$, the null hypothesis cannot be rejected: the result is insignificant. Accordingly, the probability of type 1 error is $S(z)$. The classical critical value is set such that the probability of type 1 error in one single test equals the significance level:
or equivalently $z = Z(\a)$.
The first nonclassical element of the model is the rewards accruing to significant results. To capture the facts that significant results are more likely to be published than insignificant results, and publications yield rewards to scientists, we assume that the expected rewards $v^s$ from a study with significant results are higher than the expected rewards $v^i$ from a study with insignificant results.
Scientists have ample opportunity to p-hack. However, their resources---time, money, manpower, stamina---are not infinite. Hence, they cannot systematically obtain significant results C21. We assume that it takes a random amount of resources to conduct an experiment, and the scientist must keep the cumulative resources used below a random limit $L$. Once the scientist has exhausted more resources than $L$, she must stop working on the project. The resource limit captures the many resource constraints faced by scientists: limited access to data, limited funding, limited coauthor time, limited time before publication of similar results by competing research teams, limited stamina to work on specific projects, or limited time before the opportunity to work on more promising projects arises. Following F07, we assume that the resource limit has an exponential distribution with rate $\l>0$, so $\P{L>l} = \exp{-\l l}$ for any $l>0$.\footnote{Here research is costless to the scientist. But the robust critical value is not modified if the scientist incurs a cost of doing research (appendix (ref)).}
\paragraph{Experiments} The experiments are denoted by $n = 0, 1, 2, \ldots, \infty$, with $n=0$ corresponding to not starting the research project. It takes a random amount of resources to conduct an experiment and collect a dataset. The cumulative amount of resources required to complete $1, 2, \ldots$ experiments is $D_1, D_2, \ldots$ given by a renewal process independent of the resource limit $L$. That is, the resources required for each experiment, $D_1, D_2-D_1, D_3-D_2, \ldots$, are independent and identically distributed (iid) according to a distribution independent of $L$.
\paragraph{First experiment} If resources are exhausted before the first experiment is completed, $L<D_1$, the scientist is not able to obtain any results. If the resources are not exhausted when the first experiment is completed, $L> D_1$, the scientist is able to collect a first dataset and construct a test statistic. This first test statistic is $T_1$, which is independent of the resource variables. The scientist then decides to submit this result to a journal, or to run another experiment.
\paragraph{Nth experiment} If the scientist chooses to run experiment $n\geq 2$, the scientist begins collecting a $n$th dataset of the same size and drawn from the same underlying population as previous datasets. If resources are exhausted before experiment $n$ is completed, $L<D_n$, the scientist must stop the project before obtaining the $n$th dataset and submits the best result obtained up to the previous experiment, $\max{T_1,\ldots,T_{n-1}}$. If resources are not exhausted, $L>D_n$, the scientist obtains the $n$th dataset and constructs the $n$th statistic, $T_n$, which is iid with $T_1, T_2, \ldots, T_{n-1}$.\footnote{By modeling successive test statistics as independent, we are able to derive a robust critical value that controls the probability of type 1 error across a wide variety of common p-hacking strategies that induce positive dependence---without having to specify which particular p-hacking strategy was used by the scientist (appendix (ref)).} She may then submit the best of the $n$ test statistics, $\max{T_1,\ldots,T_n}$, or she may run yet another experiment.\footnote{Here the scientist analyzes the datasets obtained from successive experiments in isolation. The scientist might instead pool the datasets and analyze the pooled data. Thankfully, the robust critical value computed here maintains the type 1 error rate below the significance level with data pooling (appendix (ref)).}
\paragraph{Infinite p-hacking} $n=\infty$ corresponds to running infinitely many experiments and never reporting any result.
Following F07, we introduce the index of the first experiment that cannot be completed before resources are exhausted: $K = \min{n \geq 1: D_n > L}$. Let $\g$ be the probability that the first experiment can be completed:
The index $K$ is independent of the test statistics $T_1$, $T_2$, \ldots, and it has a geometric distribution with success probability $1-\g$, so $\P{K>k} = \g^k$ for $k=0,1,2,\ldots$.\footnote{Here each experiment is completed with the same probability $\g$. Experiments might instead be more and more difficult to run and less and less likely to be completed. Fortunately, the robust critical value computed here maintains the type 1 error rate below the significance level with increasingly difficult experiments (appendix (ref)).}
\paragraph{No results} If the scientist does not start the research project, she receives a payoff normalized to $y_0 = 0$. If resources are exhausted before the end of the first experiment, the scientist does not obtain any result, so she receives the same payoff of $y_1 = 0$. If the scientist never concludes the research project and keeps on p-hacking forever, she also receives a payoff $y_{\infty} = 0$. In all other cases, she receives a positive payoff.
\paragraph{Exhausted resources} The scientist is not able to continue p-hacking once the project resources are exhausted. To capture this constraint, we set to zero all payoffs once resources are exhausted: $y_n = 0$ in any step $n>K$. With these payoffs, the scientist never continues past step $K$. At step $K$, the scientist cannot obtain a new test statistic, but she can submit for publication the best test statistic from the previous $K-1$ hypothesis tests, $\max{T_1,\ldots,T_{K-1}}$. If that statistic is significant, the payoff is $y_K = v^s$; if that statistic is not significant, the payoff is $y_K = v^i$.
\paragraph{Non-exhausted resources} Any experiment $n<K$ can be completed before running out of resources, so the scientist can submit the best statistic from the $n$ previous tests, $\max{T_1,\ldots,T_n}$. If that statistic is significant, the payoff is $y_n = v^s$; if not, the payoff is $y_n = v^i$.\footnote{Here the scientist does not discount the future, so a significant result yields the same payoff irrespective of when it is obtained. But the robust critical value is not modified if the scientist discounts future payoffs (appendix (ref)).}
The scientist p-hacks as long as she wishes. At each experiment, she may decide to stop and receive a payoff, or she may decide to continue to the next experiment. If she is able to complete the next experiment, she computes another test statistic. The scientist's problem, which we now solve, is to choose a time to stop p-hacking so as to maximize expected payoffs.
The stopping rule chosen by the scientist, the critical value $z$, and the random research events determine the random time $N(z)$ at which the scientist stops p-hacking. The problem of the scientist is to choose a stopping time to maximize expected payoffs.
As long as she is able to complete at least one hypothesis test, the scientist reports a random statistic $R(z)$ upon stopping. This is the best test statistic that she has been able to obtain through p-hacking. It may be significant or insignificant, and the scientist may be able to publish it or not.
An optimal stopping time $N(z)$ exists because two conditions are satisfied F07. Let $Y_n$ denote the random payoff received by the scientist when she stops at time $n$. First, $Y_n \leq v^s$ almost surely, so $\sup_n Y_n < \infty$ almost surely. Second, because the resources inevitably run out, $Y_n \as 0 = y_{\infty}$ as $n\to \infty$. Furthermore, the optimal stopping time is given by the principle of optimality of dynamic programming: it is optimal to stop as soon as the payoff is at least as high as the best payoff that can be expected by continuing.
We find the optimal stopping time by considering the various situations faced by the scientist.
\paragraph{Starting the research project} If the scientist does not start the research project, she receives $Y_0 =0$. In contrast, if she starts she earns a nonnegative payoff: $0$ if resources are exhausted before the first experiment is completed; $v^i$ if she obtains an insignificant result; or $v^s$ if she obtains a significant result. Hence it is always optimal to start the research project.
\paragraph{Continuing after insignificant results} How does the scientist behave when she still has resources to allocate to the project? A first possibility is that the result at experiment $n$ and all the results before that are insignificant. Since the best result found by the scientist is insignificant, the scientist earns $Y_n = v^i$ by stopping at experiment $n$. All possible payoffs are more than the payoff received for an insignificant result, $v^i$, so all expected payoffs are more than $v^i$. Since the scientist is expected to obtain more than $v^i$ by continuing, it is not optimal to stop without obtaining a significant result.
\paragraph{Stopping after a significant result} If the result of test $n$ is significant, the best result found by the scientist is significant, so the scientist earns $Y_n = v^s$ by stopping at experiment $n$. All possible payoffs are less than the payoff received for a significant result, $v^s$, so all expected payoffs are less than $v^s$. Hence, the scientist cannot do better by continuing. She optimally stops at experiment $n$ and reports $R(z) = \max{T_1,\ldots,T_{n}} > z$. In fact, the principle of optimality indicates that she should stop at the first occurrence of a significant result.
\paragraph{Stopping when resources are depleted} Once resources are depleted, the scientist must stop p-hacking. Hence, she stops at step $K$ if she had not stopped before. There are two possibilities. If $K=1$, resources are depleted before the first experiment, so the scientist has nothing to report. If $K>1$, the scientist submits the best test statistic that she has collected. This best result is necessarily insignificant, otherwise she would have stopped before. So she reports $R(z) = \max{T_1,\ldots,T_{K-1}} \leq z$.
\paragraph{Summary} The optimality principle gives the following results:
Based on the scientist's p-hacking strategy, we compute the critical value robust to p-hacking. This critical value ensures that the probability of type 1 error remains below the significance level even as the scientist adjusts her behavior to the critical value itself.
We compute the distribution of the optimal stopping time. Since the distribution is used to calculate the critical value, we compute it under the null hypothesis.
\paragraph{Probability of reaching significance at experiment $n$} Under the null hypothesis, the probability that the test statistic from experiment $n$ reaches the critical value $z$ is given by the test statistic's survival function: $\P(T_n > z) = S(z)$, where $\P$ denotes the probability measure under $H_0$.
\paragraph{Probability of continuing after experiment $n$} The scientist continues p-hacking after any experiment if she has not run out of resources during that experiment, which happens with probability $\g$, and the latest result is insignificant, which happens with probability $1-S(z)$. The two events are independent, so the probability that the scientist continues p-hacking is $\g [1-S(z)]$. Conversely, the probability that the scientist stops at any experiment is
\paragraph{Distribution of the stopping time} The probability of stopping at each experiment is constant, given by (ref). The optimal stopping time therefore has a geometric distribution with success probability (ref). The probability that the optimal stopping time is $n \geq 1$ is
\paragraph{Expected number of experiments} Given that the optimal stopping time has a geometric distribution with success probability (ref), we obtain the following result:
Since classical critical values are defined by (ref), we infer the following result:
\paragraph{P-hacking under the alternative hypothesis} In (ref), $1-\a$ represents the probability of obtaining an insignificant result from an experiment when the classical critical value is used to determine significance and the null hypothesis is true. When the alternative hypothesis is true instead, the probability of obtaining an insignificant result becomes $\b$, where $1-\b$ is the power of the hypothesis test. Hence, if the alternative hypothesis is true, the expected number of experiments is $1/(1-\b\g)$. In many fields, hypothesis tests are acceptable only if their power is above 80% DGK07. Setting power to $1-\b=80\%$, we find that the expected number of experiments under the alternative is $1/(1-0.2\times \g) < 1/(1-0.2) = 1.25$: there is almost no p-hacking. This is unsurprising. If the alternative hypothesis is true and the study is well powered, the null hypothesis is rejected most of the time, which makes p-hacking unnecessary. Hence, if we see a lot of p-hacking, either the alternative hypothesis is false, or the alternative hypothesis is true but tests have low power I05.
Next, we compute the probability of type 1 error as a function of the critical value.
The proof is in appendix (ref); it relies on an appropriate application of the law of total probability. Since classical critical values are defined by (ref), we infer the following:
When scientists p-hack under classical critical values, the probability of type 1 error exceeds the significance level. Hence, the standard for significance set by classical critical values is too low: significance is reached more often than purported by the test's significance level. This is problematic because hypothesis tests are only informative insofar as true null hypotheses are not rejected more often than the significance level.
\paragraph{Effects of critical value on type 1 error rate} Changing the critical value $z$ has two effects on the probability of type 1 error (equation (ref)). First, there is a mechanical effect: a higher critical value reduces the probability that a test statistic exceeds it ($S(z)$ is decreasing in $z$). Second, there is a behavioral effect: the optimal stopping time and reported test statistic are altered by the critical value. When the critical value is larger, scientists p-hack more in hope of reaching significance ($\E(N(z))$ is increasing in $z$). The behavioral effect was not taken into account by previous corrections for p-hacking A54,Lo83,G06. The novelty of this analysis is to propose a critical value that accounts for it.
\paragraph{Computing the robust critical value} The robust critical value is such that the probability of type 1 error equals the significance level $\a$ when scientists p-hack. Since the probability of type 1 error with p-hacking is given by (ref), the robust critical value $z^*$ is implicitly defined by
From this definition we obtain the following result (proof details are in appendix (ref)):
\paragraph{P-hacking under the robust critical value} The robust critical value corrects the distortion introduced by p-hacking without eliminating p-hacking. In fact, because the significance standards imposed by the robust critical value are more stringent than classical standards, scientists p-hack more under the robust critical value. Combining (ref) and (ref), we obtain the following corollary:
Our correction for p-hacking can be formulated as a nonstandard Bonferroni correction:
This relation is obtained by evaluating (ref) at $z^*$, and using $\a^* = S(z^*)$ and $S^*(z^*)=\a$. Unlike a standard Bonferroni correction, the number of experiments used for the correction is not observed, and it is not the number of experiments prevailing under a standard critical value. Rather, it is the average number of experiments under the robust critical value when the null hypothesis is true. Thanks to the model, we can link this number to the probability $\g$, which we can calibrate (section (ref)).
Finally, we discuss how the results are influenced by the completion probability $\g$, which is the main parameter of the model.
\paragraph{Higher completion probability} From equations (ref), (ref), and (ref), we obtain the following:
The corollary indicates that critical values should be higher for research teams with more resources---more time, more money, or more manpower. Research teams with more resources are less likely to be forced to interrupt a study before completion, so they can p-hack more. To control their type 1 error rate properly, a higher critical value is required. The corollary also implies that critical values should be raised when technological progress makes p-hacking easier. An example of such progress is the advent of online surveys and online experiments in social science, which have simplified the task of collecting data. Finally, the corollary implies that critical values should be higher in fields in which p-hacking is easier.
\paragraph{Completion probability of 1} From (ref), (ref), (ref), and (ref), we obtain the following results:
The corollary indicates that if scientists can complete any number of experiments, they will continue experimenting until they reach significance. Since all null hypotheses are eventually rejected, the probability of type 1 error is 1. At this limit, scientists successively experiment to reach a foregone conclusion A54. The robust critical value continues to exist, but it becomes arbitrarily large to offset the arbitrarily large amount of p-hacking.
To illustrate the amount of correction that p-hacking might require, we calibrate the completion probability $\g$ from the lifecycle of studies in the medical sciences. We then compute the resulting robust critical value.
\paragraph{Calibration method} In the model, with probability $1-\g$, the first experiment cannot be completed before running out of resources. The probability $1-\g$ therefore is the share of studies that stop before completion, while the probability $\g$ is the share of studies that are completed. We use data collected by DAA08 to calibrate $\g$ (table (ref)). DAA08 review 16 metastudies that each follow a cohort of medical studies. The studies are followed from protocol approval to publication, so we can measure the fraction of studies that were stopped before completion and thus $\g$.
\paragraph{Studies that never started} Overall the data include 6903 approved studies. We focus on the 4563 studies whose fate is known---either by surveying the scientists who conducted the studies or by searching the literature. In this pool, 658 were never started, or $658/4563 = 14.4\%$.
\paragraph{Studies that started but stopped early} In addition, not all the studies that started were completed. Of the 3905 studies that started, 228 were still ongoing when the cohort studies were written, so 3677 studies started and stopped. Of these, 243 stopped early, before any analysis could be conducted. Hence, $243/3677 = 6.6\%$ of the studies that started had to stop before completion.
\paragraph{Calibrated completion probability} Adding the studies that stopped early to those that never started, we find that $14.4\% + (1-14.4\%) \times 6.6\% = 20.0\%$ of the approved studies could not be completed. This yields a completion probability of $\g = 1 - 20.0\% = 80.0\%$.
We now compute the robust critical value using the Bonferroni correction (ref) and the completion probability observed in the medical sciences, $\g=80\%$.
\paragraph{Formula} Since the significance level $\a$ is always less than 10%, and since $\g$ is less than 1, $1-\a\g$ is close to 1, and the average number of experiments under the robust critical value is close to $1/(1-\g)$ (equation (ref)). This gives a simple Bonferroni correction to deal with p-hacking (equation (ref)). The classical significance level $\a^*$ required to correct p-hacking is approximately $1-\g$ times the desired significance level $\a$:
\paragraph{Numerical application} With $\g= 80\%$, the classical significance level required to deal with p-hacking is one fifth of the desired significance level: $\a^* = (1-0.8) \times \a = \a/5$. For instance, the critical value that achieves a significance level of 5% under p-hacking is the critical value that yields a significance level of $5\%/5 = 1\%$ under classical conditions. This rule of thumb works for any test statistic. For a $z$-test with a significance level of 5%, the robust critical value is $2.33$ instead of $1.64$ if the test is one-sided, and $2.58$ instead of $1.96$ if the test is two-sided. These robust critical values also apply to a large-sample $t$-test with a significance level of 5%.
\paragraph{Comparison with the BBJ18 proposal} To address the replication crisis in science, BBJ18 propose that scientists replace the standard significance level of $5\%$ by a lower significance level of $0.5\%$. Such tenfold reduction in the significance level is a more aggressive response to p-hacking than the fivefold reduction obtained in this numerical exercise. However, a tenfold reduction in significance level would be appropriate for a completion probability of $\g = 90\%$ (equation (ref)). In that way, our analysis provides a theoretical underpinning for proposals to reduce the significance levels used in science. It also links the proposed reductions to the amount of resources available to scientists for p-hacking.
Here we provide additional numerical results. We fix the significance level at 5%.
\paragraph{Prevailing p-hacking} The amount of p-hacking under classical critical values is given by (ref). For the completion probability of 80%, the expected number of experiments is $4.2$ (figure (ref)). Moreover, the amount of p-hacking is increasing with the completion probability. For instance, when the completion probability increases from 70% to 90%, the average number of experiments grows from $3.0$ to $6.9$.
\paragraph{Prevailing probability of type 1 error} The probability of type 1 error under classical critical values is given by (ref). For the completion probability of 80%, although the significance level is 5%, the probability of type 1 error is $21\%$ (figure (ref)). So in this case, p-hacking quadruples the probability of type 1 error. Moreover, the distortion caused by p-hacking is more severe when the completion probability is larger---because then there is more p-hacking. For instance, when the completion probability increases from 70% to 90%, the probability of type 1 error increases from $15\%$ to $34\%$.
\paragraph{Robust critical value for one-sided $z$-test} We calculate the robust critical value when the underlying test statistic has a standard normal distribution under $H_0$, as in the common $z$-test, or in a $t$-test conducted from a large sample. We begin by calculating the robust critical value for a one-sided $z$-test. The critical value is given by (ref) where $\a=5\%$ and $Z$ is the inverse survival function for the standard normal distribution: $Z(x) = \F^{-1}(1-x)$ where $\F$ is the standard normal cumulative distribution function. For the completion probability $\g = 80\%$, the robust critical value is $2.31$, almost equal to the value of $2.33$ given by the rule of thumb (ref) (figure (ref)).
\paragraph{Robust critical value for two-sided $z$-test} Next we calculate the robust critical value for a two-sided $z$-test. The critical value is now given by (ref) where $\a = 5\%$ and $Z$ is the inverse survival function for the standard half-normal distribution: $Z(x) = \F^{-1}(1-x/2)$. For the completion probability $\g = 80\%$, the robust critical value is $2.56$, almost equal to the value of $2.58$ given by the rule of thumb (ref) (figure (ref)).
\paragraph{Sensitivity to the completion probability} Robust critical values are increasing with the completion probability, but they are not very sensitive to it. For instance, as long as the completion probability remains between 70% and 90%, the robust critical value for one-sided $z$-tests remains between $2.16$ and $2.56$ (figure (ref)), and the robust critical value for two-sided $z$-tests remains between $2.42$ and $2.79$ (figure (ref)). This is reassuring: robust critical values remain close even in fields with different p-hacking intensity.
\paragraph{P-hacking under robust critical value} The average number of experiments under robust critical value is given by (ref). For the completion probability of 80%, the expected number of experiments is $4.8$ (figure (ref)). Moreover, the amount of p-hacking is increasing with the completion probability. For instance, when the completion probability increases from 70% to 90%, the average number of experiments grows from $3.2$ to $9.6$. Further, p-hacking is more prevalent under robust critical value than under classical critical value (figure (ref)). At the completion probability of 80%, the average number of experiments is $4.2$ under classical critical value but $4.8$ under robust critical value.
We conclude by summarizing our results and comparing our approach with the registration of pre-analysis plans.
Scientific journals prefer publishing significant results. Publications, in turn, determine a scientist's career path: promotions, salary, and honors. Scientists therefore have strong incentives to hunt for statistical significance. Such p-hacking reduces the informativeness of hypothesis tests, threatening the credibility of science---leading for instance to the current replication crisis. In this paper, we develop a model of hypothesis testing with p-hacking and use it to construct critical values robust to p-hacking, which guarantee that significant results occur with the desired frequency. As an illustration, we calibrate the model to the medical sciences. For a common two-sided $z$-test with significance level of 5%, the robust critical value is $2.58$---somewhat higher than the classical critical value of $1.96$.
A popular solution to p-hacking is to ask scientists to register pre-analysis plans MCC14,CM18,NED18,ADO20. Although strict adherence to pre-analysis plans prevents certain forms of p-hacking, it also prevents scientists from exploring experimental data---a keystone of scientific discovery. By contrast, robust critical values can be used exactly like classical critical values, without preventing exploration. Another concern with pre-analysis plans is that they do not prevent scientists from repeating experiments. A plan could be registered for each experiment until an experiment delivers a significant result, which the scientist would then report with its accompanying pre-analysis plan. Therefore, even when pre-analysis plans are appropriate, it might make sense to use them in conjunction with robust critical values.