Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
73,891 characters · 21 sections · 125 citation commands
Do t-Statistic Hurdles Need to be Raised?
JEL Classification: G0, G1, C1
Keywords: stock market predictability, stock market anomalies, p-hacking, multiple testing \thispagestyle{empty}\setcounter{page}{0}
\setcounter{page}{1}
Suppose a researcher proposes a “factor” behind a phenomenon. How do we determine if this factor is worth noting? At least since fisher1925statistical, researchers have used the following procedure: (1) construct a statistic that is Student's t-distributed in the case that the factor is false, and (2) declare the factor a discovery if this t-statistic exceeds 1.96 in absolute value.\footnote{In the traditional language, a “false factor” is called a “true null hypothesis” while a “true factor” is called a “false null hypothesis.” Some readers may find the traditional language confusing, hence my choice of terminology.} More recently, several papers have called for raising this t-statistic hurdle, or “t-hurdle” to guard against false discoveries (\citet*{harvey2016and}; chordia2020anomalies), including a paper with dozens of co-authors (benjamin2018redefine).
In this paper, I examine whether these calls can be empirically justified. An empirical justification is critical, as prior beliefs regarding cutting edge research are sure to vary across scholars.\footnote{For contrasting prior beliefs regarding asset pricing, see cochrane2017macro and barberis2018psychology.} Moreover, an empirical justification seems possible, as multiple testing statistics provide methods for estimating hurdles that control the false discovery rate (FDR) (benjamini1995controlling). Indeed, FDR methods are central to \citet*{harvey2016and}'s argument and the concept is also used in benjamin2018redefine.
I find that raising the t-hurdle may be difficult to justify empirically. My results differ from previous studies because I acknowledge weak identification: the problem that likelihood functions may depend little on certain model parameters (canova2009back). This problem is important because the data on academic discoveries exhibit publication bias: results that fail to meet the existing t-hurdle are often unobserved. Unobserved results need to be extrapolated, which can lead to weak identification of the key determinants of t-hurdles. I characterize this problem in a theoretical analysis that extends benjamini1995controlling to a setting with publication bias (as in hedges1992modeling). In an empirical analysis, I bootstrap t-hurdle estimates using a rich dataset of published cross-sectional stock return predictors (ChenZimmermann2021). Consistent with the theory, the empirical estimates say little about whether hurdles should be raised, stay the same, or even be lowered.
At the same time, I find other multiple testing statistics are more strongly identified. Empirical Bayes shrinkage and the FDR among published factors (chen2020publication,chen2021most) focus on the right tail of t-stats, and this portion of the distribution tends to be well-observed, in spite of publication bias. For cross-sectional predictors, I find that published t-stats are biased upward by at most 28% and that the FDR for published predictors is at most 22%, with 95% confidence. This strong identification, as well as the weak identification of t-hurdles, is also found across 10 alternative model specifications. Robustness is intuitive, as the theoretical results impose few functional form assumptions. Readers may differ on whether a worst-case FDR of 22% is satisfactory, but at least some will interpret these results as implying that $t$-hurdles need not be raised for the cross-sectional predictability literature.
How can t-hurdles not be raised? This idea seems to fly in the face of multiple testing logic. If one tests 1,000 factors, 5% will meet the 1.96 hurdle on average, even if all 1,000 factors are false. This scenario is visualized in Panel (a) of Figure (ref), which depicts the distribution of absolute t-stats implied by a standard normal. By luck, 50 of the 1,000 false factors meet the classical hurdle. But since all factors are false, the FDR is 50/50 = 100%.\footnote{Here, the FDR is the number of false and significant factors, divided by the number of significant factors, in expectation (benjamini1995controlling).} Clearly, this scenario implies that the classical hurdle needs to be raised.
The problem with this logic is that it ignores the information obtained from multiple testing. Each additional test is another data point, which brings news about the veracity of factors overall. This news can in general be positive or negative. More data does not necessarily mean more bad news.
The data on published factors looks more like Panel (b).\footnote{\citet*{mccrary2016conservative} shows the distribution of t-stats from meta-studies in political science, psychology, economics, and all fields of science. Of t-stats that exceed 1.96, roughly half also exceed 3.0. In Figure (ref), Panel (b), parameters are selected to match this moment. A similar distribution is seen in cross-sectional asset pricing (\citet*{chen2020publication}).} In this panel, there are once again 1,000 false factors (red), so 50 false factors sneak past the 1.96 hurdle. But now there are an additional 1,000 true factors (blue), and among all factors, 1,000 have t-stats that exceed 1.96. Thus, the 50 false factors represent only 5% of the 1,000 factors that meet the classical hurdle. Supposing that an FDR of 5% is sufficient, as it is in a variety of fields from genomics to functional imaging (benjamini2020selective), the classical hurdle need not be raised at all.
Panel (b) illustrates how multiple testing provides more information. With a single test, asking the question “what share of t-stats exceed 1.96?” is impossible to answer. But with multiple testing, answering this question is a straightforward counting exercise. With many tests one can also estimate the share of factors that are false (e.g. storey2002direct). If this share is small, then a t-stat of 1.5 may be due to bad luck, in which case the classical hurdle can be lowered (benjamini2000adaptive).
My theoretical results characterize exactly when FDR methods imply a raising of statistical hurdles. The theory revolves around $\pi_{F}$, the share of false factors among all factors under consideration. Assuming that an FDR of 5% is sufficient, the t-hurdle should be raised if and only if $\pi_{F}$ exceeds the share of t-stats larger than 1.96. Panel (b) of Figure (ref) shows the knife edge case in which both of these objects are equal to 50%, and thus the classical t-hurdle is sufficient.
The key to empirically justifying a higher t-hurdle, then, is to empirically justify a large $\pi_{F}$. Unfortunately, my theoretical results suggest that $\pi_{F}$ is weakly identified under publication bias. I show that, if false and true factors are in a sense distinct, then values of $\pi_{F}$ ranging from 0 to $2/3$ imply almost identical distributions in the region in which data is well-observed. As a result, empirical studies may say little about the proper t-hurdle.
To quantify this problem, I specify a parametric model of biased publication and fit it to the ChenZimmermann2021 dataset using quasi-maximum likelihood. The model includes the key parameter $\pi_{F}$, as well as additional parameters that fully describe the distribution of t-stats. The estimated parameters, in turn, provide consistent formulas for revised t-hurdles. Bootstrapped estimates show that $\pi_{F}$ is weakly identified: I re-sample the data, re-estimate the model, and find that the estimated $\pi_{F}$ ranges from 0% to 70% (90% confidence interval). t-hurdles that control the FDR at 5% range from 0 to 3.0 (90% C.I.).
More positively, I find that multiple testing statistics that target published findings are strongly identified, even in the presence of publication bias. The shrinkage and FDR for published t-stats tend to be determined by the properties of true factors. And under the same conditions that imply weak identification of $\pi_{F}$, the properties of true factors are strongly identified. Intuitively, if $\pi_{F}$ does not affect the observed distribution, it must be the properties of true factors that determine the data. Empirical estimates confirm strong identification for the Chen-Zimmermann dataset. These results are robust to a cluster-bootstrap that closely mimics correlations in the empirical data.
An alternative method for handling publication bias is to algorithmically generate a complete set of factors (yan2017fundamental,chordia2020anomalies). Standard FDR methods can then be applied and $\pi_{F}$ estimated without the identification problems introduced by publication bias. This $\pi_{F}$, however, should be interpreted as an upper bound on the $\pi_{F}$ of the literature, as expert researchers should be able to find true factors at a higher rate than a simple algorithmic procedure (chen2021most).
Code to replicate all figures and tables can be found at \url{https://github.com/chenandrewy/qml-pub-bias}.
Many papers examine multiple testing effects in cross-sectional asset pricing.\footnote{These papers include harvey2016and,yan2017fundamental,chordia2020anomalies,chen2020publication,jacobs2020anomalies,chen2021limits,chen2021most,harvey2021uncovering; and jensen2021there.} Among these papers, harvey2016and; chordia2020anomalies; and harvey2021uncovering focus on t-hurdle corrections. To these papers, I add a framework for understanding when t-hurdles need to be raised. The framework shows that most estimates in harvey2016and and chordia2020anomalies assume that t-stat hurdles need to be raised, and thus they cannot answer the question of whether t-hurdles need to be raised. It also highlights how the multiple testing algorithm perhaps most common in finance (benjamini2001control Theorem 1.3) is likely too conservative.
I also add to harvey2016and; chordia2020anomalies; and harvey2021uncovering by examining the identification problems that come with publication bias (copas1999works,hedges2005selection). These problems would be seen in standard errors on t-hurdles, which are not provided by harvey2016and or harvey2021uncovering. I revisit their estimates and find that the standard errors are so wide they say little about whether t-stat hurdles should be raised or even be lowered.
In a contemporaneous paper, harvey2021uncovering critique three assumptions that have been used in this literature: (1) cross-factor correlations are all equal, (2) the publication probability is a strict t-stat cutoff, and (3) the selected parametric models may be too restrictive. I find that none of these issues has a significant effect on the small shrinkage and FDR estimates found by chen2020publication and jensen2021there. I find similar estimates assuming (1) weak dependence across factors, (2) a smoothly increasing publication probability, and (3) many distinct parametric forms for latent effect sizes. Harvey and Liu propose instead to model latent effect sizes by randomly drawing from empirical data on published and data-mined factors. This method deviates strongly from the literature (efron2012large,andrews2019identification) and it has not been shown to recover effect sizes either theoretically or in simulations.
Identification under publication bias is also studied in andrews2019identification, who prove identification of a non-parametric model. Their proof obtains the latent distributions by solving a system of ODEs, effectively assuming an unlimited sample of published tests. My analysis is less restrictive about data availability. Also unlike Andrews and Kasy, I connect to the literature on false discovery rates and empirical Bayes shrinkage. More broadly, my paper helps bridge the literatures on publication bias (hedges2005selection,mcshane2016adjusting) and multiple testing (efron2012large) and provides guidance on which multiple testing corrections should be applied under publication bias.
There are other arguments against raising t-hurdles. Raising t-hurdles would lead to published data that is in a sense more distorted. This reasoning motivates, in part, the push for pre-analysis plans, which lower t-hurdles provided that the analysis is rigorously pre-specified (olken2015promises,kasy2023optimal). Raising t-hurdles also fails to address common misinterpretations of statistical significance and artificial dichotimization introduced by hypothesis testing (mcshane2019abandon,chen2023publication).
I describe a general model (Section (ref)), define multiple testing statistics (Section (ref)), and then prove results that illustrate weak and strong identification (Sections (ref)-(ref)). Section (ref) explains why some commonly-used FDR methods cannot answer the question of whether t-stat hurdles should be raised. Section (ref) discusses false negative rates.
A literature is generated in two steps. In the first step, researchers generate an “unbiased” set of $N$ factors that obey the assumptions in benjamini1995controlling. Factor $i$ has a t-stat $t_{i}$ that follows
where $\mu_{i}$ is the latent effect size of factor $i$ (e.g. the expected return) divided by its standard error. $\mu_{i}$ can also be thought of as the “corrected” t-stat, since it corrects $t_{i}$ for sampling error. $\mu_{i}$ depends on whether factor $i$ is false ($F_{i}$) or true ($T_{i}$)
where $g\left(\cdot|\lambda\right)$ is an arbitrary probability distribution and $\lambda$ is a vector of parameters. Last, and most importantly, factor $i$ is false with probability $\pi_{F}$.\footnote{benjamini1995controlling leave open the possibility of other distributional assumptions for Equation ((ref)), but the standard normal is commonly used in practice (efron2012large; harvey2016and). chen2021limits shows that the standard normal assumption holds for long-short portfolios from the asset pricing literature.}
I describe the factors from the first step as “unbiased” because $t_{i}|F_{i}\sim\text{Normal}\left(0,1\right),$ and thus the theory of fisher1925statistical and benjamini1995controlling applies. However, the second step of literature generation creates a bias.
In the second step, factor $i$ is published depending on its statistical significance:
where $\pub_{i}$ is the event that factor $i$ is published, $s\left(|t_{i}|\right)$ is a function with values in $[0,1]$, and $\bar{s}$ and $\tgood$ are constants. Conceptually, $s(|t_{i}|)$ captures the prevailing t-stat hurdles that were applied in the generation of the published data.
This second step implies $t_{i}|\left(F_{i},\pub_{i}\right)\nsim\text{Normal}\left(0,1\right)$, violating the assumptions of both fisher1925statistical and benjamini1995controlling. Thus, to estimate the FDR and related statistics, one needs to recover the properties of the unbiased factors generated in the first step.
Equations ((ref))-((ref)) describe a “satisficing” model of publication bias. They say that a higher t-stat implies a higher probability of publication but once the t-stat is high enough the community is satisfied and there are no additional gains. A satisficing model is consistent with the view that researchers are primarily incentivized to uncover convincing and interesting mechanisms and that statistical significance plays a secondary role. This assumption nests the functional forms used in previous estimates of publication bias (harvey2016and,andrews2019identification,chen2020publication).
I refer to t-stats that exceed $\tgood$ as “well-observed.” $\tgood$ is typically 1.96 (andrews2019identification) or a number between 1.96 and 3.0 (harvey2016and,chen2020publication). Relatively few t-stats below $\tgood$ are observed and the data moments in this region are distorted. In contrast, the likelihood conditional on $|t_{i}|>\tgood$ is not distorted, so one can make inferences about $\pi_{F}$ and $\lambda$ directly from the conditional moments in this region.
In reality, the probability of publication depends not only on $|t_{i}|$ but also on supporting evidence. However, data on supporting evidence is rarely available, so the publication bias literature focuses on Equation ((ref)), which ignores supporting evidence. The net effect of this misspecification is unclear. Omitting supporting evidence tends to bias estimates of $\pi_{F}$ and $\lambda$ in a way that implies stronger latent effects, as these parameters must absorb the larger $|t_{i}|$ values induced by the supporting evidence. On the other hand, omitting supporting evidence tends to imply weaker latent effects, as the supporting evidence is by definition a signal of strong latent effects. Appendix (ref) formalizes these issues and provides some simulation evidence suggesting that the net bias is small.
The model abstracts from dynamic issues like out-of-sample decay in effect size. More generally, the effect size depends on whether it is measured in the original sample ($\mu_{i}$) or out-of-sample ($\mu_{i}^{\text{OOS}}$). While researchers typically aim to find factors with stable effect sizes, empirical evidence in cross-sectional asset pricing finds $\mu_{i}^{\text{OOS}}\approx0.50\mu_{i}$ (mclean2016does,chen2020publication). Thus, finding that $\mu_{i}$ is close to $t_{i}$ does not necessarily imply out-of-sample robustness.
Calls for raising statistical hurdles often come down to controlling the false discovery rate with a t-stat hurdle:
where $\Fdr\left(h\right)$ is the Bayesian formulation from efron2008microarrays:
(see also efron2001empirical and storey2002direct). Equation ((ref)) finds the lowest t-stat hurdle that results in less than 5% of factors being false, where 5% is chosen for ease of exposition (the main results hold for other choices). $\Fdr\left(h\right)$ is used in both ioannidis2005most and benjamin2018redefine. jensen2021there estimates an object similar to $\Fdr\left(1.96\right)$ for cross-sectional return predictors. The chen2018publication version of chen2020publication estimates both $\text{hurdle}\left(5\%\right)$ and $\Fdr\left(h\right)$.
As shown by storey2002direct and storey2004strong, the Bayesian formulation is equivalent to the Benjamini-Hochberg (benjamini1995controlling,benjamini2000adaptive) approach under weak dependence. To see this, define the Benjamini-Hochberg FDR:
where $I\left(\cdot\right)$ is an indicator function and I assume $\sum_{i=1}^{N}I\left(|t_{i}|>h\right)>0$. Then suppose the following weak law of large numbers hold: For any $h\in\mathbb{R}_{+}$,
These expressions say that if you keep counting the shares of factors that exceed a hurdle, you'll eventually find the probability that a factor exceeds a hurdle. Rearranging Equation ((ref)) and plugging in Equations ((ref))-((ref)) leads to Equation ((ref)):
where the last equality applies Bayes rule. harvey2016and's Section 4 uses Equation ((ref)) in the constraint of Equation ((ref)) to argue that t-stat hurdles need to be raised.
As discussed in benjamini2000adaptive (Remark 2), Equations ((ref)) and ((ref)) imply that hurdles need not be raised and may even be lowered. The following proposition characterizes exactly when this occurs:
The proof is in Appendix (ref).
Proposition (ref) says the t-stat hurdle needs to be raised if and only if $\pi_{F}$ is larger than $\Pr\left(|t_{i}|>1.96\right)$. This comparison is visualized in Panel (b) of Figure (ref). $\pi_{F}$ is the share of false factors (red) relative to the total mass, while $\Pr\left(|t_{i}|>1.96\right)$ is the mass to the right of the vertical line. In this illustration, these probabilities are equal, so the multiple testing hurdle is equal to the classical one. In contrast, Panel (a) shows a setting in which the share of the red mass (100%) is much larger than the mass to the right of the vertical line (5%), implying that hurdle needs to be raised.
Some readers may have the intuition that $\pi_{F}>\Pr\left(|t_{i}|>1.96\right)$ must be the “right” case of Proposition (ref). This intuition is perhaps natural, as statistics is a conservative institution. However, there is no logical reason for why we must have $\pi_{F}>\Pr\left(|t_{i}|>1.96\right)$. Assuming that this expression holds amounts to placing additional restrictions on the model of Section (ref) and it is unclear how to motivate these restrictions. Thus, it is the role of data to tell us which case of Proposition (ref) is the right one. For example, benjamini2006adaptive reviews several methods for estimating $\pi_{F}$.
The methods in benjamini2006adaptive, however, assume that all factors are observed. Under publication bias (Equations ((ref))-((ref))) it may be difficult to identify $\pi_{F}$. Indeed, the meta-study literature on publication bias runs into flat likelihood functions and estimation problems for other parameters (copas1999works and hedges2005selection). The following proposition shows publication bias leads to identification problems for $\pi_{F}$:
The proof is in Appendix (ref).
Proposition 2 says that, under certain assumptions, $\pi_{F}$ has little effect on model predictions. In particular, Equation ((ref)) says the distribution of t-stats in the well-observed region changes by at most a factor of $2\varepsilon$ across values of $\pi_{F}$ ranging from 0 to 2/3. The bound of 2/3 is chosen for illustrative purposes. More generally, a bound of $\bar{\pi}$ for $\pi_{F}$ implies an upper bound of $\bar{\pi}/(1-\bar{\pi})\varepsilon$ on the RHS of Equation ((ref)) (see proof).
The key assumption is Equation ((ref)), which formalizes the implicit assumption in hypotheses testing that true and factors are distinct. If true and false factors are not distinct, then the pursuit of separating true from false factors is in a sense misguided. In my setting, false factors by definition have $|t_{i}|$ close to zero, so distinctiveness can be thought of as saying that true factors are more likely to have $|t_{i}|>\tgood$. Equation ((ref)) says that true factors are $1/\varepsilon$ times more likely to be found in this region than false factors, where $\varepsilon$ is presumably small.
Figure (ref) illustrates Proposition (ref) by overlaying alternative parameterizations of \citet*{harvey2016and}'s parametric model. HLZ's baseline estimate implies $\pi_{F}=0.444$ and $g\left(\cdot|\lambda\right)$ is an exponential distribution with mean of about 2. This estimate implies that raising t-hurdles is necessary. According to HLZ's Table 5, the classical t-hurdle of 1.96 needs to be raised to 2.27 to control the FDR at 5%. The distribution of t-stats implied by HLZ's baseline estimate is shown in the blue bars of Figure (ref).
But now consider an alternative model, that uses all the same parameters as HLZ's baseline but simply changes $\pi_{F}$ to zero. With no false factors, the alternative model implies that t-hurdles can be lowered, all the way to 0. The predicted distribution of t-stats is shown in the red bars of Figure (ref).
These two models differ markedly in their densities of t-stats near zero. However, these t-stats are difficult to observe and must be extrapolated. Indeed, HLZ assume $\tgood=2.57$ (vertical line), and for $|t_{i}|>\tgood$ the two distributions are nearly identical. Thus, it is unlikely that the data can tell us whether t-stats should be raised, stay the same, or even be lowered.
This identification problem should show up in standard error estimates. They imply that small perturbations to the data imply very different t-stat hurdles, and thus large standard errors. HLZ do not provide standard errors for their hurdle estimates. I fill this gap in Section (ref).
Revised hurdles are not the only way to deal with a many-factor setting. Chapter 1 of efron2012large's textbook on large scale inference examines empirical Bayes shrinkage:
where $\bar{t}\in\mathbb{R}_{+}$. Equation ((ref)) measures how much you should shrink an observed t-stat ($|t_{i}|$) toward zero in order to recover its unbiased counterpart ($E\left(\mu_{i}\left||t_{i}|=\bar{t}\right.\right)$). For example, $\text{shrinkage}\left(4.0\right)=0.25$ means that a t-stat of 4.0 should be shrunk by 25% and that the unbiased t-stat is $4.0\times(1-0.25)=3.0$. Since the sample mean return is proportional to $t_{i}$, the same shrinkage adjustment applies to the sample mean return. Asset pricing papers that study empirical Bayes shrinkage include chen2020publication, chinco2021estimating, chen2023zeroing, and jensen2021there.
Chapter 2 of efron2012large examines the local FDR:
$\fdr\left(\bar{t}\right)$ is just the probability that a factor with a t-stat of $\bar{t}$ is false. Integrating $\fdr\left(\bar{t}\right)$ from $h$ to infinity leads to Equation ((ref)).
In contrast to $\text{hurdle}\left(5\%\right)$, empirical Bayes shrinkage and the local fdr can be used to target only published findings (by selecting $\bar{t}$ to be equal to the t-stats of published factors). And in the published region, shrinkage and the local fdr depend more so on the parameter vector that governs true factors ($\lambda$). To see this, note that the key term in Equation ((ref)) can be written as
where $f_{|t||F}\left(\bar{t}\right)$ and $f_{|t||T}\left(\bar{t}\right)$ are the densities of $|t_{i}|$ given that $i$ is false or true, respectively. Intuitively, shrinkage of the full model is the same as shrinkage assuming $\pi_{F}=0$, but with a correction that increases in $\pi_{F}$. A similar argument can be made for the local FDR. Appendix (ref) provides the derivations.
For empirically relevant $\bar{t}$ and $\lambda$, $\Delta$ is often small. For example, the HLZ dataset has a mean $|t_{i}|$ of around 4.0, corresponding to $f_{|t||F}\left(4.0\right)=0.00027$. Meanwhile, their estimates imply $f_{|t||T}\left(4.0\right)=0.076$ , as true factors are much more likely to have large t-stats. So even if $\pi_{F}$ is as large as 0.9, $\Delta$ is only 0.03. Similarly large t-stats are commonly found in replications of cross-sectional predictability papers (ChenZimmermann2021). More broadly, brodeur2016star find that t-stats of 4.0 are quite common in the main hypotheses tests reported in top economics journals.
This analysis rests critically on $\lambda$, which must be estimated with published data and its associated identification problems. Fortunately, the following proposition shows that $\lambda$ can be strongly identified
The proof is in Appendix (ref).
Proposition (ref) says that, under the same condition that lead to weak identification of t-hurdles (Equation ((ref))), the model's predictions about well-observed data ($P\left(|t_{i}|<\bar{t}||t_{i}|>\tgood\right)$) are almost the same as the predictions that come from assuming all factors are true ($P\left(|t_{i}|<\bar{t}||t_{i}|>\tgood,T_{i}\right)$). As in the Proposition (ref), this proposition uses the $\pi_{F}<2/3$ for illustrative purposes. The more general case in which $\pi_{F}$ is bounded by $\bar{\pi}$ implies an upper bound of $\bar{\pi}/(1-\bar{\pi})\varepsilon$ on the RHS of Equation ((ref)) (see proof).
An implication of Proposition (ref) is that a reasonable estimate of $\lambda$ could potentially be found by estimating the model while assuming $\pi_{F}=0$. This result can be seen in Figure (ref). The red bars are essentially the model that one would estimate assuming $\pi_{F}=0$. Thus, assuming $\pi_{F}=0$ would lead to roughly the same $\lambda$ as HLZ's baseline estimate, in which $\pi_{F}=0.44$. Intuitively, if $\pi_{F}$ is not at all responsible for fitting the t-stats to the right of 1.96, it must be the other parameters in HLZ's estimate that accomplish this.
Taken with Equation ((ref)), Proposition (ref) implies that the bias in large t-stats could potentially be identified with just published data. This result is seen in the small standard errors in chen2020publication's shrinkage estimates as well as the robustness of their estimates to alternative modeling assumptions. Indeed, Chen and Zimmermann's headline bias estimate of 12% is not far from harvey2021uncovering's median estimate of 19%, even though the two papers use very different modeling assumptions and estimation methods. Similar estimates are also found in chinco2021estimating and jensen2021there, who use yet another set of methods (see discussion in chen2023publication). Section (ref) provides additional evidence that these bias estimates are strongly identified for the cross-sectional predictability literature.
The key to strong identification is flexibility. Unlike, $\text{hurdle}(5\%)$, $\text{shrinkage}\left(\bar{t}\right)$ and $\text{\fdr\ensuremath{\left(\bar{t}\right)} }$allow one to focus on the portion of the data that is well-observed. If the data are not informative about $\text{shrinkage}\left(1.0\right)$, one can instead examine $\text{shrinkage}\left(4.0\right)$. The cost of this flexibility is that $\text{shrinkage}\left(\bar{t}\right)$ and $\text{\fdr\ensuremath{\left(\bar{t}\right)}}$ do not provide a precise answer to the question, “do t-hurdles need to be raised?” They provide important evidence (e.g. published t-stats are biased upward by 12%). But ultimately an additional framework is required to convert this evidence into a new t-hurdle.
Also unlike $\text{hurdle}(5\%)$, $\text{shrinkage}\left(\bar{t}\right)$ provides economic magnitudes. Thus, it avoids the potentially artificial dichotomization that can come from hypothesis testing (mcshane2019abandon). This gain, however, comes at additional analytical costs, namely that $\text{shrinkage}\left(\bar{t}\right)$ requires selecting $g(\cdot|\lambda)$ and estimating $\lambda$.\footnote{Under publication bias, the local FDR can be bounded without estimating $\lambda$ by using a kind of worst case scenario (chen2021most).}
An important caveat of this analysis is that it relies on $\pi_{F}$ being not too close to 1.0. For $\pi_{F}=0.9$ or higher, the $2\varepsilon$ bound becomes a $9\varepsilon$ bound (or higher), and Proposition (ref) may have little bite. Intuitively, if almost all factors are false, then even the extreme right tail of the distribution will be largely determined by false factors. However, we will see that empirical evidence suggests $\pi_{F}\ge0.9$ is unlikely cross-sectional return predictors (Section (ref)).
Some FDR methods avoid estimation of $\pi_{F}$ by effectively assuming $\pi_{F}\ge1.0$. As a result, these conservative methods assume that t-hurdles need to be raised, and thus cannot answer the question of whether t-hurdles need to be raised.
These conservative algorithms include benjamini2001control's (BY's) Theorem 1.3, which is emphasized in HLZ, and is popular in the finance literature. As shown by storey2002direct, BY's Theorem 1.3 is equivalent to replacing the constraint in Equation ((ref)) with an estimator:\footnote{This is seen Equation (13) of storey2002direct. See also the Equivalence Theorem of efron2002empirical. Strictly speaking, these papers show this equivalence for the benjamini1995controlling algorithm, however, the BY algorithm simply modifies the benjamini1995controlling algorithm with a constant factor.}
where
Compare Equation ((ref)) to the result of applying Bayes rule to $\Pr\left(F_{i}\left||t_{i}|>h\right.\right)$
Thus, $h_{BY1.3}\left(5\%\right)$ effectively assumes that $\pi_{F}=\sum_{i=1}^{N}\frac{1}{i}$. In HLZ's case, where $N\approx300$, we have $\pi_{F}=6.3$. So this algorithm implies that, not only is $\pi_{F}$ is large, but it is six times larger than 1.0. Since $\pi_{F}$>1.0 is physically impossible, it is hard to justify why such a severe penalty is necessary. Indeed, $h_{BY1.3}\left(5\%\right)$ is described as “a severe penalty” and “not really necessary” (efron2012large). Even the original benjamini2001control paper says this algorithm is “very often unneeded, and yields too conservative of a procedure.”
An implication of this conservatism is that BY's Theorem 1.3 always implies that t-hurdles need to be raised. Assuming $\pi_{F}>1$, only one case of Proposition (ref) is possible:
The proof is in Appendix (ref).
Similarly, the seminal benjamini1995controlling (BH95) algorithm is equivalent to replacing $\left(\sum_{i=1}^{N}\frac{1}{i}\right)$ with 1.0 in Equation ((ref)). Thus, BH95 is isomorphic to using $\pi_{F}=1.0$ in Equation ((ref)), and Proposition (ref) implies that BH95 raises t-hurdles in all cases except for the extreme case in which $\sum_{i=1}^{N}I\left(|t_{i}|>1.96\right)=N$.
The benefit of imposing $\pi_{F}=1.0$ is simplicity. The BH95 hurdle can be computed for an arbitrary dataset without any additional specifications. The cost, however, is that BH95 discards all information in the data that can inform us about $\pi_{F}$. So while BH95 is an excellent first pass for controlling for multiple testing, it cannot tell us which case of Proposition (ref) holds, and cannot tell us whether t-hurdles must be raised.\footnote{Unlike the finance literature, the statistics literature tends to favor the BH95 algorithm over BY's Theorem 1.3 (benjamini2020selective). HLZ argue for using BY's Theorem 1.3, as it controls the FDR under arbitrary dependence. But BH95 controls the FDR under weak dependence (storey2004strong) and simulations suggest FDR control under arbitrary dependence if the tests in question use z-statistics (reiner2007fdr).}
In their first draft, Benjamini and Hochberg (1995) did not assume that t-stat hurdles should be raised. As described in benjamini2010discovering, the 1989 draft recommended a graphical method for estimating $\pi_{F}$. After many years of rejections, the authors shifted to the conservative formulation that became the seminal BH95 algorithm. The original algorithm, along with the result that t-stat hurdles could possibly be lowered, was eventually published in benjamini2000adaptive, and many statisticians went on to propose additional estimators for $\pi_{F}$ (efron2001empirical,allison2002mixture,storey2002direct,genovese2004stochastic,benjamini2006adaptive, among others).
A natural question is how publication bias affects identification of false negative rates (FNRs), also known as the type II error rate. This question can be analyzed using a Bayesian formulation of the FNR:
(see efron2012large Chapter 4.3). $\Fnr(h)$ is the probability a factor is true, given that the factor fails to meet the hurdle $h$. One can alternatively define the FNR using factor counts as in Equation ((ref)) (see genovese2002operating) but these definitions are equivalent under weak dependence.
Applying Bayes rule shows how the FNR runs into identification issues:
In the presence of publication bias, $\pi_{F}$ may be weakly identified (Proposition (ref)). Moreover, the very definition of publication bias (Equations ((ref))-((ref))) implies that the data that bears on $\Pr\left(|t_{i}|\le h\right)$ is limited. Intuitively, to estimate a false negative rate one needs to know the number of insignificant t-stats, and publication bias means that insignificant t-stats are poorly observed.
An alternative way to address the FNR is to use the FDR control that comes closest to achieving the constraint in Equation ((ref)). Achieving this constraint would typically involve forming a point estimate of $\pi_{F}$ (rather than assuming an upper bound), as is pursued in this paper.
The theoretical results illustrate weak and strong identification, but a precise description of these issues requires empirical estimates of sampling uncertainty. This section provides one such estimate by bootstrapping estimates using a dataset constructed from asset pricing publications.
My data begins with 207 published cross-sectional stock return predictors from the ChenZimmermann2021 (CZ) dataset (March 2022 release). I focus on their original predictor portfolios, which consists of long-short portfolios constructed following the procedures in the original studies.
This dataset has several advantages as a setting for studying publication bias. CZ show that the replicated t-stats closely match the originals, which rules out coding errors, fraud, or other more nefarious sources of bias. CZ also show that the pairwise correlations cluster around zero (see also \citet*{chen2021most,bessembinder2021time}), suggesting that the weak dependence assumptions underlying standard estimation methods are valid (wooldridge1994estimation), and that modeling the correlations will add relatively little to estimation efficiency. Moreover, the monthly returns in this dataset can be used to account for correlations in my bootstrapped standard errors. chen2021limits shows the distribution of t-stats is quite similar to the distribution found in HLZ, which eases comparison. Last, this dataset is publicly available at www.openassetpricing.com.
To simplify the baseline model, I drop predictors with in-sample t-stats that fall below 1.96, leading to a sample of 183 predictors. The approach of dropping t-stats $<$ 1.96 is also used in HLZ. Carefully modeling the 24 predictors with smaller t-stats requires a relatively complicated publication probability function and is likely to introduce more identification problems (copas1999works). In the robustness section, I include t-stats < 1.96 and find broadly similar results (Section (ref)).
Since all of CZ's factors are cross-sectional return predictors, I refer to them as “predictors” in what follows.
The structural model adds three assumptions to the model of Section (ref). The additional structure can be thought of as restrictions that help identify the key model objects (e.g. $\pi_{F},$ $\Pr\left(|t_{i}|>1.96\right)$).
The first is the unbiased t-stat $\mu_{i}$ is log-normal for true factors
where $\lambda_{\mu}$ and $\lambda_{\sigma}$ are mean and standard deviation of $\left(\log\mu_{i}\right)|T_{i}$, respectively. This form is arguably the simplest assumption one can make about $\mu_{i}|T_{i}$ while ensuring that true predictors have an expected return that is distinct from 0. Section (ref) shows that alternative distributional assumptions have little effect on the results.
The second assumption is a functional form for the publication probability (Equation ((ref))). I assume a staircase function
where the t-stat cutoffs of 1.96 and 2.58 correspond to the traditional 5% and 1% significance cutoffs, respectively, and $\eta$ represents a “haircut” for marginally-significant predictors. One can think of this function as the simplest intuitive model of selective publication for marginally-significant predictors. This functional form is also assumed in HLZ. Section (ref) shows that a logistic form (as in chen2020publication and harvey2021uncovering) leads to similar results.
The last assumption is that $\eta$ lies within an interval obtained from intuition
Equation ((ref)) says that at least some marginally-significant predictors are missing ($\eta\le2/3$), but at least a significant minority are reported ($\eta\ge1/3$). HLZ uses the more restrictive assumption that $\eta=1/2$, though they also examine $\eta=1/3$. Section (ref) shows that relaxing this assumption or restricting this assumption has little effect on the main results.
I also assume $\pi_{F}\in[0.01,0.99]$ for technical reasons. I use numerical integration to compute the log-likelihoods and multiple testing statistics, and I find that these methods can become poorly behaved for $\pi_{F}$ that is extremely close to 0 or 1.0. One can increase the range of $\pi_{F}$ with more complicated numerical methods but this complication is unlikely to affect the main results.
I estimate $\theta\equiv\left(\pi_{F},\lambda_{\mu},\lambda_{\sigma}\right)$ with quasi-maximum likelihood (QML). I choose $\hat{\theta}$ to maximize the mean log marginal likelihood of $|t_{i}|$ conditional on $i$ being published. I do not estimate $\bar{s}$, as it is not identified (it is cancelled out in the conditional likelihood). Under technical conditions, this QML estimator is consistent (\citet*{wooldridge1994estimation,liu2020forecasting}).\footnote{The key technical assumption is that the log marginal likelihood of a single observation satisfies the uniform weak law of large numbers.} Intuitively, the model implies that the mean derivatives of the marginal likelihoods are zero when evaluated at the true parameters, and this implication can be used as moment conditions to estimate the model.
I confirm that QML is consistent under pairwise correlations as high as 0.90 in Appendix (ref). Indeed, in simulated samples of only 200 t-stats, I find QML is essentially unbiased for $\pi_{F}$ and produces standard errors of around 0.10 to 0.15. These simulations assume no publication bias and demonstrate that the large standard errors I find for $\pi_{F}$ are not due to QML.
I measure estimation uncertainty using two bootstraps. My baseline bootstrap is simple and non-parametric: I draw a full set of t-stats from the empirical data with replacement and re-run QML. This bootstrap is transparent and also ensures that the bootstrapped data displays a similar selection bias and distributional properties as the original dataset. Moreover, cross-sectional predictor correlations are typically close to zero (mclean2016does,ChenZimmermann2021). These mild correlations suggest that this simple bootstrap will lead to similar results as one that accounts for correlations.
For robustness, I also examine a semi-parametric bootstrap that carefully accounts for correlations. In short, this second approach combines a cluster bootstrap with a parametric bootstrap. The cluster bootstrap ensures that the samples closely mimic the correlation structure in the data while the parametric bootstrap ensures that the samples mimic the distribution of observed t-stats. Details on the semi-parametric bootstrap are found in Appendix (ref)
Table (ref) shows the resulting parameter estimates. The point estimate finds that essentially all factors are true ($\hat{\pi}_{F}=0.01$), but the 90% confidence interval is huge, ranging from 0.01 to about 0.90. Even the 50% C.I. is quite large, ranging from 0.01 to about 0.70. There results are consistent with Proposition (ref) and suggest that t-hurdles are weakly identified.
Huge uncertainty about $\pi_{f}$ obtains regardless of whether I use the simple non-parametric bootstrap (shown in the table with no parentheses) or the semi-parametric cluster bootstrap (shown with parentheses). For all parameters in Table (ref), accounting for correlations has almost no effect on estimation uncertainty. This result is intuitive, as the typical predictor correlation is close to zero (ChenZimmermann2021, see also Figure (ref) in the Appendix).
In contrast to $\pi_{F}$, the parameters that govern true factors ($\lambda_{\mu}$ and $\lambda_{\sigma}$) are strongly identified. Table (ref) converts these parameters into the expected unbiased t-stat $\E\left(\mu_{i}|T_{i}\right)$ and standard deviation of the unbiased t-stat $\SD\left(\mu_{i}|T_{i}\right)$ for ease of interpretation.\footnote{The lognormal distribution implies
} The bootstrapped estimates imply that $\E\left(\mu_{i}|T_{i}\right)$ is between 2.0 and 3.8 while $\SD\left(\mu_{i}|T_{i}\right)$ is between 1.9 and 2.7, with 90% confidence. These results are consistent with Proposition (ref).
Figure (ref) provides the intuition behind these results. It compares the predictions of the point estimate with an alternative model that also uses QML, but the alternative model fixes $\pi_{F}$ at 2/3. Panel (a) shows that both models fit data to the right of $2.58$ very well. Recall that this is the subsample of the data that is well-observed (Equation ((ref))). Thus, even a slight perturbation to the observed data can move the point estimate from $\pi_{F}=0.01$ to $\pi_{F}=2/3$.
The models differ in their predictions about t-stats between 1.96 and 2.58, suggesting the data in this region can help identify $\pi_{F}$. However, many t-stats in this region may be missing, and it is difficult know a-priori just how many. Both QML estimates restrict uncertainty by assuming that the fraction missing is between 1/3 and 2/3 (Equation ((ref))). Nevertheless, this restriction is unable to identify $\pi_{F}$. Section (ref) shows that even assuming the missing fraction is 0.5 cannot identify $\pi_{F}$
Thus, the key to identification is the distribution of t-stats below 1.96. As seen in Figure (ref), the two models have very different predictions for this region. The data tell us very little about which is the right story, as there are only 24 observations in this region. Indeed, a close read of ChenZimmermann2021 shows that even these 24 observations are poorly defined, as it is unclear if these observations should even be called “predictors.”
This ambiguity leads me to drop t-stats below 1.96 in the estimation (though they are still shown in Figure (ref)). Nevertheless, Section (ref) shows that taking on this additional data still implies substantial uncertainty. Intuitively, only 7 t-stats below 1.5 are observed, and it is very hard to say how representative these t-stats are. This very small sample problem suggests that adding the standard errors of predictor mean returns to the estimation, which in principle can non-parametrically identify the publication probability (andrews2019identification), will still result in weak identification of $\pi_{F}$. Indeed, previous versions of this paper did include these standard errors in the estimation and found similar results.\footnote{Code for the previous versions can be found at \url{https://github.com/chenandrewy/t-hurdles}.}
Panel (b) of Figure (ref) illustrates strong identification. It examines the distribution of true predictors' t-stats $(t_{i}|T_{i})$ in the point estimate and the alternative model. Despite very different implications about $\pi_{F}$, both models imply similar distributions of $t_{i}|T_{i}$. Intuitively, a highly dispersed $t_{i}|T_{i}$ is required to fit the long right tail in observed t-stats, regardless of $\pi_{F}$. Since $t_{i}|T_{i}$ is just $\mu_{i}|T_{i}$ plus standard normal noise, this result implies that the parameters governing $\mu_{i}|T_{i}$ are strongly identified, leading to the relatively narrow confidence bounds in Table (ref).
Returning to Table (ref), the probability of publishing a marginally-significant t-stat ($\eta$) tends to be larger than 50% in its bootstrapped distribution, implying that the asset pricing literature does not strongly discriminate against predictors with t-stats between 1.96 and 2.58. This result is intuitive, as the asset pricing literature contains relatively little discussion of marginal significance, and instead focuses on economic mechanisms.
Table (ref) shows that the bootstrapped $\eta$ often bumps against the upper bound of $2/3$ imposed by the model (Equation ((ref))). This restriction was imposed to exclude the possibility that weak identification is due to an excessively general model. Section (ref) shows that relaxing these bounds does not change the main results.
As very similar estimates obtain using either the non-parametric or semi-parametric bootstrap, for the remainder of the paper I discuss only the non-parametric bootstrap.
With bootstrapped parameters in hand, I can finally quantify weak and strong identification of multiple testing statistics.
Figure (ref) plots the bootstrapped distribution of t-hurdles (Equation ((ref))). Panel (a) examines the more common FDR upper bound of 5%, which results in a highly dispersed t-stat hurdle of between 0 and 3.0. In other words, the data say little about how to adjust t-hurdles for multiple testing. Indeed, the classical 5% counterpart of 1.96 lies right in the middle of the distribution, implying substantial uncertainty about whether the hurdle should be raised or lowered.
The estimates imply a large mode in t-hurdles close to 0, reflective of the large mode in bootstrapped estimates of $\pi_{F}$ close to zero that is implicit in Table (ref). This mode is intuitive given the shape of the empirical distribution (Figure (ref)): This unimodal shape provides little evidence in support of a bimodal distribution in $\mu_{i}$, and thus simply assuming $\mu_{i}$ is lognormal ($\pi_{F}=0$) provides a strong fit to the data. This result is also consistent with Proposition (ref) and Figure (ref), which show that the observed distribution of the full model is approximately the same as a model with $\pi_{F}=0$.
Some readers may find $\pi_{F}=0$ to be implausible, but Figure (ref) suggests that this mode is immaterial. Excluding the mode at 0, the distribution of t-hurdle estimates is still highly dispersed, with values ranging from 1.0 to 3.0. Additionally, Section (ref) shows that restricting $\pi_{F}\ge0.20$ still results in substantial uncertainty about the right t-hurdle.
Panel (b) shows that weak identification is also seen if one selects the unusually restrictive FDR $\le$ 1% t-hurdle. The resulting distribution of t-hurdles is, once again, highly dispersed, ranging from 0 to 3.5. The classical 1% hurdle of 2.58 lies well-within the interior of this distribution.
Figure (ref) shows that the data are informative about shrinkage and the local FDR for published predictors. Panel (a) examines shrinkage by evaluating Equation ((ref)) using bootstrapped data and then taking the mean across published t-stats within each bootstrap. The panel shows the distribution of this mean across bootstraps. The vast majority of the distribution lies below 26%, which is the upper bound on statistical bias estimated in mclean2016does's out-of-sample tests. This result implies that published sample mean returns are at least 74% due to true expected returns rather than publication bias, multiple testing, or other related statistical effects.
The mode of the shrinkage distribution is close to 12%, which is the point estimate found by chen2020publication. This similarity obtains despite the fact that I assume a bi-modal distribution in t-stats, providing a counterexample to harvey2021uncovering's claim that Chen and Zimmermann's estimate is sensitive to their unimodal assumption. Indeed, Section (ref) shows Chen and Zimmermann's estimate is robust to 10 alternative models, all of which use bi-modal distributions.
Panel (b) shows similar results for the mean local FDR. Panel (b) computes the mean local FDR by evaluating Equation ((ref)) at bootstrapped parameter values and taking the mean across published t-stats within the bootstrap. Quantitatively similar to Panel (a), Panel (b) finds that published factors are at least 75% true, with high confidence. The figure displays a large mode near 0, due to the unimodal shape of the empirical data (see Section (ref)), but even if this mode were excluded the estimates still imply that the published predictors are largely true with high confidence.
For comparison, Figure (ref) also plots the mean local FDR for published predictors implied by HLZ's baseline estimates. HLZ's estimate implies that only 6% of published findings are false, close to the middle of the bootstrapped distribution. This result is surprising, given HLZ's verbal statement that “most claimed research findings in financial economics are likely false.” However, this estimate is natural given their numerical estimates.
To understand HLZ's numerical estimates, note that the mean local FDR for published factors is approximately the Bayesian expression
HLZ find that $\pi_{F}=0.444$ (Table 2) and that of 1,378 total t-stats (also Table 2), 353 t-stats that exceed 1.96 (page 28). Plugging these numbers into Equation ((ref)) leads to
This approximation is a touch higher than the 6% found by applying Equation ((ref)) to their model simulation because the simulation assumes a more stringent publication criteria (only half of t-stats between 1.96 and 2.58 are published). chen2021most finds that even HLZ's conservative estimates imply the FDR is rather small.
Figures (ref) and (ref) depend on functional form and distributional assumptions. However, Propositions (ref) and (ref) do not, suggesting that similar results may obtain under a wide variety of assumptions.
This subsection confirms this robustness for 10 alternative sets of assumptions. Table (ref) shows the 5th and 95th percentile of multiple testing statistics, computed across 500 bootstrapped estimates for each set of assumptions. Most of the 90% confidence intervals for t-hurdles span 0 and 2.9, implying substantial uncertainty about the proper correction to the classical hurdle of 1.96. In contrast, all specifications imply shrinkage is at most 29%, and the local FDR is at most 25%, with 95% confidence.
Indeed, previous versions of this paper found similar results for additional estimation and modeling assumptions, including estimations that account for standard errors and t-stats separately, and estimations that modeled the full distribution of correlations. Additional robustness can be found using the github code ({[}a github site{]}) or using the github code for the previous draft ({[}a github site{]})
The remainder of this section briefly describes each set of assumptions and their motivations.
Section (ref) simplifies modeling by dropping the 24 t-stats that fall below 1.96. Specifications (2) and (3) include these small t-stats\textemdash as long as they exceed 0.50. The three t-stats that fall below 0.50 are 0.06, 0.09, and 0.40. Including these tiny t-stats leads to shrinkage estimates that divide by tiny numbers (Equation ((ref))), leading to extreme outliers that drive the means.
To accommodate the additional 21 t-stats that fall below 1.96, I add two additional steps to the staircase publication probability:
where the additional cutoff of 1.50 is chosen to match the cutoff used in mclean2016does. 1.50 is also close to the two-sided 10% hurdle of 1.64.
Specifications (2) and (3) differ in how they restrict the parameters $\eta_{a}$, $\eta_{b}$, and $\eta_{c}$. Specification (2) makes no restrictions. Specification (3) assumes $\eta_{c}\in[1/3,2/3]$ (same as in the baseline), and imposes that the other probabilities are smaller than 1/3.
The baseline estimates show a large mode at $\hat{\pi}_{F}=0.01$ (Table (ref)), leading to large modes in Figures (ref) and (ref). Some readers may find such a small $\pi_{F}$ implausible and would like to restrict $\pi_{F}$ to some minimum value.
In Table (ref), Specifications (4) and (5) use the restrictions $\pi_{F}\ge0.10$ and $\pi_{F}\ge0.20$, respectively.
The baseline model restricts $\eta\in[1/3,2/3]$ to show that weak identification is not due to an excessively general model. Specification (6) restricts even further to $\eta=0.5$, which is HLZ's baseline assumption.
Specification (7) examines a more general $\eta\in[1/3,1.0]$. This choice is motivated by the fact that the baseline bootstrap bumps up against the upper bound of $\eta=$2/3 (Table (ref) ), suggesting that $\eta\ge2/3$ is preferred by the data.
Specification (8) assumes a logistic form for the publication probability
which is the same function used in chen2020publication (see also cochrane2005risk,harvey2021uncovering).
The baseline model assumes $\mu_{i}|T_{i}$ is log-normal, as this is the simplest assumption one can make that ensures that true predictors have a positive mean return that is distinct from 0.
In Table (ref), Specification (9) assumes $\mu_{i}|T_{i}$ is exponential following HLZ and Specification (10) assumes $\mu_{i}|T_{i}$ is a scaled t-distribution, similar to andrews2019identification and chen2020publication. Specification (11) assumes $\mu_{i}|T_{i}$ is a mixture of two normals, which can be motivated by the idea that there are two distinct types of true predictors, each with a bell-shaped distribution.
Specifications (10) and (11) often imply that $\mu_{i}|T_{i}$ can be either positive or negative, which implies that the sign of $t_{i}$ should be accounted for in the shrinkage estimation (Equation ((ref))). For these two rows, the shrinkage shown thus uses
which implicitly assumes that the journals use theory to find the appropriate sign of predictability. Most papers in the Chen-Zimmermann dataset motivate their signs using theory.
Motivated by concerns about scientific credibility, a series of papers call for statistical hurdles to be raised. I show these calls may be difficult to justify empirically. Publication bias means that the key parameter in these arguments may be weakly identified. As a result, empirical data may say little about whether hurdles should be raised, stay the same, or even be lowered.
More positively, I find that other multiple testing statistics can be strongly identified. In particular, shrinkage and the local FDR for published factors may be able to be pinned down without knowledge of t-stats near zero. For the cross-sectional return predictability literature, these statistics imply that published findings are at least 75% true, with high confidence, across many model specifications.
What “at least 75% true” says about the proper $t$-hurdle is subjective, but at least some readers will argue that a literature that is mostly true is a healthy one. Still, others may argue that a literature should strive for trust, and even for achieving “no questions asked” from their readers.
A caveat about these empirical results is that they are based on the field of cross-sectional predictability. While other fields of economics also feature very large t-stats (brodeur2016star), a rigorous analysis is required to pin down the FDR. Moreover, cross-sectional predictability is known for its clean data, standardized methods, and strong replicability (ChenZimmermann2021,jensen2021there). These features allay concerns about coding errors and the improper calculation of t-stats that may remain if shrinkage and FDR estimates were applied to other fields.
Regardless, my results demonstrate that the debate about scientific credibility should take multiple testing statistics more seriously. These statistics do not simply say that statistical hurdles should be raised, or that a large share of findings are false. Identification is important, and some multiple testing statistics are more strongly identified than others. Most important, my paper shows that the debate should focus on the more strongly identified statistics.