Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
54,864 characters · 19 sections · 109 citation commands
High-Throughput Asset Pricing
JEL Classification: G0, G1, C1
\noindentKeywords: stock market predictability, stock market anomalies, p-hacking, multiple testing \thispagestyle{empty}\setcounter{page}{0}
\setcounter{page}{1}
Data mining refers to searching data for interesting patterns. This search leads to data mining bias, if many patterns are just chance results, as is surely the case with stock return data. To address this problem, the asset pricing literature recommends restricting the search to patterns consistent with theory (\citet*{cochrane2005risk,harvey2017presidential}). However, recent empirical evidence finds this method is ineffective, even for theories published in top finance journals (\citet*{chen2022peer}).
We offer a different solution. Instead of mining data less, we recommend mining data rigorously. Rigorous data mining means conditioning interesting results on the fact that they come from searching through data. This conditioning can be achieved using empirical Bayes (robbins1956empirical,efron1973stein,efron2012large). Rigorous data mining also means that the search should be systematic, as is commonly done in high-throughput biology and chemistry (yang2021high). Ironically, systematic search implies that asset pricing should involve more data mining, not less.
We use empirical Bayes (EB) to mine for out-of-sample returns among 136,000 long-short trading strategies. The trading strategies are constructed from systematically searching data on accounting ratios, past returns, and stock tickers. Through this “high-throughput asset pricing,” we construct a portfolio with out-of-sample returns that are comparable to the returns from the best journals in finance.
Our data-mined portfolio is the simple average of the top 1% of strategies, based on EB-predicted Sharpe ratios. It earns out-of-sample returns of 5.7% per year over the 1983-2020 sample, compared to the mean return of 5.9% per year found by averaging the 200 published strategies from ChenZimmermann2021. But unlike the published strategies, which were selected with knowledge of stock return patterns that occurred in the 1980s and 1990s, our strategies can be constructed using only information available in real time.
In fact, even naively mining for the largest Sharpe ratios leads to publication-like performance. We provide a theoretical explanation for this phenomenon in Proposition (ref), which shows that under standard statistical practices (fisher1925statistical), naive data mining often selects the same set of strategies as an ideal Bayesian. However, while the naively-selected strategies may be optimal, naive performance estimates are distorted, illustrating the importance of rigorously data mining with EB.
The top 1% portfolio selected by EB provides insights into the nature of return predictability. 91.0% of strategies in this portfolio are equal-weighted accounting ratio strategies. Almost all of the remainder are equal-weighted past-return strategies. Moreover, the returns of the top 1% strategy are concentrated in the pre-2004 data. These facts are consistent with the theory that predictability is largely due to limited attention and the slow incorporation of information into stock prices (peng2005learning; Chordia2014Have).
Other facts shed light on the drivers of the recent decline in cross-sectional predictability. We find that the returns of the top 5% and top 10% of portfolios are also concentrated in the pre-2004 data. These strategies are enormous in number: the top 5% consists of 6,305 strategies, and the top 10% consists of 12,610. As many of these strategies are unlikely to be found in academic journals, this suggests that the key driver of the recent declines in predictability is improvements in information technology (Chordia2014Have), rather than investors learning from academic publications (Mclean2016Does). Consistent with this idea, we find that the top 20 strategies according to predicted Sharpe ratios using data available in 1993 have themes rarely seen in academic journals, like mortgage debt, growth in interest expense, and depreciation. Themes that were popular in academia in 1993, like book-to-market, momentum, and sales growth are missing from this list.
Overall, high-throughput asset pricing provides not only a method for dealing with look-ahead bias, but also a more rigorous method for documenting asset pricing facts. We post our strategy returns and code publicly, and encourage future researchers to use these methods.
Unlike many big data methods, EB provides a transparent intuition. In essence, EB measures the distance between the empirical t-stat distribution and the standard normal null. Ticker-based strategies have t-stats that are extremely close to the null, implying no predictability. In contrast, equal-weighted accounting t-stats are too fat tailed to be consistent with the null, implying strong predictability. Thus, just by visually inspecting the t-stat distributions, one can see where predictability is concentrated.
EB provides highly accurate predictions in pre-2004 data. We construct 120 portfolio tests using the 136,000 data-mined strategies, and compare EB-predicted returns with out-of-sample returns. In almost all of the 120 portfolios, the EB predictions are within 2 standard errors of the out-of-sample mean.
Post-2004, EB has more difficulty with accuracy, though it still captures broad patterns in out-of-sample returns. Compared to pre-2004, predicted returns are closer to zero, and only equal-weighted accounting strategies show notable predicted returns. However, out-of-sample returns are even closer to zero than predicted. This difficulty might be expected given the rise of information technology around 2004, which likely led to a structural break in predictability (Chordia2014Have; kim2021causal). Our EB predictions are constructed using a simple 20-year rolling window, and thus fail to account for this break. This difficulty suggests that a smart data miner armed with theory might have understood the implications of the internet, and could perhaps have performed much better than our theory-free EB mining process.
We also illustrate how improper use of multiple testing statistics can lead to poor data mining results. We demonstrate this possibility using \citepos{harvey2016and} recommended method for false discovery control. harvey2016and recommend applying \citepos{benjamini2001control} Theorem 1.3 to construct a t-stat hurdle that controls the false discovery rate (FDR) at the 1% level. Nearly all of our 136,000 trading strategies fail to meet this hurdle, suggesting that there are few interesting patterns in this data. But in fact, simple out-of-sample tests show there are thousands of strategies with notable out-of-sample returns. We find similar results following the recommended multiple testing control in chordia2020anomalies, which is based on romano2007control. In contrast, the storey2002direct FDR control recommended in barras2010false, captures the majority of notable portfolios.
Fortunately, this error can be avoided by rigorously studying the statistics. According to benjamini2001control, their Theorem 1.3 is “very often unneeded, and yields too conservative of a procedure.” This negative sentiment is echoed in Efron's efron2012large textbook on large scale inference. In contrast, the EB methods we use are recommended for settings like ours in Chapter 1 of efron2012large, as well as Chapters 6 and 7 of efron2016computer.\footnote{A brief explanation of why benjamini2001control Theorem 1.3 is excessively conservative is found in Section 2.5 of chen2024t.} The statistics literature has relatively little to say about the method recommended in chordia2020anomalies. We provide our own characterization, which illustrates how this method is appropriate if selecting a null strategy is catastrophic. But using the standard null, that the mean long-short return (or alpha) is zero, Chordia et al.'s method implies unneeded conservatism.
We add to yan2017fundamental and chen2022peer, who document that mining accounting data can produce substantial out-of-sample returns. Accounting data is important: Chen et al. find that mining ticker variables leads to out of sample returns of approximately zero. Thus, one needs a method for identifying the predictive power of accounting data in real time. Our empirical Bayes formulas provide one such method.
The literature on multiple testing in asset pricing features disagreement on both the methods that should be used and the empirical extent of multiple testing problems. chen2020publication; chen2022zeroing; and jensen2023there recommend empirical Bayes shrinkage. In contrast, harvey2016and; harvey2020false; and chordia2020anomalies recommend conservative false discovery controls, much more conservative than the FDR methods in barras2010false. We show how empirical Bayes shrinkage and the recommended method from barras2010false leads to much more accurate inferences. More recently, marrow2024real use empirical Bayes to study past return signals, with a focus on signal interactions and optimal weighting of more recent data.
In contrast to the intuition that simplicity is a virtue, we find that studying an enormous number of potential predictors leads to insights about the nature of return predictability. A similar theme is found in kelly2024virtue and didisheim2023complexity, who illustrate the “virtue of complexity” in the modeling of expected returns.
We describe the data (Section (ref)) and how we rigorously mine it (Sections (ref)-(ref)).
Table (ref) describes our data-mined strategies. The strategies are either based on accounting ratios, past returns, or tickers. Accounting ratio strategies are taken from chen2022peer.\footnote{We are grateful that the authors make their data publicly available.} The past return and ticker strategies are inspired by yan2017fundamental and harvey2017presidential, respectively, but we generate our own strategies in order to ensure that the number of strategies is comparable across data sources and to ensure that each type of strategy consists of many distinct strategies.\footnote{Results that mine data following yan2017fundamental and harvey2017presidential are similar and can be found in the first draft of our paper on arxiv.org or via our github site.}
A key feature of these strategies is that they are not selected based on having notable historical returns. Instead, they are constructed to systematically explore various types of data. So unlike most datasets in asset pricing (e.g. Ken French's size- and B/M-sorted portfolios; ChenZimmermann2021), ours is arguably free of data mining bias. Indeed, Table (ref) shows that the median sample mean return is close to zero for all sets of strategies.
In high-throughput research, the median measurement is relatively unimportant. What matters is that the extreme measurements show promise for, say, a pharmaceutical intervention or cancer prediction. The extreme measurements in Table (ref) suggest that accounting and past return data show promise for predicting returns. These data lead to mean returns that can exceed 5 percent per year in absolute value.
For further details on the strategy definitions, see Appendix (ref) or our github site.
The 136,000 strategies in Table (ref) contain the potential for significant data mining bias. To understand the bias, let $r_{i}$ be a performance measure for strategy $i$ (e.g. mean return, alpha) and decompose it as follows:
where $\mu_{i}$ is the actual performance and $\varepsilon_{i}$ is sampling error or luck.
Data mining involves selecting $i$ with large $r_{i}$. Suppose we set $\bar{r} \gg 0$, and search for $i^\ast \in \{1;2;\ldots;136,000\}$ such that $r_{i^\ast} = \bar{r}$. This practice is dangerous because one might think $\bar{r}$ is a good estimate of $\mu_{i^\ast}$. However, $\bar{r}$ is in fact biased upward
Selecting for large $r_{i}$ also selects for large $\varepsilon_{i}$, leading to $E\left(\varepsilon_{i^\ast} | r_{i^\ast} = \bar{r}\right) > 0$ and the bias in Equation ((ref)).
To data mine safely, one needs to remove the luck term $E\left(\varepsilon_{i}|r_{i^\ast} = \bar{r}\right)$. This term is just a conditional expectation, so it can be computed using Bayes rule, provided one has a probability model for $\mu_{i}$ and $r_{i}$.
Suppose one has a probability model, with parameter vector $\Omega$. The bias can then be removed by computing
where $\hat{\Omega}$ is a consistent (frequentist) estimate of the probability model parameters. This method, of applying frequentist estimates to Bayesian formulas is known as “empirical Bayes” (robbins1956empirical,efron1973stein).
Equation ((ref)) conditions on only one statistic regarding strategy $i$. A more optimal estimate uses more information
where $X_{i}$ is a vector of additional statistics for strategy $i$ and $\bar{X}$ is a realized value of $X_{i}$. For example, $X_{i}$ can include the standard error of $r_{i}$, the portfolio weighting (equal- or value-weighted), and the signal data source (accounting, past returns, tickers).
We use Equation ((ref)) to search our 136,000 strategies for large expected returns. We will not use economic theory to determine the probability model, and thus our search is largely atheoretical. However, we recognize the bias that comes from such a search (Equation ((ref))), and carefully correct for it. Thus, we describe our methods as “rigorous data mining.”
In empirical asset pricing, we are often interested in two questions:
If one is interested only in the first question, then there is a sense in which naively mining data, without accounting for data mining bias, is often optimal.
To understand this, we add structure to the model. First, explicitly define the additional statistics $X_i$:
where $D_i$ is the strategy “family” (e.g. equal-weighted accounting) and $\SE_i$ is the standard error of $r_i$. Actual performance follows
where $g_{D_{i}, \SE_{i}}\left(\cdot\right)$ is a distribution that depends on $D_i$ and $\SE_i$. Measured performance follows
where $f_{\mu_{i},\SE_{i}}\left(\cdot\right)$ is a distribution that depends on $\mu_{i}$ and $\SE_{i}$. This is a hierarchical structure, where the strategy family determines the actual performance, which in turn determines the measured performance.
Second, define data mining. Naive data mining chooses a hurdle $h$ and then selects strategies
In contrast, EB data mining uses the bias-adjusted measure to select strategies
where $h^\prime$ is chosen to select the same number of strategies as in naive data mining.
In general, Equations ((ref)) and ((ref)) imply different sets of strategies. However, under some natural conditions, the selections are identical:
Conditions (ref) and (ref) arise naturally when using long samples (e.g. 300 months of returns), standardized performance measures (e.g. t-statistics), and strict statistical hurdles (e.g. 5% critical levels). Under these conditions, actual performance is a strictly increasing function of only the measured performance, as proved in Appendix (ref). As a result, data-mined performance provides a reliable signal of actual performance, even if the magnitudes are distorted. The proposition assumes some exact conditions, and leading to identical selections, but approximate conditions would likely lead to similar selections.
One interpretation of Proposition (ref) is that Fisher's fisher1925statistical focus on t-statistics set future researchers up for success, even in the modern era of big data.
On the other hand, Fisher would likely have been unsatisfied with finding the best strategies. He most likely would implore us to find unbiased estimates for these best performers. Thus to rigorously mine data, one should still apply empirical Bayes.
We select as our performance measure the t-statistic on the raw long-short return, and assume that standard errors are precisely measured, implying
The latent performance is a mixture of two normals that depends on the strategy family $D_i$.
where $d$ is one of the six strategy families that comes from combining three data sources (accounting, past returns, tickers) with two portfolio formation methods (equal-weighted and value-weighted). Mixture normals are parsimonious, easy to understand, and yet allow for skewness and fat tails.
We then estimate $\Omega\equiv\left[\theta_{d,1}, \sigma_{d,1}^2, \theta_{d,2}, \sigma_{d,2}^2, \lambda_d\right]_{d=1,\ldots,6}$ using quasi-maximum likelihood. The quasi-likelihood is computed using the distr package (ruckdeschel2006s4). Optimization of $\Omega$ uses nloptr (NLopt).This estimation is done using the past 20 years of long-short returns, separately for each “forecasting year” spanning 1983-2019.
Finally, we recover EB predictions by computing Equation ((ref)) with distr, which produces an EB prediction of the expected return in units of standard errors. The EB predicted return is just Equation ((ref)) multiplied by the standard error. Similarly, the EB predicted Sharpe ratio is Equation ((ref)) multiplied by the square root of the number of periods in the sample.
For further details see Appendix (ref) or our github site.
We show that data mining leads to research-like out-of-sample returns (Section (ref)) and take a look at which kinds of strategies are identified by data mining (Section (ref)). We also provide intuition for why data mining produces such high returns (Section (ref)).
Can data mining generate out-of-sample returns? To answer this question we construct simple out-of-sample portfolio tests.
Each year, we sign strategies to have positive predicted returns, and then form portfolios that equally-weight strategies in the top $X\%$ of predicted Sharpe ratios. We use both EB predictions and standard naive predictions. We examine $X=$ 1, 5, and 10. For comparison, we also examine a portfolio that equally-weighs published strategies from the ChenZimmermann2021 dataset.
Table (ref) shows the result. Using empirical Bayes (EB Mining), the top 1% of strategies perform similarly to strategies published in top finance journals. Over the full 1983-2020 sample, the top 1% portfolio earns 5.70% per year, compared to the 5.88% return from published strategies. The Sharpe ratio from EB mining is smaller, at 1.46 vs 2.03 for published strategies. However, unlike the EB-mined strategies, which are formed using only information available in real-time, the published strategies contain look-ahead bias. Indeed, if we focus on strategies in top journals that were published pre-2004, the performance is very similar to the EB-mined strategies in terms of either mean returns or Sharpe ratios.
The EB-mined returns are robust. The top 5% and top 10% of data-mined strategies also perform well and are extremely statistically significant, indicating that the performance of the top 1% is not driven by outliers.
Panel B shows that even naive data mining produces research-like returns. Simply choosing the top 1% of strategies based on their past Sharpe ratios leads to an out-of-sample Sharpe ratio of 1.45. This is almost exactly the same as the Sharpe ratio from the top 1% using EB mining, consistent with Proposition (ref). The top 5% and top 10% of naive strategies underperform a bit relative to EB mining, but the intuition behind Proposition (ref) still goes through.
Figure (ref) takes a closer look by plotting the value of \$1 invested in each portfolio over time. The top 1% data-mined strategies have similar performance to published strategies throughout the figure. All portfolios show relatively little cyclicality during the recessions of 1991, 2009, and 2020. Indeed, the returns are fairly consistent throughout the chart, with an important caveat.
The caveat is that returns are concentrated in the pre-2004 sample. This is seen in the flattening of the solid line in Figure (ref) around 2004. The EB Mining top 1% portfolio returns 8.17% per year from 1983-2005, compared to just 2.03% from 2005-2020. A similar decay is seen across all portfolios, both data-mined and academic. This decay is consistent with Chordia2014Have and chen2022zeroing, who argue that the rise of information technology reduced return predictability.
Overall, we find that one can find long-short returns comparable to those from the best journals in finance, just by mining data, with little thought about the underlying economics. Moreover, rigorous data mining can discriminate between data sources that have no information about future returns, like stock market tickers, from data that is rich in information, like accounting ratios. Unlike the published strategy returns, our returns can be found using only information available in real-time. These results show that high-throughput methods provide a bias-free approach to studying stock market predictability. Our strategy returns and code are public, and we encourage future researchers to use these methods.
Table (ref) takes a closer look at the top 1% strategies produced by rigorous data mining. Panel A shows that 91.0% of the top 1% come from the equal-weighted accounting family and 8.6% come from equal-weighted past returns. The other strategy families comprise a negligible part of the top 1%. Ticker strategies are completely absent.
Taken with Table (ref), these results show that cross-sectional predictability is concentrated in accounting data, small stocks, and pre-2004 samples. These stylized facts offer a parsimonious description of the “factor zoo.” Theories that wish to capture the big picture of cross-sectional predictability should be consistent with these facts. For example, slow diffusion of economic information is consistent, as this diffusion would be especially slow in small stocks and before the internet era. In this way, high throughput asset pricing provides a way to not only identify out-of-sample returns, but to also provide insight into the underlying economics.
Panel B of Table (ref) shows that many of the top 1% strategies are quite far from the predictors noted in the academic literature. In 1993, academics were focused on predictors like book-to-market, 12-month momentum, and sales growth (fama1992cross; jegadeesh1993returns; lakonishok1994contrarian). None of these predictors are in the top 20 strategies based on predicted Sharpe ratios from rigorous data mining. Instead, the common themes from data mining include shorting stocks with high or growing debt, as well as buying stocks with high depreciation, depletion, and amortization. Another theme is buying stocks with high returns in quarters $t$ minus 17 and 18.
Based on textbook risk-based or behavioral asset pricing, one might expect that these data-mined predictors will average zero returns out-of-sample. But this is not the case. The realized Sharpe ratios for these strategies in the 10 years after 1993 averages around 1.0 (“SR OOS” column).
Panel B of Table (ref) focuses on 1993 because well-known predictability papers were published around that time (e.g. fama1993common). In other years, the top 20 list is different, though shorting variables related to debt growth remains a common theme. For further details see Appendix Tables (ref) and (ref).
Unlike many big data and machine learning methods, empirical Bayes has a transparent intuition. The intuition can be seen in a special case of the prediction Equation ((ref)). If $\mu_{i}\mid \left(X_i, D_i = d\right) \sim\text{Normal}\left(0,\sigma_{d}^{2}\right)$, we have
where $\widehat{\Var}\left(r_{i} | D_i = d\right)$ is an estimate of the cross-strategy variance of performance measures among strategies with data family $d$.
This expression says that rigorous mining involves shrinking performance measures $r_{i}$ toward zero at a rate of $\frac{1}{\widehat{\Var}\left(r_{i} | D_i = d\right)}$. $\widehat{\Var}\left(r_{i} | D_i = d\right)$ measures how far the data are from the null of $r_{i} \sim\text{Normal}\left(0,1\right)$, which we imposed in Equation ((ref)). If there is no predictability, then $r_{i}\sim\text{Normal}\left(0,1\right)$, $\widehat{\Var}\left(r_{i} | D_i = d\right)\approx1$, and all $r_{i}$ are shrunk to zero. But if data are far from the null, then a large $r_{i}$ is a signal of large $\mu_{i}$---even if $r_{i}$ is found from searching tens of thousands of strategies, unguided by economic theory.
Figure (ref) shows that equal-weighted accounting strategies (upper left) are far from the null using data from 1964 to 1983. Equal-weighted past return strategies (middle left) also show a notable deviation. In contrast, the other strategy families are quite close to the null. Indeed, for both families of ticker-based strategies, the null is a very good fit for the data.
Accordingly, Equation ((ref)) implies that the strategies with strong actual performance will be found in equal-weighted accounting and equal-weighted past-return strategies. This intuition is consistent with Panel A of Table (ref), which shows that the vast majority of the best data-mined strategies come from these families.
Compared to data available in 1983, all strategy families are closer to the null using data from 1985-2004, as seen in Figure (ref). All value-weighted families are very close to the null, implying that predictability in large stocks is essentially gone. The long left tail in equal-weighted past return strategies also disappears. Only equal-weighted accounting strategies are visually far from the null. These results imply that predictability is concentrated in the earlier part of the sample.
The intuition in Figures (ref) and (ref) is so simple that one might even skip the quasi-maximum likelihood estimation. Just looking at these charts, and the distance between the data and the null, one can already tell that predictability is concentrated in small stocks, accounting data, and the earlier sample. That is, one can already tell where predictability is concentrated, if one understands the intuition in Equation ((ref)).
This section takes a closer look at the EB predictions and accuracy. We see when and where EB predictions are successful and when they struggle.
To examine accuracy, we use out-of-sample portfolio sorts. For each year and each strategy family, we form 20 portfolios by sorting strategies into equal-sized groups based on the past 20 years of mean returns. We then predict the mean returns for each portfolio by averaging the EB predictions (Equation ((ref))), which are also based on the past 20 years of data. Finally, we form a portfolio that equally-weighs strategies in each group and hold for one year (the “out-of-sample” periods).
Figure (ref) shows the in-sample, predicted, and out-of-sample returns for each portfolio, averaged over the out-of-sample periods from 1983 to 2004. For all six families, there are sizable in-sample returns (dashed line) in the extreme in-sample groups. For accounting strategies, in-sample returns are as extreme as -11% per year. A naive read of this result is that one can flip the long and short legs and find +11% returns out-of-sample. Past return strategies see a similar $\pm10$ percent return in the extreme groups. Even ticker-based strategies show in-sample long-short returns of up to 4 percent per year.
However, the predicted returns are typically much closer to zero. In fact, for both ticker-based strategy families, the predicted return (solid line) is almost exactly zero for all 40 in-sample groups. This result is intuitive given how close the ticker t-stats are to the null of no predictability (Figure (ref)). This closeness implies that the extreme returns can be entirely accounted for by luck, and so shrinkage should be 100% (Equation ((ref))). Significant shrinkage is also seen in value-weighted accounting strategies (top right panel). Rigorous data mining recommends that the extreme returns of around -8% and +9% (dashed line) be shrunk down to about -3% and +2 (solid line), respectively.
Rigorous mining predicts much higher returns in equal-weighted accounting strategies (upper left panel). For these strategies, the predicted returns are actually not far from the in-sample return. This result is consistent with chen2020publication, who find shrinkage of only 12% for published anomalies, which are largely equal-weighted and based on accounting variables. Predictability is also seen in both families of past return strategies.
These predictions are borne out in out-of-sample returns (markers with error bars). The first group of EW accounting strategies returns -8 percent per year out-of-sample from 1983-2004, almost exactly the same as the EB prediction. Similar accuracy is seen throughout all 120 bins in Figure (ref).
These results show that rigorous data mining offers economic insights that are difficult to derive from theory. While theories of slow information diffusion may tell you that predictability is concentrated in small stocks, accounting signals, and pre-2004 data, they are unlikely tell you how much predictability there is. In contrast, empirical Bayes provides quantitative, accurate estimates of the precise amount of predictability.
We split our OOS tests in the mid-2000s, motivated by the idea that there was likely a structural break during this period due to the rise of information technology (Chordia2014Have). Comparing the distribution of t-stats available in 1983 vs 2004 supports the idea that the structure of financial markets changed (see Section (ref)).
This structural change can be seen by comparing Figure (ref) (EB predictions 2004-2020) to Figure (ref) (EB predictions 1983-2004). In all panels, the predicted returns shift closer to zero post-2004. Most notably, the predictability that was present in past return strategies pre-2004 is largely gone. Consistent with these predictions, the past return portfolios show a flat or even negative relationship between out-of-sample and in-sample returns post-2004. A similar weakening of EB predictions and flattening of out-of-sample returns is seen in the accounting VW family.
An exception to this pattern is the family of equal-weighted accounting ratio strategies (top left). In this chart, the shrinkage is still relatively small, with EB predictions implying returns as extreme as -9 percent per year. This prediction and others in this panel miss the mark: the out-of-sample returns are much closer to zero throughout this panel.
This poor accuracy is natural given the fact that the estimations use a rolling window consisting of the past 20 years of data. This fixed window implies that, for much of the period 2004-2020, our estimates rely on data from a time when accounting statements needed to be retrieved by traditional (snail) mail for investors without special access to the SEC reading room (bowles2023anomaly).
This result implies an important role for economic theory: when structural breaks occur, there is no way for data mining to provide a clear understanding of the economy, no matter how rigorously the mining is done. Theory is sometimes used this way in economics and finance, but this is typically not the case. Instead, theory is typically used to understand patterns found in long samples of data, spanning many decades. In our view, the future of theory is bright for theorists who study structural breaks, even in the era of big data. Indeed, a smart data miner armed with theory might have understood the implications of the internet for stock return predictability, and could perhaps have performed much better than our theory-free EB mining process.
Our main analysis corrects for data mining bias using empirical Bayes shrinkage, following chen2020publication; chen2022zeroing; and jensen2023there. An alternative approach is to use false discovery controls, following harvey2016and; barras2010false; or chordia2020anomalies. The ideal approach remains an unsettled question. Our dataset of 136,000 trading strategies provides a natural testing ground.
We examine the following false discovery controls:
For each year and each strategy family, we apply these methods using the past 20 years of data to estimate a t-statistic hurdle. We then examine whether these hurdles are able to separate strategies with high out-of-sample returns from those with low out-of-sample returns. This structure is the same as in Section (ref).
Figure (ref) shows the results. The vertical lines show the mean hurdle across all years. The markers show the mean out-of-sample returns of portfolios formed by equally weighting strategies, sorted into 20 groups based on the in-sample t-statistic. Groups of strategies that a false discovery control declares “significant” lie on the outside of the respective vertical lines.
The BY1.3 (1%) and RW (5%, 5%) methods miss out on the majority of portfolios with notable out-of-sample performance. Out of the 5 groups that have out-of-sample returns of at least 3% per year, only 1 lies outside of the solid lines corresponding to BY1.3 (1%). Only 3 of 5 lie outside the dot-dashed lines corresponding to RW (5%, 5%). The dashed line, corresponding to Storey (10%), performs much better, capturing 4 of the 5 portfolios. Similar results are found using alternative parameter choices examined by HLZ; harvey2020false; and barras2010false (see Appendix Figure (ref)).
Thus, Storey (10%) provides an easy-to-compute alternative to empirical Bayes. However, Storey cannot provide bias-adjusted performance estimates that are naturally available from empirical Bayes. Overall, our results imply that Storey forms a strong first step for rigorous data mining, while empirical Bayes is recommended for more refined estimates. These results are broadly consistent with the statistics literature, which generally recommends Storey as a preliminary examination, while suggesting empirical Bayes for greater precision (e.g. benjamini2010discovering; efron2012large).
We discuss this literature and the algorithms in more detail below.
harvey2016and (HLZ) recommend using \citepos{benjamini2001control} Theorem 1.3. Several followups to the influential HLZ paper use this method, including harvey2020false,chordia2020anomalies; and jensen2023there.
We state the theorem number 1.3 because the bulk of the original paper focuses on Theorem 1.2. Indeed, benjamini2001control describe Theorem 1.3 as “very often unneeded, and yields too conservative of a procedure” (page 1183). In his textbook on large scale inference, efron2012large agrees, stating that the theorem represents a “severe penalty” and is “not really necessary” (section 4.2). Moreover, the statistics literature uses the “BY algorithm” to refer to benjamini2005false, which is an entirely different procedure (e.g. efron2012large Chapter 11.4).
BY1.3 begins by choosing a parameter $q^{\ast}$ and then solving
where $t_i$ is the t-statistic for strategy $i$, $Z$ is a standard normal random variable,
and $N$ is the number of strategies in the year-family. \citepos{benjamini2001control} Theorem 1.3 proves that this algorithm implies a false discovery rate $\le q^{\ast}$. BY1.3 amounts to modifying the seminal benjamini1995controlling algorithm with a constant factor, $\pi_{\text{BY1.3}}$. This modification makes the algorithm more conservative.
HLZ recommend this conservative approach, claiming benjamini1995controlling “is only valid when the test statistics are independent or positively dependent” (page 21). This statement is false. storey2001estimating and storey2004strong show validity under weak dependence assumptions (see also chen2024most).
HLZ are also conservative in their choice of $q^{\ast}$. For their main results, they use $q^{\ast}=1\%$ citing the fact that the “significance level is subjective,” though they also examine $q^{\ast}=5\%$ for robustness. In contrast, the statistics literature generally recommends $q^{\ast}=5\%$ or 10% (e.g. benjamini2010discovering; efron2012large).
Given this context, it is perhaps unsurprising that BY1.3 (1%) fails to identify most out-of-sample performers in Figure (ref). BY1.3 (5%) performs somewhat better, identifying 2 out of 5 groups with out-of-sample returns of at least 3% per year (Appendix Figure (ref)).
While HLZ recommend modifying benjamini1995controlling to be more conservative, much of the statistics literature goes in the opposite direction, modifying benjamini1995controlling to be more aggressive. In finance, barras2010false take this approach.
Barras et al. recommend the storey2002direct algorithm, which can be written as
where $t_i$ is the t-statistic for strategy $i$, $Z$ is a standard normal random variable,
and the cutoff of $1.0$ is selected for ease of interpretation. storey2002direct proves that this algorithm implies a false discovery rate $\le q^{\ast}$ under independence assumptions, though storey2001estimating and storey2004strong extend this result to weak dependence.
Comparing Equations ((ref))-((ref)) to the Equations ((ref))-((ref)), we see that the only difference is the constant factor, $\pi_{\text{BY1.3}}$ vs $\pi_{\text{Storey}}$. These constants are qualtitatively different: $\pi_{\text{BY1.3}} = \sum_{i=1}^{N}\frac{1}{i} \approx 0.6 + \log N \gg 1$ , while $\pi_{\text{Storey}}\le 1.0$. As shown in storey2002direct, $\pi_{\text{Storey}}$ can be interpreted as an estimate of the probability that a strategy is null, which can be at most 1.0. Other statistics papers that recommend a constant that is at most 1.0 include benjamini2000adaptive, efron2001empirical,genovese2006false,benjamini2006adaptive.
barras2010false do not emphasize a particular choice of $q^{\ast}$, and instead examine values ranging from 5% to 20%. Figure (ref) uses $q^{\ast}=10\%$, because Barras et al. use 10% in their illustrative examples.
Once again, given the support from the statistics literature, it is perhaps unsurprising that Storey (10%) and (20%) perform well. Equations ((ref))-((ref)) are easy to implement, making it a useful alternative to our empirical Bayes method.
However, there are two downsides to using Storey. A simple, symmetric testing algorithm like Equations ((ref))-((ref)) does not handle skewed distributions well. This limitation may explain why Storey struggles to identify out-of-sample performers in past return strategies, which feature a long right tail (Figure (ref)). The second is that Storey cannot provide bias-adjusted performance estimates. Such estimates are naturally available from a more general empirical Bayes method, and would provide clean connections with portfolio choice and asset pricing questions.
chordia2020anomalies recommend combining \citepos{romano2007control} Algorithms 4.1 and 2.1, which we refer to as “RW.” This algorithm is a natural choice for asset pricing researchers, as its predecessor romano2005stepwise is motivated by data mining for CAPM anomalies. Like HLZ's method, the Romano and Wolf methods have been used in influential asset pricing papers, including chordia2020anomalies; engelberg2023do; heath2023reusing; and debodt2025competition.
Unlike Storey and BY1.3, the statistics literature has relatively little discussion of the Romano and Wolf methods. Neither Romano and Wolf (romano2005stepwise) nor Romano and Wolf (romano2007control) is found in the textbooks efron2012large and efron2016computer. The two papers are also not found in the review articles on multiple testing benjamini2010discovering and benjamini2020selective. Thus, we provide some discussion here.
The goal of RW can be written as follows: find an $h$ that ensures
where
and $p^\ast$ and $q^\ast$ are thresholds selected by the researcher. Null strategies are, typically, those with an actual performance of zero.
Figure (ref) illustrates Equation ((ref)), by simulating one of our QML estimates many times. We run 2,000 simulations, each one consisting of 29,000 strategies. For simplicity, we assume all strategiesare independent. The plot shows histograms of actual performance ($\mu_i$ in Equation (ref)) for strategies that meet the hurdle $|t_i|>3.0$, where $h=3.0$ is selected for illustrative purposes. Using this chart, we can ask whether this $h=3.0$ hurdle achieves Equation (ref), and thus understand FDP risk control.
Since the $h=3.0$ hurdle is quite stringent, the vast majority of strategies are non-null. However, there is still a risk that a strategy with $|t_i|>3.0$ has near-zero actual performance, as seen in the left tail of the histogram. The FDP characterizes this risk. It is, approximately, the share of strategies in the first bin.\footnote{More formally, one can consider the first bin to be an upper bound on the FDP (see chen2024most).} On average, the share of strategies in this bin is about 5% (bars), indicating that the FDR is approximately controlled at a 5% level.
Even though the FDP is on average about 5%, there is a risk that it is higher. This risk is seen in the lines of Figure (ref), which plot extreme order statistics across the 2,000 simulations. The 95th percentile line implies the FDP exceeds 7% in 5% of simulations. To achieve FDP risk control with $p^\ast=5\%$ and $q^\ast=5\%$, a $h>3.0$ is required. The RW method finds this $h$.
Thus, the RW method aims to control the tail risk of a tail risk. Such an algorithm is a natural choice if selecting a null strategy is catastrophic. In such a case, one may want to ensure not only that a null is highly improbable, but that the probability that a null is somewhat probable is also improbable. However, in the standard setting where the null is that the strategy has zero long-short return or zero alpha, then the RW method tends to imply extreme conservatism.
This conservatism leads to the results in Figure (ref) and Appendix Figure (ref). Choosing $p^\ast=0.05$ and $q^\ast=0.05$ or 0.10, as in chordia2020anomalies and harvey2020evaluation, leads to hurdles that many notable out-of-sample performers fail to clear.
The RW method is rather complex. It uses cluster bootstrap methods, involves testing all possible subsets of selected sets of strategies, iterating over many possible tests and sets. We describe our implementation in Appendix (ref) and provide code in our Github repo.
We show that a solution to data mining bias is to mine data rigorously. We systematically search 136,000 long-short strategies and find out-of-sample performance comparable to academic research. Simply searching for strategies with the largest t-stats leads to publication-like out-of-sample performance, a fact we explain in a Bayesian model. While naive data mining leads to distorted performance estimates, empirical Bayes provides unbiased predictions in samples without structural breaks. The forecast errors around structural breaks suggest a role for theory in the era of big data.
This high-throughput method shows that returns are concentrated in accounting signals, small stocks, and pre-2004 periods, consistent with mispricing and slow information diffusion theories. While these results could potentially be gleaned from a deep read of the anomalies literature, our method provides a scientific method for documenting these stylized facts. We provide our data and code publicly, and hope others follow in using high-throughput methods.
Our out-of-sample tests offer an intuitive method for comparing multiple testing methods. We find that methods popular in finance would lead researchers to miss out on the majority of signals with notable out-of-sample performance. In contrast, methods recommended by the statistics literature perform well.