The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
54,864 characters
High-Throughput Asset Pricing
\title{\textbf{High-Throughput Asset Pricing}}
\author{{Andrew Y. Chen}\\
{\normalsize Federal Reserve Board}\and {Chukwuma Dim}\\
{\normalsize George Washington University}}
\date{June 2025\thanks{Alternative title: ``How I learned to stop worrying and love data
mining.'' First posted to arXiv.org: November 2023. Replication code:
\protect\protect\url{https://github.com/chenandrewy/high-throughput-ap}.
Long-short strategy data: \protect\protect\url{https://sites.google.com/site/chenandrewy/}. We thank our discussants: Rudiger Weber, Julio Crego, and Morad Zekhnini for helpful comments. We also thank participants at the TBEAR Network Workshop (KIT), FutFinInfo, MFA, Iowa State, and Baylor. The views expressed herein are those of the authors and do not necessarily
reflect the position of the Board of Governors of the Federal Reserve
or the Federal Reserve System.}}
\maketitle
\begin{abstract}
\begin{singlespace}
\noindent We apply empirical Bayes (EB) to mine data on 136,000 long-short strategies constructed from accounting ratios, past returns, and ticker symbols. This ``high-throughput asset pricing'' matches the out-of-sample performance of top journals while eliminating look-ahead bias. Naively mining for the largest Sharpe ratios leads to similar performance, consistent with our theoretical results, though EB uniquely provides unbiased predictions with transparent intuition. Predictability is concentrated in accounting strategies, small stocks, and pre-2004 periods, consistent with limited attention theories. Multiple testing methods popular in finance fail to identify most out-of-sample performers. High-throughput methods provide a rigorous, unbiased framework for understanding asset prices.
\end{singlespace}
\end{abstract}
\vspace{10ex}
\textbf{JEL Classification}: G0, G1, C1
\noindent\textbf{Keywords}: stock market predictability,
stock market anomalies, p-hacking, multiple testing \thispagestyle{empty}\setcounter{page}{0}
\vspace{10ex}
\pagebreak
\section{Introduction}
\setcounter{page}{1}
Data mining refers to searching data for interesting patterns. This
search leads to data mining bias, if many patterns are just chance
results, as is surely the case with stock return data. To address
this problem, the asset pricing literature recommends restricting
the search to patterns consistent with theory (\citet*{cochrane2005risk,harvey2017presidential}).
However, recent empirical evidence finds this method is ineffective,
even for theories published in top finance journals (\citet*{chen2022peer}).
We offer a different solution. Instead of mining data less, we recommend
mining data \emph{rigorously}. Rigorous data mining means conditioning
interesting results on the fact that they come from searching through
data. This conditioning can be achieved using empirical Bayes (\citet{robbins1956empirical,efron1973stein,efron2012large}).
Rigorous data mining also means that the search should be systematic,
as is commonly done in high-throughput biology and chemistry (\citet{yang2021high}).
Ironically, systematic search implies that asset pricing should involve
\emph{more} data mining, not less.
We use empirical Bayes (EB) to mine for out-of-sample returns among
136,000 long-short trading strategies. The trading strategies are
constructed from systematically searching data on accounting ratios,
past returns, and stock tickers. Through this ``high-throughput asset
pricing,'' we construct a portfolio with out-of-sample returns that
are comparable to the returns from the best journals in finance.
Our data-mined portfolio is the simple average of the top 1\% of strategies,
based on EB-predicted Sharpe ratios. It earns out-of-sample returns
of 5.7\% per year over the 1983-2020 sample, compared to the mean
return of 5.9\% per year found by averaging the 200 published strategies from \citet{ChenZimmermann2021}. But unlike the published strategies, which
were selected with knowledge of stock return patterns that occurred
in the 1980s and 1990s, our strategies can be constructed using only
information available in real time.
In fact, even naively mining for the largest Sharpe ratios leads to publication-like performance. We provide a theoretical explanation for this phenomenon in Proposition \ref{prop:optimal-naive}, which shows that under standard statistical practices (\citealt{fisher1925statistical}), naive data mining often selects the same set of strategies as an ideal Bayesian. However, while the naively-selected strategies may be optimal, naive performance estimates are distorted, illustrating the importance of rigorously data mining with EB.
The top 1\% portfolio selected by EB provides insights into the
nature of return predictability. 91.0\% of strategies in this portfolio
are equal-weighted accounting ratio strategies. Almost all of the
remainder are equal-weighted past-return strategies. Moreover, the
returns of the top 1\% strategy are concentrated in the pre-2004 data.
These facts are consistent with the theory that predictability is
largely due to limited attention and the slow incorporation of information
into stock prices (\citet{peng2005learning}; \citet{Chordia2014Have}).
Other facts shed light on the drivers of the recent decline in cross-sectional
predictability. We find that the returns of the top 5\% and top 10\%
of portfolios are also concentrated in the pre-2004 data. These strategies
are enormous in number: the top 5\% consists of 6,305 strategies,
and the top 10\% consists of 12,610. As many of these strategies are
unlikely to be found in academic journals, this suggests that the
key driver of the recent declines in predictability is improvements
in information technology (\citet{Chordia2014Have}), rather than
investors learning from academic publications (\citet{Mclean2016Does}).
Consistent with this idea, we find that the top 20 strategies according
to predicted Sharpe ratios using data available in 1993 have themes
rarely seen in academic journals, like mortgage debt, growth in interest
expense, and depreciation. Themes that were popular in academia in
1993, like book-to-market, momentum, and sales growth are missing
from this list.
Overall, high-throughput asset pricing provides not only a method
for dealing with look-ahead bias, but also a more rigorous method
for documenting asset pricing facts. We post our strategy returns
and code publicly, and encourage future researchers to use these methods.
Unlike many big data methods, EB provides a transparent
intuition. In essence, EB measures the distance between
the empirical t-stat distribution and the standard normal null. Ticker-based
strategies have t-stats that are extremely close to the null, implying
no predictability. In contrast, equal-weighted accounting t-stats
are too fat tailed to be consistent with the null, implying strong
predictability. Thus, just by visually inspecting the t-stat distributions,
one can see where predictability is concentrated.
EB provides highly accurate predictions in pre-2004 data. We construct
120 portfolio tests using the 136,000 data-mined strategies, and compare
EB-predicted returns with out-of-sample returns. In almost all of
the 120 portfolios, the EB predictions are within 2 standard errors
of the out-of-sample mean.
Post-2004, EB has more difficulty with accuracy, though it still captures
broad patterns in out-of-sample returns. Compared to pre-2004, predicted
returns are closer to zero, and only equal-weighted accounting strategies
show notable predicted returns. However, out-of-sample returns are
even closer to zero than predicted. This difficulty might be expected
given the rise of information technology around 2004, which likely
led to a structural break in predictability (\citealt{Chordia2014Have}; \citealt{kim2021causal}).
Our EB predictions are constructed using a simple 20-year rolling
window, and thus fail to account for this break. This difficulty suggests
that a smart data miner armed with theory might have understood the
implications of the internet, and could perhaps have performed much
better than our theory-free EB mining process.
We also illustrate how improper use of multiple testing statistics can lead to poor data mining results. We demonstrate this possibility using \citepos{harvey2016and} recommended method for false discovery control. \citet{harvey2016and} recommend applying \citepos{benjamini2001control} Theorem 1.3 to construct a t-stat hurdle that controls the false discovery rate (FDR) at the 1\% level. Nearly all of our 136,000 trading strategies fail to meet this hurdle, suggesting that there are few interesting patterns in this data. But in fact, simple out-of-sample tests show there are thousands of strategies with notable out-of-sample returns. We find similar results following the recommended multiple testing control in \citet{chordia2020anomalies}, which is based on \citet{romano2007control}. In contrast, the \citet{storey2002direct} FDR control recommended in \citet{barras2010false}, captures the majority of notable portfolios.
Fortunately, this error can be avoided by rigorously studying the statistics. According to \citet{benjamini2001control}, their Theorem 1.3 is ``very often unneeded, and yields too conservative of a procedure.'' This negative sentiment is echoed in Efron's \citeyearpar{efron2012large} textbook on large scale inference. In contrast, the EB methods we use are recommended for settings like ours in Chapter 1 of \citet{efron2012large}, as well as Chapters 6 and 7 of \citet{efron2016computer}.\footnote{A brief explanation of why \citet{benjamini2001control} Theorem 1.3 is excessively conservative is found in Section 2.5 of \citet{chen2024t}.} The statistics literature has relatively little to say about the method recommended in \citet{chordia2020anomalies}. We provide our own characterization, which illustrates how this method is appropriate if selecting a null strategy is catastrophic. But using the standard null, that the mean long-short return (or alpha) is zero, Chordia et al.'s method implies unneeded conservatism.
\subsection{Related Literature}
We add to \citet{yan2017fundamental} and \citet{chen2022peer}, who
document that mining accounting data can produce substantial out-of-sample
returns. Accounting data is important: Chen et al. find that mining
ticker variables leads to out of sample returns of approximately zero.
Thus, one needs a method for identifying the predictive power of accounting
data in real time. Our empirical Bayes formulas provide one such method.
The literature on multiple testing in asset pricing features disagreement
on both the methods that should be used and the empirical extent of
multiple testing problems. \citet{chen2020publication}; \citet{chen2022zeroing};
and \citet{jensen2023there} recommend empirical Bayes shrinkage.
In contrast, \citet{harvey2016and}; \citet{harvey2020false}; and
\citet{chordia2020anomalies} recommend conservative false discovery
controls, much more conservative than the FDR methods in \citet{barras2010false}.
We show how empirical Bayes shrinkage and the recommended method from
\citet{barras2010false} leads to much more accurate inferences. More recently, \citet{marrow2024real} use empirical Bayes to study past return signals, with a focus on signal interactions and optimal weighting of more recent data.
In contrast to the intuition that simplicity is a virtue, we find
that studying an enormous number of potential predictors leads to
insights about the nature of return predictability. A similar theme
is found in \citet{kelly2024virtue} and \citet{didisheim2023complexity},
who illustrate the ``virtue of complexity'' in the modeling of expected
returns.
\section{Data and Methods\protect}\label{sec:methods}
We describe the data (Section \ref{sec:data}) and how we rigorously
mine it (Sections \ref{sec:stats:eb-overview}-\ref{sec:stats:eb-details}).
\subsection{Data on 136,000 Trading Strategies\protect}\label{sec:data}
Table \ref{tab:data-descrip} describes our data-mined strategies.
The strategies are either based on accounting ratios, past returns,
or tickers. Accounting ratio strategies are taken from \citet{chen2022peer}.\footnote{We are grateful that the authors make their data publicly available.}
The past return and ticker strategies are inspired by \citet{yan2017fundamental}
and \citet{harvey2017presidential}, respectively, but we generate
our own strategies in order to ensure that the number of strategies
is comparable across data sources and to ensure that each type of
strategy consists of many distinct strategies.\footnote{Results that mine data following \citet{yan2017fundamental} and \citet{harvey2017presidential}
are similar and can be found in the first draft of our paper on arxiv.org
or via our github site.}
\begin{center}
{[}Table \ref{tab:data-descrip}, Overview of Trading Strategies,
about here{]}
\par\end{center}
A key feature of these strategies is that they are \emph{not }selected
based on having notable historical returns. Instead, they are constructed
to systematically explore various types of data. So unlike most datasets
in asset pricing (e.g. Ken French's size- and B/M-sorted portfolios;
\citet{ChenZimmermann2021}), ours is arguably free of data mining
bias. Indeed, Table \ref{tab:data-descrip} shows that the median
sample mean return is close to zero for all sets of strategies.
In high-throughput research, the median measurement is relatively
unimportant. What matters is that the extreme measurements show promise
for, say, a pharmaceutical intervention or cancer prediction. The
extreme measurements in Table \ref{tab:data-descrip} suggest that
accounting and past return data show promise for predicting returns.
These data lead to mean returns that can exceed 5 percent per year
in absolute value.
For further details on the strategy definitions, see Appendix \ref{sec:app:data}
or our github site.
\subsection{Empirical Bayes Overview}\label{sec:stats:eb-overview}
The 136,000 strategies in Table \ref{tab:data-descrip} contain the
potential for significant data mining bias. To understand the bias, let $r_{i}$ be a performance measure for strategy $i$ (e.g. mean return, alpha) and decompose it as follows:
\begin{align}\label{eq:r=mu+e}
r_{i} & =\mu_{i}+\varepsilon_{i}
\end{align}
where $\mu_{i}$ is the actual performance and $\varepsilon_{i}$ is sampling error or luck.
Data mining involves selecting $i$ with large $r_{i}$. Suppose we set $\bar{r} \gg 0$, and search for $i^\ast \in \{1;2;\ldots;136,000\}$ such that $r_{i^\ast} = \bar{r}$. This practice is dangerous because one might think $\bar{r}$ is a good estimate of $\mu_{i^\ast}$. However, $\bar{r}$ is in fact biased upward
\begin{align} \label{eq:bias-demo}
\bar{r} &= E\left(r_{i^\ast} | r_{i^\ast} = \bar{r}\right) \\
&= E\left(\mu_{i^\ast} | r_{i^\ast} = \bar{r}\right)
+ \underbrace{
E\left(\varepsilon_{i^\ast} | r_{i^\ast} = \bar{r}\right)
}_{>0}
> \mu_{i^\ast}. \notag
\end{align}
Selecting for large $r_{i}$ also selects for large $\varepsilon_{i}$, leading to $E\left(\varepsilon_{i^\ast} | r_{i^\ast} = \bar{r}\right) > 0$ and the bias in Equation (\ref{eq:bias-demo}).
To data mine safely, one needs to remove the luck term $E\left(\varepsilon_{i}|r_{i^\ast} = \bar{r}\right)$. This term is just a conditional expectation, so it can be computed using Bayes rule, provided one has a probability model for $\mu_{i}$ and $r_{i}$.
Suppose one has a probability model, with parameter vector $\Omega$. The bias can then be removed by computing
\begin{align}\label{eq:luck-sketch}
E\left(\mu_{i^\ast}|r_{i^\ast} = \bar{r};\hat{\Omega}\right)
= \bar{r} - E\left(\varepsilon_{i^\ast}|r_{i^\ast} = \bar{r};\hat{\Omega}\right)
\end{align}
where $\hat{\Omega}$ is a consistent (frequentist) estimate of the
probability model parameters. This method, of applying frequentist estimates to Bayesian formulas is known as ``empirical Bayes'' (\citet{robbins1956empirical,efron1973stein}).
Equation (\ref{eq:luck-sketch}) conditions on only one statistic regarding strategy $i$. A more optimal estimate uses more information
\begin{align}\label{eq:ephat-def}
E\left(
\mu_{i^\ast}|r_{i^\ast} = \bar{r}, X_{i} = \bar{X};\hat{\Omega}
\right)
=
\bar{r}
- E\left(
\varepsilon_{i^\ast}|r_{i^\ast} = \bar{r}, X_{i} = \bar{X};\hat{\Omega}
\right)
\end{align}
where $X_{i}$ is a vector of additional statistics for strategy $i$ and $\bar{X}$ is a realized value of $X_{i}$. For example, $X_{i}$ can include the standard error of $r_{i}$, the portfolio weighting (equal- or value-weighted), and the signal data source (accounting, past returns, tickers).
We use Equation (\ref{eq:ephat-def}) to search our 136,000 strategies for large expected returns. We will not use economic theory
to determine the probability model, and thus our search is largely
atheoretical. However, we recognize the bias that comes from such
a search (Equation (\ref{eq:bias-demo})), and carefully correct for
it. Thus, we describe our methods as ``rigorous data mining.''
\subsection{Optimal Naive Data Mining}\label{sec:stats:naive-theory}
In empirical asset pricing, we are often interested in two questions:
\begin{enumerate}
\item What are the best strategies?
\item What is the performance of the best strategies?
\end{enumerate}
If one is interested \emph{only} in the first question, then there is a sense in which naively mining data, without accounting for data mining bias, is often optimal.
To understand this, we add structure to the model. First, explicitly define the additional statistics $X_i$:
\begin{align}
X_i &= \left[ D_i, \SE_i \right]
\end{align}
where $D_i$ is the strategy ``family'' (e.g. equal-weighted accounting) and $\SE_i$ is the standard error of $r_i$. Actual performance follows
\begin{align}\label{eq:mu-X}
\mu_{i}|X_{i} &\sim g_{D_{i}, \SE_{i}}\left(\cdot\right)
\end{align}
where $g_{D_{i}, \SE_{i}}\left(\cdot\right)$ is a distribution that depends on $D_i$ and $\SE_i$. Measured performance follows
\begin{align}
r_{i}|\mu_{i}, X_{i} &\sim f_{\mu_{i},\SE_{i}}\left(\cdot\right)
\end{align}
where $f_{\mu_{i},\SE_{i}}\left(\cdot\right)$ is a distribution that depends on $\mu_{i}$ and $\SE_{i}$. This is a hierarchical structure, where the strategy family determines the actual performance, which in turn determines the measured performance.
Second, define data mining. Naive data mining chooses a hurdle $h$ and then selects strategies
\begin{align}\label{eq:naive-selection}
\{i: r_{i} > h\}
\end{align}
In contrast, EB data mining uses the bias-adjusted measure to select strategies
\begin{align}\label{eq:eb-selection}
\{i: E\left(\mu_{i}\mid r_{i},X_{i}\right) > h^\prime\}
\end{align}
where $h^\prime$ is chosen to select the same number of strategies as in naive data mining.
In general, Equations (\ref{eq:naive-selection}) and (\ref{eq:eb-selection}) imply different sets of strategies. However, under some natural conditions, the selections are identical:
\begin{proposition}\label{prop:optimal-naive}
Consider the following two conditions:
\begin{enumerate}
\item \label{cond:normal} The performance measure satisfies
\begin{align}
r_{i} \mid \mu_{i}, X_{i}
\sim \text{Normal}\left(\mu_{i}, \SE^2\right)
\end{align}
where $\SE$ is a constant.
\item \label{cond:tail} The hurdle $h$ satisfies
\begin{align}
\Pr\left(r_{i}>h|D_{i}\right)&=0
\quad\text{if }D_{i}\in\mathcal{D} \label{eq:tail-cond1} \\
\mu_{i}|X_{i} &\sim g_{\SE_{i}}\left(\cdot\right)
\quad\text{if }D_{i}\in\mathcal{D} \label{eq:tail-cond2}
\end{align}
where $\mathcal{D}$ is a subset of the possible strategy families and $g_{\SE_{i}}\left(\cdot\right)$ is distribution with positive variance that does not depend on $D_{i}$.
\end{enumerate}
If conditions \ref{cond:normal} and \ref{cond:tail} hold, then naive data mining selects the same set of strategies as empirical Bayes.
\end{proposition}
Conditions \ref{cond:normal} and \ref{cond:tail} arise naturally when using long samples (e.g. 300 months of returns), standardized performance measures (e.g. t-statistics), and strict statistical hurdles (e.g. 5\% critical levels). Under these conditions, actual performance is a strictly increasing function of only the measured performance, as proved in Appendix \ref{sec:app:proof}. As a result, data-mined performance provides a reliable signal of actual performance, even if the magnitudes are distorted. The proposition assumes some exact conditions, and leading to identical selections, but approximate conditions would likely lead to similar selections.
One interpretation of Proposition \ref{prop:optimal-naive} is that Fisher's \citeyearpar{fisher1925statistical} focus on t-statistics set future researchers up for success, even in the modern era of big data.
On the other hand, Fisher would likely have been unsatisfied with finding the best strategies. He most likely would implore us to find unbiased estimates for these best performers. Thus to rigorously mine data, one should still apply empirical Bayes.
\subsection{Empirical Bayes Implementation \protect}\label{sec:stats:eb-details}
We select as our performance measure the t-statistic on the raw long-short return, and assume that standard errors are precisely measured, implying
\begin{align}
r_{i}|\mu_{i}, X_{i} &\sim \text{Normal}\left(\mu_{i}, 1\right)
\label{eq:t=theta+delta}
\end{align}
The latent performance is a mixture of two normals that depends on the strategy family $D_i$.
\begin{align}
\mu_{i}\mid (X_i , D_i = d)
& \sim\begin{cases}
\text{Normal}\left(
\theta_{d,1}, \sigma_{d,1}^2
\right)
& \text{with prob }\lambda_{d}\\
\text{Normal}\left(
\theta_{d,2},\sigma_{d,2}^{2}
\right)
& \text{otherwise}
\end{cases}.
\label{eq:theta~mixnorm}
\end{align}
where $d$ is one of the six strategy families that comes from combining three data sources (accounting, past returns, tickers) with two portfolio formation methods (equal-weighted and value-weighted). Mixture normals are parsimonious, easy to understand, and yet allow for skewness and fat tails.
We then estimate $\Omega\equiv\left[\theta_{d,1}, \sigma_{d,1}^2, \theta_{d,2}, \sigma_{d,2}^2, \lambda_d\right]_{d=1,\ldots,6}$ using quasi-maximum likelihood. The quasi-likelihood is computed using the distr package (\citet{ruckdeschel2006s4}). Optimization of $\Omega$ uses nloptr (\citet{NLopt}).This estimation is done using the past 20 years of long-short returns, separately for each ``forecasting year'' spanning 1983-2019.
Finally, we recover EB predictions by computing Equation (\ref{eq:ephat-def}) with distr, which produces an EB prediction of the expected return in units of standard errors. The EB predicted return is just Equation (\ref{eq:ephat-def}) multiplied by the standard error. Similarly, the EB predicted Sharpe ratio is Equation (\ref{eq:ephat-def}) multiplied by the square root of the number of periods in the sample.
For further details see Appendix
\ref{sec:app:theory} or our github site.
\section{Performance of the Best Data Mined Strategies\protect}\label{sec:best}
We show that data mining leads to research-like out-of-sample
returns (Section \ref{sec:best-oos}) and take a look at which kinds
of strategies are identified by data mining (Section \ref{sec:best-detail}).
We also provide intuition for why data mining produces such high returns (Section \ref{sec:best-intuition}).
\subsection{Out-of-Sample Returns\protect}\label{sec:best-oos}
Can data mining generate out-of-sample returns? To answer this question we construct simple out-of-sample portfolio tests.
Each year, we sign strategies to have positive predicted returns,
and then form portfolios that equally-weight strategies in the top
$X\%$ of predicted Sharpe ratios. We use both EB predictions and standard naive predictions. We examine $X=$ 1, 5, and 10.
For comparison, we also examine a portfolio that equally-weighs published
strategies from the \citet{ChenZimmermann2021} dataset.
Table \ref{tab:beststrats} shows the result. Using empirical Bayes (EB Mining), the top 1\% of strategies perform similarly to strategies published in top finance
journals. Over the full 1983-2020 sample, the top 1\% portfolio earns
5.70\% per year, compared to the 5.88\% return from published strategies.
The Sharpe ratio from EB mining is smaller, at 1.46 vs 2.03 for
published strategies. However, unlike the EB-mined strategies, which
are formed using only information available in real-time, the published
strategies contain look-ahead bias. Indeed, if we focus on strategies
in top journals that were published pre-2004, the performance is very
similar to the EB-mined strategies in terms of either mean returns
or Sharpe ratios.
\begin{center}
{[}Table \ref{tab:beststrats}, Returns of Long-Short Portfolios Data-Mined,
about here{]}
\par\end{center}
The EB-mined returns are robust. The top 5\% and top 10\% of data-mined
strategies also perform well and are extremely statistically significant, indicating that the performance of the top 1\% is not driven by outliers.
Panel B shows that even naive data mining produces research-like returns. Simply choosing the top 1\% of strategies based on their past Sharpe ratios leads to an out-of-sample Sharpe ratio of 1.45. This is almost exactly the same as the Sharpe ratio from the top 1\% using EB mining, consistent with Proposition \ref{prop:optimal-naive}. The top 5\% and top 10\% of naive strategies underperform a bit relative to EB mining, but the intuition behind Proposition \ref{prop:optimal-naive} still goes through.
Figure \ref{fig:beststrats-cret} takes a closer look by plotting
the value of \$1 invested in each portfolio over time. The top 1\%
data-mined strategies have similar performance to published strategies
throughout the figure. All portfolios show relatively little
cyclicality during the recessions of 1991, 2009, and 2020. Indeed,
the returns are fairly consistent throughout the chart, with an important
caveat.
\begin{center}
{[}Figure \ref{fig:beststrats-cret}, Cumulative Long-Short Returns,
about here{]}
\par\end{center}
The caveat is that returns are concentrated in the pre-2004 sample.
This is seen in the flattening of the solid line in Figure \ref{fig:beststrats-cret} around 2004. The EB Mining top 1\% portfolio returns 8.17\% per year from 1983-2005, compared to just 2.03\% from 2005-2020. A similar decay is seen across all portfolios, both data-mined and academic. This decay is consistent with \citet{Chordia2014Have} and \citet{chen2022zeroing}, who argue that the rise of information technology reduced return predictability.
Overall, we find that one can find long-short returns comparable to
those from the best journals in finance, just by mining data, with
little thought about the underlying economics. Moreover, rigorous
data mining can discriminate between data sources that have no information
about future returns, like stock market tickers, from data that is
rich in information, like accounting ratios. Unlike the published
strategy returns, our returns can be found using only information
available in real-time. These results show that high-throughput methods
provide a bias-free approach to studying stock market predictability.
Our strategy returns and code are public, and we encourage future
researchers to use these methods.
\subsection{The Composition of the Top 1\%\protect}\label{sec:best-detail}
Table \ref{tab:top_strat_desc} takes a closer look at the top 1\%
strategies produced by rigorous data mining. Panel A shows that 91.0\%
of the top 1\% come from the equal-weighted accounting family and
8.6\% come from equal-weighted past returns. The other strategy families
comprise a negligible part of the top 1\%. Ticker strategies are completely
absent.
\begin{center}
{[}Table \ref{tab:top_strat_desc}, Description of Top 1\% Data-Mined
Strategies, about here{]}
\par\end{center}
Taken with Table \ref{tab:beststrats}, these results show that cross-sectional
predictability is concentrated in accounting data, small stocks, and
pre-2004 samples. These stylized facts offer a parsimonious description
of the ``factor zoo.'' Theories that wish to capture the big picture
of cross-sectional predictability should be consistent with these facts.
For example, slow diffusion of economic information is consistent,
as this diffusion would be especially slow in small stocks and before
the internet era. In this way, high throughput asset pricing provides
a way to not only identify out-of-sample returns, but to also provide
insight into the underlying economics.
Panel B of Table \ref{tab:top_strat_desc} shows that many of the top 1\% strategies are quite far from
the predictors noted in the academic literature. In 1993, academics
were focused on predictors like book-to-market, 12-month momentum,
and sales growth (\citet{fama1992cross}; \citet{jegadeesh1993returns};
\citet{lakonishok1994contrarian}). None of these predictors are in
the top 20 strategies based on predicted Sharpe ratios from rigorous
data mining. Instead, the common themes from data mining include shorting
stocks with high or growing debt, as well as buying stocks with high
depreciation, depletion, and amortization. Another theme is buying
stocks with high returns in quarters $t$ minus 17 and 18.
Based on textbook risk-based or behavioral asset pricing, one might
expect that these data-mined predictors will average zero returns
out-of-sample. But this is not the case. The realized Sharpe ratios
for these strategies in the 10 years after 1993 averages around 1.0
(``SR OOS'' column).
Panel B of Table \ref{tab:top_strat_desc} focuses on 1993 because well-known predictability papers were published around that time (e.g. \citet{fama1993common}). In other years, the top 20 list is different, though shorting variables related to debt growth remains a common theme. For further details see
Appendix Tables \ref{tab:top_strat_alt1} and \ref{tab:top_strat_alt2}.
\subsection{Shrinkage Intuition\protect}\label{sec:best-intuition}
Unlike many big data and machine learning methods, empirical Bayes
has a transparent intuition. The intuition can be seen in a special
case of the prediction Equation (\ref{eq:ephat-def}). If $\mu_{i}\mid \left(X_i, D_i = d\right) \sim\text{Normal}\left(0,\sigma_{d}^{2}\right)$,
we have
\begin{align}
E\left(\mu_i \mid r_i = \bar{r}, X_i = \bar{X}, D_i = d\right)
& =\left[
1-\frac{1}{\widehat{\Var}\left(r_{i} | D_i = d\right)}
\right]
\bar{r},\label{eq:shrink-ez}
\end{align}
where $\widehat{\Var}\left(r_{i} | D_i = d\right)$ is an estimate of the cross-strategy variance of performance measures among strategies with data family $d$.
This expression says that rigorous mining involves shrinking performance measures $r_{i}$ toward zero at a rate of $\frac{1}{\widehat{\Var}\left(r_{i} | D_i = d\right)}$. $\widehat{\Var}\left(r_{i} | D_i = d\right)$ measures how far the data are from the null of $r_{i} \sim\text{Normal}\left(0,1\right)$, which we imposed in Equation (\ref{eq:t=theta+delta}). If there is no predictability, then $r_{i}\sim\text{Normal}\left(0,1\right)$, $\widehat{\Var}\left(r_{i} | D_i = d\right)\approx1$, and all $r_{i}$ are shrunk to zero. But if data are far from the null, then a large $r_{i}$ is a signal of large $\mu_{i}$---even if $r_{i}$ is found from searching tens of thousands of strategies, unguided by economic theory.
Figure \ref{fig:t-stat-1983} shows that equal-weighted
accounting strategies (upper left) are far from the null using data
from 1964 to 1983. Equal-weighted past return strategies (middle left)
also show a notable deviation. In contrast, the other strategy families
are quite close to the null. Indeed, for both families of ticker-based
strategies, the null is a very good fit for the data.
\begin{center}
{[}Figure \ref{fig:t-stat-1983}, Distribution of t-stats
in 1983, about here{]}
\par\end{center}
Accordingly, Equation (\ref{eq:shrink-ez}) implies that the strategies with strong actual performance will be found in equal-weighted accounting
and equal-weighted past-return strategies. This intuition is consistent
with Panel A of Table \ref{tab:top_strat_desc}, which shows that
the vast majority of the best data-mined strategies come from these
families.
Compared to data available in 1983, all strategy families are closer
to the null using data from 1985-2004, as seen in Figure \ref{fig:t-stat-2004}.
All value-weighted families are very close to the null, implying that
predictability in large stocks is essentially gone. The long left
tail in equal-weighted past return strategies also disappears. Only
equal-weighted accounting strategies are visually far from the null.
These results imply that predictability is concentrated in the earlier
part of the sample.
\begin{center}
{[}Figure \ref{fig:t-stat-2004}, Distribution of t-stats
in 2004, about here{]}
\par\end{center}
The intuition in Figures \ref{fig:t-stat-1983} and \ref{fig:t-stat-2004}
is so simple that one might even skip the quasi-maximum likelihood
estimation. Just looking at these charts, and the distance between
the data and the null, one can already tell that predictability is
concentrated in small stocks, accounting data, and the earlier sample.
That is, one can already tell where predictability is concentrated,
if one understands the intuition in Equation (\ref{eq:shrink-ez}).
\section{Empirical Bayes Prediction Accuracy Across the Cross-Section\protect}\label{sec:accuracy}
This section takes a closer look at the EB predictions and accuracy.
We see when and where EB predictions are successful and when they
struggle.
\subsection{EB Prediction Accuracy 1983-2004\protect}\label{sec:oos-cross-pre2004}
To examine accuracy, we use out-of-sample portfolio sorts. For each
year and each strategy family, we form 20 portfolios by sorting strategies
into equal-sized groups based on the past 20 years of mean returns.
We then predict the mean returns for each portfolio by averaging the
EB predictions (Equation (\ref{eq:ephat-def})), which are also based
on the past 20 years of data. Finally, we form a portfolio that equally-weighs
strategies in each group and hold for one year (the ``out-of-sample''
periods).
Figure \ref{fig:xpred-1} shows the in-sample, predicted, and out-of-sample
returns for each portfolio, averaged over the out-of-sample periods from
1983 to 2004. For all six families, there are sizable in-sample returns
(dashed line) in the extreme in-sample groups. For accounting strategies,
in-sample returns are as extreme as -11\% per year. A naive read of
this result is that one can flip the long and short legs and find
+11\% returns out-of-sample. Past return strategies see a similar
$\pm10$ percent return in the extreme groups. Even ticker-based strategies
show in-sample long-short returns of up to 4 percent per year.
\begin{center}
{[}Figure \ref{fig:xpred-1}, Empirical Bayes Predictions 1983-2004,
about here{]}
\par\end{center}
However, the predicted returns are typically much closer to zero.
In fact, for both ticker-based strategy families, the predicted return
(solid line) is almost exactly zero for all 40 in-sample groups. This
result is intuitive given how close the ticker t-stats are to the
null of no predictability (Figure \ref{fig:t-stat-2004}).
This closeness implies that the extreme returns can be entirely accounted
for by luck, and so shrinkage should be 100\% (Equation (\ref{eq:shrink-ez})).
Significant shrinkage is also seen in value-weighted accounting strategies
(top right panel). Rigorous data mining recommends that the extreme
returns of around -8\% and +9\% (dashed line) be shrunk down to about
-3\% and +2 (solid line), respectively.
Rigorous mining predicts much higher returns in equal-weighted accounting
strategies (upper left panel). For these strategies, the predicted
returns are actually not far from the in-sample return. This result
is consistent with \citet{chen2020publication}, who find shrinkage
of only 12\% for published anomalies, which are largely equal-weighted
and based on accounting variables. Predictability is also seen in
both families of past return strategies.
These predictions are borne out in out-of-sample returns (markers
with error bars). The first group of EW accounting strategies returns
-8 percent per year out-of-sample from 1983-2004, almost exactly the
same as the EB prediction. Similar accuracy is seen throughout all
120 bins in Figure \ref{fig:xpred-1}.
These results show that rigorous data mining offers economic insights
that are difficult to derive from theory. While theories of slow information
diffusion may tell you that predictability is concentrated in small
stocks, accounting signals, and pre-2004 data, they are unlikely tell
you how much predictability there is. In contrast, empirical Bayes
provides quantitative, accurate estimates of the precise amount of
predictability.
\subsection{EB Prediction Accuracy 2004-2020\protect}\label{sec:oos-cross-post2004}
We split our OOS tests in the mid-2000s, motivated by the idea that
there was likely a structural break during this period due to the
rise of information technology (\citet{Chordia2014Have}). Comparing
the distribution of t-stats available in 1983 vs 2004 supports the
idea that the structure of financial markets changed (see Section
\ref{sec:best-intuition}).
\begin{center}
{[}Figure \ref{fig:xpred-2}, Empirical Bayes Predictions 2004-2020,
about here{]}
\par\end{center}
This structural change can be seen by comparing Figure \ref{fig:xpred-2}
(EB predictions 2004-2020) to Figure \ref{fig:xpred-1} (EB predictions
1983-2004). In all panels, the predicted returns shift closer to zero
post-2004. Most notably, the predictability that was present in past
return strategies pre-2004 is largely gone. Consistent with these
predictions, the past return portfolios show a flat or even negative
relationship between out-of-sample and in-sample returns post-2004.
A similar weakening of EB predictions and flattening of out-of-sample
returns is seen in the accounting VW family.
An exception to this pattern is the family of equal-weighted accounting
ratio strategies (top left). In this chart, the shrinkage is still
relatively small, with EB predictions implying returns as extreme
as -9 percent per year. This prediction and others in this panel miss
the mark: the out-of-sample returns are much closer to zero throughout
this panel.
This poor accuracy is natural given the fact that the estimations
use a rolling window consisting of the past 20 years of data. This
fixed window implies that, for much of the period 2004-2020, our estimates
rely on data from a time when accounting statements needed to be retrieved
by traditional (snail) mail for investors without special access to
the SEC reading room (\citet{bowles2023anomaly}).
This result implies an important role for economic theory: when structural
breaks occur, there is no way for data mining to provide a clear understanding
of the economy, no matter how rigorously the mining is done. Theory
is sometimes used this way in economics and finance, but this is typically
not the case. Instead, theory is typically used to understand patterns
found in long samples of data, spanning many decades. In our view,
the future of theory is bright for theorists who study structural
breaks, even in the era of big data. Indeed, a smart data miner armed
with theory might have understood the implications of the internet
for stock return predictability, and could perhaps have performed
much better than our theory-free EB mining process.
\section{Comparison with False Discovery Controls\protect}\label{sec:fdr}
Our main analysis corrects for data mining bias using empirical Bayes shrinkage, following \citet{chen2020publication}; \citet{chen2022zeroing}; and \citet{jensen2023there}. An alternative approach is to use false discovery controls, following \citet{harvey2016and}; \citet{barras2010false}; or \citet{chordia2020anomalies}. The ideal approach remains an unsettled question. Our dataset of 136,000 trading strategies provides a natural testing ground.
We examine the following false discovery controls:
\begin{enumerate}
\item \textbf{BY1.3 (1\%)}: \citet{harvey2016and} (HLZ) recommend using \citepos{benjamini2001control} Theorem 1.3 at the 1\% level. HLZ is likely the most influential paper on multiple testing in empirical asset pricing.
\item \textbf{Storey (10\%)}: \citet{barras2010false}, which introduced false discovery methods to finance, study the \citet{storey2002direct} algorithm at the 10\% level.
\item \textbf{RW (5\%, 5\%)}: \citet{chordia2020anomalies} recommend combining \citepos{romano2007control} Algorithms 4.1 and 2.1. These algorithms require two parameters, both of which Chordia et al. set to 5\%.
\end{enumerate}
For each year and each strategy family, we apply these methods using the past 20 years of data to estimate a t-statistic hurdle. We then examine whether these hurdles are able to separate strategies with high out-of-sample returns from those with low out-of-sample returns. This structure is the same as in Section \ref{sec:accuracy}.
Figure \ref{fig:fd-control} shows the results. The vertical lines show the mean hurdle across all years. The markers show the mean out-of-sample returns of portfolios formed by equally weighting strategies, sorted into 20 groups based on the in-sample t-statistic. Groups of strategies that a false discovery control declares ``significant'' lie on the outside of the respective vertical lines.
\begin{center}
[Figure \ref{fig:fd-control}, False Discovery Controls, about here]
\end{center}
The BY1.3 (1\%) and RW (5\%, 5\%) methods miss out on the majority of portfolios with notable out-of-sample performance. Out of the 5 groups that have out-of-sample returns of at least 3\% per year, only 1 lies outside of the solid lines corresponding to BY1.3 (1\%). Only 3 of 5 lie outside the dot-dashed lines corresponding to RW (5\%, 5\%). The dashed line, corresponding to Storey (10\%), performs much better, capturing 4 of the 5 portfolios. Similar results are found using alternative parameter choices examined by HLZ; \citet{harvey2020false}; and \citet{barras2010false} (see Appendix Figure \ref{fig:fd-control-alt}).
Thus, Storey (10\%) provides an easy-to-compute alternative to empirical Bayes. However, Storey cannot provide bias-adjusted performance estimates that are naturally available from empirical Bayes. Overall, our results imply that Storey forms a strong first step for rigorous data mining, while empirical Bayes is recommended for more refined estimates. These results are broadly consistent with the statistics literature, which generally recommends Storey as a preliminary examination, while suggesting empirical Bayes for greater precision (e.g. \citealt{benjamini2010discovering}; \citealt{efron2012large}).
We discuss this literature and the algorithms in more detail below.
\subsection{\citet{benjamini2001control} Theorem 1.3}\label{sec:fdr-hlz}
\citet{harvey2016and} (HLZ) recommend using \citepos{benjamini2001control}
Theorem 1.3. Several followups to the influential HLZ paper use this method, including \citet{harvey2020false,chordia2020anomalies}; and \citet{jensen2023there}.
We state the theorem number 1.3 because the bulk of the original paper focuses on Theorem 1.2. Indeed, \citet{benjamini2001control} describe Theorem 1.3 as ``very often unneeded, and yields too conservative of a procedure'' (page 1183). In his textbook on large scale inference, \citet{efron2012large} agrees, stating that the theorem represents a ``severe penalty'' and is ``not really necessary'' (section 4.2). Moreover, the statistics literature uses the ``BY algorithm'' to refer to \citet{benjamini2005false}, which is an entirely different procedure (e.g. \citealt{efron2012large} Chapter 11.4).
BY1.3 begins by choosing a parameter $q^{\ast}$ and then solving
\begin{align}
h_{\text{HLZ},q^{\ast}} & \equiv\min_{h>0}\left\{ h:\left[\frac{\Pr\left(|Z|>h\right)}{\text{Share of \ensuremath{|t_{i}|>h}}}\right]\pi_{\text{BY1.3}}\le\text{\ensuremath{q^{\ast}}}\right\} \label{eq:by1.3hurdle}
\end{align}
where $t_i$ is the t-statistic for strategy $i$, $Z$ is a standard normal random variable,
\begin{align}
\pi_{\text{BY1.3}} & \equiv\sum_{i=1}^{N}\frac{1}{i},
\label{eq:by1.3penalty}
\end{align}
and $N$ is the number of strategies in the year-family. \citepos{benjamini2001control} Theorem 1.3 proves that this algorithm implies a false discovery rate
$\le q^{\ast}$. BY1.3 amounts to modifying the seminal \citet{benjamini1995controlling} algorithm with a constant factor, $\pi_{\text{BY1.3}}$. This modification makes the algorithm more conservative.
HLZ recommend this conservative approach, claiming \citet{benjamini1995controlling} ``is only valid when the test statistics are independent or positively dependent'' (page 21). This statement is false. \citet{storey2001estimating} and \citet{storey2004strong} show validity under weak dependence assumptions (see also \citealt{chen2024most}).
HLZ are also conservative in their choice of $q^{\ast}$. For their main results, they use $q^{\ast}=1\%$ citing the fact that the ``significance level is subjective,'' though they also examine $q^{\ast}=5\%$ for robustness. In contrast, the statistics literature generally recommends $q^{\ast}=5\%$ or 10\% (e.g. \citealt{benjamini2010discovering}; \citealt{efron2012large}).
Given this context, it is perhaps unsurprising that BY1.3 (1\%) fails to identify most out-of-sample performers in Figure \ref{fig:fd-control}. BY1.3 (5\%) performs somewhat better, identifying 2 out of 5 groups with out-of-sample returns of at least 3\% per year (Appendix Figure \ref{fig:fd-control-alt}).
\subsection{\citepos{storey2002direct} FDR Control\protect}\label{sec:fdr-storey}
While HLZ recommend modifying \citet{benjamini1995controlling} to be more conservative, much of the statistics literature goes in the opposite direction, modifying \citet{benjamini1995controlling} to be more aggressive. In finance, \citet{barras2010false} take this approach.
Barras et al. recommend the \citet{storey2002direct}
algorithm, which can be written as
\begin{align}
h_{\text{Storey},q^{\ast}} & \equiv\min_{h>0}\left\{ h:\left[
\frac{\Pr\left(|Z|>h\right)}{\text{Share of \ensuremath{|t_{i}|>h}}}
\right]\pi_{\text{Storey}}\le\text{\ensuremath{q^{\ast}}}\right\} \label{eq:storey-1}
\end{align}
where $t_i$ is the t-statistic for strategy $i$, $Z$ is a standard normal random variable,
\begin{align}
\pi_{\text{Storey}} & =\frac{\text{Share of \ensuremath{|t_{i}|\le1.0}}}{\Pr\left(|Z|\le 1.0\right)}=\frac{\text{Share of \ensuremath{|t_{i}|\le1.0}}}{0.68}\label{eq:storey-2}
\end{align}
and the cutoff of $1.0$ is selected for ease of interpretation. \citet{storey2002direct} proves that this algorithm implies a false discovery rate $\le q^{\ast}$ under independence assumptions, though \citet{storey2001estimating} and \citet{storey2004strong} extend this result to weak dependence.
Comparing Equations (\ref{eq:storey-1})-(\ref{eq:storey-2}) to the Equations (\ref{eq:by1.3hurdle})-(\ref{eq:by1.3penalty}), we see that the only difference is the constant factor, $\pi_{\text{BY1.3}}$ vs $\pi_{\text{Storey}}$. These constants are qualtitatively different: $\pi_{\text{BY1.3}} = \sum_{i=1}^{N}\frac{1}{i} \approx 0.6 + \log N \gg 1$ , while $\pi_{\text{Storey}}\le 1.0$. As shown in \citet{storey2002direct}, $\pi_{\text{Storey}}$ can be interpreted as an estimate of the probability that a strategy is null, which can be at most 1.0. Other statistics papers that recommend a constant that is at most 1.0 include \citet{benjamini2000adaptive, efron2001empirical,genovese2006false,benjamini2006adaptive}.
\citet{barras2010false} do not emphasize a particular choice of $q^{\ast}$, and instead examine values ranging from 5\% to 20\%. Figure \ref{fig:fd-control} uses $q^{\ast}=10\%$, because Barras et al. use 10\% in their illustrative examples.
Once again, given the support from the statistics literature, it is perhaps unsurprising that Storey (10\%) and (20\%) perform well. Equations (\ref{eq:storey-1})-(\ref{eq:storey-2}) are easy to implement, making it a useful alternative to our empirical Bayes method.
However, there are two downsides to using Storey. A simple, symmetric testing algorithm like Equations (\ref{eq:storey-1})-(\ref{eq:storey-2}) does not handle skewed distributions well. This limitation may explain why Storey struggles to identify out-of-sample performers in past return strategies, which feature a long right tail (Figure \ref{fig:t-stat-1983}). The second is that Storey cannot provide bias-adjusted performance estimates. Such estimates are naturally available from a more general empirical Bayes method, and would provide clean connections with portfolio choice and asset pricing questions.
\subsection{\citepos{romano2007control} FDP Risk Control}\label{sec:fdr-rw}
\citet{chordia2020anomalies} recommend combining \citepos{romano2007control} Algorithms 4.1 and 2.1, which we refer to as ``RW.'' This algorithm is a natural choice for asset pricing researchers, as its predecessor \citet{romano2005stepwise} is motivated by data mining for CAPM anomalies. Like HLZ's method, the Romano and Wolf methods have been used in influential asset pricing papers, including \citet{chordia2020anomalies}; \citet{engelberg2023do}; \citet{heath2023reusing}; and \citet{debodt2025competition}.
Unlike Storey and BY1.3, the statistics literature has relatively little discussion of the Romano and Wolf methods. Neither Romano and Wolf (\citeyear{romano2005stepwise}) nor Romano and Wolf (\citeyear{romano2007control}) is found in the textbooks \citet{efron2012large} and \citet{efron2016computer}. The two papers are also not found in the review articles on multiple testing \citet{benjamini2010discovering} and \citet{benjamini2020selective}. Thus, we provide some discussion here.
The goal of RW can be written as follows: find an $h$ that ensures
\begin{align}\label{eq:fdp-risk-control}
\Pr(\FDP > p^\ast) \le q^\ast
\end{align}
where
\begin{align}
\FDP \equiv \frac{\text{
Number of null strategies with $|t_i|>h$
}}{\text{
Number of strategies with $|t_i|>h$}}
\end{align}
and $p^\ast$ and $q^\ast$ are thresholds selected by the researcher. Null strategies are, typically, those with an actual performance of zero.
Figure \ref{fig:fdp-risk-demo} illustrates Equation (\ref{eq:fdp-risk-control}), by simulating one of our QML estimates many times. We run 2,000 simulations, each one consisting of 29,000 strategies. For simplicity, we assume all strategiesare independent. The plot shows histograms of actual performance ($\mu_i$ in Equation \eqref{eq:r=mu+e}) for strategies that meet the hurdle $|t_i|>3.0$, where $h=3.0$ is selected for illustrative purposes. Using this chart, we can ask whether this $h=3.0$ hurdle achieves Equation \eqref{eq:fdp-risk-control}, and thus understand FDP risk control.
\begin{center}
[Figure \ref{fig:fdp-risk-demo}, FDP Risk Control Illustration, about here]
\end{center}
Since the $h=3.0$ hurdle is quite stringent, the vast majority of strategies are non-null. However, there is still a risk that a strategy with $|t_i|>3.0$ has near-zero actual performance, as seen in the left tail of the histogram. The FDP characterizes this risk. It is, approximately, the share of strategies in the first bin.\footnote{More formally, one can consider the first bin to be an upper bound on the FDP (see \citealt{chen2024most}).} On average, the share of strategies in this bin is about 5\% (bars), indicating that the FDR is approximately controlled at a 5\% level.
Even though the FDP is on average about 5\%, there is a risk that it is higher. This risk is seen in the lines of Figure \ref{fig:fdp-risk-demo}, which plot extreme order statistics across the 2,000 simulations. The 95th percentile line implies the FDP exceeds 7\% in 5\% of simulations. To achieve FDP risk control with $p^\ast=5\%$ and $q^\ast=5\%$, a $h>3.0$ is required. The RW method finds this $h$.
Thus, the RW method aims to control the tail risk of a tail risk. Such an algorithm is a natural choice if selecting a null strategy is catastrophic. In such a case, one may want to ensure not only that a null is highly improbable, but that the probability that a null is somewhat probable is also improbable. However, in the standard setting where the null is that the strategy has zero long-short return or zero alpha, then the RW method tends to imply extreme conservatism.
This conservatism leads to the results in Figure \ref{fig:fd-control} and Appendix Figure \ref{fig:fd-control-alt}. Choosing $p^\ast=0.05$ and $q^\ast=0.05$ or 0.10, as in \citet{chordia2020anomalies} and \citet{harvey2020evaluation}, leads to hurdles that many notable out-of-sample performers fail to clear.
The RW method is rather complex. It uses cluster bootstrap methods, involves testing all possible subsets of selected sets of strategies, iterating over many possible tests and sets. We describe our implementation in Appendix \ref{sec:app:rw-details} and provide code in our Github repo.
\section{Conclusion}
We show that a solution to data mining bias is to mine data rigorously. We systematically search 136,000 long-short strategies and find out-of-sample performance comparable to academic research. Simply searching for strategies with the largest t-stats leads to publication-like out-of-sample performance, a fact we explain in a Bayesian model. While naive data mining leads to distorted performance estimates, empirical Bayes provides unbiased predictions in samples without structural breaks. The forecast errors around structural breaks suggest a role for theory in the era of big data.
This high-throughput method shows that returns are concentrated in accounting signals, small stocks, and pre-2004 periods, consistent with mispricing and slow information diffusion theories. While these results could potentially be gleaned from a deep read of the anomalies literature, our method provides a scientific method for documenting these stylized facts. We provide our data and code publicly, and hope others follow in using high-throughput methods.
Our out-of-sample tests offer an intuitive method for comparing multiple testing methods. We find that methods popular in finance would lead researchers to miss out on the majority of signals with notable out-of-sample performance. In contrast, methods recommended by the statistics literature perform well.
\newpage{}