Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
40,902 characters · 14 sections · 62 citation commands
Optimal Post-Hoc Theorizing
\newif\ifanon \anontrue \anonfalse
\ifanon \else \fi
\color{Black}JEL Classification: B41, C18, C11
\color{Black}Keywords: Publication Bias, Machine Learning, Predictivism vs Accommodation, HARKing \thispagestyle{empty}\setcounter{page}{0}
\setcounter{page}{1}
Theories formed after observing empirical results (post hoc theories), are viewed with suspicion by social scientists (e.g. kerr1998harking; harvey2017presidential). Yet some of the most successful theories in all of science were formed this way (e.g. gravity, quantum mechanics).\footnote{newton1726scholium even said “whatever is not deduced from the phenomena... ... have no place in experimental philosophy.”} Consistent with this confusion, the philosophy literature has long debated the merits of post hoc vs a priori theorizing (barnes2022prediction)
This paper provides a Bayesian model for understanding this “paradox.” It shows post hoc theory is clearly suboptimal if the sole goal of research is unbiased empirical results. Given statistics' 100-year obsession with unbiasedness (efron2001statistical), it is perhaps unsurprising that post hoc theory is viewed suspiciously.
However, the goal of research is typically more than unbiased empirical results. Another ubiquitous goal of research is to find “a good idea,” whether the idea is an investment strategy, health intervention, or model of human language. In such settings, statistical bias may matter little, as long as research provides a powerful solution.
If the goal is a “good idea,” then the optimal research method trades off a Darwinian Learning effect with a Statistical Learning effect. Darwinian Learning comes from weeding out bad theories by subjecting them to prediction competitions. Statistical Learning simply comes from theorists improving their ideas after looking at data. If Statistical Learning is stronger than Darwinian Learning, then post hoc theorizing is optimal.
In the modern world of enormous datasets and massive computing power, Statistical Learning is becoming more and more powerful. At the same time, the economic sciences have become mature, and Darwinian Learning has arguably run its course. For these reasons, I argue that post hoc theorizing is, in most cases, optimal.
For replication code and all previous versions of this paper, see \url{https://github.com/chenandrewy/Post-hoc/}.
My model is an extension of the publication bias models (hedges1984estimation; brodeur2016star; andrews2019identification; abadie2020statistical; chen2020publication; jensen2023there; kasy2024optimal). In these papers, it is unclear whether post hoc theory is harmful. In fact, the models in these papers exhibit the irrelevance result found in hempel1966philosophy; lakatos1970methodology; and elsewhere (see Section (ref)). Building on the insights of from the philosophy literature (namely maher1988prediction), I show how heterogeneous theories breaks this irrelevance.
In the philosophy literature, Maher (maher1988prediction, maher1990prediction) and Kahn, Landsburg, and Stockman (kahn1992novel; kahn1996positive) (KLS) study post hoc theorizing under heterogeneous theories. They document the selection effect that I call Darwinian Learning, and conclude that a priori theorizing is optimal, at least in normal scientific settings. Amid the centuries of debate (e.g. leibniz1678letter; newton1726scholium; keynes1921treatise), barnes1996discussion describes Maher's analysis as “the closest thing to an illuminating account of predictivism in existence.” Predictivism is the view that a priori theorizing is optimal.
My paper builds on Maher and KLS by showing how there is an offsetting effect to Darwinian Learning, namely Statistical Learning. This effect is ruled out by the assumptions in Maher and KLS. Statistical Learning is perhaps a natural extension of one of \citepos{howson1991maher} criticisms of Maher maher1988prediction and maher1990prediction, though maher1993discussion also points out flaws in \citepos{howson1991maher} criticisms. My paper provides clarity to this debate. Also unlike Howson and Franklin, I show how to connect Maher's and KLS's ideas to the models of publication bias, and the broader statistics literature on large scale inference (efron2012large).
Idea $i$ is randomly-drawn from a set $\left\{ 1,2,...,N\right\}$, and has quality $\mu_{i}$. $\mu_{i}$ is unknown but researchers can observe the measured quality
where $E\left(\varepsilon_{i}\right) = 0$. $i$ may be a real-world choice for readers (e.g. an investment strategy), in which case $\mu_{i}$ is the realized, quality of $i$ after the research is finished (“post-research”). Or $i$ may be an explanation for some phenomenon (e.g. a model of obesity in adolescents), in which case $\mu_{i}$ is the explanation's fit to the phenomenon, post-research. In either case, higher $\mu_{i}$ is better.
Using theory rules out some ideas:
where
“Theorizing” turns $S$ into a selected idea $i^\ast$, and theorizing is either a priori or post hoc:
In either case, the theory is some math or text that explains why $i^\ast$ is a good idea. In this simple model, the precise nature of the theory is not important, beyond that it argues for selecting $i^\ast$.
I assume that idea $i^\ast$ and its supporting theory are eventually re-examined with post-research data via $\mu_{i^\ast}$ (see discussion after Equation (ref)). I also require that $S$ is well-defined, and does not nest all ideas $\left\{1,2,...,N\right\}$. In other words, I assume theories are falsifiable, in the sense of Popper1959.
One may be concerned that these assumptions are inappropriate for some social sciences. Indeed, ankel2025economics provide disturbing evidence that economic theories may be immunized against refutation. If economic theories are indeed, not falsifiable, then they might as well be fairy tales. Whether fairy tales are better told a priori or post hoc is beyond the scope of this paper.
Perhaps because of kerr1998harking (“HARKing: Hypothesizing after the Results are Known”), many researchers equate post hoc theorizing with unfalsifiability. However, as seen in this model, constructing theories post hoc can be entirely consistent with Popper's notion of science.
This confusion likely stems from Kerr's loose use of language. The paper has a section titled “HARKed Hypotheses Fail Popper's Criterion of Disconfirmability.” But the text below the title clarifies, “[a] HARKed hypothesis fails this criterion, at least in a narrow, temporal sense.” In other words, the text in the section explains that the section title is not necessarily true. In fact, it seems equally reasonable to say that HARKed hypotheses fail Popper's criterion only in a narrow, temporal sense. Errors like these are found throughout kerr1998harking. See rubin2022costs for a thorough critique.
If the sole goal of research is to find an unbiased estimate of idea quality, then a priori theorizing achieves this goal. The expected $\hat{\mu}_{i}$ from a priori theorizing satisfies
where $i$ is randomly selected from $S$. In contrast, the expected $\hat{\mu}_{i}$ from post hoc theorizing is clearly biased:
Intuitively, $\hat{\mu}_{i}$ contains both $\mu_{i}$ and measurement error. Selecting on large $\hat{\mu}_{i}$ then selects for positive measurement error, leading to a biased estimate.
The preference for Equation (ref), and the fear of Equation (ref), goes back to fisher1925statistical. As described in efron2001statistical:
Taken with Lemma (ref), it is no wonder then, that economists are suspicious of post hoc theorizing.
In an ideal world, estimates from a priori theorizing are all you need. With many, many of these estimates, one eventually has estimates for every idea, including the best ideas.
But in the real world, consumers and producers of research have limited time. Consumers of research lack the time to read about every idea. Producers of research lack the time to carefully study every idea.
To introduce this real-world limitation, suppose research is restricted to reporting only a single idea, and readers are interested in the idea with the highest quality.
In this case, post hoc theorizing is actually optimal. Post hoc theorizing uses both the information in theory (Equation (ref)) and the information in the data (Equation (ref)), improving its expected quality:
Lemmas (ref) and (ref) are illustrated in Figure (ref). It simulates 200 selected ideas, with the number of potential ideas $N=100$, $\mu_i \sim \text{Normal}\left(0, 1\right)$, and $\varepsilon_i \sim \text{Normal}\left(0, 1\right)$. A priori theorizing leads to less biased estimates, seen in how the dots lie closer to the 45 degree line. However, post hoc theorizing leads to higher quality ideas, seen in how the stars tend to lie toward the right side of the chart.
The literature on stock market anomalies is an example of Lemma (ref). Readers are interested in both the magnitude of anomalies, as well as which ones are the strongest. But assuming that the magnitude meets some minimal standard, readers with limited time will just want to know which anomalies will perform the best in the future. Lemma (ref) shows that, in this case, researchers should mine the data, and report what has worked best in the past. This prescription is exactly the reverse of the conventional wisdom, that emphasizes the “dangers” of data mining (sullivan1999data; harvey2016and). However, it seems to be in-line with empirical practice, and performs quite well (chen2024does).
Large language models (LLMs) are another example. These models are tuned to perform well on common benchmarks like MMLU (Measuring Massive Multitask Language Understanding) (e.g. guo2025deepseek). Thus, the performance on these benchmarks is biased upward, just as in Lemma (ref). But in practice, this bias is not important, as long as the resulting out-of-sample performance is strong. Tuning improves out-of-sample performance, as seen in Lemma (ref).
In practice, the Fisherian ideal is impossible. Even if all researchers use theory a priori, readers with time constraints are more likely to read the research if the measured effect is large. This limited attention is arguably the raison d'etre of both peer review (klamer2002attention) and publication bias (chen2022publication)
To model limited attention, suppose a priori theory actually involves two steps. First, researchers study all ideas in $S$ and draft up their theories and empirical findings in working papers. However, not all ideas are read. Due to limited attention, only the idea with the largest measured quality becomes well-known and consumed by the public. The expected quality of this, more realistic, a priori theorizing is
which is exactly the same as the quality of post hoc theory (Lemma (ref)).
A similar irrelevance is noted in many works of philosophy (e.g. hempel1966philosophy; lakatos1970methodology; rosenkrantz1977inference; gardner1982predicting). But as noted by maher1988prediction and kahn1996positive, this irrelevance can be broken if theories are endogenous.
Let's make the model richer, with endogenous, heterogeneous theories. This richer model is a generalization of maher1988prediction and kahn1996positive. Importantly, it allows for an effect I call “Statistical Learning.” As in Section (ref), I assume that the research community has limited time, and is primarily interested in finding ideas with the highest quality.
As before, there are ideas $i\in \{1,2,...,N\}$, measured idea quality $\hat{\mu}_{i}$, and true idea quality $\mu_{i}$. But now theories come from combining a “data input” with a “theory type.”
The data input ($\mathcal{D}$ or $\mathcal{O}$) is known. $\mathcal{D}$ is the case that the data input includes all of the measured effects ($\hat{\mu}_{1},\hat{\mu}_{2},...,\hat{\mu}_{N}$). $\mathcal{O}$ is the case that the theory is given access to none of these effects. Post hoc theorizing, then, is represented by $\mathcal{D}$, while a priori theorizing is $\mathcal{O}$.
The theory type has a quality $T$ which is unknown. For simplicity, assume the quality is either good (represented by $G$) or bad ($B$). Intuitively, not all theories types are the same, and we may not know how good a particular theory type is.
Combining a theory type with a data input leads to a theory, which in turn provides a recommended idea $i^{\ast}$. As before, $i^\ast$ is a random integer with support $S$, and the theory is some math and/or text that explains why $i^\ast$ is recommended. But now I'll use conditional probability notation to account for the data input and theory type. For example, $i^{\ast}|G,\mathcal{O}$ is the recommended idea generated by a good theory type and no data (a priori).
It's reasonable to think that the good theory type leads to higher quality ideas, a priori. This can be formalized by first order stochastic dominance:
For example, one may think that while bad theory types recommend any idea in $S$ with equal probability, good theory types are twice as likely to recommend ideas from the top quartile of $\mu_{i}$ (as compared to the second-to-top quartile). An implication of Equation (ref) is that good theory types typically lead to higher measured quality $\hat{\mu}_{i^{\ast}}$ than bad theory types.
If theory is done post hoc, researchers examine measured qualities $\hat{\mu}_1,\hat{\mu}_2,...,\hat{\mu}_N$, as well as the theory type, to construct a theory that selects idea $i^{\ast}|T,\mathcal{D}$. I allow $i^{\ast}|T,\mathcal{D}$ to be general, but assume the following restriction:
that is, using bad type theories always lead researchers to select the idea with the strongest measured quality (provided the idea is consistent with some theory). This assumption can be thought of as bad theory types being unable to distinguish between ideas in $S$, and Bayesian researchers who optimize on the posterior mean based on this information and $\hat{\mu}_{i}$ (see chen2025high).
After $i^{\ast}$ is chosen, readers decide if they are interested in the theory and idea. Assume readers are uninterested unless
where $h$ is some kind of economic and/or statistical hurdle. Only theories and ideas readers are interested in are published. This assumption follows the econometric literature on publication bias (andrews2019identification).
An immediate implication of heterogeneous theories is heterogeneous measured quality:
Lemma (ref) provides an alternative way to think about the chen2022peer (CLZ) “peer review vs data mining” experiment. CLZ compare stock trading ideas from peer review to data-mined trading ideas, using post-publication returns. If we call the post-publication returns $\hat{\mu}_{i^\ast}$, neither the peer-reviewed nor data-mined ideas had access to this data, so $\mathcal{O}$ holds for both groups of ideas. Then, one can think of peer-reviewed ideas as $i^\ast|T,\mathcal{O}$, since we do not know if the theory type is $G$ or $B$. In contrast, we can think of the data-mined ideas $i^\ast |B,\mathcal{O}$. As powerfully demonstrated by novy2025ai, anyone can add text to these ideas and call it a theory.
From this framing, CLZ's empirical results are a test of whether $G$ theory types exist. If $G$ theory types comprise a significant fraction of the theories in the CLZ sample, then Lemma (ref) implies that the published strategies have higher $\hat{\mu_{i^\ast}}$. Unfortunately, CLZ find that published strategies fail to outperform, implying that $G$ theories are rare.
The CLZ experiment illustrates the Darwinian selection of theories. If we force theorists to announce their ideas before looking at the data, then the bad theory types cannot hide behind data mining. This intuition helps justify the belief that a priori theorizing provides “discipline” and that post hoc theorizing is “too easy.” The following proposition formalizes this idea:
Proposition (ref) is illustrated in Figure (ref). It shows histograms generated by parameters deliberately chosen to highlight the power of Darwinian selection.
Under a priori theorizing, published ideas mostly come from the good theory types (Panel (a), left). Naturally, good theory types are better at separating good ideas from bad ones, a priori. Post hoc, published ideas largely come from the bad theory types (Panel (b), left). This happens because bad theory types lead researchers to check far more ideas for the highest measure quality, as these bad theories cannot discriminate among ideas. As a result, bad theory types are more likely to lead to publication, despite having lower actual quality. The final result is that a priori theorizing leads to published ideas with higher actual quality (vertical dashed lines).
Proposition (ref) captures the key insight of Maher (maher1988prediction; maher1990prediction) and Kahn, Landsburg, and Stockman (kahn1992novel, kahn1996positive). If theories are heterogeneous, then forcing theorists to announce their ideas before looking at the data helps eliminate bad theories, as in Darwinian selection. In Maher's terminology, a theory is a “method,” and the theory type is “reliability,” but the idea is the same.
Maher and KLS push further. They claim that, not only does a priori theorizing produce Darwinian selection, but that the resulting hypotheses are more likely to be true. The analogue here is that $\mathcal{O}$ implies not only that $G$ is more likely, but that $\mu_{i^\ast}$ is higher. We'll see that this conclusion is not necessarily true.\footnote{barnes1996discussion revisits Maher (1988, 1990, 1993) and does not go further. His Eq (4) stops here, and considers more deeply the terms in the Bayes rule version of $P\left(G|\mathcal{O},\hat{\mu}_{i^{\ast}}>h\right)$. }
An interesting feature of Proposition (ref) is that it shows a virtue of publication bias. While requiring $\hat{\mu}_{i^{\ast}}>h$ leads to biased estimates, it helps weed out bad theories types. This result is closely analogous to Lemma (ref).
Research is not only interested in finding good theory types, but also good ideas. In fact, one can argue that finding good ideas is the ultimate goal.
Whether post hoc theory helps or hurts for finding good ideas is characterized by the following proposition:
The proof is in Appendix (ref).
The proposition says that whether post hoc or a priori theorizing leads to better ideas depends on the relative size of two effects:
Naturally, if Statistical Learning exceeds Darwinian Learning, then it's often better to look at the data---i.e. post hoc theory may be optimal.
There is no hard and fast rule for which effect is larger. There are certainly settings where Statistical Learning is miniscule (e.g. when the data is extremely noisy). And there are certainly settings where Darwinian Learning is ineffective (e.g. when all theory types are the same).
Similarly, there are contradictory historical examples. Mendeleev's prediction of elements is a shockingly impressive example of a priori theorizing. But Planck's law of radiation is a shockingly impressive example of post hoc theorizing. Proposition (ref) provides a way to understand these seemingly contradictory phenomena.
If theories are homogenous in quality, then there is no Darwinian Learning, and thus Proposition (ref) implies that post hoc theory is optimal.
Figure (ref) illustrates this phenomenon, by examining many variations of the model from Figure (ref). In Figure (ref), theory types were extremely heterogeneous: bad theory types cannot eliminate any ideas, while good theory types eliminate the worst 98% of ideas. This extreme-heterogeneity model is shown in the right most markers of Figure (ref). For this model, the improvement from post hoc theory is a negative 30%: i.e. published ideas have 30% lower quality under post hoc theory (top panel). Correspondingly, Darwinian Learning is very large, and far exceeds Statistical Learning (bottom panel).
However, reducing the heterogeneity of theories leads to post hoc theory being optimal. Moving from right to left in Figure (ref), the improvement from post hoc theory turns positive once good theory types can eliminate the worst 75% of ideas. Here, Statistical Learning is exactly equal to Darwinian Learning (bottom panel). For models with any less heterogeneity, post hoc theory is optimal.
Another implication of Proposition (ref) is that larger datasets tend to imply post hoc theory is optimal. Naturally, larger datasets imply more Statistical Learning.
To model this, one can think of measured quality $\hat{\mu}_i$ as a t-statistic, in which case a large dataset implies high $\Var(\hat{\mu}_i)$. Intuitively, as the sample size increases, so does the probability of finding statistically-significant t-stats (abadie2020statistical).
To formalize this interpretation, suppose that underlying Equation (ref) is a panel data model:
where $M$ is the number of observations for idea $i$. Moreover, suppose we fix the hurdle for readers' interest at $h=2.0$ (see Equation (ref)). Then a natural way to map Equation (ref) to Equation (ref) is to define $\hat{\mu}_i$ as the t-statistic for $\chi_i$:
where $\Var_{i}(\varepsilon_i)$ is the variance holding fixed the idea, and Equation (ref) assumes that the central limit theorem holds and $\sigma^2$ is observed. Thus, in this setting, the standard deviation of $\hat{\mu}_i$ is increasing in the sample size $M$.
Figure (ref) illustrates how large datasets affect optimal theorizing, interpreted through the panel data model (Equations (ref)-(ref)). It revisits the model from Figure (ref), but examines alternative choices for the variance of $\mu_i$. The x-axis plots $\sqrt{\Var(\hat{\mu}_i)}$, which can be interpreted as either the dispersion of t-statistics or a measure of the sample size.
The left-most markers correspond to the model from Figure (ref). This model was selected to illustrate the power of Darwinian selection. Thus, $\Var(\hat{\mu}_i)$ is close to 1.0, indicating the measured quality is close to the null distribution, and noise dominates the data. Thus, Statistical Learning is small, and a priori theorizing is optimal.
But as $\Var(\hat{\mu}_i)$ increases, so does the amount of signal, holding fixed $\Var(\varepsilon_i)$ at 1.0. The amount of Statistical Learning then increases, and post hoc theorizing, starts to become optimal at approximately $\sqrt{\Var(\mu_i)}=1.75$.
$\sqrt{\Var(\mu_i)}=1.75$ relatively small. For comparison, chen2020publication and jensen2023there estimate $\sqrt{\Var(\mu_i)}\approx 3.0$ for empirical asset pricing (see discussion in chen2022publication). For settings like this, where $\hat{\mu}_i$ provides a strong signal about the underlying $\mu_i$, Statistical Learning most likely exceeds Darwinian Learning, and thus post hoc theorizing is typically optimal.
As a field of research matures, institutions arise that standardize the many aspects of research, including the peer review process, the statistical analysis, and theory. It is reasonable to think, then, that mature fields have theories that are relatively homogeneous in quality. In fact, homogeneous theory quality is a reasonable definition of a mature field.
Economics is arguably mature. Before the 1950s, there was wild variety in the way that economists theorized. But theory began to solidify with the contributions of Arrow and Samuelson. And though behavioral economics has risen in popularity in recent decades, and the 2008 financial crisis brought on significant criticism of economic models, the basic structure of theory has been largely stable since the 1980s. It is thus reasonable to think that economic theories are fairly homogeneous in quality, and that Darwinian Learning is small.
At the same time, the modern era has seen the rise of huge datasets and enormous computing power. As discussed in Section (ref), this implies that standardized measures of idea quality are dispersed, and thus Statistical Learning is large.
Taken together, these arguments imply that post hoc theory is typically optimal in the modern era of economics.
This argument has some surprising implications. Pre-analysis plans should not be followed. Journals should favor theories that accommodate the data, post hoc. At least, these are the prescriptions for a literature that focuses on finding the best ideas, and places less emphasis on unbiasedness.
While it may feel uncomfortable to favor results over unbiasedness, this is precisely the approach taken by the computer science literature. Following this practical route, computer science has essentially taken over machine learning, which could have been the territory of statisticians. Perhaps the maturation of statistical theories, as well as the rise of big data, tilted the balance in favor of post hoc theorizing, and thus the dominance of computer scientists.
This paper presents a framework for understanding several questions about the scientific method: Why is post hoc theorizing viewed as a problem? How do we square this problem with highly-successful post hoc theories? Does the classical view of post hoc theory still hold up in the modern era of big data?
The framework shows that the distrust of post hoc theorizing is to a significant extent a relic of idealized, pre-modern statistics. With practical constraints on researchers' time, and a focus on results over unbiasedness, a priori theorizing is not always superior. Instead, there is a trade-off between Darwinian Learning, which comes from forcing theorists into prediction contests, and Statistical Learning, which arises as researchers learn from data. With modern datasets and computing power, Statistical Learning is clearly very significant. At the same time, it is unclear that Darwinian Learning still matters, in a world of mature theories.
A caveat is that a priori theorizing has benefits that are omitted from my analysis. Most important, barnes2008paradox points out that prediction contests provide an accessible, democratic way to establish what is good science. The main alternative is the peer review process, which is inscrutable to outsiders, and can potentially be abused.\footnote{Additionally, KLS argue that the choice of a priori vs post hoc theorizing may be endogenous, which can lead to additional selection effects, over and above Proposition (ref). However, the basic logic that a priori theorizing helps through inducing selection is still captured by Proposition (ref).}
A second caveat is that none of this analysis matters if economic theories are not, eventually, tested with post-research data. If economic theories are really not falsifiable, then the value of a priori and post hoc theorizing is an unscientific question, and thus beyond the scope of this paper.