EconBase
← Back to paper

Competing Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

81,172 characters · 14 sections · 85 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Competing Models

\pretitle{

flushleft} \posttitle{

} \preauthor{

flushleft} \postauthor{

} \predate{

flushleft} \postdate{

} \settowidth{\thanksmarkwidth}{*} {-\thanksmarkwidth}

\renewenvironment{abstract} { {{\abstractname}\nobreak}}

abstractDifferent agents need to make a prediction. They observe identical data, but have different models: they predict using different explanatory variables. We study which agent believes they have the best predictive ability---as measured by the smallest subjective posterior mean squared prediction error---and show how it depends on the sample size. With small samples, we present results suggesting it is an agent using a low-dimensional model. With large samples, it is generally an agent with a high-dimensional model, possibly including irrelevant variables, but never excluding relevant ones. We apply our results to characterize the winning model in an auction of productive assets, to argue that entrepreneurs and investors with simple models will be over-represented in new sectors, and to understand the proliferation of “factors” that explain the cross-sectional variation of expected stock returns in the asset-pricing literature.

\thispagestyle{empty}

\setcounter{page}{1}

Introduction

The value that individuals assign to a choice often depends on how well they believe they can predict unknown variables. How much an entrepreneur is willing to pay for a company, or whether they choose to enter a new market, depends on their belief in their own ability to predict and respond to future conditions, like market demand, costs, and competition. Households are more likely to invest in financial assets if they believe they can predict future market values.

In this paper, we study how individuals' assessments of their own predictive ability interacts with the models they use and the available sample size. Our agents are Bayesian and observe the same data, but their predictions are based on different models: some agents believe only a few covariates matter for predictions, while others believe that many more do. We ask: What are the characteristics of the model of the agent who, after observing the data, believes they have the best predictive ability, as measured by the smallest subjective posterior mean squared prediction error (subjective MSPE)? Colloquially, a candidate who believes they have the best predictive ability may be described as the most “confident.'’ Similarly, if the subjective MSPE is below the (unknown) objective MSPE of their model, they may be described as “overconfident.'’ In what follows, we refer to the agent's assessment by subjective MSPE, but in our applications we expand on this confidence/ overconfidence interpretation to deliver novel implications to the behavioral literature on overconfidence.

We show that the answer depends on the model's dimension and the sample size. With small samples, agents with the smallest subjective MSPE use a low-dimensional model, using only a few covariates, regardless of the true data generating process (DGP). In contrast, with large samples, agents with the smallest subjective MSPE use a high-dimensional model, possibly including irrelevant covariates, but never excluding relevant ones. In single-agent decision problems, this results in novel comparative statics: the dimension of agents' models and the dataset's sample size influence the value they assign to each action, holding fixed other standard considerations (e.g., risk aversion, outside options, etc.). In settings where agents compete and relative subjective expected prediction error matters, model dimension and sample size determine the winning model.

\paragraph{Our model.} As a concrete example, consider a second-price auction where a productive asset is sold to the highest bidder. The new owner of the asset will choose an action $a$, and her payoff will be given by $M - (a-y)^2$, where $M$ is a known positive quantity and $y$ is a random variable. Thus, the value of the asset depends on how well the agent can predict $y$.

There are multiple interesting economic issues in such a setting. Bidders may have different payoff functions, different actions sets, or different information. We abstract from all those issues and focus on the impact of using priors that involve a simpler as compared to a more complex relationship between explanatory variables and the variable of interest $y$. Specifically, suppose all agents agree that $y$ is a linear function of a number of covariates $\{x_j\}_{j \in \{1, \dots, k\}}$ plus a noise term, i.e., $y = \sum \beta_j x_j + \epsilon$. Both the $\beta_j$'s and the variance of $\epsilon$ are unknown, and agents may have different prior distributions on them. In particular, some agents may believe that only a subset of the covariates matters for predicting $y$.

All agents are given the same data: $n$ independent draws of $x$ and $y$, according to an unknown process. Agents are Bayesian but have different priors---as in harrison1978speculative or morris1994trade. Thus, each agent computes a posterior distribution of $\beta_j$'s and the variance of $\epsilon$, and will use these posterior distributions to solve for their optimal action. Notice that there is no winner's curse in this setup, as winning the auction has no effect on the winner's posterior distribution. Therefore, the auction has the usual equilibrium in dominant strategies where each agent bids her expected value of the asset if she becomes the owner and gets to choose $a$. The winner is the agent with the lowest subjective MSPE. As everything else is equal, this is a competition among models. We ask: What are the characteristics of the model that, after observing the data, has the lowest subjective MSPE?

Note that there is a trivial reason why certain models may have lower subjective MSPE: their priors may contain less uncertainty about the world. The most extreme case is when an agent is dogmatic and she has a (right or wrong) deterministic model. That agent would believe she has no prediction error and would always bid $M$ in the auction described above. To focus on more interesting effects, we prove all our results under the assumption that, absent data, all agents have the same expected loss.

\paragraph{Results.} Our first result, Lemma (ref), characterizes the subjective MSPE of an agent as a function of their prior and observed data. We prove a subjective/ Bayesian variant of a standard decomposition to show that subjective MSPE can be written as the sum of two components, which we term 1) model fit: the agent's posterior expectation of the variance of the regression residual, $\epsilon$; and 2) model estimation uncertainty: the agent's degree of uncertainty about the coefficients in her regression model. Crucially, we show that the latter depends on the model's dimension. This implies that, while our Bayesian agents use their posteriors to compute the best action and do not explicitly care about the dimension of their model, the dimension affects their subjective MSPE.

This characterization has two immediate implications, depending on the size of the dataset. Our first set of results pertains to the case of small samples. Here we show that “model estimation uncertainty” plays a critical role. While agents who use only a few covariates may have a lower model fit, they will also have a lower model estimation uncertainty, since they have fewer parameters to estimate. One complication with small samples is that the actual realized dataset matters, not just the agent's prior. Assuming that priors take a convenient conjugate form typical in Bayesian linear regression, we show how small sample sizes favor small models---even assuming that all agents have the same prior expectation about their prediction error. First, Proposition (ref) shows that when the dataset consists of a single datapoint, the lowest subjective MSPE is of a model that contains a single covariate (regardless of the realized data, the true DGP, or the parameters of the prior). Second, Proposition (ref) shows that for any fixed sample size and for any true data generating process, lower-dimensional models have a lower subjective MSPE with high probability, as long as the prior on the variance of $\epsilon$ is high enough. Intuitively, with small samples, uncertainty about the parameters of the model plays a crucial role. Smaller models have an advantage since uncertainty about the parameters decreases faster as data accumulates. Even if these smaller models are misspecified, if the sample is small their model fit will not be much lower, meaning that they will have the lowest subjective MSPE. Third, we show that smaller models also have a smaller subjective MSPE when agents all believe they know the variance of the error and this variance is common. Finally, we show that this is also true when agents need to compute their future expected subjective MSPE before data is realized, but knowing that a data set of size $n$, for any $n$, is going to be revealed before the action is chosen.

Next, we consider the case of large sample size. Here, model estimation uncertainty vanishes: agents will have no uncertainty about their fitted parameters, even if they are using the wrong model. Subjective MSPE is therefore based solely on model fit. Proposition (ref) then shows that models that omit a covariate that is relevant for prediction never prevail. At the same time, we also show that high-dimensional models---those that contain additional covariates irrelevant to the true DGP---may continue to win, even asymptotically. Even though these high-dimensional models will converge to the true DGP, for any finite sample they remain strictly different. We show that the probability of winning for high-dimensional models remains strictly above zero, even asymptotically. In turn, this shows that the role of priors does not vanish asymptotically: it continues to affect a model's probability of winning, even with arbitrarily large samples.

\paragraph{Applications.} In the main body of the paper, we discuss the case of an auction of a productive asset as a leading example. In Section (ref), we present two additional applications. First, we consider a simple model of entry in which returns depend on prediction error: for example, the decision of an entrepreneur to enter a new sector, or the decision of a household to invest in a risky asset. Conditional on entry, the agent must make further choices. For example, as in classic organizational economics models, the entrepreneur must choose a strategy $a$ that fits the (predicted) state of the world $y$ and the loss function is the quadratic difference between $a$ and $y$, as in, e.g., marschak1972economic, roberts1992economics or alonso2008does. Alternatively, the investor must predict price movements to profitably buy/sell the asset. Agents may have simple or complex models: they make their forecasts using few or many covariates). The results of this paper provide a novel comparative static: in new sectors/asset classes, we should observe an over representation of investors with “simple” models, even when reality is “complex.” We connect this to the literature on overconfidence of entrepreneurs and investors.

Second, we use our framework to understand the proliferation of “factors” that explain the cross-sectional variation of expected stock returns in the asset-pricing literature. We argue that the increase in the number of test portfolios used to compute the popular Fama-French cross-sectional regressions mechanically favors asset-pricing models with several factors. Our empirical analysis on the evolution of the “factor zoo” can be viewed as a particular instance of a simple model of scientific progress. Early models (when there is little data available) may be overly simple relative to the truth, and become more complicated as samples accumulate.

\paragraph{Related Literature.} Our results sit within the large and growing body of work in economic theory on agents with misspecified models; we defer a full discussion to Section (ref). One key difference is that in most of the literature, misspecified models are evaluated using their objective performance. Our paper, along with a few contemporaneous or subsequent ones eliaz2018model, levy2019misspecified, he2020evolutionarily, focuses instead on agents' subjective perception of their prediction error, the key metric in our applications.

Our results may also, at a high level, be reminiscent of model-selection methods in Statistics and Machine Learning, with one big difference: our results emerge as the outcome of competition among Bayesian decision makers using different models. By contrast, the model selection literature proposes and studies techniques to explicitly penalize high-dimensional models. The Bayesian statisticians in our paper cannot discard covariates.

\paragraph{Outline.} The remainder of the paper is organized as follows. Section (ref) outlines the formal model, and characterizes the subjective MSPE of a single agent, the foundation of our results. Section (ref) illustrates the key trade-offs under competition with a simple numerical simulation. Section (ref) then presents formal results for the case where the size of the dataset, $n$, is small, while Section (ref) considers the case where $n$ is large. Section (ref) studies the applications described above. Section (ref) discusses the related literature, and Section (ref) concludes.

Model and Single-Agent Problem

Agents want to predict a real-valued variable $y$. There are $k$ real-valued covariates (or explanatory variables) $x \in \mathbb{R}^k$.

\paragraph{Data and Data Generating Process.} Before making a prediction, agents observe a common data set, denoted $D_n$, composed of $n$ i.i.d. draws of $y$ and $x$. We denote the data as $D_n = (Y,X)$, where $Y \in \mathbb{R}^n$ and $X \in \mathbb{R}^{n \times k}$. The assumption that all agents observe the same data will be relevant for our applications (for example, in an auction setting, this avoids winner's curse).

A true Data Generating Process (DGP), denoted $\mathbb{P}$, determines the joint distribution of the random variables $y$ and $x$. Most of our results assume only that true distribution of covariates has finite moments of all orders (and a positive definite matrix of second moments).

\paragraph{Statistical Models.} Agents do not know $\mathbb{P}$ but work with a statistical model: a family of plausible joint distributions for $y$ and $x$. In particular, agents posit a linear relation between $y$ and the covariates $x \in \mathbb{R}^k$, i.e., conditional on $x$:

equation[equation omitted — 169 chars of source]

that is, agents assume $y|x$ is a homoskedastic linear regression with Gaussian errors and parameters $(\beta,\sigma^2) \in \mathbb{R}^k \times \mathbb{R}_+$.\footnote{Because covariates in $x$ can be correlated, our framework allows the agents to consider a wide family of non-linear relations. For example, the non-linear process $y=3\frac{x^3_1}{\sqrt{x_5}}+\epsilon$, can be accommodated by defining a new observable equal to $\frac{x^3_1}{\sqrt{x_5}}$. While not all non-linear processes can be expressed this way, especially since we assume finitely many covariates, good approximations can always be achieved. }$^{,}$\footnote{We use the notation $\mathcal{N}_{k}(\mu,\Sigma)$ to denote a multivariate normal distribution of dimension $k$ with mean $\mu$ and covariance matrix $\Sigma$. See p. 171 of hogg for a textbook reference on this convention.}

Agents also assume that the covariates follow a distribution $P$ which belongs to some parametric family.\footnote{A family of distributions $\mathcal{P}$ is said to be parametric, if its elements are indexed by a finite-dimensional vector. One example is $x \sim \mathcal{N}_k(0,\Sigma)$, where $\Sigma$ is an unknown positive definite matrix.} This may or may not be the correct distribution. We assume that, for any element in this family, the matrix $E_{P}[xx']$ is positive definite and that the random vector $x$ has finite moments of all orders.\footnote{Positive definiteness of the matrix rules out the case in which one covariate is a linear combination of some of the others.}

Together, $P$, $\beta$, and $\sigma^2$ fully define a joint distribution over $y$ and $x$, which we denote by $Q_{\theta}$, with parameter $\theta : = (\beta,\sigma^2,P)$.

\paragraph{Different Explanatory Variables.} Different agents may consider different explanatory variables in $x$ as relevant for their prediction. We assume that agents consider at least one explanatory variable in their models. The following notation will be useful. If $\{1,2\ldots, k\}$ label the explanatory variables in $x$, we denote by $J \subseteq \{1,2\ldots, k\}$ the subset that an agent considers possibly relevant for prediction. For a given vector $\beta$, the subvector consisting solely of the components in $J \subseteq \{1,\ldots, k\}$ denote by $\beta_J$. Let $x_J$ be the analogous subvector of $x$, and $X_J$ the corresponding submatrix of $X$.

\paragraph{Misspecification.} An agent who considers the variables in the set $J$ has the statistical model $\{ Q_{\theta_J} \}_{\theta_J \in \Theta_{J}}$. Here $\Theta_J$ is the set of parameters corresponding to variables $J$, i.e., $\theta_J:= (\beta_J, \sigma^2, P).$ This model is said to be misspecified if there is no $\theta_J \in \Theta_J$ for which $Q_{\theta_{J}} = \mathbb{P}$ (and it is correctly specified otherwise). In words, a statistical model is misspecified if it does not contain the true DGP kleijn2012bernstein. When the true DGP $\mathbb{P}$ is also a Gaussian linear regression model as in (ref) (which we will assume for some of our results), let $J_0$ denote the covariates with non-zero $\beta$s in the true DGP. Then, note that the model associated with any set of variables $J$ for which $J_0 \not\subseteq J$ is necessarily misspecified.\footnote{At first glance, it might seem reasonable to say that a model $y=\beta_1 x_1 + \epsilon$ need not be misspecified when the model is truly $y=\beta_1 x_1 + \beta_2 x_2 + \epsilon'$, as long as the distributions of $\epsilon$ and $\beta_2 x_2 + \epsilon'$ coincide. But this is ruled out by the assumption that error distributions are restricted to be Gaussian, centered at zero, and independent of covariates.}

\paragraph{Priors.} Agents are Bayesians. An agent who considers variables $J$ as relevant for prediction has a prior $\pi$ over $\Theta_J$. It will be convenient to denote by $J(\pi)$ the set of variables that an agent with prior $\pi$ considers relevant for prediction. Formally, let $\pi_j$ denote the marginal distribution over $\beta_j$ corresponding to prior $\pi$. If $\delta_0$ denotes a Dirac measure at zero, then

align*[align* omitted — 79 chars of source]

In a slight abuse of terminology, we sometimes use $J \subseteq \{1, \dots, k\}$ to refer to a model, which should be understood as the set of explanatory variables that are not exactly equal to zero under the prior $\pi$. Strictly speaking, though, a statistical model refers to the collection of distributions over data given parameters as we have defined above; see mccullagh2002statistical.

\paragraph{Actions, Utility, and Optimal Prediction.}

Agents make a prediction of $y$ given covariates $x$. Formally, they construct a prediction function $f$ that maps $x$ into $y$, i.e., $f: \mathbb{R}^k \rightarrow \mathbb{R}.$ They minimize a standard quadratic loss function, equal to the square of the difference between the true $y$ and their forecast $f$, i.e., $(y-f)^2.$ Denote by $L (f,\theta)$ the agent's loss under prediction function $f$ if the true DGP is $Q_{\theta}$, i.e.,

align[align omitted — 89 chars of source]

If $\pi$ is the agent's prior over $\theta$ and $D_n$ is the observed data, then characterizing the optimal prediction $f^{*}$ is a standard problem. The agent chooses $f$ to minimize $\mathbb{E}_{\pi}[L (f,\theta)|D_n]$, which can be rewritten as

equation[equation omitted — 156 chars of source]

The first term does not depend on $f$. The second term involves the average error incurred in predicting $x^{\prime} \beta$ using $f(x)$.\footnote{The inner expectation averages over values of $x$. The outer one averages over the values of $\beta$ and $P$.} With standard arguments (i.e., exchanging the order of integration and taking first-order conditions), we can see that this is minimized by

equation[equation omitted — 165 chars of source]

Thus, a Bayesian decision maker with a posterior $\pi|D_n$, model $J(\pi)$, and a square loss function, forecasts $y$ at $x$ as her Bayesian posterior mean of $x'\beta$. This is a standard result.

The agent's posterior loss, conditional on her using the optimal prediction function characterized above, is denoted $L^*(\pi, D_n)$. We refer to it as the subjective posterior mean-squared prediction error (subjective MSPE).

Decomposing Subjective MSPE

We now characterize an agent's subjective MSPE, our key dimension of interest. The key forces at play will already be evident from the following lemma.

restatable{lemma}{oneagentposteriorloss} Suppose that $\beta_{J(\pi)}$ is independent of $P$ under the posterior distribution. The agent's subjective MSPE can be decomposed as: \begin{align} &L^*(\pi, D_n) = \underbrace{\mathbb{E}_{\pi} \left[\sigma^2 | D_n\right]}_{Model Fit} + \underbrace{tr\left(\mathbb{V}_{\pi}[\beta_{J(\pi)}|D_n] \; \mathbb{E}_{\pi} \left[ \mathbb{E}_{P}[{x}_{J(\pi)}{x}_{J(\pi)}'] \: | \: D_n \right] \right)}_{Model Estimation Uncertainty}, \end{align} where $\mathbb{V}(\cdot)$ is the variance-covariance operator, and $\textup{tr}$ is the trace operator.

This lemma is reminiscent of standard decompositions of mean-squared prediction error in frequentist linear regression models hansen2020econometrics, except that in this case it characterizes the subjective MSPE of the agent using their own prior. The lemma shows that the agent's subjective MSPE, $L^*(\pi,D_n)$, is the sum of two components. The first, the posterior expectation of $\sigma^2_{\epsilon}$, is the agent's estimate of the irreducible noise in the system. We interpret this term as a measure of model fit, i.e., how well the model explains the data (as all unexplained variation must be ascribed to noise).

The second term, $\textup{tr}\left(\mathbb{V}_{\pi}[\beta_J|D_n]\; \mathbb{E}_{\pi} \left[ \mathbb{E}_{P}[x_Jx_J'] \: | \: D_n \right] \right)$, is the trace of the variance-covariance matrix of the coefficients of the model (adjusted by the posterior mean of $\mathbb{E}_{P}[x_Jx_J']$). We interpret this term as a measure of how uncertain the agent is in her estimation of the parameters of the model according to her own prior, capturing model estimation uncertainty. To illustrate why, consider the simpler case in which the posterior mean of $\mathbb{E}_{P}[x_J x_J']$ is the identity matrix. Then, the second term reduces to $\textup{tr}\left(\mathbb{V}_{\pi}[\beta_J|D_n]\right)$, i.e., $\sum_{j\in J} \mathbb{V}_{\pi}[\beta_j|D_n]$; this is simply the sum of the posterior variances of the parameters $\beta_j$, indeed a measure of model estimation uncertainty. In the next section we will show that this decomposition has immediate implications for which model leads to the lowest subjective MSPE.

remarkThe independence of $\beta_J$ and $P$ under the posterior distribution will hold under general assumptions. For example, it holds when agents know the distribution of covariates, or if we consider statistical models in which $P$ does not enter the parametric model of $y|x$ and $(\beta,\sigma^2)$ does not affect the distribution of $x$.

Competing Models: A First Look

Suppose agents participate in a mechanism that selects the agent with the lowest subjective MSPE. A leading example, discussed in the introduction, is that of a second-price auction of a productive asset. In this scenario, the winner of the productive asset will choose an action $a$ and her payoff will be given by $M - (a-y)^2$, where $M$ is a known positive quantity and $y$ is the variable agents aim to predict. Thus, the value of the asset depends on how well the agent can predict $y$. Assuming that different priors are the only dimension of heterogeneity among participants, and noting that there is no winner's curse, the auction will select the agent with the lowest subjective MSPE. Applying Lemma (ref), we immediately derive that this is the agent with the best trade-off between model fit and model estimation uncertainty.

Before we dive into formal results, we present a simple simulation to illustrate the key forces at play and our main findings. Suppose that there are six covariates, $\{x_1, \ldots , x_6\}$, of which only the first five are relevant for prediction in the true DGP, i.e., $y = \sum_{j=1}^5 \beta_j x_j + \epsilon$, and $\epsilon | x \sim \mathcal{N}(0,\sigma^2)$. For simplicity, assume that each nonzero regression coefficient, $\beta_j$, is equal to $1$, as is $\sigma^2$. Also assume that $x \sim N(0,\mathbb{I}_6)$ under the true DGP.

Considering all subsets of covariates, there are $63$ agents with linear regression models, one for each nonempty subset of $\{x_1, \ldots , x_6\}$. By construction, 61 are misspecified, one has the exactly correct model, and one has a model of higher dimension compared to the true DGP. For the simulation, we assume that agents' priors belong to a well-behaved family, common in Bayesian linear regression, and parametrize them so that all agents have the same subjective MSPE before seeing any data.\footnote{Specifically, the agents are assumed to have the following parametric model for their covariates: $x_J \sim \mathcal{N}_{|J|}(0, \Sigma_J)$ where $\Sigma_J$ is an unknown, positive definite matrix. Further, we assume that the priors over the $\beta$,$\sigma^2$ and $\Sigma_J$ belong to the Normal-Inverse-Gamma-Inverse-Wishart family of Definition (ref) below. In this simulation we set $a_0=2$, $b_0=10$, and $\gamma_0 = .001$.}

Figure (ref) plots the frequency of the size of the model of the agent with the lowest subjective MSPE for datasets of size $n \in \{1, \ldots, 50\}$. Two patterns emerge.

First, when $n$ is small, low-dimensional models tend to win despite being misspecified. In fact, when $n=1$, we can see in Figure (ref) that the winner is a model with a single covariate.

figure[figure omitted — 199 chars of source]

Second, as $n$ grows large, misspecified models never win. At the same time, the high-dimensional model that includes the redundant variable $x_6$ continues to win with relative frequency that appears to converge to a steady state close to 0.3---strictly above 0.

We will now show that both patterns hold more generally.

The Winner with Small $n$

We begin with the case in which the number of observations $n$ is small. For tractability, we focus on a special class of priors, widely used in Bayesian linear regression.

definitionA prior has the Normal-Inverse-Gamma-Inverse-Wishart form with hyper-parameters $(a_0, b_0,\gamma_0)$, with $(a_0, b_0,\gamma_0) \gg 0$ and $a_0 > 1$, if \begin{align*} \beta_{J} | \sigma^2 \sim \mathcal{N}_{|J|}\Bigg(0, \frac{\sigma^2}{\gamma_0 |J|} \mathbb{I}_{|J|} \Bigg), \ \ \ \ \ \ \sigma^2 \sim Inv-Gamma(a_0, b_0), \\ x_{J} \sim \mathcal{N}_{|J|}(0,\Sigma_J), \ \ \ \ \ \ \Sigma_{J} \sim Inv-Wishart(\gamma_0 |J| \mathbb{I}_{|J|}, 2|J|+1) \end{align*} where $\textrm{Inv-Gamma}(a_0,b_0)$ is the Inverse-Gamma distribution with parameters $a_0$ and $b_0$, and Inv-Wishart($\gamma |J| \mathbb{I}_{|J|}$,$2|J|+1$) is the Inverse-Wishart distribution with $J \times J$ scale matrix $\gamma |J| \mathbb{I}_{|J|}$ and degrees of freedom $2|J|+1$.\footnote{The Inverse-Gamma is a two-parameter family of distributions on the positive real line. For parameters $a_0>1, b_0 \geq 0$ it has mean $b_0/a_0-1$. The Inverse-Wishart is a two-parameter family of distributions on real-valued positive definite matrices. The first parameter (scale matrix) is a symmetric positive definite matrix. The second is a non-negative scalar at least as large as the dimension of the scale matrix.} The prior on $\Sigma_J$ is assumed independent of $(\beta, \sigma^2)$.

The Normal-Inverse-Gamma-Inverse-Wishart priors are conjugate priors for the Gaussian linear regression model and the posterior can be expressed in a closed form as a function of the data---see Appendix (ref) for details. All results in this section have analogs in the setting where the distribution $P$ on covariates is assumed to be known (not necessarily Gaussian) and $E_{P}[xx']= \mathbb{I}_k$.

One convenient feature of this family of priors, readily checked, is that they imply that all agents have the same subjective MPSE when no data is released, i.e., when $n=0$. Differences in subjective MSPE arise therefore only from the fact that the subjective MSPE evolves differently for models of different dimensions.

Using this family of priors allows us to further simplify the expression for subjective MSPE in Lemma (ref).

restatable{lemma}{unknownP} Consider a prior $\pi$ as in Definition (ref). Then \begin{align} L^*(\pi, D_n) = \underbrace{\mathbb{E}_{\pi}[\sigma^2|D_n]}_{Model Fit} + \underbrace{\mathbb{E}_{\pi}[\sigma^2|D_n] \left( \frac{|J(\pi)|} { n + |J(\pi)|} \right). }_{Model Estimation Uncertainty} \end{align}

Equation (ref) shows that the dimension of an agent's model enters explicitly into the subjective MSPE via the term $|J(\pi)| \: / \: ( n + |J(\pi)|).$ As this is increasing in $|J(\pi)|$, higher-dimensional models have a disadvantage: as they have more parameters to estimate, their model estimation uncertainty will decrease more slowly. For higher-dimensional models to have lower subjective MSPE, therefore, they must compensate for this by a sufficiently better model fit. Specifically, the ratio between the model fits must be such that

align[align omitted — 256 chars of source]

In general, however, characterizing the model fit is analytically difficult, as it depends also on the realized data $D_n$, which in small samples can vary substantially. The following results describe how lower-dimensional models have lower subjective MSPE in some cases in which model fit is simple to analyze. First, when the sample size $n=1$; we this to be the case for a model using a single covariate. Second, it holds true when prior variance of $\epsilon$ of the models is high enough relative to $n$. Third, it applies when either all agents believe they know the variance of the error term, $\sigma^2$, and have the same belief, or when subjective MSPE is computed before actual data is released but the agents know that it will be released before they must choose their action---a case that has direct applications (as we will discuss later).

One-Dimensional Models Win when $n=1$

The first result shows that when $n=1$, low-dimensional models have the lowest subjective MPSE, regardless of the realized data, the true DGP, or the parameters of the prior.

restatable{proposition}{onedataunknownP} Suppose all priors are as in Definition (ref) with shared hyper-parameters. Suppose also that for every single covariate model, i.e., every $J$ such that $|J|=1$, there is an agent with that model, and that all agents use at least one covariate. If $n=1$, then the agent with lowest subject MSPE is an agent with a single covariate model.

We prove this result by showing that, when $n=1$, model fit is minimized by some model that considers only a single variable. This means that there is no model with more than one covariate that can improve the model fit of the best one-dimensional model.

Low-Dimensional Models Win when Prior Variance Is High

The sharp characterization obtained for $n=1$ does not hold for other sample sizes. Indeed, our simulations show that models with more than one covariate may have the lowest subjective MSPE with positive probability for $n>1$. As we have discussed, the identity of this model depends on the trade-off between model fit and model estimation uncertainty.

One way to capture the advantage of small-dimensional models when $n$ is small is the following. One characteristic of few data points is that the prior continues to play a relevant role. This can be captured by making sure that the prior mean is large enough relative to the amount of data. In this case, high-dimensional models cannot improve sufficiently on the model fit term relative to lower-dimensional models (as the model fit is poor for any model). Therefore, in this setting also, the advantage that low-dimensional models have in terms of model uncertainty is the dominant factor.

restatable{proposition}{highprobwinnerUP} Let $\Pi$ be a finite set of agents' priors that satisfy Definitions (ref) with shared hyper-parameters $(a_0,b_0, \gamma)$. Let $|\underline{J}|$ be the size of the smallest model in this set. Fix the size of the dataset $n$. For any $p \in (0,1)$, there exists $b_0$ large enough so \begin{align*} \mathbb{P}\left(D_n : \exists \pi^* \in \operatornamewithlimits{argmin}_{\pi \in \Pi} L^*(\pi,D_n) s.t. |J(\pi^*)| = |J| \right) >p, \end{align*} i.e., with probability at least $p$ over datasets $D_n$, the agent with the lowest subjective MSPE has the smallest size model among all the agents.

As an illustration of this result, let us return to the simulations in Section (ref). Figure (ref) reports the winning fraction of models of size $1$, as we increase the shared hyper-parameter $b_0$ (all other simulation parameters stay the same as above). Growing $b_0$ corresponds to a larger prior mean for all agents.

figure[figure omitted — 300 chars of source]

Low-Dimensional Models Win when Error Variance Is Known or Data Are Not Released but Expected

We conclude this analysis considering two other cases in which the model fit is easy to solve analytically, which again show an advantage of lower-dimensional models. We present them as two observations since they follow directly from Lemma (ref) above.

observationSuppose all agents treat $\sigma^2$ as a known common value, but the priors on $\beta|\sigma^2$ and $\Sigma$ are as in Definition (ref). Then, for any $D_n$, $$ L^*(\pi', D_n) < L^*(\pi,D_n) \ \Longleftrightarrow \ J(\pi')|< |J(\pi)|. $$ That is, if $\sigma^2$ is believed to be known (possibly incorrectly), subjective MSPE is ranked by model dimension.

This result shows that when all agents believe they know the error variance $\sigma^2$, then model dimension induces a precise ranking between models: smaller models always have smaller subjective MSPE. This result follows directly from Lemma (ref).

We now turn to the case in which agents have not received any data, but know that they will receive $n$ data points before making their choice of action. They therefore have to compute their expected subjective MSPE, which we denote by $\mathbb{E}_{\pi'}[L^*(\pi', D_n)]$.

observationFor any $\pi$ and $\pi'$, $$ \mathbb{E}_{\pi'}[L^*(\pi', D_n)] < \mathbb{E}_{\pi}[L^*(\pi,D_n)] \ \Longleftrightarrow \ J(\pi')|< |J(\pi)|. $$

This result shows that if agents have not received any data but know that they will receive it later, again we find small-dimensional models have lower expected MSPE. This observation also follows straightforwardly from Lemma (ref): by the martingale property of beliefs, the expected model fit is constant among models, meaning that the winner must be a low-dimensional model.

The Winner with Large $n$

We now characterize the winner for large $n$. Our results will be derived for a much more general class of priors than the previous section, but it is helpful to start by recalling Lemma (ref), which assumes Normal-Inverse-Gamma-Inverse-Wishart priors. In this case, the subjective MSPE is \[ \underbrace{\mathbb{E}_{\pi}[\sigma^2|D_n]}_{\textrm{Model Fit}} + \underbrace{\mathbb{E}_{\pi}[\sigma^2|D_n] \frac{|J(\pi)|}{n + |J(\pi)|)}}_{\textrm{Model Estimation Uncertainty}}.\] From this formula, it is immediate to see that model estimation uncertainty vanishes as $n$ grows large, making the model fit the crucial aspect. This means that, because misspecified models have worse model fit than correctly specified ones, they must therefore also have worse subjective MPSE when $n$ is large enough.

The comparison is, however, less straightforward between a model that uses the exact same variables as the true DGP and another that also includes additional irrelevant covariates. For both, model fit converges to the true residual variance ($\sigma^2_0$), since both are correctly specified, and model estimation uncertainty converges to zero. Which one has lower subjective MSPE depends on how quickly these converge, which in turn depends on the realized data and on the prior. Our simulations suggest that the long-run behaviors may be such that the larger model may continue to win, with a probability bounded away from zero even at the limit. We will now show how this holds in general.

For this analysis, we do not need to assume that priors have a specific form as we did in the previous section. We simplify our analysis in the body of the paper by making two assumptions: that the true DGP is of the linear Gaussian form, which allows some models to be identical to the true DGP (Assumption (ref)); and that priors over the $\beta_i$s have full support with a smooth density (on the subset of relevant covariates $J(\pi)$), while priors on $\sigma^2_{\epsilon}$ are not degenerate (Assumption (ref)).

assumptionThere exist parameters $\theta_0 := (\beta_0, \sigma_0^2,P_0)$ such that $Q_{\theta_0} = \mathbb{P}$.
assumptionPriors are characterized by a smooth and strictly positive probability density function $\pi(\cdot)$ over $({\beta_{J(\pi)}}^{\prime},\sigma^2)^{\prime} \in \mathbb{R}^{|J(\pi)| }\times \mathbb{R}_+$. The prior over $({\beta_{J(\pi)}}^{\prime},\sigma^2)^{\prime}$ is independent of the prior over $P$.\footnote{By definition, the prior of an agent for any $\beta_{\kappa}$, $\kappa \not\in J(\pi),$ is degenerate at 0.} In addition, for each agent, there exists $n$ large enough for which $\mathbb{E}_{\pi}[\sigma^2 | D_n] < \infty$ almost surely.

Recall that $J_0$ denotes the set of covariates that are relevant in the true DGP.

restatable{proposition}{largenwinnereasy} Suppose the true DGP $\mathbb{P}$ satisfies Assumption (ref) with parameters $(\beta_0, \sigma_0^2)$. Let $\Pi$ be a finite collection of priors that satisfy Assumption (ref) and that contains $\pi^*$ with $J_0 \subseteq J(\pi^*)$. If \begin{align} tr \left( n \mathbb{V}_{\pi}(\beta_{J(\pi)} | D_n ) \mathbb{E}_{\pi} \left[ \mathbb{E}_{P}[ x_{J(\pi)} x_{J(\pi)}' ] \: | \: D_n \right] \right) = O_{\mathbb{P}}(1), \end{align} for every prior $\pi \in \Pi$, then $$\lim_{n \rightarrow \infty} \mathbb{P} \left( \exists \pi \in \operatornamewithlimits{argmin}_{\pi \in \Pi} L^*(\pi, D_n) \textrm{ s.t } J_0 \nsubseteq J(\pi) \right) = 0.$$ Moreover, for any $\pi$ for which $J_0 \subset J(\pi)$, $$ \lim_{n \rightarrow \infty} \mathbb{P}\big( L^*(\pi,D_n) < L^*(\pi_0,D_n) \big) \in (0,1],$$ where $\pi_0$ is any prior for which $J(\pi_0) = J_0$.

A crucial assumption in Proposition (ref) is (ref). This assumption will be verified whenever the posterior variance of $\beta$ decreases to zero at rate $n$. Lemma (ref) already tells us that this condition is satisfied in the special case of Normal-Inverse-Gamma-Inverse-Wishart priors. In fact, this condition holds very generally due to the Bernstein-von Mises theorem, which states that posterior distributions based on parametric models (misspecified or correctly specified) will typically behave like Gaussian distributions, with a variance that decreases at rate $n$.\footnote{See the Bernstein-von Mises theorem for misspecified parametric models of kleijn2012bernstein. This result can be thought of as richer versions of the classical results concerning posterior distributions of misspecified models in berk1970consistency.}

Proposition (ref) has two takeaways. The first part tells us that a misspecified model, because it excludes relevant variables, never wins as the sample size grows large. Any model that is not misspecified will have lower subjective MPSE with probability approaching $1$ as $n$ grows large. The second part shows that any model larger than the true one defeats the latter with a probability that is strictly positive, even asymptotically.

We have already discussed the intuition for the first result. The assumptions of the theorem, i.e. Assumptions (ref) and (ref) combined with (ref), guarantee that model estimation uncertainty converges to zero for all agents.

For the second result, from a technical perspective, our result is based on an asymptotic expansion for the posterior mean of the variance parameter in the linear regression model based on the general results in KTK:1990. This is not a fairly technical result, so, for some intuition, let us return to the case of Normal-Inverse-Gamma-Inverse-Wishart priors. Here, when $n$ is large, it is possible to approximate its distribution using asymptotic theory.

To this end, let $ \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{\beta}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{\beta}{\tmpbox} _J$ denote the OLS estimator based on the variables listed in $J$, and let $ \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{\sigma}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{\sigma}{\tmpbox} ^2_J$ be the corresponding residual variance estimator: \[ \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{\sigma}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{\sigma}{\tmpbox} ^2_J \equiv (Y-X_J' \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{\beta}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{\beta}{\tmpbox} _J)'(Y-X_J' \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{\beta}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{\beta}{\tmpbox} _J)/n.\] A key observation in our analysis---that holds in the Normal-Inverse-Gamma model, but also for more general priors---is that as the sample size grows large

equation[equation omitted — 605 chars of source]

where $O_{\mathbb{P}}(1)$ refers to a term that is bounded with high probability under $\mathbb{P}$. In the case of Normal-Inverse-Gamma-Inverse-Wishart priors, algebraic manipulations can be used to verify the approximation with a leading term equal to \[ -\gamma \beta_0'\beta_0 (|J|-|J_0|).\] Deriving an analogous result for other priors requires additional effort, given the lack of closed-form solutions for the posterior distributions.\footnote{We refer the reader to Lemma (ref), which uses the Kaas-Tierney-Kadane expansions of posterior moments in KTK:1990 to verify the approximation.}

Assumption (ref), along with standard results from regression analysis---e.g., Equation 5.28 in greene2018econometric and Theorem 5.1 therein---implies $n( \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{\sigma}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{\sigma}{\tmpbox} ^2_{J_0}- \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{\sigma}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{\sigma}{\tmpbox} ^2_J)/\sigma^2_0$ converges in distribution to a chi-squared random variable with $|J|-|J_0|$ degrees of freedom. This means that the probability that the larger model wins can be approximated by the probability of the event: \[ \chi^2_{|J|-|J_0|} -\gamma (\beta_0'\beta_0/\sigma^2_0) (|J|-|J_0|) > (|J|-|J_0|). \]

In the context of our simulations---where the true DGP only included the first five covariates, with coefficients $\beta=(1,1,1,1,1)'$, and $\sigma^2=1$---the probability that a model with six variables defeats the true DGP is roughly \[ P(\chi^2_{1} > 1 + 5 \gamma). \] When $\gamma=.001$, this probability is $0.3161$, which is close to what we see in Figure (ref). We provide a more general formula in Appendix (ref).

It is important to remark that these results hold even if the true DGP is different from (ref). For example, the distribution of errors in the true DGP may be heteroskedastic or non-normal with thin-enough tails, the distribution on covariates $P$ may be misspecified as long as it has finite second moments, and the true DGP may not be linear. Although Proposition (ref) is presented under special conditions (Assumptions (ref) and (ref)), it is possible to prove that our results continue to hold under much weaker ones.\footnote{For example, for the expansion of Lemma (ref) to hold, priors do not need to be smooth. It is sufficient that they are differentiable up to the fourth-order.} We hope that the simpler framework helps the reader understand the main forces at play in the competition among models.

\paragraph{Connection with the Akaike Information Criterion.} A different way to understand our results is to relate the model selection they induce to the Akaike Information Criterion (AIC), a well-studied model selection criterion in econometrics and statistics. In what follows, we illustrate that the loss function of an agent with Normal-Inverse-Gamma-Inverse-Wishart prior is “close” to the AIC for the linear regression model.

definition[Akaike Information Criterion] Given a dataset $D_n = (Y, X)$ with $n$ data points and $k$ possible covariates, the AIC for linear regression evaluates a model $J$ as \begin{align*} &L_{Akaike} (J,n, D_n) = \ln \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{\sigma}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{\sigma}{\tmpbox} ^2_J + \frac{2|J|}{n},\\ where & \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{\sigma}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{\sigma}{\tmpbox} ^2_J = \frac1n \min_{\beta \in \mathbb{R}^{|J|}} ({y} - {X} _J \beta)' ({y}- X_J \beta). \end{align*}

The expression $ \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{\sigma}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{\sigma}{\tmpbox} ^2_J$ is the OLS estimator of the residual variance based on a model with covariates $X_J$ in the dataset $D_n$. As is well understood, the model with lower estimated variance may not be the model with the best out-of-sample performance. This is because selecting based on average residuals favors models that have more covariates, which may overfit the data. The AIC compensates for this by adding a penalty term equal to $\frac{2|J|}{n},$ i.e., twice the ratio of the number of covariates in the model and the number of data points. Algebra shows that, if agents have an uninformative Normal-Inverse-Gamma-Inverse-Wishart prior, then the posterior loss is approximately equal to \[ \ln \Bigg( \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{\sigma}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{\sigma}{\tmpbox} ^2_J \Bigg) + \ln \left(1 + \frac{|J|}{|J|+n} \right). \]

Thus, if the sample size is large and the agents' distribution of covariates is well-specified, the posterior loss of an agent with prior $\pi$ is approximately equal to the AIC (with a penalty of $\ln (1 + |J|/(|J|+ n)) \approx |J|/n$ instead of $2|J|/n$).

The prevalence of larger models in the model competition can be then associated to the “conservativeness” of the AIC for model selection. Proposition (ref), however, makes clear that the relation is only qualitative: larger models will prevail in large samples, but the probability of a larger model being selected will continue to be affected by the prior.

Finally, it is worth reiterating that the foundations of the AIC are normative: the criterion was proposed as a way to select models to avoid overfitting. Conversely, our analyis provides a positive foundation for a solution similar to the AIC: we study the outcomes when Bayesian agents compete in a way that selects the agent with the lowest subjective MPSE.

Applications

In previous sections, we considered the auction of a productive asset as a leading example. We now discuss two additional applications. The first is a simple model of entry when returns depend on prediction error, which we connect to the literature on overconfidence. Second, we use our framework to understand the proliferation of “factors” in the asset pricing literature.

Selection of Simple Models in Entry and Investment

We begin with an application to a single-agent decision problem. An agent is faced with a risky entry choice---she has to choose between a risky option and a safe one that gives her a utility normalized to $0$. The utility of the risky option depends (in part) on the agent's ability to predict an unknown variable and take an action. This is the case of an entrepreneur who has the option to invest in a new venture, where expected returns depend in part on the ability to predict and adapt to future market demand, political situations, or trade agreements. Alternatively, this could be an individual investor considering trading an asset: returns depend on the ability of the investor to predict future price movements and trade accordingly.

Formally, suppose that the expected utility of the risky option is $$ \mathbb{E}(r)=v - L^{*}(\pi, D_n), $$ where $v$ summarizes agent-specific costs and benefits that are independent of prediction error, while $L^{*}(\pi, D_n)$ is the component that depends on prediction error---in line with our notation, the subjective MPSE. The prior $\pi$ summarizes the agent's prior belief about the relationship between the unknown variable they need to predict (e.g., market demand for an entrepreneur, price movement for an investor) and various observables they consider relevant. The agent's model is the set of variables they consider relevant for prediction. The data $D_n$ is past data about this relationship. The agent knows $v$ and $\pi$, observes $D_n$, and then chooses the risky option if its expected utility is positive.

The results of this paper directly apply. Ceteris paribus (fixing $v$), with few data points, agents with “simple models” are systematically more confident in their prediction error ( $L^{*}(\pi, D_n)$ is lower) and are, therefore, more likely to take the risky option. Crucially, this holds whether their simple models are correct or not. This has an immediate implication that provides a novel comparative static: entrepreneurs with simple models are over-represented in sectors where little data has accumulated, even when the true DGP is complex. For example, this margin of selection suggests that, ceteris paribus, the entrepreneurs more eager to invest in a country that just opened up to foreign investment, or in new technologies, will tend to be those that believe they can predict future conditions using relatively few covariates---that have a simple model---even when reality is much more complex. Similarly, investors with simpler models are more likely to enter into new asset classes (e.g., crypto-currencies).

So far we have assumed that no new data is revealed after the investment decision. In reality, however, new data accumulates after the initial investment/ choice to enter is made. Entrepreneurs or investors may take this into account in their decision, expecting to be able to refine their prediction. The discussion of Section (ref) suggests that this only strengthens the selection in favor of simple models. To illustrate, consider the setup above but suppose that entrepreneurs must decide whether or not to invest before any data is revealed, but knowing that some data will be revealed at a later stage. As we discussed in Section (ref), agents with simple models are more confident about how much they will be able to learn from the yet-to-be-released data. In this case, entrepreneurs/ investors with overly simplistic models will always be over represented.

\paragraph{Connections to Overconfidence.} These findings connect to established empirical facts on overconfidence and entry. Several studies have shown that entrepreneurs are, by various measures, overconfident koellinger2007think, COOPER198897. Similarly, a large body of evidence shows that (especially retail) investors are often overconfident about their knowledge and information odean1999investors, statman2006investor. This is commonly attributed to either incorrect beliefs, with selection favoring individuals with overly optimistic priors, or post-decision bolstering.

Our results suggest a novel margin of selection related to model complexity in relation to overconfidence: entry into areas with limited past data (e.g., “new” areas) are systematically biased towards entrepreneurs or investors with models that are “simple.” As we have seen, when the true DGP is complex, these individuals are also overconfident in their predictive ability: their subjective MPSE is on average lower than it should be. The relevant margin of selection may be the simplicity of the model, which generates both a higher likelihood of entrance and overconfidence in predictive ability. Similarly, investment in new asset classes is more common for investors who, ceteris paribus, have simple predictive models of price movement. Our results also suggest that this margin of selection is attenuated as data accumulates: entry into more established sectors/ technologies/ countries may be systematically different from entry into “new” areas in terms of complexity of the entrepreneur's model; investors in established assets/ areas may be systematically different from investors in new asset classes/ trends (e.g., crypto-currency).

Existing studies on overconfidence note also that agents often misreact, or underreact, to new information. For example, odean1999investors studies the trading patterns of retail traders, and argues that their trades are systematically incorrect: “while investors' overconfidence in the precision of their information may contribute to this finding, it is not sufficient to explain it. These investors must be systematically misinterpreting information available to them. They do not simply misconstrue the precision of their information, but its very meaning.” This is consistent with our finding that, when the true DGP is complex and involves many variables, entry into trading may favor agents with incorrect models, i.e., ones that exclude covariates that are relevant for prediction and/or include irrelevant ones.

Competing Factor Models\footnote{We thank Stefano Giglio for useful discussions in writing this section.}

A large body of work in finance has studied the cross-sectional variation in asset expected returns (i.e., why different assets earn different average returns). The classical asset-pricing framework jensen1972capital,fama1973risk posits that, at each point in time, asset returns are governed by a multi-factor model. The return of each asset---an individual stock or a portfolio---is an asset-specific linear combination of these factors (with time-invariant coefficients) plus random noise.

The search for factors to explain the cross-sectional variation of expected stock returns has produced hundreds of potential candidates. The literature has evolved from the parsimonious model of fama1993common using only three factors (market return, size premium, value premium) to the factor library in feng2020taming that contains 150 risk factors.\footnote{See Appendix (ref) for a description of these factors, how they are constructed, and the year in which they were published.}

We use our framework to understand the proliferation of factors in this literature. We argue that the increase in the number of test portfolios used to compute the Fama-French cross-sectional regressions mechanically favors asset-pricing models with several factors.

To make this point, we view different collections of factors as competing models or, more precisely, competing sets of risk factors---a terminology that has been used, incidentally, by fama1993common, fama2015five. \footnote{From fama1993common: “The average excess returns on the portfolios that serve as dependent variables give perspective on the range of average returns that competing sets of risk factors must explain.” On p.\ 13: “The wide range of average returns on the 25 stock portfolios, and the size and book-to-market effects in average returns, present interesting challenges for competing sets of risk factors.” Finally, in fama2015five “estimate the proportion of the cross-section of expected returns left unexplained by competing models.” } We then take the number of test portfolios as the number of available data points to predict the cross-section of expected returns. We consider the different factor models as different Bayesian agents competing to predict cross-sectional returns. Our results suggest that the winning model depends crucially on the sample size. With a few test portfolios, smaller factor models will be selected. Conversely, increasing the number of test portfolios favors high-dimensional factor models.

Let $i$ index an asset in the cross-section. The outcome variable $Y_i$ denotes the excess return of asset $i$ averaged over the different time periods for which data on returns and factors are available.\footnote{We follow feng2020taming and use monthly returns from July 1976 to December 2017.} Each asset $i$ has an associated 150-dimensional vector of covariates, $X_i$, containing the asset's factor loadings.\footnote{In the standard Fama-French two-pass regressions, these factor loadings are estimated from asset-by-asset time series regressions of excess returns on factors BAI201531. For simplicity, we ignore the estimation error and treat the estimated factor loadings as the true factor loadings. This allows us to stay within our simple linear regression framework. The competing models are the different subsets of the factors that different agents believe to be relevant to explain the cross-section of asset expected returns.} In principle, assets could be individual stocks or portfolios. We follow the literature and focus on portfolios.\footnote{There is some discussion in the literature about what is the right unit of observation to test asset-pricing models: see the discussion in ang_liu_schwarz_2020.} The number of portfolios under consideration $(N)$ gives the sample size for the cross-sectional regressions.

We start with the 25 (5 $\times$ 5) portfolios sorted by size and book-to-market ratio as used in fama1993common.\footnote{A 5 $\times$ 5 bivariate portfolio refers to the common practice of grouping stocks by the quintiles of the cross-sectional distribution of both size and book-to-market.} We then consider the larger collection of 2,875 (5 $\times$ 5) bivariate-sorted portfolios in feng2020taming.\footnote{The sorting is based on size and each of the 115 factors marked with asterix in the table in Appendix (ref). The dataset was obtained directly from the replication files provided by feng2020taming.}

\paragraph{Competing (Factor) Models.} We consider “competition” between three different models. The first (Fama-French) uses only the original Fama-French factors (excess market return, and the “small minus big” and “high minus low” factors). The second model (Factor Zoo) uses all 150 factors in feng2020taming. The third (FGX-DS) uses the 135 factors introduced up until 2011, plus five additional factors obtained by the “Double Selection” approach of feng2020taming.\footnote{These are the investment and profitability factors of hou2015digesting, the “robust minus weak” factor of fama2015five, the intermediary risk factor of he2017intermediary, and the “Quality Minus Junk” factor of asness2019quality.} For each model, we consider the same set of hyper-parameters $(a_0,b_0)$, chosen to maximize the marginal likelihood of the largest model with the largest dataset, and set $\gamma$ so that all models have the same prior subjective MPSE before any data. Details are provided in Appendix (ref).

Figure (ref) presents our results. It shows the subjective MSPE of each model. To make units easier to interpret, the results are presented as a percentage relative to the worst competing model (out of the three under consideration).

Consistent with our theorems, if we consider only the 25 ($5\, \times\, 5$) bivariate sorted portfolios on size and book-to-market (as fama1993common originally did), the simple three-factor model achieves the best subjective MPSE (around half of the subjective MPSE obtained with the largest models). With 2,875 ($5 \times 5$) portfolios, however, the ranking reverses, and larger models now have the lowest subjective MPSE. (Appendix (ref) presents several robustness checks with different models, sample sizes, and subsets of portfolios.)

figure[figure omitted — 161 chars of source]

\paragraph{A General Model of Scientific Progress.} This discussion suggests a simple model of scientific progress. There is public interest in predicting a variable $y$ as a function of observed covariates $x$, and there are competing scientists, each described by a prior belief about the world; these priors differ in what covariates they think are relevant for predicting $y$. Data (publicly) accumulates as i.i.d.\ draws from the unknown DGP.

Suppose each model's success also depends on its subjective MSPE. Scientists who believe their model has low prediction error would be more forceful about it, staking their career on its predictions. Others whose subjective MSPE is high may be worried about mistakes and damage to their reputation. Practitioners or politicians, who may cite scientific research to justify their actions, may be more prone to adopt models with low subjective MSPE.

With these assumptions, our results suggest the following dynamic of scientific progress. In the early stages of a field, when data is relatively scarce, overly simple models prevail---including frameworks that exclude relevant covariates. Over time, more and more data accumulates, and more nuanced models come into vogue, involving ever-increasing collections of covariates. Overly-simple models are then discarded, since they are unable to fit the data as well as larger ones; the “scientific paradigm,” understood as the collection of relevant variables, becomes more complex. This is in line with casual observation and with dynamics described in epistemology. For example, this aligns with what kuhn2012structure describes as the path of progress of “normal” science, i.e., after a dominant paradigm has been established.\footnote{Changing paradigms are outside the scope of this work---see, e.g., ortoleva2012modeling for a model of a non-Bayesian decision-maker who changes paradigm (selects a new prior) upon receiving information that is unexpected according to their current prior.}

Related Literature

A large body of literature has studied model misspecification in individual decision-making, with famous examples like overconfidence and correlation neglect. A few recent theoretical contributions to this enormous literature include heidhues2018unrealistic and ortoleva2015overconfidence, to which we refer for further references. In misspecified learning settings, “feedback loops” between the agents' misspecified beliefs and the action they take add further technical challenges---see, e.g., fudenberg2017active, fudenberg2020limits, heidhues2020convergence.

Recent works have studied the implications of agents with misspecified models in various strategic settings. For instance, bohren2016informational, bohren2017bounded, frick2019misinterpreting, and frick2019dispersed study social learning when agents have misspecified models that cause them to misinterpret other agents' actions. mailath2019wisdom study a stylized prediction market where Bayesian agents have different models (defined as different partitions of a common state space) and discuss the possibility of information aggregation.

Recent works consider specifically the outcomes when agents' models are misspecified in the sense we study here, i.e., there is a payoff-relevant dependent variable, and agents either include irrelevant independent variables or exclude dependent variables. schwartzstein2021using considers this in the context of persuasion, where competing persuaders may “overfit” the data to better persuade a receiver. levy2019misspecified study a political economy setting where there are both “simple” world views and complicated ones, and ask whether political competition disciplines overly simplistic world views, finding instead that they recur in dynamic settings. Finally, he2020evolutionarily ask whether (and what kind of) misspecifications can be evolutionarily stable. These works find reasons for why misspecified models may survive (overfit models in the former, simple models in the case of the latter two), although the exact mechanism is different than ours. Recent work also considers the possibility that agents with misspecified models may be able to realize it, and characterize the kinds of misspecifications that survive---see, e.g., fudenberg2020misperceptions, who consider an evolutionary framework, or gagnon2021channeled, who study a setting where the agent perceives their errors through the framework of their own model.

In strategic settings, esponda2016berk define a learning-based solution concept (“Berk-Nash Equilibrium”) for games in which agents' beliefs are misspecified. More broadly, solution concepts have been posited for settings where agents suffer from some sort of misspecification, including well-known examples like analogy-based equilibrium jehiel2005analogy and cursed equilibrium eyster2005cursed.

Several works consider outcomes when some agents behave in a way that can be construed as coming from a misspecified model. For instance, in spiegler2006market, spiegler2013placebo, society misunderstands the relationship between outcomes and the actions of strategic agents, which affects the actions these agents take in equilibrium and resulting outcomes (studied in the context of a market for quacks or its implications for political reforms). levy2019misspecified study a dynamic model of political competition where agents have different (misspecified) models of the world; the study uses this model to provide a foundation for the recurrence of populism. liang2018games studies outcomes in games of incomplete information where agents behave like statisticians and have limited information.\footnote{There is a larger literature that studies the outcomes when agents are modeled as statisticians or machine learners, e.g., al2009decision, al2014coarse, acemoglu2016fragility and cherry2006statistical.}

A novel approach to modeling misspecification in economic theory is the directed acyclic graph approach; see pearl2009causality. This is exploited in a single-person decision framework in spiegler2016bayesian, which studies a single decision maker with a misspecified causal model and large amounts of data. The paper shows that the decision maker may evaluate actions differently than their long-run frequencies, and exhibit artifacts such as “reverse causation” and coarse decision making. This approach is then used in eliaz2018model, which proposes a model of competing narratives. A narrative is a causal model that maps actions into consequences, including other random, unrelated variables. An equilibrium notion is defined, and the paper studies the distribution of narratives that is obtained in equilibrium.

Finally, the understanding that agents should be cognizant that their models may be misspecified has also led to new approaches in mechanism design, where the designer accounts for misspecification in various ways. The literature on robust mechanism design (beginning with the seminal bergemann2005robust) provides foundations for using stronger solution concepts. madarasz2017sellers show that an optimal mechanism may perform very poorly if the planner's model is even slightly misspecified, and they identify a class of near optimal mechanisms that degrade gracefully. Works such as chassang2013calibrated and carroll2015robustness develop optimal “robust” contracts and contrast to classical optimal contracting.

Since one natural application of our model is an auction, our results are related to atakan2014auctions, who consider the competitive sale of assets whose value depends on how they are utilized.\footnote{bond2010information study a trading environment with a similar feature.} The successful bidder chooses an action that determines, together with the state of the world, the payoff generated by the asset. They focus on a setting where bidders have a common prior but observe private signals. Their main result is the possibility of (complete) failure of information aggregation. Our model is similar in that the value of the object depends on an action taken by the agent. However, our work considers a complementary environment where all bidders observe the same information but have different priors. Information aggregation is ruled out by assumption, and our key theme is model selection.

We assume that agents have different priors and are fully aware they have different priors: that is to say, our agents agree to disagree. This assumption has been used in economic theory at least since harrison1978speculative. We refer the reader to morris1995common for a discussion of the common and heterogeneous prior traditions in economic theory. Heterogenous priors have been used in a number of applications in bargaining yildiz2003bargaining, trade morris1994trade, financial markets scheinkman2003overconfidence,ottaviani2015price, and more.

\paragraph{Relation to Model Selection.} Large literatures in statistics, econometrics, and machine learning study model selection methods and provide normative foundations; see ModelSelection08 and burnham2003model for textbook overviews. Popular approaches include, for example, the $C_p$ criterion of Mallows, the Akaike Information Criterion (AIC) of akaike1974new, and the Bayes Information Criterion (BIC) of schwarz1978. We showed that is a connection between our large data results and the AIC introduced in akaike1974new, in particular, to the asymptotic properties of the AIC characterized in the seminal paper of nishii1984asymptotic.

While some of our asymptotic results are reminiscent of the model selection literature, there are three important differences. First, the aims of this literature are very different from ours. Ours is a positive approach of studying which model emerges from a competition between Bayesian agents with misspecified models. The approach in the model selection literature is instead normative: various methods of model selection are proposed and studied with a view to avoiding over-fitting and/or selecting “good” models according to some metric. The results we are aware of speak to the asymptotic efficiency of these techniques. Second, not only are our results derived from a completely different setting, but they are also proven with different techniques. Third, the connection is limited to the large-data result. We are not aware of any analogs to our small-sample results.

Discussion and Conclusion

A variable of interest is related to a vector of covariates. Different agents have different models of this relationship: in particular they rule in/rule out different covariates as being potentially related to prediction. All agents observe a common dataset of size $n$, drawn from the true DGP. We ask: Who is the agent with the highest confidence in their own predictive ability, in the form of the lowest mean-squared prediction error according to their own subjective posterior? We study the relationship between sample size and the dimension of the winning model. This applies to all cases in which confidence in predictive ability affects selection. We show results of two kinds.

First, when $n$ is small, models that employ few covariates may take the lead, even if the true DGP is more complex. To establish this result formally, we use Normal-Inverse-Gamma-Inverse-Wishart priors. Second, when $n$ is large, misspecified models (i.e., models that rule out an observable that is relevant for prediction) never win. However, high-dimensional models that include irrelevant covariates (but do not exclude relevant ones) may continue to win. Our results show that the effect of the prior on model competition does not vanish in large samples. These results hold for a very general class of priors and true DGPs.

Finally, we give two applications. First, we apply our results to a model of entry: entrepreneurs decide whether to enter a new market, households decide whether to invest in a new asset class. We show that, insofar as prediction error of future variables is relevant for profitability, our results suggest a new margin of selection: when data is relatively scarce, agents with simpler models will be over-represented in the entry decision. Our second application is to understand the proliferation of factors that explain the cross-sectional variation of expected stock returns in the asset-pricing literature. We show how the increase in the number of test portfolios used to compute the cross-sectional regressions mechanically favors models with several factors.