Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
68,282 characters · 9 sections · 36 citation commands
On the Wisdom of Crowds (of Economists)
\newgeometry{includeheadfoot,left=0.8in,right=0.8in,top=0.8in,bottom=0.8in,bindingoffset=0.in,nohead}
\thispagestyle{empty}
\restoregeometry
\thispagestyle{empty} \setcounter{tocdepth}{3}
\setcounter{page}{1} \thispagestyle{empty}
The wisdom of crowds, or lack thereof, is traditionally and presently a central issue in psychology, history, and political science; see for example surowiecki2005 regarding wisdom, and kindleberger2023 regarding lack thereof. Perhaps most prominently, however, the wisdom of crowds is also---and again, traditionally and presently---a central issue in economics and finance, where heterogeneous information and expectations formation play a prominent role.\footnote{Interestingly, moreover, it also features prominently in new disciplines like machine learning and artificial intelligence, via forecast combination methods like ensemble averaging DSZ2023.}
In this paper we focus on economics and finance, studying the “wisdom" of “crowds" of professional economists. We focus on the U.S. Survey of Professional Forecasters (SPF), which is important not only in facilitating empirical academic research in macroeconomics and financial economics, but also---and crucially---in guiding real-time policy, business, and investment management decisions.\footnote{On real-time policy and its evaluation, see for example John Tayor's inaugural NBER Feldstein Lecture at \url{https://www.hoover.org/sites/default/files/gmwg-empirically-evaluating-economic-policy-in-real-time.pdf}.}$^,$\footnote{For an introduction to the SPF, see the materials at \url{https://www.philadelphiafed.org/surveys-and-data/real-time-data-research/survey-of-professional-forecasters}.}
In particular, we study SPF crowd behavior as crowd size grows, asking precisely the same sorts of questions of SPF “forecast portfolios" that one asks of financial asset portfolios:\footnote{Classic early work on which we build includes, for example, makridakis1983averages and BD1995. We progress much further, however, particularly regarding analytic characterization.} How quickly, and with what patterns, do diversification benefits become operative, and eventually dissipate, as portfolio size (the number of forecasters) grows, and why? Do the results differ across the key SPF variables (growth and inflation), and if so, why? What are the implications for survey size and design?
We answer the above questions using what we call “crowd size signature plots," which summarize the forecasting performance of $k$-average forecasts as a function of $k$, where $k$ is the number of forecasts in the average. We examine both “direct signature plots" (estimated directly from the SPF data) and “model-based signature plots" (estimated by fitting a simple model of forecast-error equicorrelation), and we characterize their surprisingly intimate relationship, uncovering en route a remarkably simple analytic closed-form signature plot estimator. We focus throughout on estimating and interpreting signature plots for both growth and inflation, providing individual and comparative assessments of their paths and patterns.
\paragraph{Related Literature.}
This paper contributes to the forecast combination literature by examining the performance patterns of simple $k$-average forecasts as the number of forecasters, $k$, increases. We note that this paper is not focused on the problem of optimal forecast combination (see BatesandGranger1969 for a pioneering contribution, and for example timmermann2006forecast, hsiao2014 and wang2022 for recent developments). The motivation for focusing on simple $k$-averages is twofold. First, it is a well-established empirical fact that simple averages of forecasts produce combined forecasts with surprisingly good forecasting performance despite not being optimal in general (e.g., clemen1989, stock2004combination, and genre2013).\footnote{This empirical fact has been called a “forecast combination puzzle”; work focused on studying the causes of this puzzle includes smith2009simple, Elliott2025, and claeskens2016simple.} Second, simple averages can be optimal in specific contexts. For example, EngleandKelly2012 show that when forecast errors exhibit equicorrelation, equal weights are optimal; Elliott2025 later provide precise necessary and sufficient conditions under which simple averages are optimal.
Our focus on performance patterns of simple k-averages as k increases builds on earlier empirical work. makridakis1983averages, using the M-Competition time-series dataset and combinations of forecasts across multiple forecasting methods, document that simple-average combinations become more accurate, with sharply diminishing marginal gains, as the number of component forecasts increases. We find the same qualitative pattern in the SPF when averaging across professional forecasters. We extend these results by providing a theoretical foundation: we derive an exact expression for the direct crowd size signature plot that links its level and slope to two interpretable features of SPF forecast errors (average dispersion and average dependence). This delivers an equivalent equicorrelation representation and explains both the hyperbolic shape of the signature plots and the rapid exhaustion of diversification gains.
In another related paper, chan2018some propose a framework to study the properties of a combined forecast, and derive conditions under which the simple average is optimal. In their setting, they point out that the performance of a combined forecast improves as the number of forecasts or models increases, and that a simple average will asymptotically outperform any fixed individual forecast. Although this is consistent with our results, our emphasis is more general, as we characterize the entire trajectory of performance improvements, including the practically relevant setting of a small to moderate number of forecasters.
Our results on relating direct crowd size signature plots to an equicorrelation model-based signature plot are connected to the results in clemen1985limits. In the case of normally distributed forecast errors, they show that the value of additional information for forecast combination diminishes when the additional information is correlated. This is consistent with our findings documenting diversification gains; however, our results go beyond this by not requiring the normality assumption.
Finally, we contribute to the SPF survey-forecast literature (e.g., CroushoreStark2019; clements2022) by providing evidence and theory directly relevant for survey design: how many forecasters are needed to capture most of the benefits of averaging, and why do those benefits differ across variables? Using SPF growth and inflation forecasts, we document a steeply diminishing-returns pattern in crowd-size performance, with most improvements realized by about 5-10 representative forecasters. Moreover, we show that diversification gains are considerably larger for inflation than for growth, and we provide an exact theoretical characterization linking the inflation-growth difference to the underlying variance-covariance structure of individual forecast errors.
We proceed as follows. In section (ref) we study the SPF, estimating its crowd size signature plots directly, for both growth and inflation. In section (ref) we introduce and estimate the equicorrelation model, beginning in section (ref) by characterizing its crowd size signature plots analytically for any parameter configuration, and continuing in section (ref) by estimating its parameters by minimizing divergence between direct and model-based signature plots, again providing empirical results for both growth and inflation. In section (ref) we conclude and sketch several directions for future research.
In this section we introduce the idea of crowd size signature plots, and we calculate them directly for SPF point forecasts of real output growth (“growth") and GDP deflator inflation (“inflation"), for forecast horizons $h=1,2, 3, {\rm and ~} 4$, corresponding to short-, medium-, and longer-term forecasts.\footnote{The SPF contains quarterly level forecasts of real GDP and the GDP implicit price deflator. We transform the level forecasts into growth and inflation forecasts by computing annualized quarter-on-quarter growth rates, and we compute the corresponding forecast errors using realized values as of December 2023. See Appendix (ref) for details.} Our sample period is 1968Q4-2023Q2.
The SPF was started in 1968Q4 and is currently conducted and maintained by the Federal Reserve Bank of Philadelphia.\footnote{For a recent introduction see CroushoreStark2019.} In Figure (ref) we show the evolution of the number of forecast participants, which declined until 1990Q2, when the Federal Reserve Bank of Philadelphia took control of the survey, after which it has had approximately 40 participants. Participants stayed for 15 quarters on average, with a minimum of 1 quarter and a maximum of 125 quarters.
We us now sketch the basic framework. Let $N$ refer to a set of forecasts with $N \times 1$ zero-mean time-$t$ error vector $e_t$, $t=1, ..., T$, and let $k \le N$ refer to a subset of forecasts. We consider $k$-forecast averages, and we seek to characterize $k$-forecast mean-squared forecast error ($MSE$). For a particular $k$-forecast average corresponding to group $g^*_k$, the forecast error is just the average of the individual forecast errors, so we have
For any choice of $k$, however, there are $N \choose k$ possible $k$-forecast averages. We focus on the $k$-average $\widehat{MSE}^*_T(k)$ given by equation (ref), averaged across all groups of size $k$,
where $g_k$ is an arbitrary member of the set of groups of size $k$.
Among other things, one may be interested in:
The “ratio" signature plots, especially the $\widehat{MSE}^{avg}_{R,NT}(k)$ plot defined in equation (ref), deserve special mention. Scaling by $\widehat{MSE}^{avg}_{NT}(1)$ (the benchmark MSE corresponding to no averaging) makes $\widehat{MSE}^{avg}_{R,NT}(1) \equiv 1$, which facilitates $\widehat{MSE}^{avg}_{NT}(k)$ comparisons across different variables like growth and inflation. For that reason we use it extensively in the graphical presentations that follow.
$\widehat{MSE}^{avg}_{NT}(k)$ and the related objects above are readily calculated in principle, but complications arise in practice. In particular, calculation is infeasible unless $N$ is very small, due to the huge number of different $k$-average forecasts. For example, for $N=40$ and $k=20$ we obtain ${N \choose k} = O(10^{11})$. Hence we proceed by approximating $\widehat{MSE}^{avg}_{NT}(k)$ as follows:
As $B \rightarrow \infty$, the above approximation will converge to $\widehat{MSE}^{avg}_{NT}(k)$, and in this paper we set $B=30,000$. The approximation works, moreover, for unbalanced as well as balanced panels, which is important because the SPF panel is unbalanced due to entry and exit of forecasters.
We show direct crowd size signature plots for growth and inflation in Figure (ref) ($\widehat{MSE}^{avg}_{R,NT}(k)$), Figure (ref) ($\widehat{DMSE}^{avg}_{R,NT}(k)$), and Figure (ref) ($\widehat{F}^{avg}_{R,NT}(k)$). Several features are apparent, as distilled in the following remarks:
Figures (ref) and (ref), together with Remarks (ref) and (ref), are closely linked and merit additional and joint discussion. Figure (ref) and Remark (ref) show a markedly stronger consensus in growth forecasts than in inflation forecasts (compare, for example, the growth and inflation forecast distributions in Figure (ref) for $k=1$). This pattern accords with the broader empirical macroeconomics literature, in which low-ordered autoregressions typically provide hard-to-beat descriptions of short-term U.S. growth SW2002,CP2013, whereas inflation dynamics are more complex, warranting consideration of a wider set of models SW2007,PR2007.\footnote{A similar situation appears in the agent-based modeling literature, where simulated agents struggle to learn a stable inflation model and often adopt heterogeneous specifications ABK2013,LS2025.} In this light, Figure (ref) and Remark (ref) make perfect sense: the greater heterogeneity in inflation forecasts produces greater gains from inflation forecast diversification, yielding a strikingly lower $\widehat{MSE}^{avg}_{R,NT}(k)$ signature-plot asymptote for inflation (about 60%) than for growth (about 80%).
Having empirically characterized crowd size signature plots directly in the SPF data, we now proceed to characterize them analytically in a simple covariance-stationary equicorrelation model, in which $e_{t} \sim (0, \Sigma)$, where $0$ is the $N \times 1$ zero vector and $\Sigma$ is an $N\times N$ forecast-error covariance matrix displaying equicorrelation, by which we mean that all variances are identical and all correlations are identical.
A trivial equicorrelation example occurs when $\Sigma = \sigma^2 I$, where $I$ denotes the $N\times N$ identity matrix, so that all variances are equal ($\sigma^2$), and all correlations are equal (0). Of course the zero-correlation case is unrealistic, because, for example, economic forecast errors are invariably positively correlated due to overlap of information sets, but it will serve as a useful benchmark, so we begin with it.
Simple averaging is the fully optimal forecast combination in the zero-correlation environment, which is obvious since the forecasts are exchangeable. More formally, the optimality of simple averaging (equal combining weights) follows from the multivariate BatesandGranger1969 formula for $MSE$-optimal combining weights,
where $\iota$ is a $k$-dimensional column vector of ones. For $\Sigma = \sigma^{2} I$ the optimal weights collapse to $$ \lambda^* = (\sigma^{-2} N)^{-1} \sigma^{-2} \iota= \frac{1}{N} \iota. \nonumber $$
Analytical asymptotic results are straightforward for simple averages in the zero-correlation environment. Let $MSE^{avg}_{N \infty}(k; \sigma) = plim_{T \rightarrow \infty} \left ( \widehat{MSE}^{avg}_{NT}(k; \sigma) \right )$; then we have
Moreover,
and
Notice that $\sigma$ cancels in the $R^{avg}_{N \infty}(k; \sigma)$ calculation, so we simply write $R^{avg}_{N \infty}(k)$.
We now move to a richer equicorrelation case with equal but nonzero correlations, so that instead of $\Sigma = \sigma^{2} I$ we have
where
and $\rho \in \left ] \frac{-1}{N-1}, 1 \right [$.\footnote{$R$ is positive definite if and only if $\rho \in \left ] \frac{-1}{N-1}, 1 \right [$. See Lemma 2.1 of EngleandKelly2012.} Recent work, in particular EngleandKelly2012, has made use of equicorrelation in the context of modeling multivariate financial asset return volatility.
Importantly, the optimality of simple averaging under zero correlation is preserved under equicorrelation. That is, equicorrelation is sufficient for the optimality of simple averaging -- an immediate implication of the results of Elliott2025, who shows that a necessary and sufficient condition for optimality of simple averaging is that row sums of $\Sigma$ be equal. Equicorrelation is one such case, although there are of course others, obtained by manipulating correlations in their relation to variances to keep row sums equal, but none are nearly so compelling and readily interpretable as equicorrelation.
To understand the optimality of simple averaging under equicorrelation in the context of BatesandGranger1969, just as we did earlier under zero correlation, consider the inverse covariance matrix in the Bates-Granger expression for the optimal combining weight vector, (ref). In the equicorrelation case we have
where\footnote{See Lemma 2.1 of EngleandKelly2012.}
Then, using equation ((ref)), the first part of the optimal combining weight (ref) is
and the second part is
Inserting equations ((ref)) and ((ref)) into equation (ref) again yields
establishing the optimality of equal weights.
Having now introduced equicorrelation and shown that it implies optimality of simple average forecast combinations, it is of interest to motivate it -- despite its stark simplicity -- in terms of economic considerations. First, obviously but importantly, the information sets of economic forecasters are quite highly overlapping, so it is not necessarily unreasonable to suppose that the forecast error variances are roughly equal, and that pairs of forecast errors are roughly equally correlated.
Second, less obviously but very importantly, equicorrelation is closely linked to factor structure, which is a great workhorse of modern macroeconomic and business-cycle analysis SW2. In particular, equicorrelation arises when forecast errors have single-factor structure with equal factor loadings and equal idiosyncratic shock variances, as for example in:
$$ z_{t} = \phi z_{t-1} + v_{t}, $$ where $ w_{it} \sim iid (0, \sigma_{w}^{2})$, $v_{t} \sim iid (0, \sigma_{v}^{2})$, and $w_{it} \perp v_{t}$, $\forall i, t$, $i = 1, ..., N$, $t = 1, ..., T$.
Finally, a large literature from the 1980s onward documents and exploits the routine outstanding empirical performance of simple average forecast combinations, despite the fact that simple averages are not optimal in general clemen1989, genre2013, ET2016, DS2019. As we have seen, however, equicorrelation is sufficient (and almost necessary) for optimality of simple averages, so that if simple averages routinely perform well, then the equicorrelation model is routinely reasonable -- and a natural model to pair with the simple averages embodied in the SPF.
Analytical results for $MSE^{avg}_{N \infty}(\cdot)$, $DMSE^{avg}_{N \infty}(\cdot)$ and $R^{avg}_{N \infty}(\cdot)$ are easy to obtain under equicorrelation, just as they were under zero correlation. Immediately,
In addition, for $MSE$ changes and ratios we have
and
One can immediately verify that if $\rho=0$ then the equicorrelation signature plot results (ref)-(ref) collapse to the corresponding earlier zero-correlation results (ref)-(ref).
In Figure (ref) we show $MSE^{avg}_{R,N \infty}(k; \rho)$ as a function of $k$, for various equicorrelations, $\rho$. Several remarks are in order:
Here we estimate the equicorrelation model by choosing its parameters $\rho$ and $\sigma$ to make the equicorrelation $MSE^{avg}_{N \infty}(k; \rho, \sigma)$ as close as possible to the SPF $\widehat{MSE}^{avg}_{NT}(k)$. This estimation strategy is closely related to, but different from, GMM estimation. Rather than matching model and data moments, it matches more interesting and interpretable functions of those moments, namely model-based and direct crowd size signature plots -- as per the “indirect inference" of Smith1993 and Gourierouxetal1993. Henceforth we refer to it simply as the “matching estimator."
Specifically, we solve for $(\hat{\rho}, \hat{\sigma})$ such that
where
and the minimization is constrained such that $\sigma>0$ and $\rho \in \left] \frac{-1}{N-1},1 \right [$.
Calculating the minimum in equation (ref) is very simple, because the bivariate minimization can be reduced to a univariate minimization in $\rho \in \left] \frac{-1}{N-1},1 \right [$. To see this, recall that $MSE_{N \infty}^{avg}(k; \rho, \sigma)=\frac{\sigma^2}{k}[1+(k-1)\rho]$, so that the first-order condition for $\sigma^2$ is
and the first-order condition for $\rho$ is
Combining equations (ref) and (ref) yields
where $c_1=\sum_{k=1}^N \frac{\widehat{MSE}^{avg}_{NT}(k)}{k}$, $c_2=\sum_{k=1}^N \frac{1}{k^2}$ and $c_3=\sum_{k=1}^N \frac{k-1}{k^2}$. Hence, at the optimum and conditional on the data, there is a deterministic inverse relationship between $\rho$ and $\sigma$ (see Figure (ref)), enabling one to restrict the parameter search to the small open interval $\rho \in \left] \frac{-1}{N-1},1 \right [$, as well as to explore the objective function visually as a function of $\rho$ alone (see Figure (ref)).
In Table (ref) we show the complete set of estimates (for ${\sigma}^2$ and ${\rho}$, for growth and inflation, for $h=1, ...4$). The $\hat{\rho}$ estimates accord with the earlier-discussed asymptotes of the direct signature plots, and the $\hat{\sigma}^2$ estimates increase with forecast horizon, reflecting the fact that the distant future is harder to forecast than the near future, and implying that fitted model-based equicorrelation signature plots should shift upward with horizon. In Figure (ref) we show side-by-side direct $\widehat{MSE}^{avg}_{NT}(k)$ (left column) and equicorrelation model-based ${MSE}^{avg}_{R,N \infty}(k; \hat{\rho})$ (right column) signature plots, which reveal a remarkably good equicorrelation model fit -- the direct and model-based plots are effectively identical.
Here we present a closed-form solution for the direct crowd size signature plot. The result is significant in its own right and reveals why our numerical matching estimates for the equicorrelation model produce fitted signature plots that align so closely with direct signature plots. To maintain precision it will prove useful to state it as a formal theorem.
Several remarks are in order:
Let us now provide a characterization of $DMSE_{N \infty}^{avg}(k)$ (and $DMSE^{avg}_{R,N \infty}(k)$), which of course follows immediately from our earlier characterization of $MSE_{N \infty}^{avg}(k)$ (and $MSE^{avg}_{R,N \infty}(k)$). Immediately,
As we have seen, equicorrelation ultimately emerges as a device for convenient calculation and understanding of direct signature plots, rather than as a separate “model" producing separate signature plots. In particular, nothing we have done requires that the equicorrelation be “true." Nevertheless, it may be of interest in some contexts to assess whether forecast errors are truly equicorrelated, which can be done (under stronger assumptions than those invoked thus far) by maximum-likelihood estimation of a dynamic-factor model, followed by likelihood-ratio tests of the equicorrelation restrictions.
Consider in particular a standard dynamic single-factor model (DFM),
where $ w_{it} \sim iid \, N(0, \sigma_{wi}^{2})$, $v_{t} \sim iid \, N(0, \sigma_{v}^{2})$, and $w_{it} \perp v_{t}$, $\forall i, t$. The implied forecast-error covariance matrix $\Sigma$ fails to satisfy equicorrelation; that is,
because the forecast-error variances generally vary with $i$, and their correlations generally vary with $i$ and $j$. In particular, simple calculations reveal that
where $var(z_t) = \frac{\sigma_v^2}{1 - \phi^2}$, and
Nevertheless, certain simple restrictions on the DFM (ref) produce certain forms of equicorrelation. First, it is apparent from equation (ref) that $\rho_{ij} = \rho,~ \forall i \ne j$, if and only if
so that imposition of the constraint (ref) on equation (ref) produces a “weak" form of equicorrelation with identical correlations ($\rho$) but allowing for potentially different idiosyncratic shock variances ($\sigma_1^2, ..., \sigma_N^2$). That is,
Second, it is also apparent from equation (ref) that if we impose the stronger restriction,
which of course implies the weaker restriction (ref), then we obtain (“strong") equicorrelation as we have defined it throughout this paper, with identical correlations ($\rho$) and identical idiosyncratic shock variances ($\sigma^2$). That is,
The DFM (ref) is already in state-space form, and one pass of the Kalman filter yields the innovations needed for exact Gaussian likelihood evaluation and optimization, and it also accounts for missing observations associated with survey entry and exit. One may also impose the weak or strong equicorrelation restrictions (ref) or (ref), respectively, and assess them using likelihood-ratio tests. Unsurprisingly, such formal tests (not reported) reject equicorrelation -- a stylized “toy" data-generating process if ever there was one -- for both growth and inflation, so we will not pursue formal testing.
More interestingly, we will provide a preliminary assessment of the size of deviations from equicorrelation, for both growth and inflation, quite apart from assessing whether they happen to be exactly zero. We construct a balanced panel with no entry or exit, 2010Q1-2019Q4 (40 quarters), containing ten forecasts for growth and inflation, and we examine their covariance matrices, as shown in Figure (ref).\footnote{The forecasts are ordered from lowest to highest error variance for each variable.}
First consider growth forecast errors, in the upper panel of Figure (ref). No shading means that the absolute deviation of the object in the cell (a variance or covariance) from the median across all cells (all variances or all covariances) is less than 10%. For example, the forecast with ID 3 has a variance of growth forecast error that is -2.2% below the median variance, so its cell is unshaded. Similarly, light red shading means that the absolute deviation from the median is 10-20%, and darker red shading means that the absolute deviation from the median is 20-30%. All growth covariance cells are white, as are all but two variance cells. Hence, at least by our intentionally rough metric, equicorrelation appears to be not too bad an approximation for growth forecast errors.
Now consider inflation forecast errors, in the lower panel of Figure (ref). The shading convention is the same as earlier, but now there is a fourth, dark red, category corresponding to ${>}30\%$. The situation is sharply different, with many red cells, for both variances and covariances. This echoes our earlier discussion of the relative lack of consensus in inflation forecasts due to the difficulty of determining an appropriate inflation model, as per Figures (ref) and (ref), together with Remarks (ref) and (ref).
We have studied the properties of macroeconomic survey forecast response averages as the number of survey respondents grows, characterizing the speed and pattern of the “gains from diversification" and their eventual decrease with “portfolio size" (the number of survey respondents) in both (1) the key real-world data-based environment of the U.S. Survey of Professional Forecasters (SPF), and (2) the theoretical model-based environment of equicorrelated forecast errors. We proceeded by proposing and comparing various direct and model-based “crowd size signature plots," which summarize the forecasting performance of $k$-average forecasts as a function of $k$, where $k$ is the number of forecasts in the average. We then estimated the equicorrelation model for growth and inflation forecast errors by choosing model parameters to minimize the divergence between direct and model-based signature plots.
The results indicate near-perfect equicorrelation model fit for both growth and inflation, which we explicated by showing analytically that, under conditions, the direct and fitted equicorrelation model-based signature plots are identical at a particular model parameter configuration, which we characterize. We find that the gains from diversification are greater for inflation forecasts than for growth forecasts, but that both the inflation and growth diversification gains nevertheless decrease quite quickly (which we also explain analytically), so that fewer SPF respondents than currently used may be adequate.
Several directions for future research appear promising, including, in no particular order: