Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
46,829 characters · 7 sections · 97 citation commands
No-Regret Forecasting with Egalitarian Committees
\thispagestyle{empty}
\fontsize{12}{18pt}\selectfont \setcounter{page}{1}
Big data has been a buzzword in social science in recent years, and its popularity is witnessed in surveys such as Varian2014 and EinavLevin2014a in economics, LazerRadford2017 in sociology, and Brady2019 in political science. The importance of data in empirical studies is self-evident. However, as argued by Lovell1983 almost forty years ago, “it is by no means obvious that reductions in the costs of data mining have been matched by a proportional increase in our knowledge of how the economy actually works.” The advance in theory is also important and indispensable for scientific progress. Data science --- thus named perhaps because it emphasizes dialogues between theory and data --- suggests an effective way to improve knowledge.
A dialogue between theory and data has been exemplified by the forecast combination puzzle in econometrics. This puzzle refers to the phenomenon that the equal-weight scheme, which is theoretically suboptimal in general, often outperforms the forecast combination with BatesGranger1969's (BatesGranger1969) optimal weights as well as other sophisticated combination methods in empirical studies. The early dialogue had sparkled thought-provoking works on the combination of forecasts, as documented in Clemen1989. More recent surveys on the forecast combinations are provided by Timmermann2006 as well as ElliottTimmermann2016a. Indeed, the dialogue on such a puzzle continues to this day. For example, acknowledging the equal-weight scheme as a high benchmark, DieboldShin2019 propose the egalitarian ridge regression, which is the ridge regression with shrinkage toward the equal-weight scheme.\footnote{ Alternatively, DieboldPauly1990 propose empirical Bayes forecasting procedures with shrinkage toward the equal-weight scheme. This Bayesian approach has been applied in, for example, StockWatson2004 and AiolfiTimmermann2006. } Their idea has heuristic appeal but leaves open the question of an appropriate computational method.
In this paper, we fortify the theoretical and computational foundations of DieboldShin2019's (DieboldShin2019) shrinkage approach by developing a real-time forecasting algorithm under the decision-making framework given that the growing literature on forward-looking models in economics highlights the role of forecasting in decision making.\footnote{ As indicated in ClaridaGaliEtAl2000 and Mavroeidis2010, a forecast-based interest rate rule in response to future macroeconomic conditions can provide a guideline for a monetary policy maker. Forecasting matters in not only a public policy but a private agent's decision as well. TanakaBloomEtAl2020 build a simple model concerning a firm's decisions on inputs under uncertainty to rationalize the empirical evidence that its GDP forecast accuracy is a predictor of profitability and productivity. } A decision maker first organizes egalitarian committees, that is, committees providing their forecasts, respectively, via a ridge regression with shrinkage toward the simple average of individual forecasts selected by mixed integer quadratic programming (MIQP). The application of MIQP and partition of a parameter space solve DieboldShin2019's computational difficulty in egalitarian ridge regression with simultaneous selection of individual forecasts. Next, the decision maker pools committee forecasts by applying a variant of FreundSchapire1997's (FreundSchapire1997) hedge algorithm. This variant depends on an estimate of the maximal committee loss for the duration of its implementation. The decision maker's two-stage implementation of real-time forecasting is referred to as hedge egalitarian committees algorithm, henceforth abbreviated to HECA.
We establish non-asymptotic upper bounds on the average regret attained by HECA, which is the decision maker's average forecasting loss in excess of the smallest average forecasting loss accomplished by these egalitarian committees. First, these upper bounds indicate the decision maker's own acumen of business cycles could pay off because given the committee forecasts, a more precise estimate of the maximal committee loss ceteris paribus yields a tighter upper bound on the average regret. Furthermore, these upper bounds show that HECA has no-regret property; that is, the decision maker's long-run performance should be at least as good as the best long-run performance accomplished by the egalitarian committees. This result is in line with the findings in the online learning literature. An excellent overview of this literature is recently provided by Cesa-BianchiOrabona2021. More importantly, this no-regret property implies the superiority of HECA relative to the equal-weight scheme in the long run. It is such `theoretical' superiority that makes HECA eligible for the competition with the equal-weight scheme; however, its `empirical' superiority remains to be examined.
To examine whether HECA outperforms the equal-weight scheme in an empirical study, we focus on the quarterly one-year-ahead forecasts of Euro-area real GDP growth in Survey of Professional Forecasters (SPF), which is conducted by the European Central Bank (ECB). The purpose of selecting this dataset is twofold. On the one hand, it generates the equal-weight scheme that has particularly hard-to-beat forecasting performance, as demonstrated in GenreKennyEtAl2013, ConflittiDeMolEtAl2015, and DieboldShin2019. On the other hand, it involves not-so-big data such that theoretical parts of data science (inclusive of domain knowledge, statistical methods, and computational techniques) are crucially important. Our empirical results show that HECA keeps pace with the equal-weight scheme before the outbreak of COVID-19 but wins the competition during the COVID-19 recession. Despite the superiority of HECA relative to the equal-weight scheme, HECA suffers from an upsurge in forecasting loss around the onset of COVID-19 pandemic. This pattern is consistent with previous research, as indicated in ChauvetPotter2013. We also find that the formation of egalitarian committees gives HECA an advantage over FreundSchapire1997's hedge algorithm during the COVID-19 recession.
In addition to the pursuit of forecasting performance, we are dedicated to credible forecasting in a spirit that only credible assumptions are maintained, as emphasized in Manski2013a. To achieve this goal, we treat the data generating process (DGP) of target variables and their individual forecasts as a black box; that is, no assumption on such DGP is imposed. Instead, the proposed HECA is a data-driven and adaptive approach: At each round, it outputs a combined forecast with more weights on forecasts provided by committees that have performed well in the past; furthermore, the built-in updating mechanism enables HECA to adapt to the environment in the presence of structural breaks that may make forecasters' relative performance unstable over time. The unstable performance is called model instability in literature, and an excellent survey of this issue is provided by Rossi2013. Additionally, the committee forecasts, as inputs of HECA, are obtained by the rolling egalitarian ridge regression scheme. The rolling scheme is used to guard against possible parameter drift, as indicated by West2006, whereas the shrinkage toward the simple average of selected individual forecasts is supported by empirical evidence in literature.
HECA embodies interdisciplinary research, which is another marked characteristic of data science.\footnote{ This interdisciplinary characteristic is vividly illustrated in Drew Conway's data science Venn diagram, which can be found at \url{http://drewconway.com/zia/2013/3/26/the-data-science-venn-diagram}. } Knowledge in econometric literature, statistical methods, numerical and computational techniques, and online learning modeling are woven into HECA for the decision maker's real-time forecasting. The empirical findings in econometrics treat the equal-weight scheme as a high benchmark, to which the weights should shrink. In spite of alleviating numerical instability, the ridge regression in general fails to achieve the simultaneous selection of individual forecasts, which is implemented by mixed integer optimization. The hedge algorithm further allows for the adaptability to sequential data in real time. Equipped with these designs, HECA complements, but does not replace, existing forecasting methods. It is particularly useful in the situation where the decision maker has limited access to predictors. For surveys of data-rich methods, the reader is referred to StockWatson2006 and ChauvetPotter2013.
Throughout this paper, we write $\Norm{z}_{1}$, $\Norm{z}_{2}$, and $\Norm{z}_{\infty}$ for the one-norm, two-norm, and infinity-norm of a generic column vector $z$ in an Euclidean space, respectively. We denote the collection of positive integers and the collection of real numbers by $\mathbb{N}$ and $\mathbb{R}$, respectively.
The structure of the remaining paper is as follows. Section (ref) describes features of data collected from the ECB SPF and Eurostat. Section (ref) presents the organization of egalitarian committees, the decision maker's HECA, and the theoretical upper bounds on the average regret of HECA. Section (ref) discusses the empirical results of applying HECA to real-time forecasting of the year-on-year growth rate of euro area. Section (ref) concludes. Technical proofs are deferred to the appendix.
The forecast target variables in this paper are the year-on-year Euro-area GDP growth estimates collected from Eurostat, the European Statistical Agency.\footnote{ These estimates are available at \url{https://ec.europa.eu/eurostat/web/national-accounts/data/other}. } Due to data revisions, several estimates for a given quarter are released by Eurostat. Following GenreKennyEtAl2013, we focus on the $t+45$ flash estimates, which are published about 45 days after the associated quarter, for our empirical study in Section (ref). The evaluation sample runs from the first quarter of 2012 to the third quarter of 2020.
As in GenreKennyEtAl2013, ConflittiDeMolEtAl2015, and DieboldShin2019, we focus on the quarterly one-year-ahead forecasts of Euro-area real GDP growth in the ECB SPF. These one-year-ahead forecasts, however, are actually six to eight months ahead. For example, in the questionnaire for the third quarter of 2018, macroeconomic experts participating in the SPF are asked for the expected year-on-year real GDP growth for the first quarter of 2019 and provided with the GDP growth for the first quarter of 2018 as a reference.
One noticeable feature in the SPF is the frequent entry, exit, and reentry of experts so that an unbalanced panel arises. As pointed out in GenreKennyEtAl2013, such an unbalanced panel may yield sampling distortions. To lessen the extent of undesirable distortions, we exclude experts who did not reply in two consecutive quarters during the evaluation period spanning from the first quarter of 2012 to the third quarter of 2020. After this removal, there remain $21$ experts. Hereafter, we focus on the forecasts provided by this filtered panel of experts. In Figure (ref), we mark a slot by the notation x if a forecast is provided by an associated expert for a specific quarter in the SPF; otherwise, we leave it blank. We further replace each missing value of a forecast with the simple average of the rest of the reported forecasts for the same quarter. For example, Expert 038 provides forecasts throughout the evaluation period except the one for the third quarter of 2015; this missing forecast is filled in with the simple average of forecasts provided by the other 19 experts, as the forecast associated with Expert 110 is also unreported.
Another well-known feature in the SPF is the forecast combination puzzle: It is hard for other sophisticated schemes to improve on the performance of equal-weight scheme, as indicated by GenreKennyEtAl2013 and ConflittiDeMolEtAl2015. Some rationales behind this puzzle are proposed in literature. From a theoretical perspective, as shown in Timmermann2006, the equal-weight scheme is optimal if individual forecast errors have the same variance and identical pairwise correlation. From a practical perspective, as noted in SmithWallis2009 and ConflittiDeMolEtAl2015, the finite-sample error and numerical instability may make the estimated optimal weights inferior to the equal-weight scheme in terms of forecasting performance.
Both perspectives are crucial in an empirical study using the evaluation sample from the SPF. First, Table (ref) and Figure (ref) show the sample variances and pairwise correlation coefficients, respectively, of individual forecast errors for the $21$ experts in the filtered panel. Since the sample variances are similar to each other whereas the sample correlation coefficients are centered around $0.995$, the hard-to-beat performance of equal-weight scheme is unsurprising. In addition, the asymptotic approximation of estimated weights is arguably imprecise because the overall sample size --- $35$ quarters --- is obviously small. More importantly, the numerical instability in the estimated optimal weights is severe; for example, the condition number associated with the ordinary least square regression using the evaluation sample is $31,112$.\footnote{ The (2-norm) condition number of a matrix $A$ is defined as
and often used to evaluate the stability of a linear system in numerical analysis. As argued in BelsleyKuhEtAl1980, “moderate to strong relations are associated with condition indexes of $30$ to $100$.”} This finding yields a clue as to the application of ridge regression, a classical approach in literature to alleviating numerical instability, to the estimation of optimal weights. Recognizing the remarkable performance of equal-weight scheme, DieboldShin2019 propose the egalitarian ridge regression, which is the ridge regression with shrinkage toward the equal-weight scheme. Following their approach, we further develop an algorithm in the next section that can select experts in each quarter and achieve some satisfactory objective in hindsight.
The fundamental importance of economic forecasting for forward-looking private agents and public policy makers motivates us to propose an algorithm that incorporates features of the SPF forecasts and outperforms the equal-weight scheme under the decision-making framework. Roughly speaking, we consider the situation where a single decision maker is allowed access to forecasts provided by anonymous experts using either quantitative models or model-free judgments, and such forecasts generating processes are unknown to the decision maker.\footnote{ These experts' potentially strategic behaviors are also ignored by the decision maker. For the strategic forecasting, we refer the reader to MarinovicOttavianiEtAl2013 and references cited therein. } This decision maker is assumed to minimize the cumulative squared loss without discounting. The squared loss can be replaced with other loss functions, for example those documented in Section 2.2 of ElliottTimmermann2016a, in the rest of this paper. We refrain from this replacement because the squared loss is used in common with the literature on forecast combination puzzle.
To elaborate on the proposed method, we now introduce notation. Suppose that there are $M$ experts providing a forecast of the target variable $y_{t}$, respectively, before its realization. These individual forecasts are denoted by $f_{t}\equiv(f_{t,1},\dots,f_{t,M})^{\top}$, where $f_{t, m}$ stands for the forecast of $y_{t}$ provided by expert $m\in\{1,\dots,M\}$.\footnote{ In our empirical analysis in Section (ref), the vector $f_{t}$ of forecasts in the SPF are six to eight months prior to the release of the $t + 45$ flash estimate $y_{t}$ by Eurostat. } Accessing the data encompassing current forecasts, realized target variables, and their corresponding forecasts, the decision maker announces his or her own forecast of $y_{t}$ by a two-stage method: At the first stage, the decision maker imagines $M$ committees $\{\mathcal{C}_{c}\}_{c=1}^{M}$, where committee $\mathcal{C}_{c}$ consists of $c$ members selected among $M$ experts; subsequently, each committee provides a forecast $\hat{y}_{t,c}$, which is a combination of forecasts provided by its $c$ members. At the second stage, this decision maker applies the hedge algorithm to $\{\hat{y}_{t,c}\}_{c=1}^{M}$ and then yields his or her own forecast of $y_{t}$. To complete the two-stage method, we explain how these committees $\{\mathcal{C}_{c}\}_{c=1}^{M}$ are formed at the first stage and how the hedge algorithm works at the second stage in the subsections below.
The imaginary committee $\mathcal{C}_{c}$ is organized by solving the following optimization problem for a fixed rolling window $r\in \mathbb{N}$ and every $\lambda$ in a pre-specified set $\Lambda$ of grids:\footnote{ As indicated in ElliottTimmermann2016a (ElliottTimmermann2016a, p.\ 378), the length of estimation window can be selected by the cross-validation method, which is however rarely done. }
where $\Norm{b}_{0}$ is the number of nonzero elements in $b\equiv(b_{1},\dots,b_{M})^{\top}$ and $\bm{1}$ is the $M$ dimensional column vector of ones. The rolling scheme is adopted to guard against possible parameter drift. Since all individual forecasts are measured on the same scale, they are not standardized in this ridge-type regression. In addition, the lag term can be set to be either $l=1$ or $l=2$ for real-time forecasting with the SPF forecasts. Let $\hat{\beta}_{t,c}(\lambda)$ be a solution to problem (P1) associated with $\lambda$, and $\iota_{m}\in \mathbb{R}^{M}$ be a unit vector with $m$-th element equal to one. The tuning parameter $\hat{\lambda}_{t,c}$ is selected by setting
where $r_{\lambda}\in \mathbb{N}$ denotes the number of periods for validation. The set
is called the egalitarian committee with $c$ members, for problem (P1) can be viewed as a subproblem of partial egalitarian ridge regression in DieboldShin2019. To see this, let
where $\hat{\beta}_{t,c}(\lambda)$ is a minimizer of (P1) for each $c\in\{1,\dots,M\}$ and a fixed $\lambda$. Although the objective function is discontinuous due to $\Norm{b}_{0}$, we have
Phrased differently, the partition of a parameter space according to the value of $\Norm{b}_{0}$ allows us to recover $\tilde{\beta}_{t}(\lambda)$. Therefore, DieboldShin2019's (DieboldShin2019) `one-step' partial egalitarian ridge regression can be equivalently implemented as long as problem (P1) is successfully solved for every $c$.\footnote{ Although DieboldShin2019's (DieboldShin2019) partial egalitarian ridge regression concerns the inclusion of Tibshirani1996's (Tibshirani1996) one-norm regularization in the objective function rather than in the constraints, the idea of partitioning a parameter space still works mutatis mutandis. }
To solve problem (P1), we recast it as the following MIQP:
If $\epsilon$ is the smallest machine-representable positive real number,\footnote{ As defined in Judd1998 (Judd1998, p.\ 30), a machine zero is referred to as a quantity equivalent to zero on a machine. The positive real number $\epsilon$ is not a machine zero, but every positive real number less than $\epsilon$ is a machine zero. } then problems (P1) and (P2) are computationally equivalent. The intuition is that under the constraints $d_{j}\epsilon\leq b_{j}\leq d_{j}$ and $d_{j}\in\{0,1\}$, the dummy variable $d_{j}=\mathbbm{1}_{[b_{j}>0]}$ indicates whether $b_{j}$ is positive; that is, expert $j$ is selected in the committee $\mathcal{C}_{c}$ with size $\Norm{b}_{0}$, which is equal to the sum of $d_{j}$'s. We summarize the discussion in the following proposition.
A conceptually simple method of solving problem (P2) is exhaustive enumeration. To see this, note that there are $\binom{M}{c}$ feasible choices of $d\equiv(d_{1},\dots,d_{M})^{\top}$ in problem (P2). For any given feasible $d$, this optimization problem is essentially the constrained ridge regression with the unknown parameter $b$. Implementing these $\binom{M}{c}$ ridge regressions thus suffices to solve problem (P2). We call this approach complete subset ridge regressions by analogy with complete subset regressions in ElliottGarganoEtAl2013. This exhaustive approach, however, may be computationally inefficient because every egalitarian committee is asked to provide a forecast $\hat{y}_{t,c}\equiv f_{t}^{\top}\hat{\beta}_{t,c}(\hat{\lambda}_{t,c})$; consequently, there are $\binom{M}{1}+\binom{M}{2}+\dots+\binom{M}{M}=2^{M}-1$ ridge regressions to be carried out for every $\lambda\in\Lambda$.
Instead of such an exhaustive search for $\{\hat{\beta}_{t,c}(\lambda)\}_{c=1}^{M}$ in the parameter space, the modern solver Gurobi can be used to implement the MIQP in problem (P2). The practical tractability of moderate-size MIQP, though NP-hard in nature, can be attributed to the rapid advances in computation power. According to BertsimasDunn2019, the overall speedup, inclusive of solvers and hardware, is approximately two trillion between 1991 and 2016. The advance of mixed integer optimization has sparked recent studies in statistics and econometrics, for example BertsimasKingEtAl2016 and ChenLee2018, among others. Details about algorithmic developments of mixed integer optimization can be found in JuengerLieblingEtAl2010, ConfortiCornuejolsEtAl2014, and references cited therein.
In econometrics and statistics, assumptions about the DGP of $(y_{t}, f_{t}^{\top})$ are usually imposed to establish theoretically nice properties of a forecasting method. ElliottTimmermann2016a (ElliottTimmermann2016a, p.\ 320), however, put it this way:
Additionally, the knowledge about how individual forecasts are generated by experts is unknown in principle to the decision maker. A practical example given in Diebold2015 is that forecasts are purchased from a vendor using proprietary models, which are not revealed to the decision maker. Recognizing such limits to knowledge, the decision maker attempts to neither model nor assume the DGP of $(y_{t}, f_{t}^{\top})$ and is dedicated to credible forecasting.\footnote{ As argued in Manski2013a, “[t]he fundamental difficulty of empirical research is to decide what assumptions to maintain.” He further argues that “[s]tronger assumptions yield conclusions that are more powerful but less credible.” } This decision maker thus deals with the real-time forecasting problem by the hedge algorithm based on the past performance of the egalitarian committees, as described in the next subsection.
After receiving forecasts $\hat{y}_{t}\equiv(\hat{y}_{t,1},\dots,\hat{y}_{t,M})^{\top}$ made by all egalitarian committees, the decision maker announces his or her own forecast, which is a weighted average of $\{\hat{y}_{t,c}\}_{c=1}^{M}$. Subsequently, the nature announces the realization of the target variable $y_{t}$. The sequence of target variables can be generated as in statistical models. For example, it can represent business cycles undulating along a trend, either deterministic or stochastic; it can exhibit structural breaks with changing points, either known or unknown; and it can describe switching among different states, either observed or unobservable. Further examples about statistical modeling can be found in Pesaran2015 and PenaTsay2021. Alternatively, this sequence of target variables can be adversarially generated as in game-theoretic models, where the nature attempts to maximize the decision maker's forecasting loss. Details about game-theoretic analysis can be found in Cesa-BianchiLugosi2006 and SchapireFreund2012. Briefly, the following happen in order for each round $t$:
Adopting the aforementioned strategy (i.e., announcement of $\hat{\hat{y}}_{t}$ in each round) the decision maker ex ante aims to obtain small average regret
by selecting the sequence $\{\pi_{t}\}_{t=1}^{T}$ of distributions. To achieve this goal, this decision maker selects $\{\pi_{t}\}_{t=1}^{T}$ by HECA, whose pseudocode is shown as Algorithm (ref).
HECA, adapted from the hedge algorithm in FreundSchapire1997, incorporates features of the decision maker's real-time forecasting based on the SPF forecasts. Suppose that the decision maker announces his or her forecast $\hat{\hat{y}}_{t}$ immediately after receiving the SPF forecasts. In this case, the decision maker does not know the realized forecasting loss $\{\ell_{t,c}\}_{c=1}^{M}$ until round $t+2$. Hence, in the first two rounds, the decision maker has no information about any committee's performance and thus uses the uniform distributions $\pi_{1}$ and $\pi_{2}$. In the third and ensuing rounds, the decision maker observes every committee's performance $\{\ell_{(t-2),c}\}_{c=1}^{M}$ of two rounds prior, thereby updating the distribution $\pi_{t}\equiv(\pi_{t,1},\dots,\pi_{t,M})^{\top}$. Note that HECA requires an estimate of the maximal committee loss throughout the $T$ rounds. The decision maker assumes $B_{1}$ to be this maximal loss, which might be a biased estimate, in the first two rounds, and updates it subsequently round by round.
HECA reflects the presumption embraced by the decision maker: A committee with relatively better performance (i.e., smaller $\ell_{(t-2)}$) would maintain the momentum to perform relatively well in the current round; therefore, its weight $\pi_{t,c}$ in the combined forecast $\hat{\hat{y}}_{t}$ should relatively increase. The idea of performance-based pooling of forecasts has been used in econometrics, for example forecasts weighted by inverse mean squared error in StockWatson1999 and CapistranTimmermann2009, and aggregated forecast through exponential reweighting in Yang2004 and WeiYang2012, among others. The built-in updating mechanism makes HECA adaptive to the environment. The adaptability of HECA is in sharp contrast to the constancy of equal-weight scheme even in the possibly ever-changing environment.
The following theorem gives upper bounds on the decision maker's average regret. These upper bounds are non-asymptotic; that is, they hold for every finite $T\in\mathbb{N}$.
The non-asymptotic upper bounds in Theorem (ref) hold without requirement of any assumption on the DGP. The decision maker may underestimate or overestimate the maximal loss $\bar{B}_{T}$. The biased estimation, however, is a contributing factor of an upper bound on the average regret. A sharper upper bound can be obtained if the magnitude of underestimation (overestimation), measured by $\gamma_{\text{u}}$ ($\gamma_{\text{o}}$), is smaller. It turns out that given the same committee forecasts and duration of implementing HECA, decision makers having different evaluations of business cycle may achieve different forecasting performances even if they use HECA. Calculation of these upper bounds is also easy ex post. In contrast, upper bounds in Yang2004 and WeiYang2012 involve nuisance parameters of the underlying DGP and need further evaluation.
Moreover, if the sequence $\{\bar{B}_{T}\}_{T=1}^{\infty}$ is bounded above,\footnote{ The monotonicity of $\{\bar{B}_{T}\}_{T=1}^{\infty}$, together with its boundedness, implies that $\lim_{T\to\infty}\bar{B}_{T}<\infty$. } then HECA exhibits no regret because these upper bounds on the average regret all shrink to zero as $T$ tends to infinity. Phrased differently, the performance of HECA is at least close to that of the best egalitarian committee in the long run. The no-regret property is common in the online learning literature. We refer the reader to Cesa-BianchiLugosi2006's (Cesa-BianchiLugosi2006) monograph for early findings and Cesa-BianchiOrabona2021's (Cesa-BianchiOrabona2021) survey for recent advances. From a pragmatic standpoint, the assumption about boundedness of $\{\bar{B}_{T}\}_{T=1}^{\infty}$ could be inconsequential. To see this, note that
by Jensen's inequality; additionally, as indicated by ElliottTimmermann2016a (ElliottTimmermann2016a, p.\ 17), “[i]n practice, forecasts are usually bounded and extremely large forecasts typically get trimmed as they are deemed implausible.”
More importantly, Theorem (ref) establishes the intuition that the decision maker's two-stage method should outperform the equal-weight scheme whenever $T$ is so large that the upper bounds are small. Theorem (ref) implies that in the long run, HECA should perform at least as well as the best egalitarian committee. In addition, the best egalitarian committee would dominate the egalitarian committee $\mathcal{C}_{M}$, which should in turn be weakly better than the equal-weight scheme. It follows from these arguments that HECA would outweigh the equal-weight scheme in terms of long-run forecasting performance.
If the decision maker regularly postpones announcing $\hat{\hat{y}}_{t}$ until the realization of $y_{(t-1)}$, then the information on $\{\ell_{(t-1),c}\}_{c=1}^{M}$ could be exploited. Because of the updated information, the decision maker's forecasting performance is expected to improve. Indeed, HECA with delayed announcements (Algorithm (ref)) allows for tighter upper bounds on the average regret, as shown in the following theorem.
Convinced of the theoretically asymptotic performance of HECA, we are now concerned with its forecasting performance for the evaluation sample mentioned in Section (ref).
We first concentrate on the competition among the equal-weight scheme, HECA (Algorithm (ref) with $l=2$) and HECA with delayed announcements (Algorithm (ref) with $l=1$). The latter two algorithms involve the first-stage MIQP, which is implemented by the Gurobi Python interface, and the tuning parameters are set by $r=16$, $r_{\lambda}=1$, $\Lambda=\{0.01g\}_{g=1}^{200}$, and $\epsilon=5\times 10^{(-324)}$.\footnote{The number $5\times 10^{(-324)}$ is equal to the product of sys.float_info.min and sys.float_info.epsilon in Python 3.7, and the mixed integer optimization are carried out by Gurobi 9.0.3, which is available at \url{https://www.gurobi.com/}. } We set $B_{1}$ to be the maximum of individual forecasting losses (observed by the decision maker) from the first quarter of 2012 to the quarter prior to $t=1$. HECA, in comparison to its counterpart with delayed announcements, requires two extra rounds for `in-sample' estimation of $\hat{\beta}$'s and validation of $\hat{\lambda}$'s. Thus, we consider $t=1$ in HECA to be the fourth quarter of 2016 and $t=1$ in HECA with delayed announcements to be the second quarter of 2016.
Table (ref) reports their forecasting losses and associated differences per round for this competition. As can be seen, HECA and the equal-weight scheme are nearly neck and neck until the first quarter of 2020. Keeping abreast of the equal-weight scheme, which is a well-known high benchmark, HECA also performs well. HECA further outperforms the equal-weight scheme since the second quarter of 2021, in which “the fall in economic activity was unprecedented in depth, speed and scope”.\footnote{ This description is given in the news published on 29th March 2021 by Euro Area Business Cycle Dating Committee. Further details are available at \url{https://eabcn.org/sites/default/files/eabcdc_findings_29_march_2021.pdf}. } HECA with delayed announcements even achieves better forecasting performance by exploiting updated information. Given the small-sample survey data from the SPF, we do not pursue statistical testing for the superiority of HECA because existing tests accounting for in-sample estimation error, for example the tests developed in DieboldMariano1995 and GiacominiWhite2006, rely on out-of-sample asymptotic approximation to determine the critical value.
Moreover, Figure (ref) implies that the number of experts in the committee performing best in a single round is not constant but time-varying. Despite this instability, the theoretical results in Section (ref) suggest that HECA could perform as well as the best committee over the entire evaluation period in hindsight. As can be seen from Table (ref), the average regret is relatively small given the substantial impact of COVID-19 pandemic on the euro area economy. It is worth noting that the best egalitarian committee in the fourth quarter of 2019 and in the third quarter of 2020 are identical, and the average regret is less than $0.03$ if HECA with/without delayed announcements terminates in the fourth quarter of 2019.
Finally, we turn the spotlight on the cousin and ancestor of HECA. As shown in Table (ref), the exponential fictitious play with updating mechanism ((ref)) performs almost the same as HECA; similarly, the exponential fictitious play with updating mechanism ((ref)) very much resembles HECA with delayed announcements in terms of forecasting ability. The close resemblance gives a hint on the no-regret property of exponential fictitious play. Table (ref) shows that although neither FreundSchapire1997's (FreundSchapire1997) hedge algorithm nor HECA dominates each other before the fourth quarter of 2019, HECA wins the competition since the first quarter of 2020. Thus, combining the results in Tables (ref) and (ref), we find that during the COVID-19 recession, the formation of egalitarian committees gives HECA a competitive edge over the hedge algorithm, which beats the equal-weight scheme by adaptability.
The proposed HECA should be in the data scientist's toolkit for three reasons as follows. From a theoretical perspective, HECA outputs credible forecasting because it relies on practically convincing assumptions and meanwhile achieves an asymptotically negligible upper bound on the average regret. From an empirical perspective, HECA outweighs the equal-weight scheme after the outbreak of COVID-19 in euro area, whereas the equal-weight scheme only outperforms HECA by a margin, if any, before such an outbreak. From a methodological perspective, HECA differs from other data-rich methods in that it is applicable in the context where no extra predictor of the target variables, except for the forecasts provided by the experts, is available for the decision maker.
We do not deal with the optimal timing of implementing HECA. Our empirical results seem to suggest that compared with the equal-weight scheme, HECA would be suitable for forecasting around business cycle turning points. It is also unclear whether the duration of implementing HECA should be determined at the very beginning. We delegate these fascinating issues for future work.