EconBase
← Back to paper

Macroeconomic Data Transformations Matter

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

79,818 characters · 10 sections · 76 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Macroeconomic Data Transformations Matter

\and Maxime Leroux$^2$ \and Dalibor Stevanovic$^2$. D\'{e}partement des sciences \'{e}conomiques, UQAM.} \and St\'{e}phane Surprenant$^2$ }

abstractIn a low-dimensional linear regression setup, considering linear transformations/combinations of predictors does not alter predictions. However, when the forecasting technology either uses shrinkage or is nonlinear, it does. This is precisely the fabric of the machine learning (ML) macroeconomic forecasting environment. Pre-processing of the data translates to an alteration of the regularization -- explicit or implicit -- embedded in ML algorithms. We review old transformations and propose new ones, then empirically evaluate their merits in a substantial pseudo-out-sample exercise. It is found that traditional factors should almost always be included as predictors and moving average rotations of the data can provide important gains for various forecasting targets. Also, we note that while predicting directly the average growth rate is equivalent to averaging separate horizon forecasts when using OLS-based techniques, the latter can substantially improve on the former when regularization and/or nonparametric nonlinearities are involved.

JEL Classification: C53, C55, E37

Keywords: Machine Learning, Big Data, Forecasting.

Introduction

Following the recent enthusiasm for Machine Learning (ML) methods and widespread availability of big data, macroeconomic forecasting research gradually evolved further and further away from the traditional tightly specified OLS regression. Rather, nonparametric non-linearity and regularization of many forms are slowly taking the center stage, largely because they can provide sizable forecasting gains when compared with traditional methods (see, among others, kim2018mining,medeiros2019forecasting,gclss2020,goulet2020macroeconomy), even during the Covid-19 episode gcms2021. In such environments, different linear transformations of the informational set $X$ can change the prediction and taking first differences may not be the optimal transformation for many predictors, despite the fact that it guarantees viable frequentist inference. For instance, in penalized regression problems -- like Lasso or Ridge --, different rotations of $X$ imply different priors on $\beta$ in the original regressor space. Moreover, in tree-based models algorithms, since the problem of inverting a near singular matrix $X'X$ simply does not happen, making the use of more persistent (and potentially highly cross-correlated regressors) much less harmful. In sum, in the ML macro forecasting environment, traditional data transformations -- such as those designed to enforce stationarity mccracken2016fred -- may leave some forecasting gains on the table. To provide guidance for the growing number of researchers and practitioners in the field, we conduct an extensive pseudo-out-of-sample forecasting exercise to evaluate the virtues of standard and newly proposed data transformations.

From the ML perspective, it is often suggested that a "feature engineering" step may improve algorithms' performance kuhn2019feature. This is especially true of Random Forests (RF) and Boosted Trees (BT), two regression tree ensembles widely regarded as the most performing off-the-shelf algorithms within the modern ML canon hastie2009elements. Among other things, both successfully handle a high-dimensional $X$ by recruiting relevant predictors in a sea of useless ones. This implies the data scientist leveraging some domain knowledge can create plausibly more salient features out of the original data matrix, and let the algorithm decide whether to use them or not. Of course, an extremely flexible model, like a neural network with many layers, could very well create those relevant transformations internally in a data-driven way. Yet, this idyllic scenario is a dead end when data points are few, regressors are numerous, and a noisy $y$ serves as a prediction target. This sort of environment, of which macroeconomic forecasting is a notable example, will often benefit from any prior knowledge one can incorporate in the model. Since transforming the data transforms the prior, doing so properly by including well-motivated rotations of $X$ has the power to increase ML performance on such challenging data sets.

Macroeconomic modelers have been thinking about designing successful priors for a long time. There is a wide literature on Bayesian Vector Autoregressions (VAR) starting with litterman1984. Even earlier on, the penalized/restricted estimation of lag polynomials was extensively studied almon1965distributed,shiller1973distributed. The motivation for both strands of work is the large ratio of parameters to observations. Forty years later, many more data points are available, but models have grown in complexity. Consequently, large VARs banbura2010large and MIDAS regression ghysels2004midas still use those tools to regularize over-parametrized models. ML algorithms, usually allowing for sophisticated functional forms, also critically rely on shrinkage. However, when it comes to nonlinear nonparametric methods -- especially Boosting and Random Forests -- there are no explicit parameters to penalize. Nevertheless, in the case of RF, the ensuing ensemble averaging prediction benefits from ridge-like shrinkage as randomization allows each feature to contribute to the prediction, albeit in a moderate way hastie2009elements,mentch2019randomization. Just like rotating regressors changes the prior in a Ridge regression (see discussion in GC2019), rotating regressors in such algorithms will alter the implicit shrinkage scheme -- i.e., move the prior mean away from the traditional zero. This motivates us to propose two rotations of $X$ that implicitly implement a more time-series-friendly prior in ML models: moving average factors (MAF) and moving average rotation of $X$ (MARX). Other than those motivated above, standard transformations are also being studied. This includes factors extracted by principal components of $X$ and the inclusion of variables in levels to retrieve low frequency information.

We are interested in predicting stationary targets through a direct (in opposition to iterated) forecasting approach. There are at least two ways one can construct direct forecasts of the average growth rate of a variable over the next $h>1$ months -- an important quantity for the conduct of monetary policy and fiscal planning. A popular approach is to forecast the final object of interest by projecting it directly on the informational set $X$ (e.g., stock2002forecasting). An alternative is the path average approach where every step until the final horizon is predicted separately. A potential benefit of fitting the whole path first and then constructing the final target is to allow for the selected predictors, the harshness of regularization, and the type of nonlinearities to fully adapt when different relationships arise among the variables during the path.\footnote{An obvious drawback is that implies estimating and tuning $h$ models rather than one.} Since those three modeling elements are wildly nonlinear operations in the original input, averaging the path before or after ML is performed can produce very different results.

To evaluate the contribution of data transformations for macroeconomic prediction, we conduct an extensive pseudo-out-of-sample forecasting experiment (38 years, 10 key monthly macroeconomic indicators, 6 horizons) with three linear and two nonlinear ML methods (Elastic Net, Adaptive Lasso, Linear Boosting, Random Forests, and Boosted Trees), and two standard econometric reference models (autoregressive and factor-augmented autoregression).

Main results can be summarized as follows. First, combining non-standard data transformations, MARX, MAF and Level, minimizes the RMSE for 8 and 9 variables out of 10 when respectively predicting at short horizons 1 and 3-month ahead. They remain resilient at longer horizons as they are part of best RMSE specifications around 80% of time. Second, their contribution is magnified when combined with nonlinear ML models -- 38 out of 47 cases\footnote{There are 47 cases where at least one of these transformations is used.} -- with an advantage for Random Forests over Boosted Trees. Both algorithms allow for nonlinearities via tree base learners and make heavy use of shrinkage via ensemble averaging. This is precisely the algorithmic environment we conjectured could benefit most from non-standard transformations of $X$. Third, traditional factors can help tremendously. The overwhelming majority of best information sets for each target included factors. On that regard, this amounts to a clear takeaway message: while ML methods can handle the high-dimensional $X$ (both computationally and statistically), extracting common factors remains straightforward feature engineering that pays off. \textbf{Fourth}, the path average approach is preferred to the direct counterpart for almost all real activity variables and at most horizons. Combined with high-dimensional methods that use some form of regularization improves predictability by as much as 30%.

The rest of the paper is organized as follows. In section (ref), we present the ML predictive framework and detail the data transformations and forecasting models. In section (ref), we detail the forecasting experiment and in section (ref) we present main results. Section (ref) concludes.

Machine Learning Forecasting Framework

Machine learning algorithms offer ways to approximate unknown and potentially complicated functional forms with the objective of minimizing the expected loss of a forecast over $h$ periods. The focus of the current paper is to construct a feature matrix susceptible to improve the macroeconomic forecasting performance of off-the-shelf ML algorithms. Let $H_t = \left[ H_{1t}, ..., H_{Kt} \right]$ for $t=1,...,T$ be the vector of variables found in a large macroeconomic dataset and let $y_{t+h}$ be our target variable that is supposed stationary. The corresponding prediction problem is given by

equation[equation omitted — 50 chars of source]

To illustrate the data pre-processing point, define $Z_t \equiv f_Z(H_t)$ as the $N_Z$-dimensional feature vector, formed by combining several transformations of the variables in $H_t$.\footnote{Obviously, in the context of a pseudo-out-of-sample experiment, feature matrices must be built recursively to avoid data snooping.} The function $f_Z$ represents the data pre-processing and/or featuring engineering whose effects on forecasting performance we seek to investigate. The training problem for $f_Z = I()$ is

equation[equation omitted — 153 chars of source]

The function $g$, chosen as a point in the functional space $\mathcal{G}$, maps transformed inputs into the transformed targets. $\text{pen()}$ is the regularization function whose strength depends on some vector/scalar hyperparameter(s) $\tau$. Let $\circ$ denote the function product and $\tilde{g} := g \circ f_Z$. Clearly, introducing a general $f_Z$ leads to

align*[align* omitted — 362 chars of source]

which is, simply, a change of regularization. Now, let $g^*(f_Z^*(H_t))$ be the "oracle" combination of best transformation $f_Z$ and true function $g$. Let $g(f_Z(H_t))$ be a functional form and data pre-processing selected by the practitioner. In addition, denote $\hat{g}(Z_t)$ and $\hat y_{t+h}$ the fitted model and its forecast. The forecast error can be decomposed as

equation[equation omitted — 201 chars of source]

While the intrinsic error $e_{t+h}$ is not shrinkable, the estimation error can be reduced by either adding more relevant data points or restricting the domain $\mathcal{G}$. The benefits of the latter can be offset by a corresponding increase of the approximation error. Thus, an optimal $f_Z$ is one that entails a prior that reduces estimation error at a minimal approximation error cost. Additionally, since most ML algorithms perform variable selection, there is the extra possibility of pooling different $f_Z$'s together and let the algorithm itself choose the relevant restrictions.\footnote{More concretely, a factor $F$ is a linear combination of $X$. If an algorithm pick $F$ rather than creating its own combination of different elements of $X$, it is implicitly imposing a restriction.}

The marginal impact of the increased domain $\mathcal{G}$ has been explicitly studied in gclss2020, with $Z_t$ being factors extracted from the stationarized version of FRED-MD. The primary objective of this paper is to study the relevance of the choice of $f_Z$, combined with popular ML approximators $g$.\footnote{There are many recent contributions considering the macroeconomic forecasting problem with econometric and machine learning methods in a big data environment kim2018mining, kotchoni2019macroeconomic. However, they are done using the standard stationary version of FRED-MD database. Recently, mccracken2020fred studied the relevance of unit root tests in the choice of stationarity transformation codes for macroeconomic forecasting with factor models.} To evaluate the virtues of standard and newly proposed data transformations, we conduct a pseudo-out-of-sample (POOS) forecasting experiment using various combinations of $f_Z$'s and $g$'s.

Finally, a question often overlooked in the forecasting literature is how one should construct the forecast for average growth/difference of the level variable $Y_t$, which is the popular target in macroeconomic applications. The usual approach -- and also the least computationally demanding -- is that of fitting the model on $y_{t+h}=\sfrac{\sum_{h'=1}^h \Delta Y_{t+h'}}{h}$ directly and using $\hat{y}_{t+h}^{\text{direct}}$ as prediction, where $\Delta Y_{t+h'} = Y_{t+h'}-Y_{t+h'-1}$ is the simple growth/difference of the variable of interest. Another approach, requiring the estimation of $h$ different functions, is the path average approach where each $\Delta Y_{t+h'}$ is fitted separately and the forecast for $y_{t+h}$ is obtained from $\hat{y}_{t+h}^{\text{path-avg}}=\sfrac{\sum_{h'=1}^h \widehat{\Delta Y}_{t+h'}}{h}$.

The common wisdom -- from OLS -- is that such strategies are interchangeable. But the equivalence does not hold when regularization and nonparametric nonlinearities are involved. For instance, it breaks in the simplest possible departure from OLS, a ridge regression, where

align[align omitted — 127 chars of source]

and only if $\lambda_{h'}=\lambda \enskip \forall h'$ then

align[align omitted — 154 chars of source]

This setup naturally includes the known equivalence in the OLS case ($\lambda_{h'}=0 \enskip \forall h')$. We get even further from the equivalence with Lasso, Random Forests, and Boosted Trees which all imply the nonlinear hard-thresholding operation of variable selection -- and basis expansion creation for the last two. With those, we get even further from the equivalence by having a different $Z_{h'}^* \subset Z$ in each prediction function.

Of course, the path average approach can be rather demanding since it implies $h$ estimation (and likely cross-validation) problems --- with the benefit of providing a whole path rather than merely $y_{t+h}$. The second question address then concerns whether those benefits could additionally include forecasting gains. To investigate this and how this choice interacts with the optimal $f_Z$, we conduct the whole forecasting exercise using both schemes.

Old News

Firstly, we consider more traditional candidates for $f_Z$.

\vskip 0.2cm

{\sc Including Factors.} Common practice in the macroeconomic forecasting literature is to rely on some variant of the transformations proposed by mccracken2016fred to obtain a stationary $X_t$ out of $H_t$. Letting $X = \left[ X_t \right]_{t=1}^T$ and imposing a linear latent factor structure $X = F\Lambda + \epsilon$, we can estimate $F$ by the principal components of $X$. The feature matrix of the autoregressive diffusion index (FM hereafter) model of stock2002forecasting, stock2002macroeconomic can be formed as

align[align omitted — 89 chars of source]

where $L$ is the lag operator and $y_t$ is the current value of the target. In gclss2020, factors were deemed the most reliable shrinkage method for macroeconomic forecasting, even when considering ML alternatives. Furthermore, the combination of factors (and nothing else) with nonlinear nonparametric methods is (i) easy, (ii) fast, and (iii) often quite successful. Point (iii) is further re-enforced by this paper's results, especially for forecasting inflation, which contrasts with the results found in medeiros2019forecasting.

\vskip 0.2cm

{\sc Including Levels.} In econometrics, debates on the consequences of unit roots for frequentist inference have a long history\footnote{See for example, phillips1991criticize, phillips1991optimal, sims1988bayesian, sims1990inference, sims1991understanding.}, just as does the handling of low frequency movements for macroeconomic forecasting Graham2006. Exploiting potential cointegration has been found useful to improve forecasting accuracy under some conditions (e.g., christoffersen1998cointegration, engle1987forecasting, hall1992cointegration). From the perspective of engineering a feature matrix, the error correction term could be obtained from a first step regression \`{a} la engle1987co and is just a specific linear combination of existing variables. When it is unclear which variables should enter the cointegrating vector -- or whether there exist any such vector -- one can alternatively include both variables in levels and differences into the feature matrix. This sort of approach has been pursued most notably by cook2017macroeconomic who combine variables in levels, first differences and even second differences in the feature matrix they provide to various neural network architectures in the forecasting of US unemployment data.\footnote{Another approach is to consider factor modeling directly with nonstationary data baing2004,PENA20061237,BANERJEE2014589.}

From a purely predictive point of view, using first differences rather than levels is a linear restriction (using the vector $[1, -1]$) on how $H_t$ and $H_{t-1}$ can jointly impact $y_t$. Depending on the prior/regularization being used with a linear regression, this may largely decrease the estimation error or inflate the approximation one.\footnote{A similar comment would apply to all parametric cointegration restrictions. For recent work on the subject, see for example chan2015nonlinear.} However, it is often admitted that in a time series context (even if Bayesian inference is left largely unaltered by non-stationarity sims1988bayesian), first differences are useful because they trim out low frequencies which may easily be redundant in large macroeconomic data sets. Using a collection of highly persistent time series in $X$ can easily lead to an unstable $X'X$ inverse (or even a regularized version). Such problems naturally extend to Lasso lee2018lasso. In contrast, tree-based approaches like RF and Boosted Trees do not rely on inverting any matrix. Of course, performing tree-like sample splitting on a trending variable like raw GDP (without any subsequent split on lag GDP), is almost equivalent to split the sample according to a time trend and will often be redundant and/or useless. Nevertheless, there are numerous $H_t$'s where opting for first differencing the data is much less trivial. In such cases, there may be forecasting benefits from augmenting the usual $X$ with levels.

New Avenues

When regressors outnumber observations, regularization, whether explicit or implicit, is necessary. Hence, the ML algorithms we use all entail a prior which may or may not be well suited for a time series problem. There is a wide Bayesian VAR literature, starting with litterman1984, proposing prior structures that are thought for the multiple blocks of lags characteristic of those models. Additionally, there is a whole strand of older literature that seeks to estimate restricted lag polynomials in Autoregressive Distributed Lags (ARDL) models almon1965distributed,shiller1973distributed. While the above could be implemented in a parametric ML model with a moderate amount of pain, it is not clear how such priors framed in terms of lag polynomials can be put to use when there is no explicit lag polynomial. A more convenient approach is to (i) observe that most nonparametric ML methods implicitly shrink the individual contribution of each feature to zero in a Ridge-ean fashion hastie2009elements,elliott2013complete and (ii) rotating regressors implies a new prior in the original space. Hence, by simply creating regressors that embody the more sophisticated linear restrictions, we obtain shrinkage better suited for time series.\footnote{A cross-section RF-based example is rodriguez2006rotation who propose "Rotation Forest" that build an ensemble of trees based on different rotations of $X$.} A first step in that direction is goulet2020macroeconomy who proposes Moving Average Factors to specifically enhance RF's prediction and interpretation potential. A second is to find a rotation of the original lag polynomial such that implementing Ridge-ean shrinkage in fact yields shiller1973distributed approach to shrinking lag polynomials.

\vskip 0.2cm

{\sc Moving Average Factors.} Using factors is a standard approach to summarize parsimoniously a panel of heavily cross-correlated variables. Analogously, one can extract a few principal components from each variable-specific panel of lagged values, i.e.

align[align omitted — 186 chars of source]

to achieve a similar goal on the time axis. Define a moving average factor as the vector $M_{k}$.\footnote{While we work directly with the latent factors, a related decomposition called singular spectrum analysis works with the estimate of the summed common components, i.e. with $M_k \Gamma_k'$. Since this decomposition naturally yields a recursive formula, it has been used to forecast macroeconomic and financial variables hassani2009forecasting, hassani2013predicting, usually in an univariate fashion.} Mechanically, we obtain weighted moving averages, where the weights are the principal component estimates of the loadings in $\Gamma_k$. By construction, those extractions form moving averages of the $P_{MAF}$ lags of $X_{t,k}$ so that it summarizes most efficiently its temporal information.\footnote{$P_{MAF}$ is a tuning parameter analogous to the construction of the panel of variables (usually taken as given) in a standard factor model. We pick $P_{MAF}=12$. We keep two MAFs for each series and they are obtained by PCA.} By doing so, the goal to summarize information in $X_{t,k}^{1:P_{MAF}}$ is achieved without modifying any algorithm: we can use the MAFs which compresses information ex-ante. As it is the case for standard factors, MAF are designed to maximize the explained variance in $X_{t,k}^{1:P_{MAF}}$, not the fit to the final target. It is the learning algorithm's job to select the relevant linear combinations to maximize the fit.

\vskip 0.2cm

{\sc Moving Average Rotation of $X$.} There are many ways one can penalize a lag polynomial. One, in the Minnesota prior tradition, is to shrink all lags coefficients to zero (except for the first self-lag) with increasing harshness in $p$, the order of the lag. Another is to shrink each $\beta_p$ to $\beta_{p-1}$ and $\beta_{p+1}$ rather than to zero. Intuitively, for higher-frequency series (like monthly data would qualify for here) it is more plausible that a simple linear combination of lags impacts $y_t$ rather than a single one of them with all other coefficients set to zero.\footnote{This is basically a dense vs sparse choice. MAFs go all the way with the first view by imposing it via the extraction procedure.} For instance, it seems more likely that the average of March, April, and May employment growth could impact, say, inflation, than only May's. Mechanically, this means we expect March, April, and May 's coefficients to be close to one another, which motivated the prior ${\beta_{p}}\sim N(\beta_{p-1},\sigma_u^2 I_K)$ and more sophisticated versions of it in other works shiller1973distributed. Inputting in the ML algorithm a transformed $X$ such that its implicit shrinkage to zero is twisted into this new prior could generate forecasting gains. The only question left is how to make this operational.

The following derivation is a simple translation of GC2019's insights for time-varying parameters model to regularized lag polynomials \`{a} la shiller1973distributed.\footnote{Such reparametrization schemes are also discussed for "fused" Lasso in tibshirani2015statistical and employed for a Bayesian local-level model in koop2003bayesian.} Consider a generic regularized ARDL model with $K$ variables

align[align omitted — 196 chars of source]

where $\beta_p \in {\rm I\!R}^K$, $X_t \in {\rm I\!R}^K$, $u_p \in {\rm I\!R}^{K \times P}$, and both $y_t$ and $\epsilon_t$ are scalars.\footnote{We use $P$ as a generic maximum number of lags for presentation purposes. In Table (ref) we define $P_{MARX}$.} While we adopt the $l_2$ norm for this exposition, our main goal is to extend traditional regularized lag polynomial ideas to cases where there is no explicitly specified norm on $\beta_p - \beta_{p-1}$. For instance, elliott2013complete prove that their Complete Subset Regression procedure implies Ridge shrinkage in a special case. Moving away from linearity makes formal arguments more difficult. Nevertheless, it has been argued several times that model/ensemble averaging performs shrinkage akin to that of a ridge regression hastie2009elements. For instance, random selection of a subset of eligible features at each split encourage each feature to be included in the predictive function, but in a moderate fashion.\footnote{Recently, MSoRF argued that ensemble averaging methods \`{a} la RF prunes a latent tree. Following this view, the need for cleverly pre-assembled data combinations is even clearer.} The resulting "implicit" coefficient is an average of specifications that included the regressor and some that did not. In the latter case, the coefficient is always zero by construction. Hence, the ensemble shrinks contributions towards zero and the so-called mtry hyperparameter guides the level of shrinkage like a bandwidth parameter would olson2018making.

To get implicit regularized lag polynomial shrinkage, we now rewrite problem (ref) as a ridge regression. For all derivations to come, it is less tedious to turn to matrix notations. The Fused Ridge problem is now written as

align*[align* omitted — 172 chars of source]

where $\boldsymbol{D}$ is the first difference operator. The first step is to reparametrize the problem by using the relationship $\beta_k = C \theta_k$ that we have for all $k$ regressors. $C$ is a lower triangular matrix of ones (for the random walk case) and define ${\theta_k} = [{u_k} \quad {\beta_{0,k}}]$. For the simple case of one parameter and $P=4$: \[

bmatrix[bmatrix omitted — 71 chars of source]

=

bmatrix[bmatrix omitted — 94 chars of source]
bmatrix[bmatrix omitted — 59 chars of source]

. \] For the general case of $K$ parameters, we have $$ \boldsymbol{\beta}= \boldsymbol{C} \boldsymbol{\theta}, \quad \boldsymbol{C} \equiv I_K \otimes C $$ and $\boldsymbol{\theta}$ is just stacking all the $\theta_k$ into one long vector of length $KP$. Using the reparametrization $\boldsymbol{\beta}= \boldsymbol{C} \boldsymbol{\theta}$, the Fused Ridge problem becomes

align*[align* omitted — 184 chars of source]

Let $\boldsymbol{Z} \equiv \boldsymbol{XC}$ and use the fact that $\boldsymbol{D} = \boldsymbol{C}^{-1}$ to obtain the Ridge regression problem

align[align omitted — 194 chars of source]

We arrived at destination. Using $\boldsymbol{Z}$ rather than $\boldsymbol{X}$ in an algorithm that performs shrinkage will implicitly shrink $\beta_p$ to $\beta_{p-1}$ rather than to 0. This is obviously much more convenient than modifying the algorithm itself and is directly applicable to any algorithm using time series data as input. One question remains: what is $\boldsymbol{Z}$, exactly? For a single polynomial at time $t$, we have $Z_{t,k} = X_{t,k} C$. $C$ is gradually summing up the columns of $X_{t,k}$ over $p$. Thus, $Z_{t,k,p}=\sum_{p'=1}^{P} X_{t,k,p'}$. Dividing each $Z_{t,k,p}$ by $p$ (just another linear transformation, $\tilde{Z}_{t,k,p}$ ), it is now clear that $\tilde{\boldsymbol{Z}}$ is a matrix of moving averages. Those are of increasing order (from $p=1$ to $p=P$) and the last observation in the average is always $X_{t-1,k}$. Hence, we refer to this particular form of feature engineering as Moving Average Rotation of $X$ (MARX).

\vskip 0.2cm

{\sc Recap.} We summarize our setup in Table (ref). We have five basic sets of transformations to feed the approximation of $f_Z^*$: (1) single-period differences and growth rates following mccracken2016fred ($X_t$ and their lags), (2) principal components of $X_t$ ($F_t$ and their lags), (3) variables in levels ($H_t$ and their lags), (4) moving average factors of $X_t$ ($MAF_t$), and (5) sets of simple moving averages of $X_t$ ($MARX_t$). We consider several forecasting models in order to approximate the true functional form: Autoregressive (AR), Factor Model (FM, \`{a} la stock2002forecasting), Adaptive Lasso (AL), Elastic Net (EN), Linear Boosting (LB), Random Forest (RF), and Boosted Trees (BT). Lastly, we apply those specifications to forecasting both direct and path-average targets. The details on forecasting models are presented in Appendix (ref).

Furthermore, most ML methodologies that handle well high-dimensional data perform some form or another of variable selection. For instance, RF evaluates a certain fraction of predictors at each split and selects the most potent one. Lasso selects relevant predictors and shrinks others perfectly to zero. By rotating $X$, we can get these algorithms (and others) to perform restriction/transformation selection. Thus, one should not refrain from studying different combinations of $f_Z$'s.\footnote{Notwithstanding, some authors have noted that a trade-off emerges between how focused a RF is and its robustness via diversification. borup2020targeting sometimes get improvements over plain RF by adding a Lasso pre-processing step to trim $X$.} As a result, all the combinations of $f_Z$ thereof are admissible and 16 of them are included in the exercise. Moreover, there is a long-standing worry that well-accepted transformations may lead to some over-differenced $X_k$'s mccracken2020fred. Including MARX or MAF (which are both specific partial sums of lags) with $X$ can be seen as bridging the gap between a first difference and keeping $H_k$ in levels. Hence, interacting many $f_Z$ is not only statistically feasible, but econometrically desirable given the sizable uncertainty surrounding what is a "proper" transformation of the raw data choi2015almost.

table[table omitted — 2,622 chars of source]

Forecasting Setup

In this section, we present the results of a pseudo-out-of-sample forecasting experiment for a group of target variables at monthly frequency from the FRED-MD dataset of mccracken2016fred. Our target variables are the industrial production index (INDPRO), total nonfarm employment (EMP), unemployment rate (UNRATE), real personal income excluding current transfers (INCOME), real personal consumption expenditures (CONS), retail and food services sales (RETAIL), housing starts (HOUST), M2 money stock (M2), consumer price index (CPI), and the production price index (PPI). Given that we make predictions at horizons of 1, 3, 6, 9, 12, and 24 months, we are effectively targeting the average growth rate over those periods, except for the unemployment rate for which we target average differences. These series are representative macroeconomic indicators of the US economy, as stated in kim2018mining, which is also based on gclss2020 exercise for many ML models, itself based on kotchoni2019macroeconomic and a whole literature of extensive horse races in the spirit of SW1998comparison. The POOS period starts in January of 1980 and ends in December of 2017. We use an expanding window for estimation starting from 1960M01. Following standard practice in the literature, we evaluate the quality of point forecasts using the root Mean Square Error (RMSE). For the forecasted value at time $t$ of variable $v$ made $h$ steps before, we compute

align[align omitted — 116 chars of source]

The standard dieboldmariano (DM) test procedure is used to compare the predictive accuracy of each model against the reference factor model (FM). RMSE is the most natural loss function given that all models are trained to minimize the squared loss in-sample. We also implement the Model Confidence Set (MCS) that selects the subset of best models at a given confidence level MCS2011.

Hyperparameter selection is performed using the BIC for AR and FM and K-fold cross-validation is used for the remaining models. This approach is theoretically justified in time series models under conditions spelled out by bergmeir2018note. Moreover, gclss2020 compared it with a scheme which respects the time structure of the data and found K-fold to be performing as well as or better than this alternative scheme. All models are estimated every month while their hyperparameters are reoptimized every two years.

Results

Table (ref) shows the best RMSE data transformation combinations as well as the associated functional forms for every target and forecasting horizon. It summarizes the main findings and provide important recommendations for practitioners in the field of macroeconomic forecasting. First, including non-standard choices of macroeconomic data transformation, MARX, MAF and Level, minimize the RMSE for 8 and 9 variables out of 10 when respectively predicting 1 and 3-month ahead. Their overall importance is still resilient at longer horizons as they are part of best specifications most of the variables. Second, their success is often paired with a nonlinear functional form $g$, 38 out of 47 cases, with an advantage for Random Forests over Boosted Trees. The former is used for 26 of those 38 cases. Both algorithms make heavy use of shrinkage and allow for nonlinearities via tree base learners. This is precisely the algorithmic environment that we precedently conjectured to be where data transformations matter.

footnotesize\begin{ThreePartTable} \begin{longtable}{l|lll|llll|lll} \caption{Best model specifications - with target type} \\ \hline & INDPRO & EMP & UNRATE & INCOME & CONS & RETAIL & HOUST & M2 & CPI & PPI \\ \hline \endhead \hline \endfoot H=1 & RF{\color{ForestGreen}$\medbullet$}{\color{Cyan}$\medbullet$}{\color{RubineRed}$\medbullet$}{\color{blue}$\medbullet$} & RF{\color{ForestGreen}$\medbullet$}{\color{Cyan}$\medbullet$}{\color{RubineRed}$\medbullet$}{\color{blue}$\medbullet$} & BT{\color{ForestGreen}$\medbullet$}{\color{RubineRed}$\medbullet$} & RF{\color{RubineRed}$\medbullet$} & FM{\color{ForestGreen}$\medbullet$} & FM{\color{ForestGreen}$\medbullet$} & EN{\color{ForestGreen}$\medbullet$}{\color{blue}$\medbullet$} & RF{\color{Cyan}$\medbullet$}{\color{blue}$\medbullet$} & AL{\color{RubineRed}$\medbullet$} & EN{\color{ForestGreen}$\medbullet$}{\color{RubineRed}$\medbullet$} \\ H=3 & RF{\color{RubineRed}$\medbullet$} & RF{\color{ForestGreen}$\medbullet$}{\color{RubineRed}$\medbullet$} & RF{\color{ForestGreen}$\medbullet$}{\color{Cyan}$\medbullet$}{\color{RubineRed}$\medbullet$}{\color{blue}$\medbullet$} & RF{\color{ForestGreen}$\medbullet$}{\color{RubineRed}$\medbullet$} & RF{\color{ForestGreen}$\medbullet$}{\color{blue}$\medbullet$} & BT{\color{ForestGreen}$\medbullet$}{\color{Cyan}$\medbullet$}{\color{RubineRed}$\medbullet$} & EN{\color{ForestGreen}$\medbullet$}{\color{blue}$\medbullet$} & AL{\color{Cyan}$\medbullet$}{\color{blue}$\medbullet$} & RF{\color{ForestGreen}$\medbullet$} & EN{\color{RubineRed}$\medbullet$} \\ H=6 & \underline{RF}{\color{RubineRed}$\medbullet$} & \underline{BT}{\color{ForestGreen}$\medbullet$}{\color{RubineRed}$\medbullet$} & \underline{RF}{\color{ForestGreen}$\medbullet$}{\color{RubineRed}$\medbullet$} & \underline{RF}{\color{ForestGreen}$\medbullet$}{\color{Cyan}$\medbullet$}{\color{RubineRed}$\medbullet$} & \underline{RF}{\color{ForestGreen}$\medbullet$}{\color{blue}$\medbullet$} & AL{\color{ForestGreen}$\medbullet$}{\color{RubineRed}$\medbullet$} & \underline{RF}{\color{ForestGreen}$\medbullet$}{\color{Cyan}$\medbullet$}{\color{RubineRed}$\medbullet$} & RF{\color{ForestGreen}$\medbullet$}{\color{blue}$\medbullet$} & RF{\color{ForestGreen}$\medbullet$} & RF{\color{ForestGreen}$\medbullet$} \\ H=9 & \underline{RF}{\color{RubineRed}$\medbullet$} & \underline{BT}{\color{ForestGreen}$\medbullet$}{\color{RubineRed}$\medbullet$} & \underline{LB}{\color{ForestGreen}$\medbullet$}{\color{Cyan}$\medbullet$}{\color{RubineRed}$\medbullet$}{\color{blue}$\medbullet$} & \underline{RF}{\color{ForestGreen}$\medbullet$}{\color{RubineRed}$\medbullet$} & \underline{RF}{\color{Orange}$\medbullet$} & BT{\color{ForestGreen}$\medbullet$}{\color{Cyan}$\medbullet$}{\color{RubineRed}$\medbullet$}{\color{blue}$\medbullet$} & BT{\color{ForestGreen}$\medbullet$}{\color{Orange}$\medbullet$} & RF{\color{ForestGreen}$\medbullet$}{\color{blue}$\medbullet$} & RF{\color{ForestGreen}$\medbullet$} & RF{\color{ForestGreen}$\medbullet$} \\ H=12 & \underline{RF}{\color{RubineRed}$\medbullet$} & \underline{BT}{\color{ForestGreen}$\medbullet$}{\color{RubineRed}$\medbullet$} & \underline{LB}{\color{ForestGreen}$\medbullet$}{\color{Cyan}$\medbullet$}{\color{RubineRed}$\medbullet$}{\color{blue}$\medbullet$} & \underline{RF}{\color{ForestGreen}$\medbullet$}{\color{RubineRed}$\medbullet$} & \underline{RF}{\color{ForestGreen}$\medbullet$}{\color{Orange}$\medbullet$} & BT{\color{ForestGreen}$\medbullet$}{\color{Cyan}$\medbullet$}{\color{blue}$\medbullet$} & RF{\color{ForestGreen}$\medbullet$} & BT{\color{ForestGreen}$\medbullet$}{\color{blue}$\medbullet$} & RF{\color{ForestGreen}$\medbullet$} & RF{\color{ForestGreen}$\medbullet$} \\ H=24 & RF{\color{ForestGreen}$\medbullet$}{\color{blue}$\medbullet$} & \underline{BT}{\color{Orange}$\medbullet$} & BT{\color{ForestGreen}$\medbullet$}{\color{Orange}$\medbullet$} & \underline{RF}{\color{ForestGreen}$\medbullet$}{\color{Cyan}$\medbullet$}{\color{RubineRed}$\medbullet$} & \underline{RF}{\color{ForestGreen}$\medbullet$}{\color{Orange}$\medbullet$} & BT{\color{ForestGreen}$\medbullet$}{\color{Cyan}$\medbullet$}{\color{Orange}$\medbullet$} & RF{\color{ForestGreen}$\medbullet$} & \underline{RF}{\color{ForestGreen}$\medbullet$}{\color{blue}$\medbullet$} & \underline{RF}{\color{Cyan}$\medbullet$} & BT{\color{ForestGreen}$\medbullet$}{\color{blue}$\medbullet$} \end{longtable} \end{ThreePartTable}

{\flushleft

scriptsize{ \singlespacing Note: Bullet colors represent data transformations included in the best model specifications: {\color{ForestGreen}F}, {\color{RubineRed}MARX}, {\color{Cyan}X}, {\color{blue}\emph{\textbf{L}}} and {\color{Orange}\emph{\textbf{MAF}}}. Path average specifications are underlined.}

}

Without a doubt, the most visually obvious feature of Table (ref) is the abundance of green bullets. As expected, transforming $X$ into factors is probably the most effective form of feature engineering available to the macroeconomic forecaster. Factors are included as part of the optimal specification for the overwhelming majority of targets. Furthermore, including factors only in combination with RF is the best forecasting strategy for both CPI and PPI inflation for the vast majority of horizons. This is in line with findings in gclss2020 but in contrast with the results found in medeiros2019forecasting. The major difference with the latter is that they estimate and evaluate models on the basis of single month inflation rate, which is only the intermediary step in our path average strategy. In addition, we explore the possibility that $F$ alone could be better than $X$, rather than always both together. As it turns out, the winning combination is RF using factors as sole inputs to directly target the average growth. Finally, the omission of factors from optimal specifications for industrial production growth 3 to 12 months ahead is naturally surprising. This points out that current wisdom based on linear models may not be directly applicable to nonlinear ones. In fact, alternative rotations will sometimes do better.

There is plentiful of red bullets populating the top rows of Table (ref). Indeed, our most salient new transformation is MARX. In combination with nonlinear tree-based models, it contributes to improve forecasting accuracy for real activity series such as industrial production, employment, unemployment rate, and income, while they are best paired with penalized regressions to predict the CPI and PPI inflation rates. The dominance of MARX is particularly striking for real activity series as the transformation is included in every best specification for those variables at all horizons ranging from one month to a year. We further investigate how those RMSE gains materialize in terms of forecasts around key periods in section (ref). While MAF performance is often positively correlated with MARX, the latter is usually the better of the two, except for longer-run forecasts -- like those 2-years where MAF is featured for four variables.

Considering levels is particularly important for the M2 money stock as it is included in the best model for all horizons. For other variables, its pertinence is rather sporadic, with at least two horizons featuring it for INDPRO, UNRATE, CONS, and RETAIL.

The preference for $\hat{y}_{t+h}^{\text{direct}}$ vs $\hat{y}_{t+h}^{\text{path-avg}}$ mostly go on a variable by variable basis. However, there is clear consensus $\hat{y}_{t+h}^{\text{path-avg}} \succ \hat{y}_{t+h}^{\text{direct}}$ for all variables which strongly co-move with the business cycle (INDPRO, EMP, UNRATE, INCOME, CONS) with the notable exception of retail sales and housing starts. When it comes to nominal targets (M2, CPI, PPI), $\hat{y}_{t+h}^{\text{path-avg}} \prec \hat{y}_{t+h}^{\text{direct}}$ is unanimous for horizons 6 to 12 months, and so are the affiliated data transformations as well as the $g$ choice (all tree ensembles, with 8 out of 9 being RF). The quantitative importance of both types of gains on both sides is studied in section (ref), while section (ref) looks at implied forecasts to understand when and why $\hat{y}_{t+h}^{\text{path-avg}} \succ \hat{y}_{t+h}^{\text{direct}}$, or the reverse.

These findings are particularly important given the increasing interest in ML macro forecasting. They suggest that traditional data transformations, meant to achieve stationarity, do leave substantial forecasting gains on the practitioners' table. These losses can be successfully recovered by combining ML methods with well-motivated rotations of predictors such as MARX and MAF (or sometimes by simply including variables in levels) and by constructing the final forecast by the path average approach.

The previous results were desirably expeditive. The detailed results on the underlying performance gains and their statistical significance are presented in Appendix (ref).

Marginal Contribution of Data Pre-processing

In order to disentangle marginal effects of data transformations on forecast accuracy we run the following regression inspired by CARRIERO20191226 and gclss2020:

equation[equation omitted — 97 chars of source]

where $R^2_{t,h,v,m} \equiv 1 - \frac{e^2_{t,h,v,m}}{\frac{1}{T} \sum_{t=1}^T (y_{v,t+h} - \bar{y}_{v,h})^2}$ is the pseudo-out-of-sample $R^2$, and $e^2_{t,h,v,m}$ are squared prediction errors of model $m$ for variable $v$ and horizon $h$ at time $t$. $\psi_{t,v,h}$ is a fixed effect term that demeans the dependent variable by “forecasting target,” that is a combination of $t$, $v$, and $h$. $\alpha_{\mathcal{F}}$ is a vector of $\alpha_{\mathit{MARX}}$, $\alpha_{\mathit{MAF}}$, and $\alpha_{\mathit{F}}$ terms associated to each new data transformation considered in this paper, as well as to the factor model. $H_0$ is $\alpha_f =0 \quad \forall f \in \mathcal{F} = [\mathit{MARX}, \ \mathit{MAF}, \ \mathit{F}]$. In other words, the null is that there is no predictive accuracy gain with respect to a base model that does not have this particular data pre-processing. While the generality of ((ref)) is appealing, when investigating the heterogeneity of specific partial effects, it will be much more convenient to run specific regressions for the multiple hypothesis we wish to test. That is, to evaluate a feature $f$, we run

align[align omitted — 118 chars of source]

where $ \mathcal{M}_f$ is defined as the set of models that differs only by the feature under study $f$.

figure[figure omitted — 626 chars of source]

{\sc MARX.} Figure (ref) plots the distribution of $\alpha_{\mathit{MARX}}^{(h,v)}$ from equation ((ref)) done by $(h,v)$ subsets. Hence, we allow for heterogeneous effects of the MARX transformation according to 60 different targets. The marginal contribution of MARX on the pseudo-$R^2$ depends a lot on models, horizons, and series. However, we remark that at the short-run horizons, when combined with nonlinear methods, it produces positive and significant effects. It particularly improves the forecast accuracy for real activity series like industrial production, labor market series and income, even at larger horizons. For instance, the gains from using MARX with RF achieve 16% when predicting INDPRO at the $h=3$ horizon, and 14% in the case of employment if $h=6$. When used with linear methods, the estimates are more often on the negative side, except for inflation rates and M2 at short horizons, and a few special cases at the one and two-year ahead horizons.

figure[figure omitted — 745 chars of source]

{\sc Direct vs Path Average.} Figure (ref) reports the most unequivocal result of this paper: $\hat{y}_{t+h}^{\text{direct}}$ can prove largely suboptimal to $\hat{y}_{t+h}^{\text{path-avg}}$. For every method using a high-dimensional $Z_t$ shrunk in some way, i.e., not the OLS-based AR and FM, $\hat{y}_{t+h}^{\text{path-avg}}$ will do significantly better than the direct approach, with $\alpha_{\mathit{\text{path-avg}}}^{(h,v)}$ sometimes around 30% and highly statistically significant. As mentioned earlier, those gains are most prevalent for the highly cyclical variables {and longer horizons}. Cases where $\hat{y}_{t+h}^{\text{path-avg}} \prec \hat{y}_{t+h}^{\text{direct}}$ are rare and usually not statistically significant at the 5% level, except for AR and FM which are both fitted by OLS.

How to explain this phenomenon? Aggregating separate horizon forecasts allows to leverage the "bet on sparsity" principle of hastie2015statistical. Presume the model for $\widehat{\Delta Y}_{t+h'}$ is sparse for each $h'$, yet different. This implies that the direct model for $\hat{y}_{t+h}^{\text{direct}}$ is dense, and a much harder problem to learn. RF, BT, and Lasso will all perform better under sparsity, as every model struggle in a truly dense environment (unless it has a factor structure, upon which it becomes sparse in rotated space). An implication of this is that one should, as much as possible, try to make the problem sparse. Yet, whether sparsity will be more prevalent for $\hat{y}_{t+h}^{\text{path-avg}}$ or $\hat{y}_{t+h}^{\text{direct}}$ depends on true DGP. The evidence from Figure (ref) suggests that DGPs favoring $\hat{y}_{t+h}^{\text{path-avg}}$ are more prevalent in our experiment. What do those look like?

We find it useful to connect this question to recent works on forecasts aggregation, like Bermingham2014 who forecast the year on year inflation and compare two strategies: forecasting overall inflation directly vs forecasting individual elements of the consumption basket and using a weighted average of forecasts. They find that using more components and aggregating individual forecasts improves performance.\footnote{In a similar vein, marcellino2003macroeconomic found that forecasting inflation at the country level and then aggregating the forecasts increases does better than forecasting at the aggregate level (Euro).} They provide a simple example to rationalize their result: forecasting an aggregate variable made of two series with differing levels of persistence using only past values of the aggregate will be misspecified. In ML forecasting context, where $Z$ contains "everything" anyway, this problem translates from misspecification into making once sparse problems into a dense one, which is harder to learn. Consider a toy multi-horizon problem

equation[equation omitted — 329 chars of source]

where one needs to select a single predictor for each horizon. In this simple analogy to a high-dimensional problem, unless $k^*(1)=k^*(2)$, that is, the optimally selected regressor is the same for both horizon, the direct approach implies a "denser" problem -- estimating two coefficients rather than one for separate regressions. A scaled-up version of this is that if each horizon along the path implies 25 non-overlapping predictors, then the average growth rate model should have $25 \times h$ predictors, a much harder learning problem.

Of course, the $\hat{y}_{t+h}^{\text{direct}}$ approach might work better, even in a ML environment. For instance, the "aggregated" error term in (ref) could have a lower variance if $\mathrm{Corr}(\epsilon_{t+1},\epsilon_{t+2}) < 0$. Note that this would not imply substantial differences in the OLS paradigm since such errors would rather average out at the aggregation step in $\hat{y}_{t+h}^{\text{path-avg}}$. However, if a regularization level must be picked by cross-validation (like Lasso's $\lambda$), an environment where there is a strong common component across $h'$'s for the conditional mean could favor $\hat{y}_{t+h}^{\text{direct}}$. The reason for this is that choosing a regularization level optimized for a single horizon $h'$ could be different than what may be optimal for the final averaged prediction -- as examplified by our ridge regression case of equations (ref) and (ref). This observation is closely related to that of Granger1987 who shows that the behavior of the aggregate series can easily be dominated by a common component even if it is unimportant for each of the microeconomic unit being aggregated. Translated to our ML-based multi-horizon problem, this means we want to avoid having overly harsh regularization throwing out negligible effects for a given $h'$ whose accumulation over all $h'$'s makes them in fact non-negligible. Thus, if the noise level is much higher for single horizons forecasts, an overly strong $\lambda_{h'}$ for each $h'$ may be chosen whereas $\lambda_{h}$ for $\hat{y}_{t+h}^{\text{direct}}$ could be milder and allow for otherwise neglected signals to come through.

These potential explanations are illustrated using variable importance (VI) in Figure (ref). As shown earlier, the path average approach has outperformed the direct one when predicting real activity variables. VI measures in top panels show how models for $\hat{y}_{t+h}^{\text{path-avg}}$ use a much more polarized set of variables whereas those aiming for $\hat{y}_{t+h}^{\text{direct}}$ using a very diverse set of predictors in case of Income and Employment. This shed light on our bet-on-sparsity conjecture, i.e. that $\hat{y}_{t+h}^{\text{path-avg}}$ will have the upper hand if ${\Delta \hat{Y}_{t+h'}}$ predictive problems are quite heterogenous. In both cases, horizon 1 is quite different from 2-3-4, which also differ from the 5-12 block. It is noted in Figures (ref) and (ref) that $\hat{y}_{t+h}^{\text{path-avg}}$ visibly demonstrate a better capacity for autoregressive behavior (even at $h=12$) which provides it with a clear edge over $\hat{y}_{t+h}^{\text{direct}}$ during recessions. Interestingly, the foundation for this finding is also visible in Figure (ref) for real activity variables: $\hat{y}_{t+h}^{\text{path-avg}}$ reliance on plain AR terms is more than twice that of $\hat{y}_{t+h}^{\text{direct}}$.

The bottom panels show VI measures for CPI inflation and M2 growth. Recall that $\hat{y}_{t+h}^{\text{path-avg}} \prec \hat{y}_{t+h}^{\text{direct}}$ was unambiguous for those variables. Here again, results are in line with the above arguments. The retained predictors' sets are much more similar across the two approaches, which results from the presence of a strong common component over horizons (i.e., persistence which constitutes about 75% of normalized VI), which favors $\hat{y}_{t+h}^{\text{direct}}$.

figure[figure omitted — 1,360 chars of source]

{\sc MAF.} Figure (ref) plots the distribution of $\alpha_{\mathit{MAF}}^{(h,v)}$, conditional on including $X$ in the model. The motivation for that is that MAF, by construction, summarizes the entirety of $[X_{t-p}]_{p=1}^{p=P_{MAF}}$ with no special emphasis on the most recent information.\footnote{Of course, one could alter the PCA weights in MAF to introduce priority on recent lags \`{a} la Minesota-prior, but we leave that possibility for future research.} Thus, it is better-advised to always include the raw $X$ with MAF, so recent information may interact with the lag polynomial summary if ever needed. MAF contributions are overall more muted than that of MARX, except when used with Linear Boosting method. Nevertheless, it is noticed that it shares common gains with the latter as short horizons ($h=3,6$) of real activity variables also benefit from it. More convincing improvements are observed for retail sales at the 2-year horizons for nonlinear methods.

figure[figure omitted — 604 chars of source]

{\sc Traditional Factors.} It has already been documented that factors matter -- and a lot stock2002forecasting, stock2002macroeconomic. Figure (ref) allows us to evaluate their quantitative effects. Including a handful of factors rather than all of (stationary) $X$ improves substantially and significantly forecast accuracy. The case for this is even stronger when those are used in conjunction with nonlinear methods, especially for prediction at longer horizons. This finding supports the view that a factor model is an accurate depiction of the macroeconomy, as originally suggested in the works of Sargent-Sims(1977) and Geweke1977 and later expanded in various forecasting and structural analysis applications stock2002forecasting,Bernanke-Boivin-Eliasz(2005). In this line of thought, transforming $X$ into $F$ is not merely a mechanical dimension reduction step. Rather, it is meaningful feature engineering uncovering true latent factors which contains most, if not all, the relevant information about the current state of the economy. Once $F$'s are extracted, the standard diffusion indexes model of stock2002macroeconomic can either be upgraded by using linear methods performing variable selection, or nonlinear functional form approximators such as Random Forests and Boosted Trees.

figure[figure omitted — 571 chars of source]

Case Study

In this section we conduct "event studies" to highlight more explicitly the importance of data pre-processing when predicting real activity and inflation indicators. Figure (ref) plots cumulative squared errors for three cases where specific transformations stand out. On the left, we compare the performance of RF when predicting industrial production growth three months ahead, using either F, X or F-X-MARX as feature matrix. The middle panel shows the same exercise for employment growth. On the right, we report one-year ahead CPI inflation forecasts. Industrial production and employment examples clearly document the merits of including MARX: its cumulatively summed squared errors (when using RF) are always below the ones produced by using F and X. The gap widens slowly until the Great Recession, after which it increases substantially. As discussed in section (ref), using common factors with RF constitutes the optimal specification for CPI inflation. Figure (ref) illustrates this finding and shows that the gap between using \textit{F} or \textit{X} widens during the mid-80s, the mid-90s, and just before the Great Recession. To provide a statistical assessment of the stability of forecast accuracy, we consider the fluctuation test of Giacomini-Rossi(2010) in Appendix (ref).

figure[figure omitted — 200 chars of source]
footnotesize\flushleft { \singlespacing Notes: Cumulative squared forecast errors for INDPRO and EMP (3 months) and CPI (12 months). All use the Random Forest model and the direct approach. CPI and EMP have been scaled by 100. }

In Figure (ref), we look more closely at each model's forecasts during last three recessions and subsequent recoveries. Specifically, we plot the 3-month ahead forecasts for the period covering 3 months before, and 24 months after a recession, for industrial production and employment. The forecasting models are all RF-based, and differ by their use of either F, X or F-X-MARX. On the right side, we show the RMSE ratio of each RF specification against the benchmark FM model for the whole POOS and for the episode under analysis. In the case of industrial production, the F-X-MARX specification outperforms the others during the Great Recession and its aftermath, and improves even more upon the benchmark model compared to the full POOS period. We observe on the left panel that forecasts made with F-X-MARX are much closer to realized values at the end of recession and during the recovery. The situation is qualitatively similar during the 2001 recession but effects are smaller. Including MARX also emerges as the best alternative around the 1990-1991 recession, but the benchmark model is more competitive for this particular episode.

figure[figure omitted — 1,241 chars of source]

In the case of employment showcased in Figure (ref) in Appendix (ref), MARX again supplants F or X in all three recessions. For instance, around the Dotcom bubble burst, it displays an outstanding performance, surpassing the benchmark by 40%. However, during the Great Recession, it is outperformed by the traditional factor model. Finally, the F-X-MARX combination provides the most accurate forecast during and after the credit crunch recession of the early 1990s.

Figure (ref) illustrate the relative performance of the two target transformations for employment and income 12 months ahead. Again, we focus on the three most recent recession episodes. $\hat{y}_{t+h}^{\text{path-avg}}$ dramatically improves performance over $\hat{y}_{t+h}^{\text{direct}}$ and much of that edge visibly comes from adjusting itself more or less rapidly to new economic conditions. In contrast, $\hat{y}_{t+h}^{\text{direct}}$ is extremely smooth and report something close to the long-run average. Since the last three recessions were characterized by a slow recovery, $\hat{y}_{t+h}^{\text{path-avg}}$ procures much more credible forecasts of employment and income simply by catching up sooner with realized values. This behavior is understandable through the lenses of Figure (ref) where early horizons of $\hat{y}_{t+h}^{\text{path-avg}}$ make a pronounced use of autoregressive terms for both employment (and income, see Figure (ref) in Appendix (ref)).

figure[figure omitted — 1,323 chars of source]

Extraneous Transformations

We evaluate four additional data transformation strategies in combination with direct and path average targets. First, we accommodate for the presence of error correction terms (ECM) by considering the Factor-augmented ECM approach of BANERJEE2014589 and include level factors estimated from $I(1)$ predictors. Second, we consider volatility factors and data inspired by nggorodnichenkoVol, where both factors from $X^2$ and $X^2$ itself are included as predictors. Third, we evaluate the potential predictive gains from including fhlrDFM's dynamic factors in $Z$.

Figure (ref), in Appendix (ref), reports the distribution of average marginal effects of adding level factors in the predictors' set $Z$. Their impact is generally small and not significant at short horizons, while it depends on methods and forecasting approach at longer horizons. In the case of the direct average approach, as depicted in panel (ref), adding level factors generally deteriorates the predictive performance except for M2 with nonlinear methods. The effects are qualitatively similar when the target is achieved by the path average approach, as shown in (ref).

Adding volatility data and factors is generally harmful with linear methods and has almost no significant impact when random forest and boosted trees are used, see Figure (ref).\footnote{The very weak contribution of volatility terms to BT or RF is expected given that those transformations are locally monotone (i.e, for all points where $X_{k,t}>0$ or $X_{k,t}<0$) and trees are invariant to monotone transformations.} Hence, letting ML methods generate nonlinearities proves to be more resilient than to include simple power terms. This also suggests that volatility or other uncertainty proxies may not be the major sources of nonlinearities for macroeconomic dynamics since they would otherwise be an indispensable form of feature engineering which variable selection algorithms build their predictions from.

Finally, Figures (ref) and (ref) evaluate the marginal predictive content of dynamic factors as opposed to MAF and static factors (PCs) respectively. Considering dynamic factors as opposed to MAF improves the predictability at longer horizons when used to construct $\hat{y}_{t+h}^{\text{direct}}$, while their effects are rather small with $\hat{y}_{t+h}^{\text{path-avg}}$. When it comes to the choice between dynamic and static factors, the results are in general quantitatively small but suggest that standard principal components are preferred, especially in combination with nonlinear methods, which is analogous to the findings of boivinng2005 in linear environments.

Conclusion

This paper studies the virtues of standard and newly proposed data transformations for macroeconomic forecasting with machine learning. The classic transformations comprise the dimension reduction of stationarized data by means of principal components and the inclusion of level variables in order to take into account low frequency movements. Newly proposed avenues include moving average factors (MAF) and moving average rotation of $X$ (MARX). The last two were motivated by the need to compress the information within a lag polynomial, especially if one desires to keep $X$ close to its original -- interpretable -- space. {In addition to the aforementioned transformations focusing on $X$, we considered two pre-processing alternatives for the target variable, namely the direct and path average approaches.}

To evaluate the contribution of data transformations for macroeconomic prediction, we have considered three linear and two nonlinear ML methods (Elastic Net, Adaptive Lasso, Linear Boosting, Random Forests and Boosted Trees) in a substantive pseudo-out-of-sample forecasting exercise was done over 38 years for 10 key macroeconomic indicators and 6 horizons. With the different permutations of $f_Z$'s available from the above, we have analyzed a total of 15 different information sets. The combination of standard and non-standard data transformations (MARX, MAF, Level) is shown to minimize the RMSE, particularly at shorter horizons. Those consistent gains are usually obtained when a nonlinear nonparametric ML algorithm is being used. This is precisely the algorithmic environment we conjectured could benefit most from our proposed $f_Z$'s. Additionally, traditional factors are featured in the overwhelming majority of best information sets for each target. Therefore, while ML methods can handle the high-dimensional $X$ (both computationally and statistically), extracting common factors remains straightforward feature engineering that works.

{The way the prediction is constructed can make a great difference. The path average approach is more accurate than the direct one for almost all real activity variables (and at various horizons). The gains can be as large as 30% and are mostly observed when the path average approach is used in conjunction with regularization and/or nonparametric nonlinearity.}

As the number of researchers and practitioners in the field is ever-growing, we believe those insights constitute a strong foundation on which stronger ML-based systems can be developed to further improve macroeconomic forecasting.

\onehalfspace

\setlength\bibsep{5pt}