Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
123,891 characters · 25 sections · 63 citation commands
Forecasting inflation using disaggregates and machine learning
\singlespacing
\onehalfspacing
Economists and econometricians aim to provide as accurate inflation forecasts as possible by utilizing the most efficient approaches available. An important question is whether considering disaggregated inflation in different markets or economic classifications can enhance the forecasting performance for aggregate inflation. At first, this approach could capture trend dynamics, seasonality, and short-term changes more effectively espasa2002. In other words, using subcomponents would allow the econometric models to capture the heterogeneity underlying the aggregate variable better bermingham2014. Given an unknown data-generating process, whether direct or indirect forecasting through aggregating disaggregated forecasts can improve or not forecast accuracy is strictly an empirical question lutkepohl1984, hendry2011, faust2013. Nevertheless, a critical challenge that emerges is the increase in estimation uncertainty. To mitigate this problem, we implement a disaggregated analysis using machine learning (ML) methods that can deal with the bias-variance trade-off. Studies such as inoue2008, garcia2017, and medeiros2021 point out the benefits of these techniques for inflation forecasting. Our paper employs these techniques in the context of disaggregated analysis, something scarcely explored in the literature.
The broad literature on inflation forecasting documents that the predictive performance of survey-based forecasts is challenging to beat, especially in the short-term horizons -- current and immediate next months thomas1999, ang2007, croushore2010. faust2013 argue that “purely subjective forecasts are in effect the frontier of our ability to forecast inflation” because, besides private sectors and central banks having access to econometric models, they add expert judgment to these models. Consequently, “a useful way of assessing models is by their ability to match survey measures of inflation expectations” faust2013. A potential explanation for this phenomenon is that forecasters are likely to have a richer information set than the econometrician employing a standard set of macroeconomic variables as predictors for inflation delnegro2011. Thus, including revealed expectations among the predictors is a way to exploit an information set that is not available. bacsturk2014, altug2016, garcia2017, fulton2021, and banbura2021 find evidence favorable to the incorporation of survey-based forecasts into forecasting econometric models.
This paper examines the effectiveness of various forecasting methods for predicting aggregate inflation, focusing on aggregating disaggregated forecasts -- also known in the literature as the bottom-up approach. Using the Brazilian case as an example, we compare the predictive performance of the bottom-up approach with traditional approaches in the literature, including survey-based forecasts and direct forecasting based exclusively on the aggregate. We explore different levels of disaggregation to assess how forecasts based on disaggregate price levels fare relative to those that rely solely on aggregate. Granularity is a potential advantage of considering disaggregates. Besides the specific effects of the traditional macro-variables related to money, economic activity, government, and external sector, we include lagged and crossed effects between disaggregates. When we compute our forecasts, we also consider available survey-based expectations as a predictor to add information not captured by other variables. Finally, we employ a range of traditional time series techniques, as well as linear and nonlinear ML techniques to deal with a larger number of predictors. More specifically, we consider these modeling possibilities:
The Brazilian case is interesting for several reasons. First, the Broad Consumer Price Index (Índice de Preços ao Consumidor Amplo -- IPCA), which serves as the official Brazilian price index, is available monthly and boasts a rich structure to be explored. The index contains several disaggregation levels and all time-varying weights of goods and services in the representative consumption basket are readily available. Second, the Central Bank of Brazil conducts the Focus survey, an extensive daily survey of expectations for multiple forecast horizons for some variables, including inflation. This survey reflects experts' opinions, mainly financial market professionals, and may contain private information that is not available to the econometrician. Beyond its utility as a predictor for generating model-based forecasts, Focus' inflation expectations can be used as a benchmark to assess whether improving survey-based forecasts for a given horizon is possible. Third, due to Brazil's inflationary history, in addition to the official price index, the country has several price indexes that may be used as predictors for inflation. Hence, it is pertinent to examine whether this information is valuable for forecasting Brazilian inflation over future horizons.
\paragraph{Findings.} Among the main results of this paper are:
\paragraph{Contributions for the literature.} We can summarize the main contributions of this paper in two fields. First, this paper advances the literature on inflation forecasting via aggregation of disaggregated forecasts by considering many predictors for each disaggregate, as well as several statistical and econometric methods underexplored in this literature. Many papers employ traditional time series models and a limited number of predictors. In this context, some papers find evidence favoring the bottom-up approach for the Euro Area espasa2002, espasa2007 and various countries bruneau2007, moser2007, capistran2010, aron2012, carlo2016, fulton2021. On the other hand, some papers find that aggregating forecasts by components does not necessarily improve aggregate inflation forecasting benalal2004, hubrich2005, hendry2011. In turn, duarte2007, ibarra2012, and bermingham2014 highlight the benefits of aggregating a large number of disaggregates. By employing a large number of predictors, florido2021 points out the benefits of the disaggregated analysis in inflation nowcasting, while araujo2023 do not find good results by using a disaggregation in multi-period forecasting. We show that the bottom-up approach can generate multi-period forecasts as accurately as survey-based expectations and direct forecasts.
Second, our analysis extends the literature on machine learning (ML) benefits to forecasting inflation by showing a useful application of these methods considering the aggregation of disaggregated forecasts in a data-rich environment. The employ of ML methods to directly forecast inflation started with factor and principal component models stockwatson1999, stockwatson2002b, forni2003, bai2008, ibarra2012, and neural network models moshiri2000, nakamura2005, choudhary2012. Several other papers expanded the list of methods to shrinkage-based models (e.g., Ridge and LASSO), Bayesian methods, bagging, boosting, random forest (RF), and complete subset regressions (CSR), but keeping focus on forecast inflation directly from the aggregate inoue2008, medeiros2016br, garcia2017, zeng2017, baybuza2018, medeiros2021, araujo2023. florido2021 considers ML techniques, disaggregated inflation, and a broad set of predictors in inflation nowcasting, finding good results. araujo2023 consider the inflation disaggregated into administered prices, services, industrial goods, and food at home to generate multi-horizon inflation forecasts employing several ML models. However, their results are not favorable to the bottom-up approach. In contrast, our combination between disaggregated analysis, ML, and many predictors yields promising results and opens up new possibilities for further exploration.
We also point out five other minor contributions. First, we corroborate the findings of bacsturk2014, altug2016, garcia2017, fulton2021, and banbura2021 regarding the benefits of incorporating a survey-based expectation as a predictor when econometrician computes their forecasts. The presence of this variable is relevant to improve predictive accuracy even for some disaggregation levels. Second, when estimating a factor-augmented autoregression model using a method that allows predictor selection, we find that the factor that summarizes most of the predictors' variability is not relevant for predicting inflation. Hence, using an estimation method such as the adaLASSO or the approaches of bai2008, bai2009 instead of least squares may be beneficial. Third, our paper is one of the first to employ the FarmPredict, a model proposed by farmpredict that combines factor and sparse linear regressions. We adapt it to allow the simultaneous estimation via adaLASSO of a final model containing lags, common factors, and idiosyncratic components. Fourth, as in duarte2007, ibarra2012, and bermingham2014, we also indicate the potential benefits of considering a high level of disaggregation. However, there are caveats about how to improve the bottom-up approach by considering different models predicting different disaggregates. Finally, our analysis underscores the importance of examining sub-periods and emphasizing the benefits of model-based forecasts in volatile periods, as also pointed out by altug2016 and medeiros2021.
\paragraph{Outline.} This paper has five more sections in addition to this Introduction. Section (ref) presents the forecasting methodology, and Section (ref) describes the models, estimation, metrics, and test to compute and assess the results. Section (ref) displays the data and setup. Section (ref) presents the results and provides an economic discussion about them. Finally, Section (ref) concludes. Appendixes from (ref) to (ref) offer supplementary information and complementary results.
Let $\pi_t$ be the (aggregate) inflation at period $t$. We compute the inflation from the percentage change in a price index based on a typical consumption basket. For forecasting purposes, assume there are $J$ predictors for inflation. Let $\boldsymbol{z}_t$ be a $J$-dimensional vector of these explanatory variables observed at $t$, that is, the information set available to the econometrician to perform the forecasting. Notice that $\boldsymbol{z}_t$ can contain both the last available realizations of the predictor variables as well as lags of these variables. Lastly, let $\mathcal{M}_{t,h}$ be a time-varying mapping between explanatory variables (predictors) and inflation $h$ periods ahead. As the estimation is based on moving windows, the mapping is dependent on time, which we indicate by the subscript $t$.
There are several possibilities to estimate the mapping $\mathcal{M}_{t,h}$. Initially, we choose between linear or non-linear specifications. In a rich-data environment, we can consider dimensionality reduction or shrinkage with or without selecting predictors. Whatever the choices, we must be careful to avoid overfitting. Finally, an $h$-period-ahead forecast is given by
where hats indicate estimation.
Besides the general price index and their percentage change, the aggregate inflation $\pi_t$, now consider the availability of $N^d$ disaggregated price indexes (subcomponents of the original price index) indexed by $i = 1,\dots, N^d$. The letter $d$ indicates the disaggregation level. Let $\pi_{it}^d$ be the percentage change of the disaggregate $i$ in disaggregation level $d$ at period $t$. Let $\omega_{it}^d$ be the weight of the disaggregate $i$ at disaggregation level $d$ in the general price index at period $t$. Note that these weights are time-varying since the composition of the representative consumption basket may change over time. The relationship between inflation and price changes in disaggregated indexes is given by
that is, aggregate inflation is a weighted average of “disaggregated inflations” (price changes).
Let $\boldsymbol{z}_{t}^d$ be a $J^d$-dimensional vector of all explanatory variables for price changes observed at $t$ by the econometrician at disaggregation level $d$. Note that it is expected that $J^d > J$ since in the disaggregated case we potentially have more information: in addition to all the other explanatory variables available in the aggregated case, we can use the lagged price changes of the other disaggregates as predictors for a specific disaggregate. The question arises as to whether capturing and exploring the crossed dependence between disaggregated prices could enhance inflation forecasting. Let $\mathcal{M}_{i,t,h}^d$ be a time-varying mapping between predictors and $h$-period-ahead price variation of each disaggregate $i = 1, \dots, N^d$ at disaggregation level $d$. Following (ref), an $h$-period-ahead forecast for aggregate inflation is given by
where $\widetilde{\omega}_{it}^d$ is the weight of the disaggregate $i$ at disaggregation level $d$ in aggregate index observed at $t$ by the econometrician, that is, the last available weight at period $t$ and not the weight evaluate for the period $t$ -- which we previously indicate simply by $\omega_{it}^d$.
We employ a direct forecast approach considering expanding windows for (monthly) horizons $h \in \{0, 1, \dots, 11\}$. We take the time-adjusted predictors to fit the mapping between them and inflation in this approach. For example, suppose we want to generate a forecast for the current period ($h=0$), which is called nowcasting. In that case, we consider the most recently available information to estimate the desired mapping. Conversely, when computing a one-month-ahead forecast ($h=1$), we use the information available up until the preceding period in which the forecast is estimated. We continue this way until we calculate the forecast for $h = 11$, utilizing information available ten periods prior. Figure (ref) illustrates the exercise. Following the computation of forecasts based on a given period, we advance the time window by one period and repeat the estimation procedure for each forecast horizon, subsequently calculating new forecasts.
To enhance clarity in presenting the following forecast methods, we omit the superscript $d$ that indicates the level of disaggregation, whenever applicable.
\paragraph{Random walk (RW).} Considering the aggregated case, the forecast of the $h$-period-ahead inflation at period $t$ is given by current inflation, that is, $\widehat{\pi}_{t+h\,|\,t}^{\,\text{RW}} = \pi_t$.
\paragraph{Historical mean.} Also for the aggregated case, a prediction for $h$ periods ahead is given by historical average inflation computed at $t$, that is,
where $S$ is the number of previously observed inflation measures (expanding window length).
\paragraph{Autoregressive model -- AR($p$).} For both aggregated and disaggregated cases, in the direct forecast approach, for each horizon $h$, we can be written a $p$-order AR model as
where $\varepsilon_t$ is an error term. The order $p$ can be previously fixed or selected via some information criterion (e.g., BIC). Thus, a $h$-period-ahead inflation forecast is given by
where $\widehat{\mu}$ and $\widehat{\phi}$'s are least squares (OLS) estimates.
\paragraph{Augmented autoregressive model.} Including seasonal dummies and inflation expectation, we can write the model
where $\pi_{t\,|\,t-h}^{e}$ is the inflation expectation for the period $t$ available at $t-h$, $d_{mt}$ is a seasonal dummy that assumes value 1 for month $m$, and $\delta_{m}$ is a coefficient associated with seasonal dummy $d_{mt}$. In this framework, we estimate the coefficients via OLS, and a $h$-period-ahead forecast is given by
\paragraph{(Empirical) Hybrid New Keynesian Phillips curve (HNKPC).} Following and adapting price-setting models such as those presented in galigertler1999 and blanchard2007, we employ a forecasting model for the aggregate inflation based on a hybrid Phillips curve given by
where $g_{t-h}$ is some economic activity measure observed at $t-h$, and $\Delta s_{t-h}$ is an exchange rate measure observed at $t-h$. We compute the forecast by
where $\big(\widehat{\mu},\,\widehat{\boldsymbol{\phi}},\,\widehat{\eta},\,\widehat{\psi}_1,\,\widehat{\psi}_2\big)$ are OLS estimates.
\paragraph{Ridge (with incomplete information).} For disaggregated cases, we consider the augmented AR model (ref) with the addition of other lagged disaggregates:
where $N^d$ is the number of subcomponents in the disaggregation level indicated by $d$. We consider four disaggregation levels in this paper: aggregate inflation, economic categories defined by the BCB, and groups and subgroups from IPCA (IBGE).
We estimate the coefficients employing the Ridge estimator:
where $\lambda_i$ is a regularization parameter, $\boldsymbol{z}_{t-h}$ is a vector with all predictors, and $\boldsymbol{\beta}_i$ is a vector of coefficients. Chosen $\lambda$ via information criteria (e.g., BIC), a prediction for $h$ periods ahead is given by
\paragraph{adaLASSO (with full information).} For all cases, consider the model with full information given by
where $\boldsymbol{x}_{t-h} \in \mathbb{R}^{J \cdot p}$ is an expanded vector of potential predictors for $\pi_{it}$. We estimate this model employing the adaptive LASSO (adaLASSO). Introduced by zou2006, this method selects predictors and their optimization problem is given by
where $\xi$ is a regularization parameter, $\boldsymbol{z}_{t-h} \in \mathbb{R}^V$, $V = N^d \cdot p + 12 + J \cdot p$, is a vector of all predictors, and $\boldsymbol{\zeta} = (\zeta_1, \dots, \zeta_V)$ is a vector of weights obtained previously employing LASSO -- a estimator that assumes $\zeta_{ij} = 1$, for all $j$. More precisely, we compute the weights via
where we add $T^{-1/2}$ to allow a variable that is not selected in the first stage to have a chance of being selected in the second stage.
As before, a $h$-period-ahead forecast is $\widehat{\pi}_{i,t+h\,|\,t}^{\text{adaLASSO}} = \widehat{\mu}_i + \widehat{\boldsymbol{\beta}}_{\,\text{adaLASSO},i}\,\boldsymbol{z}_t$.
\paragraph{(Augmented) factor model.} Consider that all regressors are normalized for both aggregate and disaggregate cases. Thus, for $i = 1, \dots, N^d$, a factor-augmented autoregression model is described by
from which we compute common factors $\widehat{\boldsymbol{f}}_{t} = \left(\widehat{f}_{1t},\dots,\widehat{f}_{Kt}\right)$ and factor loadings $\boldsymbol{\lambda}_k = \left(\lambda_{1k},\dots,\lambda_{Jk}\right)$ by combining principal component analysis (PCA) and OLS. Finally, we compute $\big(\widehat{\mu}_i, \, \widehat{\boldsymbol{\phi}}, \, \widehat{\eta}_i, \, \widehat{\boldsymbol{\beta}}\big)$ via adaLASSO.
For identification purposes, we assume that
The number of factors $K$ is selected via information criterion $\text{IC}_{p2}$ of bai2002, and the forecast $h$ periods ahead is given by
where $\widehat{f}_{kt}$ is the $k$-th factor evaluated at $t$.
\paragraph{Target factor model.} Proposed by bai2008, in this “hard thresholding” version, this approach controls for the participation of normalized explanatory variables in the factor construction. In a previous stage, for each predictor indexed by $j = 1, \dots, J$, and disaggregate indexed by $i = 1, \dots, N^d$, we estimate
and run the hypothesis test \, $\theta_{ij} = 0 \, \times \, \theta_{ij} \neq 0$ \, for some significance level $\alpha$. If $\theta_{ij}$ is statistically different from zero, we employ $x_j$ in the factor estimation. Let $\boldsymbol{x}_t(\alpha, i)$ be the set of selected variables for $i$-th disaggregation. Finally, we proceed as in the traditional factor-augmented autoregressive model: we perform
which $\widehat{\boldsymbol{f}}_{t}$ and $\widehat{\boldsymbol{\lambda}}_k$ are computed via PCA and OLS. Then we estimate the augmented (target) factor model via adaLASSO and compute the forecast as before.
Some idiosyncratic errors of the factor model, that is, some $\boldsymbol{u}_t$ entries in Equation (ref), can impact the price variation, which the common factor structure does not capture. Defining $\widehat{\boldsymbol{u}}_{t} = \boldsymbol{x}_t - \sum_{k=1}^{K} \widehat{\boldsymbol{\lambda}}_{k} \, \widehat{f}_{k,t}$, a $J$-dimensional vector, we can introduce lags of $\boldsymbol{u}_t$ on the factor model:
This model is a specific form of a general model called FarmPredict proposed by farmpredict. Here, we estimate the “final equation” (ref) with all regressors simultaneously employing the adaLASSO. Next, we compute the forecast.
Introduced by elliott2013, elliott2015, this ensemble method combines estimates from all (or several) possible linear regression models, keeping the number of predictors fixed. Let $p$ be the total available predictors and $k \leqslant p$ be the number of “selected” predictors (complete subsets). The CSR involves the estimation of $\frac{\textstyle k!}{\textstyle (k-p)!k!}$ linear models. Variables when “non-selected” has their coefficients set to zero. The final CSR estimate is the average of all estimates. Thus, subset regression has a shrinkage interpretation since when averaging parameters that sometimes assume zero value, this average generates shrunken estimates of the coefficients, which can contribute to more accurate forecasts. Due to the high computational cost arising from a large number of predictors, we (pre-)select $\widetilde{p} \leqslant p$ predictors based on a ranking of t-statistics in absolute value, as in garcia2017 and medeiros2021. This procedure is similar to that used in the target factor model. So, instead of considering all available $p$ predictors, we run the CSR considering $\widetilde{p}$ pre-selected predictors.
breiman2001 introduces the random forest (RF), a model that combines several based-tree regressions using bagging. A regression tree is a nonparametric model that approximates an unknown nonlinear function with local predictions via recursive partitioning, as illustrated in Figure (ref).
Formally, a regression tree model can be written as follows:
where $\mathcal{I}_k(\boldsymbol{x}_{t-h} \in R_k)$ is an indicator function that assumes the value 1 when $\boldsymbol{x}_{t-h}$ belongs to the $k$-th region $R_k$, and $c_k$ is the average of $\pi_t$ in this region. We have to set the minimum number of observations per region. Then, we obtain $B$ trees by implementing a double draw: we draw on the observation dimension using block bootstrap, and we draw variables to incorporate in the estimation of the tree. The idea is that this double draw will ensure the variability of the trees. Let $K_b$ be the number of regions of the $b$-th tree, $b = 1, \dots, B$. Lastly, the final forecast is given by the average of the forecasts obtained by each tree evaluated in the original data, that is,
where $R_{k,b}$ is the $k$-th region of the $b$-th tree.
Methods may perform differently for distinct disaggregates or even over time for the same disaggregate. To mitigate instabilities associated with some method for some disaggregate or at some point in time, for each disaggregation, we will compute a combined forecast given by the average of forecasts generated by all methods applied to this disaggregation and Focus expectations available when the econometrician computes their forecasts, that is,
where $d$ indicates one of four possible disaggregations levels addressed in this paper (aggregate inflation, BCB categories, IBGE groups, and IBGE subgroups), $m$ indicates a method, $M^d$ is the number of methods employed to forecast the inflation for the disaggregation $d$, $N^d$ is the number of disaggregates in the disaggregation $d$, and $\widetilde{\omega}_{it}^d$ is the weight of disaggregate $i$ of the disaggregation level $d$ in the aggregate index observed at $t$ by the econometrician. The idea is to investigate whether this simple combination leads to improvements in forecast performance.
\paragraph{Metrics.} We use out-of-sample root mean squared error (RMSE) as the main metric to evaluate the forecast performance. For each horizon $h$, this metric is described by
where $\widehat{\pi}_{t+h\,|\,t}^{\,m,\,d}$ indicates a forecast generated by the model $m$ considering the disaggregation level $d$. The smaller the $\text{RMSE}_{h,\,m}$, the better the model's predictive performance. For the Diebold-Mariano test, we consider the mean squared error (MSE) defined by
\paragraph{Test.} To assess the results, we consider the widely employed test developed by dmtest. Let $\widehat{v}_{t+h\,|\,t}^{\,m} = \pi_{t+h} - \widehat{\pi}_{t+h\,|\,t}^{\,m}$ be a forecast error of the model $m$. Here, we omit the disaggregation level $d$. Let $g(\cdot)$ be a metric to be applied to $\widehat{v}_{t+h\,|\,t,\,m}$ (e.g., MSE). The Diebold-Mariano (DM) test statistic is given by
where $m'$ indicates another model, a competitor (i.e., a benchmark model or specific forecast, for example). We will consider that the normality of DM statistics is likely a trustworthy approximation, including for model-based forecasts.
\paragraph{Data.} We analyze the period from January 2004 to June 2022, totalizing 18,5 years of monthly data. For aggregate inflation, we employ the IPCA, the official Brazilian price index computed by the Brazilian Institute of Geography and Statistics (Instituto Brasileiro de Geografia e Estatística -- IBGE). For disaggregations, we consider all groups and subgroups of the IPCA. There are nine groups and 19 subgroups throughout the period analyzed. Subgroups are subdivisions and, in some cases, the group itself. For definition of groups and subgroups, and their respective average weights in the IPCA, see Table (ref) in Appendix (ref). In addition, we use a disaggregation defined by the Central Bank of Brazil (BCB) based on IBGE data. The BCB disaggregation consists of administrated, non-tradables, and tradables items. The use of this last disaggregation is interesting because, in principle, it presents more economic intuition, which can contribute to better forecast performance.
We consider inflation expectations of the Central Bank of Brazil's Focus survey and lags of the predicted variables among the admissible predictors. To forecast a disaggregate, we consider lags of other disaggregates in the same disaggregation, which allows capturing potential lagged “cross-effects”. The Focus survey has a daily frequency and contains inflation expectations formed by many economic agents (experts) for several horizons (months) ahead. Reflecting the opinion of experts, the Focus may contain private information that is not available to the econometrician -- hence the importance of considering this variable in our information set. We consider the latest available inflation expectation for the horizon of interest when we generate our forecast. Moreover, there are eighty-nine other predictor variables (and their lags) divided into ten categories: prices and money (17), commodities prices (4), economic activity (19), employment (5), electricity (4), confidence (3), finance (12), credit (4), government (12), and exchange and international transactions (9). In Appendix (ref), Table (ref) presents a description of these variables, the delay for each to become available and transformations implemented to guarantee the stationarity. We structure our dataset to closely approximate real-time data for the Brazilian case. The main challenge lies in our lack of access to the first releases of some economic activity variables, including the IBC-Br (a Brazilian Economic Activity Index) and industrial production (and their subcomponents).
\paragraph{Setup.} The reference day to compute our forecasts is the last business day of each month. For the results shown in the following section, we consider three lags for all predictive variables, including variables mentioned above, factors in factor models, idiosyncratic components in FarmPredict, and lags of aggregate and all disaggregates. The only exception is the factors in the target factor model for which we employ only one (target) factor. As mentioned in Subsection (ref), the main results are generated based on expanding windows. In this setup, we generate 114 forecasts for each horizon. The regularization parameters ($\lambda$'s) of the Ridge, LASSO, and adaLASSO are obtained via Bayesian Information Criterion (BIC). We restrict the number of possible selected variables by the ceiling of $\sqrt{T}$ to enforce discipline. The number $K$ of latent factors in factor models is selected via bai2002 information criterion $IC_{p2}$. For CSR, we set $\widetilde{p} = 20$ (number of pre-selected predictors) and $p = 4$ (number of selected variables by CSR). For pre-selecting of both target factor and CSR models, we adopt the 5% significance level ($\alpha = 0.05$). In its turn, for the RF models, we allow the trees to grow until five observations by leaf. We set the proportion of selected variables in each split to 1/3 and the number of bootstrap samples to 500 ($B = 500$). All settings are similar to those adopted by garcia2017 and medeiros2021. Finally, to estimate the empirical Phillips curve, we use the Central Bank of Brazil's economic activity index (IBC-Br) and BIS' real effective exchange rate (REER) as a proxy for economic activity and exchange rate, respectively.
Table (ref) exhibits the results of forecast performance in terms of root mean squared errors (RMSE) for different models and horizons ranging from nowcasting ($h=0$) to eleven months ahead ($h=11$), as well as for 12-month accumulated inflation. We normalize every RMSE to relative terms by computing their ratio to the RMSE of the Focus consensus -- the median expectation of the available Focus survey. Thus, a value lower than one indicates that a model numerically outperforms the Focus consensus, while a value greater than one suggests underperformance compared to the same benchmark. At this first moment, the results consider the entire period for which we compute predictions, from January 2014 to June 2022. Each panel of Table (ref) considers a group of competitors. In panel A, we have the available and ex-post Focus, the latter released by the Central Bank in the following week reflecting the experts' opinions on the same day we compute our forecasts. We note virtually no difference between the available and ex-post Focus for longer horizons. However, in the short term ($h \leqslant 3$), there is evidence that ex-post Focus statistically outperforms available Focus. Despite being only a few days apart, the informational gain is considerable for shorter horizons, which does not occur for more distant periods since it is unlikely that very relevant information about them will emerge within a few days.
Panels B to E of Table (ref) show the results for each model considering different levels of disaggregation: aggregate inflation, disaggregations from the Central Bank of Brazil (BCB), and disaggregations into IBGE groups and subgroups, respectively. Perhaps not surprisingly, it is hard to outperform the Focus survey in nowcasting. Exceptions are due to models that forecast aggregate inflation directly. However, such models perform better only than the available Focus. Considering the whole period, no alternative beats the ex-post Focus, which delivers almost 7% RMSE reduction compared to the available survey. However, it is appreciable that some models are competitive with the ex-post Focus. Specialists who report their expectations to the BCB often have access to information unavailable to econometricians, such as private data. Since models do not have this additional information and other advantages, such as including personal judgments, as pointed out by faust2013, their ability to outperform available expectations and get closer to ex-post survey-based expectations is a great result. We note that models forecasting disaggregates do not deliver good performance for nowcasting. Lastly, the combinations of models in each level of disaggregation, whose results are shown in Panel F, also do not generate forecasts better.
For other horizons, the contribution of the models becomes more effective. Despite the challenge of surpassing survey-based expectations for short-term horizons such as $h=1$ and $h=2$, several models for aggregate inflation (Panel B) achieve good results for these horizons. Like occurred for $h=0$, the hybrid Phillips curve, adaLASSO, factor model, FarmPredict, and, additionally, the complete subset regression (CSR), deliver the best forecast performances for one and two months ahead. All are statistically superior to the available Focus according to the Diebold-Mariano (DM) test considering at least the more slack significance level (i.e., 10%). Furthermore, these models also numerically outperform the ex-post Focus. On the other hand, the adaLASSO using BCB disaggregation (Panel C) is the only model employing any disaggregation among the best models. However, this model is not statistically superior to the available Focus by the DM test. Finally, regarding these shorter horizons, it is worth highlighting the performance of the average forecast of the models for aggregate inflation, which achieves the highest accuracy for $h = 2$ by presenting a statistically significant reduction of almost 5% in RMSE.
Models considering some disaggregation for the inflation yield better results starting from the 4-month horizon. The adaLASSO, factor model, and FarmPredict, all using the BCB disaggregation, perform well for forecast horizons ranging from fourth to seventh months. These models are statistically superior to available or ex-post Focus at various periods. Regarding the use of disaggregated inflation data in groups from the IBGE, it is worth mentioning the good performance of the CSR, which achieves the best result among all the options for $h = 7$. Another highlight is the combination of forecasts generated by models that directly forecast the aggregate inflation, which achieves a statistically significant reduction of 6% in RMSE from 6 to 9 months ahead. For $h \geqslant 8$, there is broad dominance of the random forest (RF), whether using aggregate inflation or some disaggregation. Frequently, for these more distant horizons, the RF registers a statistically significant reduction in RMSE ranging from 7% to 10% in comparison to the survey-based expectations. This result highlights that the RF, employing IBGE group disaggregation, achieves the best performance among all competitors for inflation accumulated over 12 months (see last column of Table (ref)), closely followed by the adaLASSO using BCB disaggregation, which achieves a similar RMSE reduction.
\paragraph{Remarks.} Considering the forecast performance of various models from January 2014 to June 2022, we observe that different approaches are more effective at different times. In the short term, machine learning models that deal directly with aggregate inflation perform better, whereas for intermediate horizons of 4 to 7 months, considering the BCB disaggregation lead to significant benefits. For the period between 6 and 9 months ahead, the average of forecasts obtained from models that used only aggregate inflation also perform well. Finally, for longer horizons of 8 months or more, regardless of the approach, the RF delivers the best forecast performances. While garcia2017 points to the superiority of the CSR in several horizons, we only verify the prevalence of the CSR for $h = 1$ and other isolated good performances. The RF's performance in predicting inflation had already been pointed out by medeiros2021 when analyzing the case of the United States and highlighting the benefits of this method for dealing with non-linearities. The advent of the COVID-19 pandemic in Brazil changes the price dynamics considerably from 2020 onwards. Because of that, in what follows, we divide our analysis into two sub-periods: (i) before the pandemic, from January 2014 to February 2020, and (ii) after the pandemic, from March 2020 onwards.
Table (ref) shows that the sub-period between January 2014 and February 2020 is quite challenging for model-based forecasts. Almost no RMSE ratios are below 1, with exceptions mainly in the short term. Nevertheless, no model is able to beat the ex-post Focus or be statistically superior to the available Focus for nowcasting. For $h = 1$, only the factor model and FarmPredict are statistically superior to the available Focus at the 10% significance level, with these two tying with the predictive performance of the ex-post Focus. For the 3-month forecast, the hybrid Phillips curve for aggregate inflation is subtly superior to the Focus in numerical terms but without statistical significance. For the other horizons, no model performs better than the survey-based expectations. There are some potential explanations for this poor performance of the models. First, since we are analyzing the first sub-period, a small sample may have affected the estimates, contributing to the models' poor forecast performance. Second, the instabilities in the Brazilian economy in 2014 and 2015 that resulted in a sharp increase in inflation in 2015, as well as the rapid disinflation that occurred from the second half of 2016, are challenging events to anticipate, especially without extensive historical data. Conversely, from 2017 until the begging of the COVID-19 pandemic, Brazilian inflation remained reasonably controlled and close to the inflation target, leaving limited opportunities for models to enhance survey-based expectations. These dynamics of the Brazilian inflation can be observed from Figure (ref) later on.
Since the COVID-19 pandemic, most models perform statistically better than the Focus survey, as seen in Table (ref). Models that look directly at aggregate inflation or use some inflation disaggregation tend to perform very well at all horizons, including nowcasting. Specifically, the adaLASSO, factor model, and FarmPredict working directly with the aggregate inflation achieve RMSE 10% lower than the RMSE of the available Focus and smaller RMSE than the ex-post Focus, something challenging to imagine before the pandemic. Meanwhile, augmented AR and ridge, both employing BCB disaggregation, also perform well. One month ahead, even some models employing the highest level of disaggregation (i.e., subgroups) deliver good results. From the results in this second sub-period, we understand the good performance of the RF for the entire period. Considering aggregate inflation, the method already registers the best performances from $h = 3$, with similar results when using any disaggregation. Frequently, the RF can reduce the RMSE of the available and ex-post Focus by up to 30% for longer horizons, regardless of whether considering aggregate inflation or some disaggregation. For 12-month accumulated inflation, the RF obtains the best performance when using the BCB disaggregation: a 35% reduction compared to the RMSE of the main benchmarks. Finally, regarding model combinations, each is statistically superior to Focus for all $h \geqslant 1$, often at a 1% significance level. However, the combinations do not beat some individual models with good predictive performance in the sub-period.
Figure (ref) presents the temporal evolution of actual inflation, Focus survey expectations, and the best aggregate- and disaggregates-based models for each horizon. Looking at the projections for $h = 0$, we partially understand why it is not easy to outperform the survey in the very-short term. The Focus consensus is very close to the actual values. Furthermore, as we will see in Subsection (ref), available inflation expectations are the primary predictor for model-based nowcasting. The survey contains much relevant information unavailable to the econometrician, so we already expect this result. When analyzing the other horizons ($h \geqslant 1$), we note that it is challenging for the survey and models to predict peaks and valleys of inflation. Already at $h = 1$, we observe outstanding forecasting errors. One consequence of COVID-19, which start is highlighted by vertical dashed lines on each plot, is that the Focus survey initially overestimated inflation and afterward systematically underestimated one. Despite the challenges of generating accurate forecasts in such an uncertain period, several models perform better than expert forecasts for all horizons. A punctual example is the adaLASSO that, using both aggregates and BCB disaggregations, as well as other models, achieves a great result in forecasting the peak observed in December 2020 at $h = 1$, a point at which the available inflation expectation is far from the actual value. Thus, other variables besides the available inflation expectations are fundamental for the performance of model-based forecasts.
\paragraph{Remarks.} The good performance of the RF is mainly due to its ability to capture the higher level of future inflation from the second half of 2020. The model generates forecasts closer to the actual inflation than the Focus expectations, which systematically underestimate inflation in that period. Since the pandemic, models for disaggregated inflation tend to provide more accurate forecasts than models for aggregate inflation, except for nowcasting. For each $h \geqslant 2$, we note that adaLASSO, factor model, FarmPredict, and CSR using any disaggregation of inflation deliver forecasts with lower RMSE than the respective models using aggregate inflation, with a few exceptions for the CSR using groups and subgroups that do not outperform the CSR using aggregate inflation. In turn, the RF performs well regardless of the target variable. These findings underscore the use of models in inflation forecasting, including the junction between disaggregated analysis and machine learning techniques, particularly during periods of higher economic instability, such as a pandemic. Previous studies such as altug2016 and medeiros2021 have also shown that models perform well during more volatile periods. Next, we will analyze the forecasts for disaggregates and identify the predictors selected by adaLASSO and FarmPredict, methods that allow variable selection.
\paragraph{Predictive performance.} Now we consider the predictive performance of the models using different disaggregations, starting with the disaggregation from the BCB. Since we lack long time series for survey-based expectations for disaggregated inflation, we use the AR model as a benchmark. Of all the BCB disaggregations, monitored prices are the most challenging to forecast since they are subject to many unexpected changes resulting from government decisions that often do not freely follow supply and demand movements. According to the results displayed in Table (ref), other models manage to beat the AR only in short or more distant horizons for this disaggregate (Panel A). In nowcasting, other methods perform significantly better than the AR, with augmented AR and Ridge standing out by obtaining more than a 30% reduction in RMSE. For $h = 1$, augmented AR achieves a 10% reduction in RMSE, the only statistically significant at any level. Some models present minor RMSE for intermediate horizons, but the results are not statistically significant according to the DM test. For ten and eleven months ahead, Ridge delivers reductions of 4% and 6% in RMSE, respectively, compared to the AR, and, specifically for $h = 11$, adaLASSO, factor model, and FarmPredict also statistically outperform the AR, with RMSE reductions ranging from 3% and 4%. Putting all horizons together, Ridge, adaLASSO, FarmPredict, and factor model generate more accurate forecasts than the AR by delivering RMSEs 2% to 4% lower than the AR model, all statistically significant at the 1% level. The good result of the augmented AR model is restricted to the short term.
Looking at the other BCB disaggregates, namely non-tradable and tradable items, we notice that machine learning models deliver better results than the traditional AR model (Panels B and C of Table (ref)). For non-tradables, once again augmented AR and Ridge stand out in nowcasting with RMSE reductions of 18% and 16%, while adaLASSO and FarmPredict achieve the best performances between one and five months ahead, with RMSE reductions oscillating between 20% and 27%. For all $h \geqslant 5$, the random forest dominates by delivering the lowest RMSE. Aggregating all horizons, all ML models perform statistically better than AR for non-tradable items. In turn, the results for tradables are similar, with the ML methods yielding subtly smaller improvements. Also for tradables, there is a predominance of the RF: this method obtains the best or second-best performance for all $h \geqslant 1$, with RMSE reductions ranging from 10% to 19%. Other methods that stand out are adaLASSO, which obtains the best result in nowcasting (almost 24% reduction in RMSE), and target factor, which registers the best performance one, seven, and eight months ahead. By gathering the forecasts for all periods, the RF obtains an average reduction of 15% in RMSE compared to the AR model, with adaLASSO coming close behind, with a reduction of 11% in RMSE. All models, except the augmented AR, outperform the AR at the 1% significance level. These findings suggest that models that include more predictors, impose restrictions on parameters, or assume other functional forms can be more advantageous in inflation forecasting than traditional time-series models such as the AR model. Lastly, we note that the improvements due to ML methods in a data-rich environment are not just observed for the pandemic period, as seen in Tables (ref) and (ref) in Appendix (ref).
\paragraph{Variable selection.} To explore potential economic intuitions for the results from aggregate inflation and BCB disaggregation, we compare what is behind the two approaches in terms of variable selections by adaLASSO and FarmPredict, as shown in Figures (ref) and (ref), respectively. In both Figures, panel A brings the predictors that each method selected to forecast the aggregate inflation directly. Meanwhile, panels B, C, and D display the predictors selected to predict price variation of administrated, non-tradable, and tradable items. To make the presentation viable, we restricted ourselves to the variables chosen at least 20% and 11.5% of the time in at least one forecast horizon, respectively. Variables definitions are shown in Table (ref) in Appendix (ref). The prefix “u_” in some variables in Figure (ref) indicates that the variable had the common factors “discounted” and, therefore, only its idiosyncratic component is left. These variables are indicated by $\widehat{u}_j$ in Equation (ref).
We can summarize the results of the variable selections in the following topics:
Now we address the predictive performance for each IBGE group. The identification of each group, as well as their respective participation in the IPCA, are available in Table (ref) in Appendix (ref). From Table (ref), we note that ML methods perform statistically better than the AR model and numerically better than the augmented AR for all components by stacking the horizons. An exception is the target factor, which does not perform well for some groups. Furthermore, from Table (ref) in Appendix (ref), which presents detailed results of each disaggregate by forecast horizon, we notice that there are infrequent horizons that do not have an ML model performing statistically better than the benchmark.
Different models perform better for different disaggregates. Education (inf.g8) is the group most benefited from using other techniques. However, given the good performance of the augmented AR, we infer that predictive improvement is mainly due to the inclusion of the February dummy. Except for augmented AR and target factor, the models also achieve predictive improvement for communication (inf.g9), including all horizons individually. Transportation is the group for which the models beat the AR model by stacking the horizons with the smallest margin (inf.g5). In addition, the ML methods are not statistically superior to AR in half of the forecast horizons for this disaggregate. The transportation group comprises public transport fares and expenses with own vehicle and fuel, mostly items whose prices are administered by the government, which are difficult to forecast. However, except once again for augmented AR and target factor, the models deliver a statistically significant average reduction of at least 3% in RMSE compared to the AR model, with Ridge's predictive gain around 10% between 6- and 9-month-ahead. Lastly, it is worth highlighting the good performance of the RF to forecast the price variation of foods and beverages (inf.g1) at all horizons.
Lastly, we examine the predictive performance for each IBGE subgroup. The results stacking all horizons are shown in Table (ref), and results for each horizon are in Table (ref) in Appendix (ref). Descriptions of subgroups are in Table (ref) in Appendix (ref). Similar to what happens with the disaggregation into groups, ML models achieve more accurate forecasts in comparison to AR models for each subgroup, but with no single model emerging as a dominant predictor for different disaggregates. The reduction of RMSE reaches 71% in the case of courses, reading, and stationery (inf.sg18), 40% for communication (inf.sg19), and 35% for household operations (inf.sg7) and personal services (inf.sg16). RF stout out for delivering the best predictive performances considering all forecast horizons for food at home (inf.sg1), appliances (inf.sg6), household operations (\texttt{inf.sg7}), and fabrics (\texttt{inf.sg10}). The Ridge performs well at all horizons for domestic fuels and energy (\texttt{inf.sg4}) and jewelry (\texttt{inf.sg10}), and the adaLASSO performs well for communication (\texttt{inf.sg19}) at all horizons as well. Furthermore, we show outstanding performances for specific horizons: adaLASSO performs well in the long-term for food away from home (\texttt{inf.sg2}), target factor in short- and intermediate-term for pharmaceutical and optical products (\texttt{inf.sg13}), while the RF performs well at short term for that one and at more distant horizons for the latter.
Some subgroups are equivalent to groups. It is the case of transportation (inf.sg12 and inf.g5), courses, reading, and stationary (inf.sg18) that is equivalent to education (inf.g8), and communication (inf.sg19 and inf.g9). Interestingly, the predictive accuracy obtained from disaggregation into subgroups is superior to that based on groups, both stacking the horizons and considering them individually. In particular, the improvements are remarkable for transportation and communication. The difference between both approaches is that, while one includes lags from other subgroups and not groups, the other includes lags from different groups and not subgroups. Note that, stacking the forecast horizons, while CSR reduces the RMSE of the AR model for \textit{group} transportation (\texttt{inf.g5}) by 5%, the RF obtains an average reduction of 14% for \textit{subgroup} transportation (\texttt{inf.sg12}). For transportation, while the ML methods using disaggregation into groups do not statistically outperform the AR model for horizons 1 to 4, 10, and 11 months ahead, some of these methods employing disaggregation into subgroups generate statistically significant improvements at the most demanding level ( i.e., 1%) for all these horizons. For communication, disaggregation into subgroups can improve the already good performance of methods when employing disaggregation into groups.
We find that ML models considering several predictors beat the AR model by a wide and statistically significant margin for all disaggregates considered. Comparing the use of group and subgroup disaggregations, the predictive benefits associated with a higher level of disaggregation appear to align with the results of duarte2007, for Portugal, ibarra2012, for Mexico, and bermingham2014, for the US and Euro Area. However, upon revisiting the results of the aggregation of disaggregated forecasts (Tables (ref) to (ref)), we note that the use of disaggregation into subgroups does not generate more accurate forecasts than the direct approach or aggregation from less profound disaggregations. A potential explanation for this result is that we do not consider the aggregation of disaggregated forecasts generated by different models in this paper. As we have seen, a single model does not dominate the forecasts of all disaggregates and often does not even exhibit dominance over time for the same disaggregate (see Figures from (ref) to (ref) in Appendix (ref)). Additionally, there is a possibility of inaccurate disaggregated forecasts occurring at some point in time, which can lead to a deterioration of the predictive accuracy of the aggregation of disaggregated forecasts. Importantly, this deterioration cannot be attributed to the use of the most recent available weights for each item in the consumption basket. When we conduct the forecasting exercises again using the actual weights, the results demonstrate no significant changes.
In this paper, we investigate the use of inflation disaggregations to forecast the aggregate via aggregation of disaggregated forecasts -- what became known as the bottom-up approach. We innovate by considering multi-horizon forecasts of several inflation disaggregates in a data-rich environment (i.e., considering many predictors), which is only possible by employing machine learning methods. Analyzing Brazilian inflation and exploring different levels of disaggregation for inflation, we conduct forecasting exercises from both direct and bottom-up forecast approaches. We highlight the relevance of considering the combination of disaggregated analysis and machine learning methods in the econometrician's toolbox.
For many forecast horizons, the aggregation of disaggregated forecasts performs as well as survey-based expectations and models that generate forecasts directly from the aggregate. Our results reinforce the benefits of using models in a data-rich environment for inflation forecasting, including aggregating disaggregated forecasts generated from machine learning techniques, mainly during volatile periods. During the COVID-19 pandemic, model-based forecasts, including those based on disaggregated data, tend to provide more accurate forecasts than survey-based expectations. For example, the random forest model based on both aggregate and disaggregated inflation delivers great results for intermediate and longer horizons. The selection of predictors obtained by the adaLASSO and FarmPredict indicates the importance of considering a broad and diversified set of variables when forecasting inflation. Regarding the prediction of individual disaggregates, we find that ML models considering several predictors beat the AR model by a wide and statistically significant margin for all disaggregates and the vast majority of horizons.
This paper can be extended in many ways in future research. Firstly, it is possible to replicate the procedures analyzed here for other developed and emerging countries. Secondly, an important possibility is the formulation and implementation of a methodology that combines different models predicting different disaggregates. As we have seen, there is no dominant technique. Thus, to fully exploit the potential of disaggregated forecasting, it is necessary to combine different forecasts to obtain the aggregate forecast. Beyond combining different models in the dimension of disaggregates, one can also combine different models over horizons to improve the forecast of time-accumulated inflation (e.g., 12-month cumulative inflation). Thirdly, combining different levels of disaggregation could be valuable since increasing the level in some branches may be advantageous while for others, it does not. For example, in the Brazilian price index employed in this paper, some groups may be worth keeping in the final combination, while others may benefit from further breakdown into subgroups or deeper disaggregation. Finally, a fourth possibility not explored in this paper is to use breakeven inflation as a predictor and benchmark in inflation forecasting.
\onehalfspacing