EconBase
← Back to paper

Nowcasting the Portuguese GDP with Monthly Data

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

39,165 characters · 6 sections · 22 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Nowcasting the Portuguese GDP with Monthly Data

\selectlanguage{english}

abstractIn this article, we present a method to forecast the Portuguese gross domestic product (GDP) in each current quarter (nowcasting). It combines bridge equations of the real GDP on readily available monthly data like the Economic Sentiment Indicator (ESI), industrial production index, cement sales or exports and imports, with forecasts for the jagged missing values computed with the well-known Hodrick and Prescott (HP) filter. As shown, this simple multivariate approach can perform as well as a Targeted Diffusion Index (TDI) model and slightly better than the univariate Theta method in terms of out-of-sample mean errors. Keywords: Time series; Macroeconomic forecasting; Nowcasting; Error correction models; Combining forecasts. \\

Introduction

Macroeconomic forecasting models are typically based on quarterly or annual data. Nevertheless, a lot of rich data are available on monthly, weekly or even daily basis including price indexes, economic sentiment/business climate indicators, industrial production, cement sales, car sales, unemployment, exports or imports among other indicators. In particular, monthly data already available for the current quarter could be combined with lagged quarterly data to produce better forecasts, either for the current quarter (nowcasting) or for subsequent quarters, typically one or two periods ahead.

The estimation of current quarterly models from high frequency data was introduced in the late 1960's by the German-American economist Otto Eckstein [1927-1984] at Data Resources, Inc., a company integrated, meanwhile, in IHS Markit Ltd., now a part of S&P Global. Though, it was the Nobel Laureate Lawrence Robert Klein [1920-2013] that formulated the two main approaches to the problem Klein1989.

On the one hand, the analyst can consider the main entries on the expenditure side of national accounts and then establishes empirical bridge equations by aggregating high-frequency (monthly) indicators into quarters and correlating those with quarterly components of the gross domestic product (GDP). On the other hand, the forecaster can collect as much high frequency data as are available in the current quarter and then she extracts the leading principal components of the quarterly averages of these several indicators and regresses the GDP on them.

Bridge equation methods, also known as “tracking models” Higgins2014, explores the interrelationships between expenditure (or income) components of GDP and monthly available data. For example, the private consumption of durable goods might be highly correlated with car sales, and other relationships could be found for national quarterly accounts (NQA) entries like gross private residential and non-residential investment, inventory changes, government expenditure and exports and imports of goods and services Klein1989.

Where possible, bridge equations are estimated from monthly data on both high frequency indicators and NQA components. However, national accounts are not reported monthly in most cases, so quarterly bridge equations must be build by aggregating or averaging the monthly indicators into quarters. Typically, end-of-sample jagged values for those indicators are forecasted with monthly univariate models, either to complete the current quarter or to forecast one or two quarters beyond.

The first approach, originally proposed by Klein1989, builds separate bridge equations for nominal GDP components and price deflators, noting that real expenditure estimates can be obtained by dividing the estimated nominal value of each NQA component by an appropriate deflator, which may be highly correlated with consumer price index. Typically, a bridge equation relates each NQA entry with the quarterly averages of one or two highly correlated variables. Finally, a national accounting entity is used to sum up the estimated (real) NQA components into GDP.

The second approach to nowcasting is founded mainly on the contributions of Stock1989,Stock2002,Stock2011. The premise of their models is that a few latent dynamic factors drive the comovements of a time series like GDP, which is also affected by zero-mean idiosyncratic disturbances. These disturbances arise from measurement error and from specific features of the data. The latent factors follow a time series process, typically a vector autoregression (VAR). This kind of “medium data” or “data-rich” forecasting methods without expenditure components Higgins2014 can cover more than 200 monthly macroeconomic indicators like the application of Giannone2008.

In Portugal, the central bank (Banco de Portugal, BdP) has a long experience in forecasting GDP using dynamic factor models with autoregressive components, the so-called Diffusion Index (DI) approach Stock2002. In this kind of models, a large number of predictors is summarized using a small number of indexes constructed by principal component analysis, and a dynamic factor model serves as the statistical framework for the estimation of the indexes and forecasting. As stressed by Dias2015, these models requires the previous determination of the factors which reflect the top-ranked principal components, that is, the ones that encompass the largest share of the common comovement in the dataset; all other lower-ranked factors are entirely disregarded independently of their possible informational content for forecasting the variable of interest, typically, the GDP growth.

In order to avoid this limitation of DI approach, Dias2010 develop the targeting principle discussed by Bai2008. Their procedure considers a synthetic regressor which is computed as a linear combination of all the factors of the dataset weighted such that each factor takes into account both the relative size of the variance captured by it and its correlation with the variable of interest at the relevant forecast horizon. Recent findings for the Portuguese real GDP growth Dias2016 suggest that this Targeted Diffusion Index (TDI) approach could reduce the out-of-sample mean squared error (MSE) by 63% using a simple autoregressive model as benchmark, while the gain with a DI model is about 50%.

The combination of the tracking and data-rich approaches is also possible. In fact, bridge equations can include common factors and may be applied either to the aggregate GDP or to its components. The GDPNow model from the Federal Reserve Bank of Atlanta Higgins2014 is a good attempt in that direction, despite its complexity and detail level.

In this paper, we present an approach that combines bridge equations of real GDP on several covariates available on a monthly basis, including coincident indicators, with forecasts for the jagged missing values computed with the well-known Hodrick1997 filter. As described below, this multivariate approach can perform as well as the TDI model developed by the Portuguese central bank, and slightly better than the univariate Theta method.

Methodology

Our general methodology is similar to the approach proposed by Miller1996. Firstly, we use a simple model to predict each relevant seasonally adjusted monthly series $Y_{t:i}$ for the current quarter $t$, where $i = 1, 2, 3$ is the number of months of data available from it. For example, when two months of data are available, the forecast for the third month of $t$ is

equation[equation omitted — 73 chars of source]

where $g_{t:2} \equiv X_{t:2} - X_{t:1}$ is the first difference of the trend $X_{s:i}$ of the series $Y_{s:i}$, $s=1, .., t$, $i = 1, 2, 3$, which is chosen to minimize either the sum of the square residuals $\varepsilon_{s:i}=Y_{s:i}-X_{s:i}$ or its smoothness Hodrick1997:

equation[equation omitted — 193 chars of source]

where $\lambda = 14400 $ is a penalty for the square of the acceleration (second difference) of the trend.

For raw series $Y_{t:i}$ particularly noisy such as the industrial production index (IPI), cement sales (CEM) or card transactions in point of sales or automated teller machines (ATM), we compute the following alternative moving average (MA) forecast for the same example:

equation[equation omitted — 142 chars of source]

The number of months $i$ of data available depends upon the time series and the day of the current quarter. For instance, at day 60, two months of the European Commission's Economic Sentiment Indicator (ESI) are available, but only the first month of the industrial production index is ready to use, that is, the IPI is delayed one month. The table (ref) indicates the selected series available at days 0, 30, 60 and 90 of the current quarter $t$, as well as at the day 10 of the next quarter $t+1$, here designated by day 100 for convenience.

table[table omitted — 1,534 chars of source]

The monthly series of exports (EXGS) and imports (IMGS) of goods and services are delayed two months with only one month available, even 10 days after the end of the current quarter. However, two months of the series of exports and imports of goods (without services, denoted respectively EXG and IMG) are available at that time. Thus, we perform a simple linear regression to forecast an additional figure $\widehat{EXGS}_{t:2}$ (or $\widehat{IMGS}_{t:2}$) for the exports (imports) of goods and services:

equation[equation omitted — 226 chars of source]

where the lower cases $exgs$ and $exg$ are the percentage growth rates compared to the same quarter of previous year for the corresponding upper case variables, namely, $exg_{t:2} = (EXG_{t:2}/EXG_{t-4:2}-1) \times 100$, where $EXG_{t:2}$ is the data on exports of goods, available at the day 100 of the current quarter $t$, and $\hat{\alpha}$ and $\hat{\beta}$ are the ordinary least squares (OLS) coefficients on these year-over-year (y-o-y) changes in order to avoid serial correlation and residual seasonality issues. Similar equations were defined for imports. As usual, the third month of exports (or imports) of goods and services is estimated using formula ((ref)) and the Hodrick-Prescott filter with the additional figure $\widehat{EXGS}_{t:2}$ (or $\widehat{IMGS}_{t:2}$).

Secondly, we average monthly data and forecasts for each quarter. This procedure is necessary because our variable of interest (GDP) is issued quarterly. So, in the same example of two months of data available in quarter $t$, we compute:

equation[equation omitted — 87 chars of source]

where the generic $\hat{Y}_{t:3}$ can be replaced by $\tilde{Y}_{t:3}$ for noisy series. The exports and imports of goods and services are still a special case in the sense that their quarterly figures are estimated using a specific multiple regression model which takes into account, as dependent variable, the exports EXP (or imports IMP) in volume provided by the NQA\footnote{Exports and imports are available only 60 days after the end of the respective quarter.} and, as independent variables, the monthly exports (or imports) of goods and services previously averaged with formula ((ref)), the lagged price deflator (DEF) of exports (or imports) and the Brent oil price (OIL). This procedure is required because the monthly series of exports and imports of goods and services are provided in current prices, that is, in (nominal) values instead of (real) volumes. As in ((ref)), the lower case points out the percentage change over the same quarter of the previous year (y-o-y) of the correspondingly level variable in upper case:

equation[equation omitted — 274 chars of source]

Thirdly, the current quarter's GDP growth ($gdp_t$) is predicted using several bridge equations which explore different combinations of the monthly time series listed in table (ref). Here, we describe six linear models from a pool of more than thirty models currently in use.

The first of these models, denoted with the superscript $^{(1)}$, includes, as independent variables, the sum of the last three quarter-over-quarter (q-o-q) percentage growth rates of real GDP ($sum_{t-1}$)\footnote{Before 2020, the last GDP growth rate was issued only 45 days after the end of the respective quarter, so it cannot be summed up to nowcast the current (next) quarter until then. Thus, the forecasts at days zero and 30 of the current quarter $t$ were made using the sum of the last two available q-o-q changes of GDP, concerned with periods $t-2$ and $t-3$.}, the quarterly averages of the Economic Sentiment Indicator ($\widehat{ESI}_t$) previously regularized (by subtracting the historical average 100) and Economic Climate Indicator ($\widehat{ICE}_t$), as well the y-o-y changes of industrial production index ($\widehat{ipi}_t$) and cement sales ($\widehat{cem}_t$):

equation[equation omitted — 241 chars of source]

where $\hat{\alpha}_0, \dots, \hat{\alpha}_5$ are the coefficients estimated with ordinary least squares (OLS).

The second model adds the international trade, that is, the real y-o-y changes of exports and imports in addition to the independent variables already considered in equation ((ref)):

equation[equation omitted — 309 chars of source]

The third and fourth models replaces ICE in equation ((ref)) with the y-o-y changes of car sales and card transactions, respectively. Cards transactions are expressed in real terms, that is, they were previously deflated with the consumer price index (CPI).

equation[equation omitted — 241 chars of source]
equation[equation omitted — 241 chars of source]

The fifth model is based in the last model, but includes the international trade:

equation[equation omitted — 312 chars of source]

Finally, the model 6 replaces exp and imp in equation ((ref)) with the Euro-coin Real Time Indicator of the Euro Area Economy (CEPR), an alternative measure of the external outlook:

equation[equation omitted — 277 chars of source]

For each model $j = 1, \dots, 6$, we compute an alternative GDP growth estimate with a simple error correction mechanism that incorporates the last observed error $\epsilon_{t-1}^{(j)}$:

equation[equation omitted — 195 chars of source]

Then, we apply the median operator X to obtain a consensus among the corrected and uncorrected (simple) forecasts for each model $j$ and also for the six models such that

equation[equation omitted — 203 chars of source]

We also perform a consensus among the simple forecasts without error correction:

equation[equation omitted — 140 chars of source]

As benchmark, we calculated one period and two periods look ahead forecasts with the Theta method proposed by Assimakopoulos2000 and explained by Hyndman2001. Briefly, this method starts with the estimation of the following regression in level:

equation[equation omitted — 101 chars of source]

for $\theta = 0$ and $\theta = 2$, where $t=1,...,n$ is the time index. Note that, when $\theta = 0$, $\hat{a}_0$ and $\hat{b}_0$ are simply the parameters of the linear time trend fitted to the GDP quarterly series. Then, for each value of $\theta$, a new series $GDP_{t,\theta}$ is constructed by doing:

equation[equation omitted — 121 chars of source]

The $h$-step ahead forecast is obtained by averaging the GDP forecasts for $\theta=0,2$:

equation[equation omitted — 120 chars of source]

where $\widehat{GDP}_{t,0}(h)$ is obtained by extrapolating the linear time trend:

equation[equation omitted — 95 chars of source]

and $\widehat{GDP}_{t,2}(h)$ is obtained using simple exponential smoothing (SES) on series $\{GDP_{t,2}\}$:

equation[equation omitted — 118 chars of source]

where $\widehat{GDP}_{1,2}=GDP_{1,2}$ (starting value) and $\gamma=0.3$ (smoothing parameter). Thus, the SES forecasts are equivalent for all $h$ Hyndman2001.

Main findings

In this application, we considered the quarter-over-quarter (chain) percentage changes of real GDP, exports and imports provided by Statistics Portugal (INE) since the first quarter of 1996 till the fourth quarter of 2019 (total of 96 observations), fully available 60 days after the end of the respective quarter, and complemented by the monthly data indicated in table (ref) (above).\footnote{Most recent data were not considered in this paper because the COVID-19 pandemic and lockdowns had originated a structural break in the GDP series, namely, for Portugal in the first quarter of 2020 that had required a different forecasting strategy and method, based on differences in levels instead of growth rates.}

An out-of-sample rolling forecasting exercise was performed since the first quarter of 2002 to assess the relative performance of the described nowcasting models. Each current quarter growth was estimated with the data available at day 0, 30, 60, 90 and 100. For days 0 and 30, the benchmark is the 2-steps ahead forecast given by the Theta method, recalling that the GDP change of the previous quarter was available only at day 45 of the current quarter (and the exports and imports at day 60) before 2020. For days 60, 90 and 100, we used the one period look ahead Theta forecast as benchmark.

The estimated coefficients for the full sample (1996Q1 - 2019Q4) are presented in table (ref). In general, these coefficients are 1% or 5% statistically significant through the six models with few exceptions concerned, namely, with industrial production index (IPI). Particularly relevant is the coefficient associated with the last q-o-q changes of real GDP (sum), suggesting the auto-correlated nature of the dependent variable. The six models have high, significant F statistics and adjusted R$^{2}$ above 90%, specially the models 2 and 5 with exports and imports.

Truncating the data may have a limited impact on these estimates. In fact, the coefficients presented in table (ref) for roughly an half of the sample (40 observations from 1996Q1 to 2005Q4) are similar from those condensed in table (ref). Nevertheless, IPI loosened its significance in all models and, sometimes, the Economic Sentiment Indicator (ESI) and the CEPR indicator.

Figures (ref) to (ref) illustrate the out-of-sample mean absolute error (MAE) for each model, progressively updated with more data over the current quarter. For each model and quarter from 2002Q1 to 2019Q4, we computed the direct (simple) forecast without last error correction, the forecast with that kind of correction (as described previously) and the median of these two estimates. A first evidence is that the simple mechanism of picking the last out-of-sample error might not be effective in reducing the MAE, except in some cases for model 3. Nevertheless, the intermediate point between the simple and corrected forecasts always performs better. This is an expected result in the sense that median has been proven very powerful for attenuating or even removing noise in time series Wen1999.

figure[figure omitted — 381 chars of source]
figure[figure omitted — 381 chars of source]

A second evidence is that the MAE becomes smaller and smaller from day 0 to day 100, that is, the incorporation of more information concerned with the current quarter improves the quality of the nowcasting exercise. This result was observed in all models and it is particularly evident from day 30 to day 60, recalling that the national accounts (GDP, exports and imports) for the previous quarter become fully available only at day 60 of the current quarter. Thus, that information should be extremely important to improve the accuracy of the GDP growth estimates in addition to the readily available data.

figure[figure omitted — 381 chars of source]

Figure (ref) compares the cumulative absolute out-of-sample error of the median of the simple forecasts given by the six models ((ref)) with the overall median with and without error correction ((ref)). The reduction of the absolute error by including the last error correction in the median is evident, especially for day 60 and beyond. Additionally, this figure confirms the relevance of using, progressively, more and more data over the current quarter. In fact, monthly data are still important in the sense that the cumulative absolute error reduces from day zero to 30 and even more from day 60 to days 90 and 100. Our forecasts are being published typically at day 100 with a sightly error reduction from day 90.

figure[figure omitted — 443 chars of source]

The gain in terms of out-of-sample MAE between day zero and day 60 is 0.08 percentage points (pp) from 0.53 to 0.45, as suggested by the last column of table (ref). With last error correction, the gain is slightly better, 0.09 pp, from 0.51 to 0.42, see table (ref). The inclusion of more high-frequency data concerned with the current quarter can improve the MAE additionally in 0.06 pp towards a final mark of 0.36 for day 100 with last error correction. This kind of improvement is also visible in mean squared error (MSE) and its root, which is directly comparable with MAE.

As suggested by figures (ref) and (ref) (and the same tables), our approach performs quite better than the Theta method. We found also that the last correction error may increase the MAE of the estimates computed with that method. Thus, error correction mechanisms such the one employed here should be avoided, or used with care, as far as the Theta method is concerned.

figure[figure omitted — 494 chars of source]
figure[figure omitted — 504 chars of source]

The Targeted Diffusion Index (TDI) model estimated by Dias2016 used a comparable dataset but for a shorter period (2002Q1-2015Q4). For this sub-sample, our approach gave a MAE of 0.42 pp which is close to 0.41 from TDI model. In root MSE, the difference between the two methods is meaningless at days 60 and 90 and our approach may perform better at day 100 (see table (ref)).

Conclusion

This article shows how relatively simple it is to design robust statistical models to nowcast macroeconomic variables of interest, like GDP, based on readily available indicators. To this end, it is important to mobilize either qualitative or quantitative high-frequency indicators correlated with GDP. In the former case, we have the coincident indicators of business sentiment or economic climate, and in the second case, quantitative indicators like industrial production, cement and car sales, or the value of exports of goods and services.

Both types of indicators (qualitative and quantitative) can be aggregated quarterly and used in simple multivariate linear regressions in the year-on-year (or chain) variation of GDP. In this quarterly aggregation, the monthly jagged values not yet available for the current quarter can be estimated in a very simple way, using a tendency-cycle decomposition such as the proposed by Hodrick1997, correcting any excessive noise with moving averages and using the median to get consensus among different estimates, eventually corrected from the last observed forecast error.

In fact, these rather practical procedures, although laborious and demanding, can lead to predictions with an average error similar to the one associated with more sophisticated methods like the Targeted Diffusion Index (TDI) which involves the treatment of tens or even hundreds of time series. The method described throughout the article has proved to be effective against the recognized Theta method, which is particularly good in its simplicity, and it can be replicated to a country rather than Portugal straightforwardly.

Acknowledgments

This work was supported by Fundação para a Ciência e Tecnologia (FCT), Lisbon, Portugal under a post-doctoral grant with the reference CUBE-MACROECO-BPD4 from Católica Lisbon Research Unit in Business & Economics (UID/GES/00407/2020).

Annex

table[table omitted — 2,634 chars of source]
table[table omitted — 2,548 chars of source]
table[table omitted — 731 chars of source]
table[table omitted — 763 chars of source]
table[table omitted — 817 chars of source]