Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
111,279 characters · 19 sections · 69 citation commands
Nowcasting and aggregation: Why small Euro area countries matter
\def\spacingset#1{ {#1}} \spacingset{1}
\if00 \fi
\if10 \fi
{\it Keywords:} hierarchical nowcasting, high-dimensional panels, aggregation, mixed-frequency data, textual data \thispagestyle{empty} \setcounter{page}{0}
\spacingset{1.8}
A much-researched example of nowcasting is that of US real GDP growth. Traditional methods rely on dynamic factor models which treat quarterly GDP growth as a latent process and use standard (monthly) macroeconomic data releases to obtain within-quarter estimates. There are several limitations of this approach. First, in the era of big data, it is challenging to expand dynamic factor models to include many high-frequency, non-traditional data increasingly used by macroeconomists to gauge the state of the economy. Second, going beyond the nowcasting of a single series to nowcasting the GDP growth for many countries simultaneously---or to nowcasting firms' earnings, which has many similarities with nowcasting GDP---is also challenging with the traditional models and methods. To that end, babii2022machine developed machine learning mixed-data sampling (MIDAS) regressions for single series nowcasting using potentially large data sets, and babii2022machinepanel extend these methods to machine learning panel data MIDAS regressions.
In this paper, we are interested in nowcasting Euro area-19 (EA-19) GDP growth (see, e.g., marcellino2003macroeconomic), which means that we aim to potentially nowcast GDP growth for multiple countries simultaneously. One can think of several ways to proceed, namely:
Option (a) is similar to the case of US GDP growth, with typically around 30 monthly series used to produce nowcasts. In particular, we use a larger set of monthly and weekly standard data for the EA-19 compared to the series used for the US, as will be explained in Section (ref). Even for this case, babii2022machine show that machine learning methods outperform traditional dynamic factor models using exactly the same data. Moving to option (b), the challenge of dealing with large data sets emerges. If we keep the same 30 series, but collect them for each individual country, we have potentially 30 $\times$ 19 = 570 predictors. Still, the target is a single series, i.e.\ aggregate EA-19 GDP growth. The more interesting options are (c) and (d), which are the novel contributions of the paper. In both cases, the combination of nowcasting and aggregation comes into play. The former involves nowcasting for all countries, regardless of their size and importance in the overall economic outlook of the EA. Option (d) has been entertained by cascaldigarcia2021 among others, who propose a multi-country nowcasting model to simultaneously predict the GDP of the Euro area and its three largest countries---namely Germany, France, and Italy---using up to 16 predictor variables per country.
Cases (c) and (d) pertain to nowcasting in a data-rich environment using high-dimensional mixed-frequency panel data models. This is a relatively new and unexplored research area. khalaf2020dynamic study low-dimensional mixed-frequency panels, but do not study forecasting. fosten2019panel study nowcasting with a mixed-frequency vector autoregression (VAR) but not in the data-rich environment. babii2022machinepanel introduce structured machine learning panel data regressions, sampled at various frequencies using the sparse-group LASSO (sg-LASSO) regularization. They derive the oracle inequalities and debiased inference for this method in the panel data setting with dependent fat-tailed data; see also babii2022machine and babii2020inference for the time-series setting. babii2022panelpe apply the method to nowcast a large cross-section of earnings data of US firms. In this paper, we apply the same framework for nowcasting panels of Euro area GDP growth based on real-time data vintages for standard macro data and news series extracted from a large set of newspaper articles. Our theoretical contribution pertains to the comparison of various aggregation schemes in high-dimensional data settings.
The question of aggregate versus disaggregate estimation and forecasting has a long history in econometrics; see pesaran1989econometric and references therein. lutkepohl2006forecasting and hendry2011combining discuss the theoretical comparison when the true data generating process is a VARMA process. Empirically, marcellino2003macroeconomic find that aggregating individual time-series forecasts of EU inflation and GDP is superior to aggregate forecasting. However, the previous literature has not studied the aggregation problem for high-dimensional machine learning regressions, which is the focus of our work and is novel in the literature. We formally show that the aggregation of individual components forecasted with pooled panel data regressions is superior to direct aggregate forecasting due to lower estimation error. Importantly, our theoretical comparison is performed under weak assumptions on the underlying DGP, allowing for heterogeneous coefficients as in pesaran1989econometric, pesaran1995estimating, and pesaran2024forecasting. Under heterogeneous coefficients, there is a trade-off between the overfitting of individual time-series regressions estimating a large number of coefficients and the misspecification error of pooled panel data regressions. The pooled panel data regressions may outperform the individual time-series regressions when the latter overfit due to large estimation error. Finally, simulations reported in the paper support the new theoretical results.
In the empirical application, we use standard macroeconomic series and non-standard textual series. In recent years, the use of newspaper data for nowcasting has been explored by various authors. See, for example, thorsrud2018words, larsen2021news, ellingsen2022news, barbaglia2022forecastingUS, and boss2025nowcasting among others. SCOTTI20161, baker2016, barbaglia2022forecastingEA, and ashwin2021nowcasting are recent examples that use newspaper data for economic analysis in multiple countries. Although the construction of country-specific indicators brings an additional level of complexity, the presumption is that the resulting indices provide an early signal of the local economic conditions that is more precise than a news-based general indicator. We follow barbaglia2022forecastingEA, who propose country-specific text-based indicators for five European countries using local news translated to English. In this paper, we extend their analysis to all EA-19 countries and show how the inclusion of news-based indices about smaller countries can improve the nowcasting performance.
We use the MIDAS sparse-group LASSO regression approach of babii2022machine, suitable for large-dimensional data environments, to test whether adding different levels of heterogeneity to panel models improves the quality of nowcasts. Moreover, we test whether specific data sources are informative when used in nowcasting settings. In particular, we look at whether news data can improve nowcasts over more traditionally used macro and financial series. Lastly, we apply several weighting schemes to compute the Euro area aggregate nowcast based on country-level GDP nowcasts. We show that this strategy leads to more accurate nowcasts irrespective of the weighting scheme, suggesting that information in all European countries contributes to overall predictions. In addition, we test whether smaller panels that include only the Big Four countries---France, Germany, Italy, and Spain---suffice to compute the Euro area nowcasts. Our results suggest that smaller countries indeed matter and models that incorporate all EA-19 produce higher quality nowcasts. The efficient estimation of panel models using large data sets benefits from the timely inclusion of information about the current state of the economy from smaller countries.
In sum, the paper makes multiple contributions. First, we extend the existing literature on forecasting/nowcasting and aggregation to high-dimensional data settings. Second, we study different aggregation schemes and find that a projection method appears to work best. Third, we showcase the use of non-traditional data in nowcasting, and more specifically news-related data, a topic that has received much attention recently. Fourth, we highlight the role played by the small Euro area countries due to their timely data releases and interconnectedness of their economies with the larger countries. Finally, our sample includes the pandemic, and we analyze its impact on nowcasting and model performance. The latter leads us to investigate nowcast combination schemes across different models.
The paper is organized as follows. Section (ref) presents the theoretical comparison of various aggregation schemes in high-dimensional data settings. Section (ref) presents the results of a Monte Carlo simulation study. Section (ref) considers the empirical problem of nowcasting EU output and aggregation. Finally, Section (ref) concludes. The Appendix contains additional empirical results and details on the construction of the news-based indicators.
The theoretical developments in this section use a standard linear panel data model setting used for the purpose of forecasting. Hence, we simplify the notation by ignoring the mixed frequency nature of the data and treat nowcasting as a special case of a generic forecasting problem. More specifically, we work with panel data $(y_{i,t},x_{i,t})\in\ensuremath{\mathbf{R}}\times\ensuremath{\mathbf{R}}^p$, where $t=1,2,\dots,T$ denotes time and $i=1,2,\dots,N$ is the cross-section. The vector of covariates $x_{i,t}$ can include $1$ for the intercept, the lags of $y_{i,t}$, and the (higher-frequency observations of) covariates. The objective is to predict the aggregate outcome:
using the covariates $\mathbf{x}_{t}\in\ensuremath{\mathbf{R}}^{N\times p}$, where $\mathbf{y}_t = (y_{1,t},\dots,y_{N,t})^\top$ and $\omega\in\ensuremath{\mathbf{R}}^N$ is a vector of aggregation weights. The simplest aggregation scheme is averaging, in which case the weights are $\omega_i=1/N$ for all units $i=1,\dots,N$. For a quadratic loss function, the optimal forecasts are given by conditional expectations, hence: $\mathbb{E}[Y_{t+1}|\mathcal{F}_t]$ = $\sum_{i=1}^N\omega_i\mathbb{E}[y_{i,t+1}|\mathcal{F}_t],$ where $\mathcal{F}_t$ is the information available at time $t$. Formally, $\mathcal{F}_t$ is the $\sigma$-field generated by $(y_{i,s},x_{i,s})$ for $i=1,\dots,N$ and $s\leq t$. This allows the disaggregate series $x_{i,t},i=1,\dots,N$ to be used when forecasting the aggregate $Y_{t+1}$. If one had perfect knowledge of the data-generating process (DGP), forecasting the aggregate outcome would be equivalent to linear aggregation of individual forecasts. However, in practice forecasts are estimated using data and the interplay between the estimation and specification errors becomes important.
We focus on linear forecasting models with one of the following two approaches:
It is assumed that the DGP is as follows:
with $i=1,\dots,N, \ t=1,\dots,T$. For $k\in\{\mathrm{A,AC,TS,P}\}$, let $\hat Y_{T+1}^k$ = $\sum_{i=1}^{N}\omega_i x_{i,T}^{\top}\hat\beta_i^k,$ and $Y_{T+1}$ = $\sum_{i=1}^{N}\omega_i y_{i,T+1}$ be the forecasted and the realized aggregate outcomes. Using the above equation, we can write the forecast error as a sum of estimation error, heterogeneous coefficients bias, and irreducible error:
where $\beta_i^k$ is a vector of model-specific linear projection coefficients and $H_{N,T}^k$ is the heterogeneity bias. Combining the identity in Equation (ref) with $\mathbb{E}[u_{i,T+1}\mid\mathcal{F}_T]=0$ in Equation (ref), the mean-squared forecasting error is decomposed as:
where $\Sigma_h = \mathrm{Var}(\mathbf{u}_{T+1}|\mathcal{F}_T)$. In the aggregate-on-aggregate regressions the aggregate outcome is regressed on aggregate covariates to obtain the direct aggregate forecast. Let $\hat\beta^{\rm A}$ be the estimated slope coefficient and let $\beta^{\rm A}$ := $\arg\min_{b\in\ensuremath{\mathbf{R}}^p}\mathbb{E}\left|Y_{t+1} - \sum_{i=1}^N\omega_ix_{i,t}^\top b\right|^2$ be the corresponding population projection coefficients. The forecasted aggregate outcome is $\hat Y_{T+1}^{\rm A}=\sum_{i=1}^N\omega_ix_{i,T}^\top\hat\beta^{\rm A}$.
In what follows, we focus on estimators that correspond either to individual time-series regressions or to pooled panel data regressions with sparse-group LASSO regularization.
Assumption (ref) (i)-(iv) corresponds to babii2022machine, Assumptions 3.1-3.4, where a technically precise statement can be found. The assumption on weights holds for the averaging aggregation scheme.
\paragraph{Aggregate-on-Components Regression:} In the aggregate-on-components regressions the aggregate outcome is regressed on all individual component covariates. Let $\hat\beta^{\rm AC}\in\ensuremath{\mathbf{R}}^{pN}$ be the vector of estimated coefficients and $\beta^{\rm AC}$ = $\arg\min_{\{b_i\}_{i=1}^N}\mathbb{E}\left|Y_{t+1} - \sum_{i=1}^N\omega_ix_{i,t}^\top b_i\right|^2$ be the corresponding population projection coefficients. The forecasted aggregate outcome is $\hat Y_{T+1}^{\rm AC} = \sum_{i=1}^N\omega_ix_{i,T}^\top\hat\beta_i^{\rm AC}$.
\paragraph{Pooled Panel Regressions:} In the pooled panel data regressions, the entire panel is used to estimate a single homogeneous regression. Let $\hat\beta^{\rm P}\in\ensuremath{\mathbf{R}}^p$ be the vector of estimated coefficients and let $\beta^{\rm P}$ = $\arg\min_{b\in\ensuremath{\mathbf{R}}^p}\mathbb{E}\left|y_{i,t+1}-x_{i,t}^\top b\right|^2$ be the corresponding population projection coefficient. The forecasted aggregate outcome is $\hat Y_{T+1}^{\rm P}$ = $\sum_{i=1}^N\omega_ix_{i,T}^\top\hat\beta^{\rm P}$.
\paragraph{Individual Time-Series Regressions:} In the individual time-series regressions, each individual time series is forecasted separately. Let $\hat\beta_i^{\rm TS}\in\ensuremath{\mathbf{R}}^p$ be the vector of estimated coefficients for unit $i$ and let $\beta_i^{\rm TS}$ = $\arg\min_{b\in\ensuremath{\mathbf{R}}^p}\mathbb{E}\left|y_{i,t+1}-x_{i,t}^\top b\right|^2$ be the corresponding population projection coefficient. The forecasted aggregate outcome is $\hat Y_{T+1}^{\rm TS} = \sum_{i=1}^N\omega_ix_{i,T}^\top\hat\beta_i^{\rm TS}$.
\paragraph{Data generating process:} In the simulation study, we consider the panel DGPs for the number of units (e.g., countries) $i=1,\dots,N,$ and time units $t=1,\dots,T$, where $N\in\{10,20\}$ and $T\in\{35, 100\}$, respectively. We generate $j=1,\dots,p$ regressors, $p\in\{50,500\}$. In our setup, the signal is $j=1$ while $j\ge2$ are noisy regressors. To simplify the DGP, we only consider the same frequency regressors as the outcome variable, hence we apply the LASSO estimator. The model we simulate from has the following form
For each $j\in[p]$, the regressor $x_{i,t,j}$ is an AR(1) process simulated around unit-specific locations, i.e., $x_{i,t,j}$ = $\mu_{i,j} + \xi_{i,t,j},$ and $\xi_{i,t,j}$ = $\rho_{i,j}\,\xi_{i,t-1,j} + \sigma_{i,j}\sqrt{1-\rho_{i,j}^2}\,\nu_{i,t,j},$ with $\nu_{i,t,j}\sim\mathcal N(0,1)$ for the Gaussian design and $\nu_{i,t,j}\sim\text{student}$-$t(5)$ for the heavy-tailed case, $\rho_{i,j}\sim\text{Uniform}(0,0.95)$, $\sigma_{i,j}>0$. We simulate $\sigma_{i,j}^2 \sim (1 + \chi^2_1)/2$. Lastly, $ \mu_{i,j} = (z_{i,j}^2 - 1) / \sqrt{2}$ where $ z_{i,j}\sim \mathcal N(0,1)$ or $ z_{i,j}\sim\text{student}$-$t(5)$.
\paragraph{Heterogeneity:} First, we impose a homogeneous AR coefficient, i.e, we set it to $\gamma_i = 0.688,\forall i\in[N].$ The heterogeneous coefficients are the intercept $\alpha_i$ and the slope $\beta_i$. We assume the following structure for $\alpha_i$ and $\beta_i$, respectively, $\alpha_i$ = $\alpha_0$ + $\phi\,\mu_{i,1}$ + $\sigma_\eta\,\eta_i,$ and $\beta_i$ = $\beta_0$ + $\pi\,\mu_{i,1}$ + $\sigma_\zeta\,\zeta_i,$ with $\eta_i,\zeta_i\sim\mathcal N(0,1)$ for Gaussian or $\eta_i,\zeta_i\sim\text{student}$-$t(5)$ for heavy-tailed scenarios, and $\phi=\rho_{\alpha x}\mathbf{\sigma}$, $\pi=\rho_{\beta x}\mathbf{\sigma}$, $\sigma_\eta^2=\sigma^2-\phi^2$, $\sigma_\zeta^2=\sigma^2-\pi^2$ and $\alpha_0=0$, $\beta_0=0.5$. Lastly, the error term is simulated from $\varepsilon_{it}$ = $\sigma_{\varepsilon,i}\,(\zeta_{it}^2-1)/\sqrt{2},$ with $\zeta_{it}\sim\mathcal N(0,1)$ or $\sim\text{student}\text{-}t(5),$ $\sigma_{\varepsilon,i}>0.$ The other free parameters are fixed to the following values: $\rho_{\alpha x}=0.1$, $\rho_{\beta x}=0.5$. Throughout the experiment, we vary the parameter $\sigma\in\{0, 0.2, 0.4, 0.6, 0.8\}$, which controls the level of heterogeneity. For $\sigma=0$, $\alpha_i = \alpha_0 = 0$ and $\beta_i = \beta_0 = 0.5$, while $\sigma\neq0$ implies heterogeneous $\alpha_i$ and $\beta_i$.
We compute predictions using four models which we estimate by applying the LASSO estimator. The first is the pooled regression model for all $N$ units, which we denote as P. Next, we estimate individual regressions for each $i$, denoted TS. The third case considers a model where we predict the aggregate outcome using all $N$ individual regressors. We call this model AC. Lastly, we consider aggregate-on-aggregate approach, denoted A, where we regress the aggregate outcome on aggregate regressors. Each aggregation assumes fixed weights set to $1/N$.
In Table (ref) we report results for the individual time-series (TS), aggregate-on-components (AC), and aggregate-on-aggregate (A) approaches, all relative to the pooled model (P) out-of-sample mean squared errors (MSE). We compute MSEs for all approaches for the target out-of-sample aggregate outcome. For pooled and individual cases, we use $1/N$ weights to aggregate our predictions.
The Monte Carlo simulations show that pooled regressions strike a balance between estimation error and heterogeneity bias. Compared to aggregate-on-aggregate and aggregate-on-components regressions, the pooled model generally achieves lower mean squared errors because it benefits from the larger effective sample size, which reduces estimation noise. However, when heterogeneity across units is strong, pooling imposes a homogeneity restriction that creates bias. This trade-off explains why pooled regressions outperform individual time-series regressions when the latter overfit in small samples, but can be worse when unit-specific variation matters. In short, pooled regressions dominate aggregate-based methods and are more robust than TS in finite samples, yet they may underperform when heterogeneity is substantial.
The simulation results also reveal that pooled regressions behave differently under distributional and dimensional changes. Moving from Gaussian to student-$t(5)$ innovations, the pooled model remains relatively stable, while competing methods such as TS and AC become more sensitive to heavy tails and deteriorate in accuracy. This robustness highlights the pooled model’s advantage in heavy-tailed environments where heavy-tailed observations inflate estimation errors elsewhere. Increasing dimensionality from $p=50$ to $p=500$ introduces additional noise, which tends to amplify the overfitting of TS and AC methods, whereas pooling, again leveraging the combined sample size, better controls variance. Overall, across both the distributional shift from Gaussian to student-$t(5)$ and the dimensional increase, pooled regressions exhibit the most reliable performance relative to the other three approaches and the results are in line with our theoretical results.
In the section we describe the machine learning MIDAS panel data models that we use for nowcasting. Then, we present the aggregation schemes used to obtain EA-19 GDP growth nowcasts. Finally, we discuss the data and the empirical results.
Our approach builds on babii2022machine, who introduced machine learning MIDAS (ML MIDAS) regressions with an application to single series nowcasting, and babii2022machinepanel, who extended the method to panel data settings. Moreover, babii2022panelpe consider an application of nowcasting price-earnings ratios for a large set of US firms. In Online Appendix Section (ref) we provide a detailed discussion of these types of models.
These models fit into a generic linear panel regression setting, albeit with regularization to take account of the high-dimensional nature of the regressors. We follow the notation in the aforementioned papers. Define $\mathbf{y}_i$ = $(y_{i,1+h},\dots,y_{i,T+h})^\top$, with $h$ the forecasting/nowcasting horizon, $\mathbf{\tilde y}_{i,q}$ = $(y_{i,1-q},\dots,y_{i,T-q})^\top$ for $q\in[Q]$ lagged dependent variables, $\mathbf{\tilde y}_i = (\mathbf{\tilde y}_{i,1}, \dots, \mathbf{\tilde y}_{i,Q})$, and $\mathbf{u}_i$ = $(u_{i,1},\dots,u_{i,T})^\top.$ Stacking time series observations in vectors, the regression equation for each $i\in[N]$ with pooling of covariates $\mathbf{x}_i$ is:
where $\iota_T$ is a size $T$ vector of ones, $\rho_i\in\ensuremath{\mathbf{R}}^{Q}$ coefficients of autoregressive lags, and $\beta\in\ensuremath{\mathbf{R}}^{LK}$ is a vector of regression slopes. We also define the vector of all outcomes $\mathbf{y} = (\mathbf{y}_1^\top,\dots, \mathbf{y}_N^\top)^\top$, regressors $\mathbf{X}=(\mathbf{x}_1^\top, \dots, \mathbf{x}_N^\top)^\top$, and errors $\mathbf{u} = (\mathbf{u}_1^\top,\dots,\mathbf{u}_N^\top)^\top$. Then stacking all observations together, we obtain
where $\mathbf{\tilde Y}$ is a diagonal matrix with elements $\mathbf{\tilde y}_i,i=1,\dots,N$, and $\rho=(\rho_1^\top,\dots,\rho_N^\top)^\top.$ To deal with the large number of predictors, we use the sg-LASSO regularization that was used successfully for individual time series machine learning regressions in babii2022machine. The MIDAS approach reduces efficiently the dimensionality of high-frequency lag coefficients. An alternative approach, known as the U-MIDAS, see foroni2015unrestricted, would estimate all individual coefficients associated with high-frequency covariate lags hoping that machine learning would pick up relevant lags. This strategy is not appealing because we always pay a price for the model selection which can be substantial with heavy-tailed time series data, typically leading to worse predictions compared to regularized MIDAS schemes; see babii2022machine, babii2022machinepanel for further discussion and details. The pooled panel data estimator with heterogeneous autoregressive dynamic and a variation of sparse-group LASSO solves
where $\|.\|_{NT}^2 = |.|^2/(NT)$ is the scaled $\ell_2$ norm squared and $\Omega_\gamma(b,c)$ = $\gamma\left[|b|_1 + |c|_1\right] + (1-\gamma)\left[\|b\|_{2,1} + \|c\|_{2,1}\right]$ is a penalty, which is a linear combination of LASSO and group LASSO penalties. The weight parameter $\gamma\in[0,1]$ determines the relative importance of the $\ell_1$ (sparsity) and the $\ell_{2,1}$ (group sparsity) norms. The amount of regularization is controlled by the regularization parameter $\lambda\geq 0$. Recall also that, for a group structure $\mathcal{G}$ described as a partition of $[p]=\{1,2,\dots,p\}$, the group LASSO norm is $\|u\|_{2,1}=\sum_{G\in\mathcal{G}}|u_G|_2$, where $u$ is a generic vector and $u_G$ are the elements of $u$ corresponding to group $G$. We assume that the group structure is observed, which in our setting corresponds to: a) country-specific autoregressive lags, and b) time series lags of covariates. It is also feasible to combine covariates of a similar nature into groups.
In addition to the pooled panel models, as in Section (ref), we also look at variations involving individual country models. In the discussion of the empirical results, we will refer to the following models:
We use the following schemes to combine country-level GDP growth nowcasts $y_{i,t|\tau}$ for country $i$ in quarter $t$, given the high-frequency information available up to $\tau$, and compute the aggregate Euro area level GDP nowcasts, denoted $y_{\text{ea},t|\tau}$:
For all weighting schemes we use $t-1$ GDP data, since at quarter $t$ this is the information available in real-time. The Euro area aggregate nowcast is computed as $ y_{\text{ea},t|\tau} = \sum_{i=1}^N W^{(q)}_{i|t} y_{i,t|\tau},$ for $q\in\{1,2,3,4\}$, corresponding to the set of weights.
Finally, we also consider what we call the Euro area model. In this model, we nowcast Euro area GDP growth based on aggregate Euro area data. For this, we use the machine learning MIDAS setup of babii2022machine.
We use real-time vintages of standard macro monthly releases from several sources (see the Appendix for the details of each series). GDP growth is quarterly and is available in real-time in our sample. The first vintage is January 2015, for which we have GDP vintages for all EA-19 countries. As predictor variables, we collect 64 traditional macroeconomic series: 57 are monthly series, three are weekly, and four are daily. Monthly series are hard data such as the unemployment rate and industrial production; 47 are available at the country level while the remaining 10 are Euro area aggregates; weekly series are oil products which are available at the country level; four daily series are financial markets data covering stock market, gold, foreign exchange, and interest rates.
It is worth mentioning that some macro series are available with a month, two, or even three months of delay relative to the nowcasting month. Since we use real-time data vintages, we naturally take into account such delays. This is particularly important when we analyze the additional gain in nowcasting accuracy when using news data---which we describe below---since such data are timely and available without delays.
We collect news data from Dow Jones Factiva. The data set contains daily printed and online full-text articles from three different sources dealing with economic and financial issues, namely The Economist, Reuters News, and The Wall Street Journal. The final data set consists of approximately $2.5$ million articles and 1 billion words from January 2005 to December 2022. We construct news-based indicators following the fine-grained aspect-based sentiment (FiGAS) by consoli2022fine. This is a rule-based algorithm that provides sentence-level sentiment scores for textual information in the English language. The sentiment is (i) aspect-based, meaning that it analyzes only the words in a sentence that relate to a specific topic of interest based on linguistic dependencies, and (ii) fine-grained, that is, the sentence-level sentiment score comes from a human-annotated dictionary tailored for economic and financial applications and is defined in $[-1,+1]$.
We compute news-based indicators for the following three topics covering different aspects of economic and financial activity: economy, employment, and inflation. Each topic is associated with a set of additional keywords that we look for in the articles: for instance, the economy topic also includes related terms, like GDP, output, or economic growth. We compute country-specific indicators for all EA-19 member states by filtering only on sentences where there is a direct mention of the country in the analysis. We refer to the online appendix for additional details about the news-based sentiment indicators, the selection of the topics' keywords, and the construction of country-specific measures.
The output of FiGAS consists of daily news-based indicators of sentiment and volume for each topic and country. The sentiment measures are obtained by summing the sentence-level sentiment scores for all articles within that day, while the volume corresponds to the number of sentences containing an explicit mention of the topic. Most importantly, the news-based sentiment indicators are real-time and we include them as additional regressors with no publication delay. In the remainder, our baseline model will include news-based indicators about the economy, employment, and inflation.
We apply the machine learning regressions described in Section (ref) to assess whether the aggregate Euro area GDP growth nowcasts are more accurate than panel data models which are based on individual country-level data and several weighting schemes of individual country nowcasts to construct an aggregate nowcast.
\paragraph{Aggregate versus panel data regressions:} Table (ref) reports the nowcast comparison between the (aggregate) Euro area model and the panel data models, namely the pooled and the HetAR panel models, and the country-specific MIDAS regressions. The forecast accuracy is measured as root-mean-squared errors (RMSEs) at three nowcast horizons (i.e., 2- and 1-month ahead and end of the quarter). The first row in Panel A of Table (ref) reports the RMSEs of the aggregate Euro area benchmark model, while the other rows report the RMSEs of the proposed panel data models relative to the benchmark: values below unity signal a better performance of the proposed model relative to the benchmark. For each panel data model, we document the performance of the four different weighting schemes ($W^{(1)}$-$W^{(4)}$) to aggregate individual country forecasts as described in Section (ref). Columns 1 to 3 report the results on the full sample, while columns 4 to 6 consider only the pre-COVID period.
Panel A reports the results for the time series regressions when we nowcast directly Euro area aggregates. While the first row regressions include only aggregated information about the Euro area as an explanatory variable (option (a) in the Introduction), the other rows in the top panel add respectively the aruoba2009real (ADS) index and information from individual EA-19 countries as additional regressors (option (b) in the Introduction). A few patterns are worth highlighting. First, the model's accuracy is largely impacted by the inclusion of the COVID period observations, with the full sample results showing much larger errors than the pre-COVID sample. Second, the inclusion of timely information about the US business cycle proxied by the ADS index does not bring any systematic improvement with respect to the benchmark model. Third, an even worse performance results from the inclusion of information about individual EA-19 countries into the time series regression model compared to the benchmark model only using aggregate EA data. Hence, option (a) is better than (b), and therefore using only the aggregate data in machine learning models suffices.
Turning to the performance of the panel data machine learning models in Panels B and C, we obtain mixed results. For the pre-COVID sample, we note that the panel data models under perform vis-\`a-vis the benchmark. In contrast, the panel data models always outperform the benchmark when looking at the full sample. This result is robust with respect to the choice of the weighting scheme and the nowcasting horizon. The proposed models exploit the additional country-specific information included in the panel data structure: this information turns out to be redundant during normal times, while it proves relevant during the COVID pandemic. Comparing the two panel data models, the HetAR model generally achieves better performances than the pooled models: the inclusion of heterogeneity in the form of country-specific lags seems to be a better choice than pooling all coefficients.
Panel D of Table (ref) reports the results with country-specific MIDAS regressions -- hence not exploiting the panel data structure, but instead estimating single regressions per country -- which are then aggregated using again the same weighting schemes. Hence, Panel D differs from Panels B and C only with respect to the panel structure, while the information set, the MIDAS structure and the weighting schemes are the same. Country-specific regressions aggregated following weighting schemes $W^{(3)}$ and $W^{(4)}$ attain good results both when looking at the full sample and in the pre-COVID period. On the one hand, Panel D shows gains with respect to the Euro area benchmark during the pre-COVID sample, when panel models perform poorly. The gains attained by country-specific regressions are large and up to 45 percentage points. On the other hand, the performance in the full sample, although still better than the benchmark, does not improve with respect to panel models which perform best when including the COVID pandemic in the analysis.
\paragraph{The Big Four:} We test whether nowcasting the four largest Euro area countries---France, Germany, Italy, and Spain (i.e., the “Big Four")---separately may give an advantage in producing more accurate predictions for the Euro area GDP. See also ashwin2021nowcasting, barbaglia2022testing, cascaldigarcia2021, among others. We compare the performance of the HetAR and pooled panel models, as well as the country-specific regressions as in the previous section. To compute the aggregated Euro area nowcasts we use weighting scheme $W^{(4)}$ modified to include only the Big Four.
Table (ref) reports the results of the four-country models relative to the benchmark model appearing in the first row of Table (ref). All point to a better performance of the model involving nowcasting only the largest European economies. This result is robust across horizons, model specifications and, interestingly, the full and pre-COVID samples. Indeed, panel data models with the Big Four attain a more accurate performance even when considering the pre-COVID sample, whereas that was not the case in Table (ref). However, looking at the full sample, panel data models with all Euro area countries always outperform the Big Four models. Hence, rather than focusing only on the indicators from the largest European countries, our evidence suggests that there is a value-added in the nowcasting model for all EA-19 countries when looking at the full sample and including the COVID in our model. Looking at the pre-COVID sample, the Big Four models are outperformed at all horizons by the country-specific regressions using all Euro area countries reported in Table (ref). Overall, this suggests the best results are obtained with the inclusion of information from all EA-19 countries both in the full sample (with panel models) and pre-COVID (with country-specific regressions).
To explain the relative improvement in nowcasting performance achieved by using the full panel of Euro area countries as opposed to just the Big Four, several factors come into play. First, from a modeling perspective, machine learning panel regression models are estimated more accurately when $N$, i.e.\ the number of countries, is large - see babii2022machinepanel. More accurate parameter estimates lead to higher quality nowcasts. Second, the real-time flow of data and information varies among countries. For instance, Belgium, a nation falling outside the Big Four category, consistently boasts superior and more timely survey data, significantly contributing to the accuracy of Euro area GDP predictions, see basselier2018nowcasting. From an economic perspective, given Belgium's strong economic ties with Germany and France due to its geographical proximity, its economic news serves as a reliable signal for the broader economies, therefore influencing the overall Euro area GDP projection. By incorporating the entire panel of countries, we can effectively capture and harness these effects. The Online Appendix provides additional insights about the drivers of the nowcasting performance and a country-level evaluation.
\paragraph{Nowcast aggregation and combination:} We now focus on how to aggregate and combine individual nowcasts, starting from the nowcasting performance of the four proposed weighting schemes. Looking at the full sample results of Table (ref), we observe that $W^{(4)}$ achieves the lowest RMSEs for the pooled panel model. This result does not hold for the HetAR model, where the $W^{(3)}$ weighting scheme produces slightly more accurate forecasts than $W^{(4)}$. Regarding country-specific regressions, $W^{(3)}$ and $W^{(4)}$ attain similar performance in both the full sample and the pre-COVID period, and substantially outperform the other weighting schemes. Figure (ref) illustrates the four estimated weighting schemes, with $W^{(3)}$ and $W^{(4)}$ notably exhibiting denser characteristics than the other schemes. Consequently, nowcasts employing denser weighting schemes yield more accurate results, further reinforcing the argument in favor of utilizing the entire panel of European countries to improve nowcasting precision. In the Online Appendix Section (ref) we further investigate the nowcasting performance across weighting schemes, where we compare the weights obtained by $W^{(4)}$ (i.e., projections on GDP) and by $W^{(3)}$ (i.e., proportion of GDP level). Compared to $W^{(3)}$, we observe that $W^{(4)}$ assigns smaller weights to Germany and to a lesser extent the Netherlands, while it inflates the relative importance of some small- and medium-sized countries, namely Austria, Belgium and Luxembourg. Although the size of these economies within the Euro area is relatively small, information about economic developments in those countries plays a relevant role in attaining more accurate nowcasts.
Finally, we explore the relative importance of smaller countries against the Big Four by plotting in Figure (ref) the sum of weights of the 15 small Euro area countries (Small15) relative to the sum of the weights of the Big Four for weighting schemes $W^{(3)}$ (dashed line) and $W^{(4)}$ (solid line). While the relative importance of smaller countries with respect to the Big Four is stable when looking at the $W^{(3)}$ weights (indeed, proportions of GDP level are steady in time and vary only at a slow pace), we observe large variability when considering $W^{(4)}$. Smaller countries are relatively more important in 2016-17 and, most notably, after 2020, when the total weight assigned to smaller countries doubles, going from 0.3 to approximately 0.6. The information coming from smaller Euro area countries plays an important role in attaining an accurate nowcast in the years following the COVID-19 pandemic, hence explaining the good performance of panel models on all EA-19 countries in the full sample reported in Table (ref).
Overall, our results highlight a heterogeneous nowcast performance of the analyzed models: panel models perform best when considering the full sample and information from all Euro area countries, country-specific MIDAS regressions outperform all other models in the pre-COVID sample, while nowcasting models on the Big Four attain a good performance across all time samples, although not the best one. Given this heterogeneity in performance, we combine forecasts obtained from individual models following the linear combination method by stock2004. In particular, we combine the forecasts obtained from the following models: the Pooled panel model and Country-specific regressions on all EA-19 countries (Panels B and D, respectively, of Table (ref)), and the HetAR model on the Big Four (Panel B of Table (ref)). For all individual forecasts, we consider the aggregated forecasts obtained with $W^{(4)}$ weights. We select forecasts aggregated with $W^{(4)}$ weights since they provide the best nowcasting performance overall. Regarding the selection of the individual nowcasting models, we include all model types analyzed in the paper (Pooled and HetAR panels, as well as country-specific MIDAS regressions) taking into account the time period and information set (i.e., all EA-19 or only the Big Four countries) where they achieve their best performance. We have also experimented with other model subsets for which we obtained similar results.
The nowcast combination results are reported in Table (ref). We start with Panel A which covers models with news data. Combining the nowcasts consistently delivers large gains which range between a 20 to 30% reduction with respect to the benchmark RMSEs. Moreover, the gains are robust across horizons and, importantly, full versus pre-COVID samples. Indeed, combinations attain more accurate nowcasts both in the full sample and in the pre-COVID period, reaching a performance that is close to the best-performing individual model. Panel B of Table (ref) explores one final aspect of our analysis, namely the added value of the news indicators. Compared to Panel A showing the nowcasting performance of models including news indicators, Panel B excludes them. The inclusion of news delivers nowcasting gains, even though these gains are not large, ranging between 1-5%. Therefore, although only marginally, the inclusion of news indicators positively impacts the accuracy of the nowcasts.
The paper studies the Euro area nowcasting using MIDAS machine learning panel data regression models. These models offer the flexibility needed to analyze extensive datasets gathered from diverse sources, sampled at varying frequencies, and available in both cross-sectional and time series dimensions. Through this research, we introduce several innovative empirical findings that carry significant relevance for policymakers. Our findings reveal a significant enhancement in the accuracy of nowcasts for the Euro area aggregate GDP when incorporating data from smaller countries. These improvements are substantial, reaching up to 30%, and remain robust across nowcasting horizons. We attribute these gains to two primary factors.
Firstly, in highly turbulent times such as the COVID-19 pandemic, the parameter estimates of the machine learning panel data regression models are estimated more precisely when a broader cross-sectional dataset is employed. Secondly, smaller countries, such as Austria, Belgium, or Luxembourg, in contrast to larger Euro-area countries like Germany or France, tend to possess higher-quality and more timely survey data. Given the high level of economic interconnectivity, especially among neighboring countries in the Euro area, the inclusion of data from these smaller nations results in more accurate signals, ultimately leading to enhanced model predictions.
In addition to our primary findings, we also contribute to the expanding body of literature concerning the application of alternative data, such as information extracted from newspaper articles, to enhance the accuracy of economic forecasts. The incorporation of such data is particularly advantageous due to their real-time availability and the lack of any publication delay for these data. Our research demonstrates that news data improve the accuracy of Euro area GDP nowcasts.
\if00 {
We received helpful comments from Peter Reusens, Paolo Paruolo, Wouter Van der Veken, and Raf Wouters, as well as participants at a National Bank of Belgium seminar, the Nowcasting workshop at the Paris School of Economics, and the 5$^{th}$ Conference on “Nontraditional Data, Machine Learning, and Natural Language Processing in Macroeconomics." The views expressed are purely those of the authors and should not, in any circumstances, be regarded as stating an official position of the European Commission. Eric Ghysels and Jonas Striaukas gratefully acknowledge the financial support of the National Bank of Belgium. Jonas Striaukas also acknowledges the financial support of the F.R.S.\---FNRS PDR under project Nr. PDR T.0044.22. } \fi
\if10 {
} \fi
\setcounter{page}{1} \setcounter{section}{0} \setcounter{equation}{0} \setcounter{table}{0} \setcounter{figure}{0}
We construct news-based indicators following the fine-grained aspect-based sentiment (FiGAS) by consoli2022fine. This method provides sentence-level sentiment scores for textual information in the English language. FiGAS has two main characteristics. First, it is aspect-based, that is, it computes a sentiment score about a specific topic, rather than the overall sentiment of all sentences in a text. Second, it is fine-grained, meaning that words are assigned a sentiment score in [-1,+1] taken from a dictionary developed for economic and financial applications.
If a sentence contains a direct mention of the topic of interest, then FiGAS tags each word in the sentence with its part-pf-speech and its dependency. Based on these tags, the algorithm checks whether the sentence corresponds to one of the linguistic rules mapped in FiGAS, which aims to capture semantic constructions that characterize the topic of interest. Examples of rules are the presence of an adjectival modifier or clause, or the topic of interest being followed by an object predicate. We refer to consoli2022fine for a detailed description of the method.
In our application, we compute news-based indicators for three topics covering different aspects of economic and financial activity: economy, employment and inflation. We associate a list of additional keywords to each topic to better represent the news coverage of that economic concept as detailed in Table (ref). The selection of the keywords starts form the list of related terms extracted from the global knowledge graph from the Global Data set of Events, Language and Tone (GDELT) and it cuts off the most frequent terms based on their presence in GDELT news items. Finally, we subset the list of additional keywords by selecting only the terms that share the same polarity: for instance, we do not include in the same topic a term that has a positive connotation (e.g., employment) with a term with a negative one (e.g., unemployment).
The news-based indicators are country-specific. We rely on one additional feature of FiGAS which allows filtering sentences with a direct mention of a geographic location. FiGAS uses named-entity recognition to identify entities in the full-text Mohit2014. If it identifies an entity, it assigns the most relevant location to the whole text, unless the sentence mentions directly another location. Table (ref) reports the list of keywords that we employ to filter the location for each Euro Area country.
We now report two examples to show how FiGAS works. Consider the sentence “the Italian economy is expected to grow more than 6% this year after the record slump recorded in 2020" that appeared in Reuters News on December 22$^{nd}$ 2021. FiGAS identifies one particular linguistic rule, more specifically, this sentence consists of the topic of interest (i.e., the Italian economy) connected to a verb (i.e., is expected) followed by an open clausal complement (i.e., to grow). Furthermore, the algorithm detects a specific mention of a location and assigns it to Italy. The sentence-level sentiment score is 0.6, representing the positive outlook for the Italian economy. Now consider this sentence “the mighty automobile industry increased its output by 1.9%" that was published in Reuters News on September 7$^{th}$ 2021. FiGAS identifies the combination of two linguistic rules that characterize the topic of interest (i.e., the automobile industry): in detail, the topic of interest is associated with two adjectival modifiers, either in the form of an adjective (i.e., mighty) or as a verb (i.e., increased). Given that there is no direct mention of any location in this sentence, FiGAS assigns the most relevant location detected in the full article, which is Germany. The overall sentence-level sentiment score is 0.49 and captures the positive connotation in which the text is describing the German automobile industry.
For each country and topic, we obtain two daily news-based indicators for sentiment and volume. The first one is obtained by summing the sentence-level sentiment scores, while the second one counts the number of sentences that discuss that topic in a day. As an example, Figure (ref) plots the economy sentiment indicators for all Euro Area 19 countries. The time series are aggregated monthly by summing and smoothed with a yearly rolling moving average. There is some clear heterogeneity across the countries showing that each indicator captures country-specific economic developments. In most cases, we notice a good fit with the business cycle fluctuations: for all major economies in the Euro Area, the economy sentiment deteriorates clearly in correspondence with the three recessionary periods in our sample.
In Section (ref) of the paper we simplified the notation by ignoring the mixed frequency nature of the data. Here we introduce the MIDAS panel data models used in the empirical nowcasting study. More specifically, we work with panel data $(y_{i,t},x_{i,t})\in\mathbb{R}\times\mathbb{R}^p$, where $t=1,2,\dots,T$ denotes time and $i=1,2,\dots,N$ is the cross-section. The vector of covariates $x_{i,t}$ can include $1$ for the intercept, the lags of $y_{i,t}$, and the (higher-frequency observations of) covariates.
Our goal is to nowcast a set of variables $y_{i,t+h}$ with $i\in[N]$ at horizon $h$ measured at some low frequency, e.g.\ quarterly, $t\in[T]$, where $[p]=\{1,2,\dots,p\}$ for a positive integer $p$. For simplicity, we assume equally spaced data at different frequencies. In particular, $n_k^H$ denotes the total number of high-frequency observations for the $k^{\rm th}$ covariate for each low-frequency period $t$, and $n^L_k$ is the number of low-frequency periods used as lags.\footnote{For example, in our application a quarter of high-frequency lags used as covariates corresponds to $n^L_k=1$ while $n_k^H=3$ indicates that three months of data are used in each quarter. Note that we can have a mix of quarterly, monthly, and weekly data, and for each covariate indexed by $k$, $n_k^H$ represents different high-frequency sampling frequencies and associated lags $n_k^L n_k^H.$} The information set consists of $K$ predictors, i.e.\ $\left\{x_{i,t-j/n_k^H,k}:\;i\in[N],t\in[T],j=0,\dots,n^L_kn_k^H-1,k\in[K]\right\},$ measured potentially at some higher frequency, e.g., monthly/weekly/daily with real-time updates into the quarter.
Consider the following panel data regression for the low-frequency panel target $y_{i,t|\tau},$ using information up to $\tau$:
where
and (a) $\alpha$ is the intercept constant across all $i$; (b) $\rho_{i,q}$ are autoregressive lag coefficients possibly different across $i$; (c) $k_{\rm max}$ is the maximum lag length which could depend on the covariate $k,$ and for each high-frequency covariate $x_{i,\tau,k}$ we have the most recent information at time $\tau.$\footnote{All parameters depend on $\tau$, but we suppress this detail to simplify notation. Furthermore, we will not keep track of the dependence of $x_t$, $k_{\max}$, and $K$ on $\tau$.} Note that this also implies the most recent vintage of data is used, i.e.\ past data revisions are incorporated as well. When $\tau$ $\leq$ $t - 1,$ we are dealing with a forecasting situation, hence, our analysis applies to forecasting as well. Note also that $k_{\max}$ is the maximum lag length which may depend on the covariate $k,$ and for each high-frequency covariate $x_{i,\tau,k}$ we have the most up-to-date information available at time $\tau.$ For some high-frequency regressors that have not been updated yet, this could be stale information; see babii2022panelpe for further discussion. In our empirical analysis, we consider the following two variations of this general formulation. For $i,j\in[N]$: (a) Pooled panel, denoted {\it pooled}, restriction: $\rho_{i,j} = \rho_j$. (b) Country-specific AR lag coefficients, denoted {\it HetAR}. Note that in the HetAR case, we model country-specific effects, therefore we investigate whether more flexible models can improve the quality of nowcasts relative to the pooled panel model by including country-specific autoregressive coefficients in the model. Lastly, in the pooled, we group pooled autoregressive lags in time series dimension rather than the cross-section.
The large number of predictors $K$ with potentially large number of high-frequency measurements $k_{\max}$ can be a rich source of predictive information, yet at the same time, estimating $N\times Q + 1 + \sum_{k=1}^K k_{\max}$ parameters, where $N$ is the size of the cross-section, $Q$ is number of autoregressive lags (constant across $i$) and $k_{\max}$ is the max of high-frequency lags for each covariate (constant across $i$), is costly and may reduce the predictive performance in small samples. In addition, observing predictors at different frequencies leads to the frequency mismatch problem due to the missing data. To solve both problems, we follow the MIDAS (machine learning) literature and instead of estimating a large number of individual slope coefficients, we consider a weight function\footnote{We use $q=1$ for nowcasting without low-frequency lags. More generally, $q-1$ denotes the number of low-frequency lags.} $\omega:[0,q]\to\ensuremath{\mathbf{R}}$, parameterized by $\beta_k\in\ensuremath{\mathbf{R}}^L$
where $[0,q]\ni s\mapsto$ $\omega(s;\beta_k)$ = $\sum_{l=0}^{L-1}\beta_{l,k}w_l(s),$ and $(w_l)_{l\geq 0}$ is a set of $L$ functions, called the dictionary. One can use either splines or Legendre polynomials as a dictionary; see babii2022machinepanel for further discussion on different choices of the dictionary. The attractive feature of this approach is that we can map the MIDAS regression to the simple linear regression framework (as in Section (ref) of the paper). To that end, assuming that $k_{\max}$ is the same for all covariates and $\mathbf{x}_i = (X_{i,1}W_k,\dots,X_{i,K}W_k)$, where for each $k\in[K]$, $X_{i,k} = \{x_{i,\tau-j/n_k^H,k},j = 0, \ldots, {k}_{\max} - 1\}_{\tau\in[T]}$ is a $T\times {k}_{\max}$ matrix of predictors and $W_k{k}_{\max} $ = $(w_l(j/n_k^H; \beta_k))_{0\leq l\leq L-1, 0 \leq j\leq {k}_{\max}}$ is a ${k}_{\max} \times L$ matrix corresponding to the dictionary. Note that the matrices $(W_k)_{k\in[K]}$ solve the frequency mismatch problem.
Say the training and validation set for a given nowcast is $1,\dots,N$ in the cross-sectional dimension and $1,\dots, T_0$ in the time dimension. To select the tuning parameters of the sparse group LASSO, we use 5-fold cross-validation adapted to the panel setting. Specifically, the time dimension is partitioned into 5 contiguous folds, and for each fold the model is re-estimated on the remaining periods and validated on the excluded fold. The fold structure is replicated across all cross-sectional units so that entire time blocks, rather than individual observations, are left out. This ensures that dependence across countries at a given date is preserved and avoids information leakage across folds. The cross-validation error is then computed for each candidate penalty parameter, i.e, ($\gamma$, $\lambda$), and the optimal value is chosen for the minimum cross validation error as is standard in cross-validation implementations. We set the grid for the $\gamma = (0, 0.05, 0.1, \dots, 0.95, 1)$ while the grid for $\lambda$ is based on the usual path construction for the LASSO, i.e., for a given $\gamma$, we find the largest $\lambda$ such that the parameter vector is zero. We then construct the grid for $\lambda$ between the largest and the $\lambda$ value multiplied by a factor $10^{-2}$ in the log-space. Once the optimal combination of both tuning parameters are chosen, we re-estimate the model on the full data set using the optimal tuning parameters. We then use the estimated coefficients to predict the outcome at time $T_0+1$. We repeat the process by using expanding window in time dimension until all out-of-sample outcome data points are exhausted.
Lastly, we note that the cross-validation procedure is implemented in the open-source R package midasml function cv.panel.sglfit. The same procedures where used in the articles babii2022machinepanel, babii2022panelpe — see both articles for further details.
Mixed-Frequency Vector AutoRegressive (MF-VAR) models have become valuable tools in the nowcasting framework, enabling the joint modelling of multiple time series observed at different frequencies while capturing the dynamic interdependencies among macroeconomic variables. Beyond improving nowcasting accuracy, MF-VAR models also support structural analysis through the interpretation of estimated coefficients, offering insights into the transmission mechanisms of economic shocks. In a single-region context, two prominent contributions are schorfheide2015real, who develop a Bayesian MF-VAR framework to forecast the U.S.\ business cycle using a mix of quarterly and monthly indicators, and chan2023high, who extend this approach by incorporating weekly data. This standard setting has been further expanded by koop2020regional, who apply the MF-VAR model to nowcast regional economic activity in the United Kingdom, and by koop2025monthly, who focus on producing monthly growth estimates for individual U.S.\ states.
We include the MF-VAR as an additional benchmark in our analysis. Similar to schorfheide2015real and koop2020regional, we address the frequency mismatch between quarterly and monthly variables using the triangular decomposition method discussed in mariano2003new, mariano2010coincident. Given the inclusion of over 50 variables for each of the EA-19 countries and the relatively short time sample, the dataset is high-dimensional. Estimating the MF-VAR in this context poses challenges related to multicollinearity and the risk of overfitting. Similar to schorfheide2015real, we model GDP growth rates using a selected subset of monthly and quarterly indicators.\footnote{The MF-VAR is estimated using the following variables: GDP growth; Industrial Production (Manufacturing); Harmonised Index of Consumer Prices; Long-term Government Bond Yield; Unemployment Rate; Trade Balance; and Real Effective Exchange Rate.} We employ a Minnesota prior and incorporate stochastic volatility, conducting inference based on 10,000 replications. The real-time nowcasting framework mirrors that of the main paper. We evaluate the performance of the MF-VAR by benchmarking it against an EA-19 model that relies solely on aggregate data.
Table (ref) reports the performance of two different specifications of the MF-VAR. In Panel A, we fit the MF-VAR only on the EA-19 aggregate time series, while, in Panel B, we fit the MF-VAR on the individual country data and obtain the aggregate EA-19 forecast by using the weights $W^{(1)}-W^{(4)}$ defined in the main paper. Looking at Panel A, we observe that the MF-VAR on aggregated data improves the nowcasting performance against the benchmark in all cases expect for the EoQ in the pre-COVID period. Moving to Panel B, a similar pattern observed in the main paper with $W^{(3)}$ and $W^{(4)}$ achieving a better performance that the other weighting schemes. The MF-VAR on EA-19 aggregated data (Panel A) performs better than all specifications using country-level data (Panel B) in the Full sample, while we observe the opposite in the Pre-COVID period.
All in all, the MF-VAR appears to provide reliable nowcasting performance, with gains that are generally robust across both the pre-COVID and full samples. While its effectiveness is confirmed when compared to the forecast combination results in Table 4 of the main paper, the linear combination panel MIDAS models generally outperform the MF-VAR across most horizons. The only exception is the 2-month-ahead nowcast in the full sample, where the MF-VAR yields a more accurate forecast. These findings highlight the MF-VAR’s competitiveness, particularly in specific forecasting windows, while also underscoring the advantages of combining forecasts in a panel MIDAS framework.
This section further investigates the nowcasting performance discussed in the main paper. It begins by exploring which variables drive the nowcast, and then evaluates the model’s performance at the country level to assess the added value of panel models without aggregating to the Euro area level.
\paragraph{Variable importance:} Figure (ref) displays the sparsity pattern resulting from the variable selection performed by the Pooled panel model in the out-of-sample exercise. A colored tile indicates that the variable is selected by the sparse group penalty; a blank tile indicates otherwise. Overall, the selection pattern appears fairly persistent over time.
Several noteworthy patterns emerge. First, news-based variables are frequently selected, whereas other high-frequency financial indicators—such as exchange rates and stock indices—are never chosen. Second, survey data generally provide useful signals. Among these, consumer-related surveys tend to be, on average, more informative than surveys on other topics (e.g., retail, construction, industry). Third, inflation-related variables (both news-based and HICP) are often selected. In contrast, energy- and fuel-related variables (e.g., Import Prices: Mining, Mfg & Energy, gasoline, gas) are rarely picked, possibly because their information is already captured by HICP dynamics. Finally, it is interesting to notice that labor-related variables are seldom selected (e.g., news-based employment indicators, unemployment rate).
Overall, these results are consistent with the main findings in the literature. First, we confirm the importance of soft data (e.g., surveys) in nowcasting Euro area economic developments, in line with cascaldigarcia2021, who highlight that their timeliness can offset the long delays in the publication of hard data. Second, we emphasize the relevance of news-based variables in enhancing the predictive performance of nowcasting models tailored to the European context ashwin2021nowcasting, barbaglia2022forecastingEA. Finally, our findings corroborate the relatively greater importance of inflation-related variables compared to unemployment indicators, consistent with boss2025nowcasting.
\paragraph{Country-level performance:} In Table (ref), we summarize the nowcasting performance at the country level. The table reports the RMSEs of the Pooled panel model relative to those obtained from country-specific regressions. This comparison allows us to directly evaluate the contribution of the panel structure to nowcasting accuracy, independently of the weighting schemes used in the main analysis.\footnote{To conserve space, we present results only for the Pooled panel model. Comparable findings are observed for the HetAR panel model. Additional results are available upon request.}
Looking at the full sample (columns 2 to 4 of Table (ref)), we observe that the panel model generally outperforms country-specific regressions. In over three-quarters of the cases, the relative RMSE falls below one, indicating superior nowcasting accuracy. On average, the panel model achieves an improvement of more than 10% in forecasting performance. Notable gains are observed in large and medium-sized countries such as France, Italy, Spain, and Portugal. However, the outlook changes when focusing on the pre-COVID sample (columns 5 to 7 of Table (ref)). In this subsample, the panel model performs better in only about 40% of the cases, with particularly weaker results in countries like France and Spain.
While this additional analysis provides valuable insights into the performance of panel models at the individual country level, it does not alter the main conclusions of the paper. On average, panel models deliver more accurate nowcasts when evaluated over the full sample, compared to the pre-COVID subsample. This finding holds both at the country level and when aggregating to the Euro area, suggesting that the influence of the weighting schemes used in the main analysis is only marginal.
To further investigate the nowcasting performance across weighting schemes, we turn to Figure (ref) where we compare the weights obtained by $W^{(4)}$ (i.e., projections on GDP) and by $W^{(3)}$ (i.e., proportion of GDP level). We take the difference between $W^{(4)}$ and $W^{(3)}$ weights: positive values indicate a larger weight given by $W^{(4)}$ with respect to $W^{(3)}$. Note that by construction the $W^{(3)}$ weights are a direct measure of the size of each national economy within the Euro area. Compared to $W^{(3)}$, we observe that $W^{(4)}$ assigns smaller weights to Germany and to a lesser extent the Netherlands, while it inflates the relative importance of some small- and medium-sized countries, namely Austria, Belgium and Luxembourg. Although the size of these economies within the Euro area is relatively small, information about economic developments in those countries play a relevant role in attaining more accurate nowcasts.