Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
55,137 characters · 16 sections · 77 citation commands
Forecasting Macroeconomic Tail Risk in Real Time:Do Textual Data Add Value?
\thispagestyle{empty}
We examine the incremental value of news-based data relative to the FRED-MD economic indicators for quantile predictions of employment, output, inflation and consumer sentiment in a high-dimensional setting. Our results suggest that news data contain valuable information that is not captured by a large set of economic indicators. We provide empirical evidence that this information can be exploited to improve tail risk predictions. The added value is largest when media coverage and sentiment are combined to compute text-based predictors. Methods that capture quantile-specific non-linearities produce overall superior forecasts relative to methods that feature linear predictive relationships. The results are robust along different modeling choices.
{KEYWORDS:} Quantile Regression, Textual Data, Topic Models, Quantile Regression Forests, Gaussian Processes \\
{JEL:} C53, C55, E27, E37
\doublespacing
Periods of economic stress such as the Global Financial Crisis or the COVID-19 pandemic have highlighted the importance of tail risk predictions. Quantile predictions of macroeconomic time series provide a more nuanced picture than point forecasts, as the former allow the predictive relationship between the target variable and the covariates to vary across quantiles. Policymakers and central banks are particularly interested in the tails, which are associated with high economic risk. For this reason, the literature on macroeconomic forecasting has paid increasing attention to quantile predictions (see, e.g., Manzan2015,korobilis2017,Adrian2019,Carriero2020,Adams2021,clark2022,Pruser2023).
Another recent development in macroeconomic forecasting is the use of textual data, which provide timely information that may embed complementary signals to (hard) economic indicators (see e.g., larsen2019, Ellingsen2022, Kalamara2022, shapiro2022, vandijk2023, bybee2021). By extracting information from textual data, researchers seek to quantify the predictive relationship between narratives and economic outcomes. As argued by Shiller2017, narratives may even have the power to causally shape economic outcomes. Therefore, quantifying changing narratives based on textual data seems to be a worthwhile endeavour.
Our study takes a data-driven approach to assess whether textual data, relative to a large set of economic indicators, add value to macroeconomic quantile predictions in high-dimensional settings. We focus on monthly tail risk forecasts (nowcasts and one month ahead) of employment, industrial production, inflation and consumer sentiment between 1999:08 and 2021:12. To distinguish between the benefits of text-based predictors for tail predictions compared to the center of the distribution, we look at a broad range of quantiles (5%, 10%, 25%, 50%, 75%, 90%, 95%). As data revisions are an important issue in out-of-sample (OOS) macroeconomic forecasting, we use real-time vintages of economic predictors contained in the prominent McCracken2016 FRED-MD database.
For the computation of text-based predictors, we opt for an unsupervised machine learning approach that is agnostic about the dominant themes in a corpus of documents, namely the Correlated Topic Model (CTM) by blei2007. We use the CTM to map nearly 800,000 newspaper articles from The New York Times and The Washington Post into a set of numerical text-based predictors. The most frequent words within each topic (word cluster) characterize the topic, and the proportion indicates the degree of media coverage of a particular topic at a particular point in time. However, the topic proportions do not indicate whether the tone of the topic is positive or negative. We therefore also use tone-adjusted (sentiment) topic proportions. To investigate whether textual data can be used to improve tail risk forecasts, we use three different sets of predictors: (i) macroeconomic predictors only, (ii) macroeconomic predictors plus unadjusted text-based predictors, and (iii) economic predictors plus unadjusted text-based predictors plus text-based predictors with tone adjustment.\footnote{Text-based predictors are a viable alternative to survey forecasts and other “soft” data for quantile predictions. While survey forecasts were found to be valuable in improving point forecasts of macroeconomic time series, they were not found to be helpful in producing more accurate estimates of variance. The reason is that professional forecasters tend to be overconfident, hence providing too narrow density forecasts; see, e.g., galvao2021,banbura2021. Given our focus on quantiles and, in particular, on the tails of the distribution, textual data appear more promising. For other “soft” data such as Google Trends, the available data history is considerably shorter than that of our text-based predictors.}
In terms of forecasting models, two key insights guided our choice. First, mechanisms to prevent overfitting are necessary for OOS forecasting in high-dimensional settings such as ours. We therefore use three Bayesian quantile regressions (QRs) with different shrinkage priors as carriero2022 have shown the importance of using shrinkage for QR in empirical macroeconomics. Second, recent studies point to the importance of capturing non-linear predictive relationships. For example, Goulet2022 consider capturing non-linearities as the “true game changer” of machine learning methods for macroeconomic forecasting. This is corroborated by Medeiros2021 and clark2022, who find strong empirical support for tree-based methods. To capture non-linear relationships between the quantile-specific target variables and the covariates, we use non-parametric Bayesian Gaussian Process Regressions Williams2006 and QR Forests Meinshausen2006. Entertaining both forecasting methods that feature linear and non-linear predictive relationships enables us to evaluate the empirical differences between the model classes.
While most studies analyze the benefits of text-based predictors for macroeconomic point forecasts, the role of text-based predictors for tail forecasts is largely unexplored. Few exceptions are Barbaglia2022, Sharpe2023 and Filippou2023 who use sentiment-based approaches to determine the relevance of textual data (newspaper articles and FOMC speeches) for macroeconomic quantile forecasts. Our approach differs from these studies in three important ways. First, we investigate the added value (if any) of textual data in a setting where the forecaster is working in a data-rich environment using the FRED dataset, which contains a large amount of information about the economy. We ask whether, in this setting, a large set of text-based predictors can provide additional value for tail risk predictions. Second, we consider not only forecasting models that can deal with high-dimensional predictor sets, but also two methods that can deal with non-linear predictive relationships. Third, we combine unsupervised topic models to extract the content of newspaper articles via estimated topic proportions and further tone-adjust these topic proportions by multiplying them with sentiment scores based on a pre-defined dictionary. By using both unadjusted and tone-adjusted text-based predictors, we can disentangle the added value of textual predictors that embed only information about media coverage from textual data that also includes information about the tone of newspaper articles. Our approach to learning from the news data is agnostic about which themes are relevant for predicting quantiles of macroeconomic times series. This contrasts with approaches such as that of Barbaglia2022, who only use text in an article that is semantically dependent on a term of interest (e.g. inflation).
Our key findings can be summarized as follows. First, textual data contain valuable predictive information that is not embedded in the large set of FRED-MD economic indicators. Our results suggest that text-based predictors can improve tail risk forecasts of macroeconomic variables, but much less for the center of their probability distributions, consistent with the narrative that news signals are most helpful in extreme economic situations. Second, adding tone-adjusted text-based predictors results in more accurate forecast results than using only unadjusted predictors. Third, methods that can capture non-linear predictive relationships produce better predictions than those that cannot. We have conducted a number of robustness checks that confirm our main findings. Among other things, we find that our results are robust to longer forecast horizons and a range of different choices to produce text-based predictors. \color{black}
The paper proceeds as follows. Section 2 lays out our methodology. Section 3 outlines our forecasting setup, introduces our predictors, and presents our forecast results. Section 4 concludes. Additional material is relegated to the appendix.\color{black}
In settings with many regressors, Bayesian shrinkage alleviates overfitting and thus noisy forecasts. Recently, carriero2022 showed the importance of using shrinkage for QR in empirical macroeconomics. For Bayesian estimation, we use the mixture representation established by yu2001. For a given variable $y$ of interest that is to be predicted for quantile $\tau $ at horizon $h$, the Bayesian QR can be stated as
where $\left \{ x_{t}\right \} _{t=1}^{T}$ is a column vector of predictors, $\beta _{\tau }$ denotes a vector of quantile-specific regression coefficients and $\varepsilon _{\tau ,t+h}$ follows an asymmetric Laplace distribution, see yu2001. The dimension of $x_{t}$ depends on the setting, which is one of the following three possibilities: (i) FRED predictors only, (ii) FRED predictors plus unadjusted text-based predictors, and (iii) FRED predictors plus unadjusted text-based predictors plus text-based predictors with tone adjustment.
Posterior inference requires specifying a likelihood and eliciting priors for the coefficients. As Markov Chain Monte Carlo estimation is slow in high dimensions, we use fast variational Bayes approximations for posterior inference. In the interest of brevity we do not give full details of the estimation, but note that, by introducing auxiliary latent variables, the likelihood is conditionally Gaussian and the errors are conditionally heteroscedastic, see yu2001 and Pruser2023. The shrinkage priors we consider fall into the class of global-local shrinkage priors, where the prior variance comprises one global term pertaining to all coefficients and another coefficient-specific term. Our chosen shrinkage priors can be written in the general form
where $J$ denotes the number of predictors, $\lambda _{\tau }$ denotes a quantile-specific global shrinkage parameter and $\psi _{\tau j}$ are quantile-specific local scaling parameters that control the coefficient-specific shrinkage intensities. Different shrinkage priors are generated by choosing different mixing densities via the functions $u$ and $\pi $. We consider three shrinkage priors that are popular choices in macroeconomic forecasting; see, e.g., huber2019, cross2020, and pruser2023DB:
Gaussian Process Regression Williams2006 is a non-parametric method that was recently used for inflation forecasting by clark2022f. It elicits a process prior on the function $g_{\tau }\left({x}_{t}\right):$
where we set the mean function $\mu _{\tau }\left( {x}_{t}\right) $ to zero. The kernel function $\mathcal{K}\left( {x}_{t},{x}_{ \mathfrak{t}}^{^{\prime }}\right) $ describes the relationship between $ {x}_{t}$ and ${x}_{\mathfrak{t}}$, for $t$, $\mathfrak{t=} 1,\ldots,T$.
As ${x}_{t}$ is observed at discrete points in time, ${g} _{\tau }=\left( g_{\tau }\left( {x}_{1}\right),\ldots,g_{\tau }\left( {x}_{T}\right) \right) ^{\prime }$:
where ${K}\left( {w}\right) $ refers to a $T\times T$-dimensional matrix with $\left( t,\mathfrak{t}\right) $-th element $ \mathcal{K}\left( {x}_{t},{x}_{\mathfrak{t}}\right) $.
The type of kernel determines the estimated function. We choose a squared exponential kernel
where we follow chaudhuri2017 for setting the hyperparameters ${w=} \left( w_{1},w_{2}\right) ^{^{\prime }}$, which govern the smoothness of the function.
Above we have outlined the function-space view of the GP regression. In the alternative weight-space view, which is convenient for estimation, the GP regression can be expressed as
where ${y}$ denotes the stacked dependent variables, ${Z}$ represents the lower Cholesky factor of ${K}$, and $\mathbf{ \gamma }_{\tau }\sim \mathcal{N}\left( {0}_{T},{I}_{T}\right) $.
Our last forecasting method is a frequentist non-parametric method, namely QR Forests, an extension of Random Forests breiman2001 for conditional point estimation to conditional quantile estimation based on an ensemble of trees Meinshausen2006. Random Forests and QR Forests capture non-linear predictive relationships and, especially due to this feature, have been found to perform well in macroeconomic forecasting; see, e.g., Medeiros2021,clark2022).
Random Forests grow a large number of trees by using $t$ observations:
where $Y$ is the variable of interest and $X$ is a (possibly high-dimensional) predictor variable. For ease of notation we drop time subscripts. For each tree and node, Random Forests uses a random subset of predictors to split on.\footnote{In our empirical work, for a $p$-dimensional predictive variable, at each node we use the default choice of $\sqrt{p}$ randomly selected predictors.} The intuition of this random selection is to de-correlate the trees and thus to decrease the variance of the forecasts. In Random Forests, the conditional mean prediction of $Y$, given $X=x$, is generated as the weighted sum over all observations:
where the weights $w_{i}\left( x\right) $ are computed over the collection of trees. In each tree, the conditional mean prediction is the simple average of all observations that fall into the same leaf when dropping down $x$; the remaining observations are neglected.
Meinshausen2006 extended the Random Forests to QR Forests. The conditional distribution function of $Y$, given $X=x$, is
Instead of approximating the conditional mean $\mathbbm{E}\left( Y|X=x\right)$ in case of Random Forests, for QR Forests, $\mathbbm{E}\left( \mathbbm{1}_{\left \{ Y\leq y\right \} }|X=x\right) $ is approximated by the weighted mean over the observations $\mathbbm{1}_{\left \{ Y_{i}\leq y\right \}}$,
where $w_{i}\left( x\right) $ are the same weights as for Random Forests. Based on $\widehat{F}\left( y|X=x\right) $, we can estimate the desired conditional $\alpha $-quantile $Q_{\alpha }\left( x\right) $ as
An important difference between QR Forests and Random Forests is that, for each node in each tree, Random Forests keep only the mean of the observations that fall in that node, whereas QR Forests keep the values of all the observations in that node to compute the conditional distribution.
We use a probabilistic topic model to create text-based predictors. Topic models are data-driven and use a probabilistic approach to identify themes within a large set of written documents. Unlike dictionary-based and Boolean approaches, topic models are initially context-agnostic, finding word clusters (topics) solely based on the co-occurrence of words.
The most prominent topic model is latent Dirichlet allocation (LDA) by blei2003. It posits that documents are generated by a stochastic process where each text is a mixture of latent topics and each topic is a probability distribution over the same vocabulary, but with different probabilities for each word. Despite its prominence and advantages, Latent Dirichlet Allocation (LDA) cannot account for the fact that certain topics tend to co-occur together within documents (e.g., inflation and commodity prices). We therefore use the more sophisticated CTM by blei2007, which was developed to address this limitation.\footnote{Economic studies have predominantly relied on LDA by blei2003 to extract news-based predictors. These predictors have then been used, for instance, to investigate the value of news data for modeling macroeconomic dynamics larsen2019, bybee2021, to construct a daily business cycle index thorsrud2020, to predict US macroeconomic variables Ellingsen2022, and to nowcast US GDP Babii2021 as well as Chinese GDP and inflation expectations zheng2023. The CTM and further extensions have been used to produce topic proportions which have then been used, for example, to analyze news coverage of China roberts2016, to investigate the impact of presidential tax speeches on economic activity dybowski2018, to forecast the equity premium adaemmer2020, and to analyze European Central Bank communication dybowski2020, bohl2023. \color{blue} \color{black}}
Figure (ref) shows a graphical model of the generative process assumed by the CTM, where edges denote dependencies, nodes are random variables, and plates are replicated variables. The only observable variables of the model are the words ($w$).
The key difference between LDA and the CTM is the assumption for the topic proportions, $\theta_d$. The CTM assumes that topic proportions follow a logistic normal distribution---which has a mean vector ($\mu$) and a covariance matrix ($\Sigma$)---as opposed to LDA, which assumes that topic proportions come from a Dirichlet distribution. Consistent with the more realistic assumptions of correlating topics, the CTM yields higher values for statistics that measure the predictive performance regarding unseen documents blei2007. The generative process of the CTM can be written as follows (local mathematical notation):
We use the partially collapsed variational EM algorithm by roberts2016 to estimate the CTM. The approach is implemented in the R-package stm by roberts2019. We are particularly interested in $\theta_d$, namely the topic proportions of each document, whose aggregates serve as our text-based predictors; see Section (ref).
We generate monthly quantile forecasts (nowcasts and one month ahead) for employment, inflation (Total CPI), industrial production, and consumer sentiment. We choose employment, inflation and industrial production as our target variables because they are key macroeconomic variables used in several forecasting studies; see, e.g. Kalamara2022, Barbaglia2022. Consumer sentiment is used to assess whether news data are important in forming household expectations larsen2021. Depending on the setting, we incorporate FRED predictors, text-based predictors, or both together. We detail our predictors in Section (ref). In addition, all model specifications include 12 lags of the respective (transformed) variable of interest.
For our target variables and the FRED predictors, we use vintage data from the McCracken2016 database. We use 98 indicators from the FRED-MD database and transform the data to achieve stationarity McCracken2016.\footnote{The transformation of the variables can be found in Appendix (ref). We use all indicators from the FRED-MD database that are available both at the beginning and at the end of our out-of-sample period.} Our estimation sample starts in 1980:06, when the text-based predictors start. We run recursive estimations based on an expanding window. Our evaluation period runs from 1999:08 to 2021:12.
We compute our predictions at the end of a given month. Due to publication lags, we include macroeconomic predictors from the previous month, while we include text-based and financial predictors from said month.\footnote{See Appendix (ref) for which variables are classified as financial predictors. Note that few indicators are lagged by more than one month, and we adjust for this.} For example, if we are at the end of December and produce a nowcast for December, we use the macroeconomic predictors from November released in December, and the financial and text-based predictors from December. Similarly, for one-month-ahead forecasts, if we are at the end of December and produce a prediction for January, we use the macroeconomic predictors from November released in December and the financial and text-based predictors from December.
Our five forecasting models are three versions of Bayesian QRs with different shrinkage priors (Horseshoe, Ridge and Lasso), QR Forests and Gaussian Process Regressions. \color{black}
We evaluate our forecasting models with the quantile score (QS), which is computed as
where $y_{t+h}$ is the actual outcome of the variable of interest in $t+h$, $Q_{\tau ,t+h}$ denotes the forecast of quantile $\tau$ for $t+h$. The indicator function $\mathbbm{1}_{\left \{ y_{t+h}\leq Q_{\tau ,t+h}\right \}}$ takes on a value of $1$ if the outcome is not higher than the quantile forecast, and $0$ otherwise.
We select all macroeconomic predictors from the McCracken2016 FRED-MD database that are available at both the end and the beginning of the sample. The predictors can be classified into eight categories: (i) output and income, (ii) labor market, (iii) housing, (iv) consumption, orders, and inventories, (v) money and credit, (vi) interest and exchange rates, (vii) prices, and (viii) stock market. For the classification of the variables, see \url{https://www.ssc.wisc.edu/ bhansen/econometrics/FRED-MD_description.pdf}.
We used the legal database LexisNexis to download 793,013 economically related newspaper articles from The New York Times and The Washington Post between 1980:06 and 2021:12. We then conducted Part-of-Speech-Tagging benoit2020 to remove anything but nouns from the articles. Our reason for this choice is that topic models aim to summarize the content of documents, which is predominantly described by nouns, in contrast to sentiment analysis which aims to describe the documents' tone. In addition, martin2015 found highest values of semantic coherence when using a nouns-only approach, a metric that strongly correlates with human judgement.
We tokenized the documents into single words, removed punctuation, numbers, symbols, stopwords, etc., and constructed a document-term matrix (dtm). A dtm counts how often a certain word (column) occurs within a certain document (row). We then computed term-frequency inverse-document-frequency (tf-idf) values for each word to extract the 10,000 most relevant terms until 1999:08. The final dtm served as the input for the CTM. In addition, the number of topics $K$ has to be given as an input by the researcher. larsen2019, thorsrud2020, Ellingsen2022 use LDA to compute topic proportions for different inference and prediction tasks. In each study, they choose 80 topics. We follow them and use $K=80$ as our baseline model.\footnote{In Appendix (ref), we provide robustness checks for the number of topics, which are based on topic interpretability roberts2019,adaemmer2020 and an automated search approach by Mimno2014. Overall, our main findings are robust to alternative choices of the number of topics.}
We estimated for each document on each day 80 topic proportions based on the newspaper articles until 1999:08. After then, we computed the topic proportions OOS to exclude any look-ahead bias, similar to Ellingsen2022. Finally, for a given month, we computed the simple average over all documents' estimated topic proportions. Note that $\theta_{d,k}$ denotes the (estimated) topic proportion of the $k$-th topic for the $d$-th document. For a given month $m$ and for each topic $k$, we compute the average topic proportion as follows:
where $D_{m}$ denotes the number of documents that fall into the $m$-th month. Applying ((ref)) to all K (estimated) topic proportions produces a $K\times1$ vector of (estimated) topic proportions for the $m$-th month, that is $\theta^{(m)}=(\theta_{1}^{(m)},\ldots,\theta_{K}^{(m)})'$.
Topic proportions only reveal information about the content of the newspaper articles, not about the tone that accompanies them. However, the tone may contain additional predictive information. We therefore also consider tone-adjusted (sentiment) topic proportions as predictors. To do so, we rely on the dictionary compiled by Barbaglia2022, where each word contains a sentiment value between -1 and 1. The more positive (negative) the sentiment score, the more positive (negative) is the tone of the respective word. Note that we use all the words in the sentiment dictionary to calculate a document's sentiment score, unlike the nouns-only approach for calculating the topic proportions. The sentiment score (Sent) for the $d$-th document is computed as follows:
where $Score_i$ is the sentiment score for the $i$-th word in the respective document. The number of words in each article is denoted by $N$. The sum of a document's sentiment score is normalized by the number of words with an associated sentiment score in the respective text. Figure (ref) shows our monthly sentiment time series, computed as $Sent^{(m)} = \frac{1}{D_m}\sum \limits_{d=1}^{D_m}Sent_{d}$. Gray-shaded areas indicate NBER-dated recessions. The time series captures well economic periods of booms and busts.
We use the sentiment score of each document to adjust each topic proportion:\footnote{Ellingsen2022 also tone-adjust their topic proportions, but in a different manner. In addition, they use the Harvard IV-4 Psychological Dictionary, which only distinguishes between positive and negative words and was not designed for economic and financial applications.}
We then compute the average over all documents of a given month m:
The first row of Figure (ref) shows the trajectories of three selected topics without any tone adjustment. A high value for a certain topic at a given point in time indicates high media coverage of said topic at that point in time.\footnote{ The labels are based on the top five words of each topic. We emphasize that topic labeling is notoriously subjective, which is why we show the top five words of all topics in Appendix (ref).} The figures show that our topics capture important political and economic events such as several debt crises in Mexico, Asia and Europe (Topic 9), the beginning of the Gulf War in 1991 (Topic 27) and the Great Recession which emanated in the housing market (Topic 46).
The second row of Figure (ref) shows the tone-adjusted topic proportions, which reflect the severity of several events. For example, the tone in articles about debt crises (9) was overall negative. The series of SentTopic 28 drops sharply at the beginning of the Gulf War in 1991 and SentTopic 46, which mainly covers issues related to the housing market, clearly changes in sentiment from positive to negative as the collapse of the US housing market became apparent during the Great Financial Crisis. We use both the monthly averages of unadjusted topic proportions ((ref)) and with tone adjustment ((ref)) as text-based predictors in our empirical analyses.
We first focus on the results for the linear Bayesian QR models with different shrinkage priors, namely Horseshoe, Lasso and Ridge. Our primary interest here is to assess the added value (if any) of the text-based predictors. We then assess how methods that use the same sets of predictors but allow for non-linear predictive relationships (QR Forests and Gaussian processes) perform against the overall most successful linear model.
For Horseshoe, Lasso and Ridge, Figure (ref) summarizes the results for nowcasts and one month ahead forecasts, respectively. Each column denotes a target variable and each row a (linear) model. Forecast accuracy is assessed by quantile scores across different quantiles $\tau$: $\tau$ = 5%, 10%, 25%, 50%, 75%, 90%, 95%. A quantile AR(1) model serves as our benchmark.\footnote{More precisely, we compute a (frequentist) quantile AR(1) model. For example, in case of a nowcast at the end of December (for December), the target variable's latest available value from November is used as the single predictor in the quantile AR(1): $y_{December}=\beta_{0,\tau}+\beta_{1,\tau}y_{November}+\varepsilon_{December}$.} Relative quantile scores below (above) one indicate more (less) precise forecasts compared to the benchmark model.
The combination of macroeconomic predictors and text-based predictors leads to accuracy gains in many cases, especially in the tails, compared to the setting with macroeconomic predictors alone. The addition of tone-adjusted topic proportions leads to even lower overall quantile scores compared to economic predictors and unadjusted topic proportions (FRED & Topics). In those cases where the addition of text-based predictors does not improve forecast performance, it does not have a materially adverse effect either. The incremental value from adding text-based predictors is pronounced in the tails. Improved tail forecasting with textual data suggests a scenario where text-based predictors are particularly helpful in extreme economic environments, where timely news data can reflect the overall economic picture and be useful predictors of macroeconomic variables. In times such as the Great Recession or the COVID-19 pandemic, uncertainty and political change may have been better captured than by hard economic indicators.
Within the class of the three linear models, the Horseshoe shrinkage prior underperforms the Lasso and Ridge in the tails. Ridge, as a purely global shrinkage prior, outperforms Horseshoe and Lasso. Thus, the richer shrinkage patterns offered by Horseshoe does not pay off in our analysis. These results are consistent with a dense representation of the prediction problem, where many weak predictors are relevant, rather than a sparse structure, where a few important predictors can be selected and the others are discarded.
Ridge produces more precise nowcasts for the left tails of employment and production and for both tails of inflation and consumer sentiment relative to a quantile AR(1) model when using FRED & Topics & SentTopics. Ridge is also successful in one month ahead forecasting left-tail outcomes of production and the right-tail outcomes of inflation and employment when using FRED & Topics & SentTopics. For consumer sentiment (nowcasts and one month ahead forecasts), quantile scores markedly improve in the left tail when (in particular tone-adjusted) text-based predictors are added. This result is interesting from an information processing aspect, supporting the notion that news data play an important role for households in forming their expectations larsen2021.
Next, we look at the performance of non-linear methods, namely Gaussian Processes and QR Forests. As we are interested in the added value of non-linear models over linear ones, we use Ridge with the full predictor set (FRED & Topics & SentTopics) as our benchmark. We do this because Ridge was the overall most successful linear model, particularly in the tails (see Figure (ref)). Figure (ref) shows the relative predictive performance of both non-linear methods with the various predictor sets. Relative quantile scores below (above) one indicate more (less) precise forecasts compared to Ridge with all predictors (FRED & Topics & SentTopics).
As indicated by the predominantly hump-shaped relative quantile scores, with many values below one in Figure (ref), both non-linear methods were overall better at tail forecasting than Ridge. This is particularly true for the right-tail outcomes of employment and both tails of production, especially when using the FRED & Topics & SentTopics predictor set. Hence, the combination of non-linear methods and a mix of economic and text-based predictors seems to be a promising approach for macroeconomic tail forecasting. That said, the marginal difference in quantile scores from adding text-based predictors is smaller overall for non-linear methods than for linear ones. One explanation for this finding may be that text-based predictors may partially compensate for missing non-linear predictive relationships in the linear models.
\color{black}
We conducted a number of robustness checks and further analyses to explore how alternative settings affect our findings in this paper.
Generating text-based predictors. The computed topic proportions obtained from the CTM are influenced by a number of choices. Here we explore the impact of alternative choices and examine their potential impact on forecast accuracy. First, we computed $K=68$ topics, which is based on the random search approach by Mimno2014. More specifically, we ran their algorithm 100 times and chose the average estimated number of topics across all runs, which was 68. Second, following adaemmer2020, who use the CTM and a similar corpus, we computed $K=100$ instead of K=80 topics. The authors show empirically that $K=100$ strikes a good balance between prediction of unseen documents and topic interpretation. Third, the CTM is nested within the more sophisticated structural topic model (STM) by roberts2016, which further allows to include covariates. We computed an STM to account for the two different newspapers. Fourth, we estimated our baseline model with 15,000 unique words instead of 10,000 unique words based on the highest tf-idf scores. Fifth, for our reported results, we used a training sample for the CTM from 1980:06 to 1999:08 (see also Section (ref)), and computed OOS topic proportions for the evaluation period from 1999:08 onwards. This approach avoids any look-ahead bias, but implicitly assumes that the vocabulary remained constant over time. In reality, however, this was certainly not the case. For example, words like “Brexit” or “COVID” were not part of the vocabulary in the estimation sample of the topic proportions. To investigate whether---and if so how much---a fixed vocabulary lowers forecast accuracy, we estimated the CTM using the vocabulary of the whole sample (i.e., with a look-ahead bias). Interestingly, we did not get lower quantile scores overall. These results are consistent with those of vandijk2023, who accounted for shifts in the vocabulary used to nowcast GDP, but did not find gains in forecast accuracy either. Sixth, we computed separate topic proportions for the last week of the month and for each topic. However, we found no added value from this alternative specification based on more recent news. Overall, the robustness checks reveal that the empirical results are not sensitive to the exact specification used to generate the topic proportions. Figure (ref) and Figure (ref) in Appendix (ref) summarize the results for the different choices. The predictor set used in each setting is FRED & Topics & SentTopics. For comparison, we also added the results for our baseline model ($K=80$). Overall, the results are robust across different choices. We have taken a conservative approach in this paper by avoiding fine-tuning the parameters that control the generation of the text-based data. More sophisticated procedures for selecting these parameters may further increase the gains in macroeconomic quantile prediction, but with an increased risk of overfitting.
Alternative forecast horizons. Figure (ref) and (ref) in Appendix (ref) summarize the results for the 6-month and the 12-month forecast horizon. In summary, the results confirm our findings for the nowcast and one month ahead forecasts. The marginal improvement from adding (tone-adjusted) text-based predictors seems even higher at these horizons, especially for the non-linear models and for tail predictions associated with economic downside risk (that is, low employment, high inflation and low production).
Quantile scores over time. Figure (ref) and (ref) in Appendix (ref) show the cumulative quantile scores at the 10% quantile relative to the quantile AR(1) benchmark for the nowcasts and one month ahead forecasts, respectively; Figure (ref) and (ref) in Appendix (ref) do the same for the 90% quantile. Note that the presentation of results starts at 2005:01 rather than 1999:08, as earlier starting dates would distort the scaling due to the few observations at the beginning of the OOS period. Overall, across the different target variables and methods, gains relative to the quantile AR(1) benchmark are most pronounced during the Great Recession and in the COVID-19 pandemic. Thus, our results are consistent with the typical finding in the economic forecasting literature that the performance of forecasting models is unstable over time; see, e.g. Rossi2013. To assess the marginal gains of the text-based predictors, Figure (ref) and (ref) in Appendix (ref) show the cumulative quantile scores at the 10% quantile relative to the FRED-only dataset as the benchmark for the nowcasts and one month ahead forecasts, respectively; Figures (ref) and (ref) in Appendix (ref) do the same for the 90% quantile. Across the different target variables and methods, we observe heterogeneous patterns and, overall, find that text-based predictors (both unadjusted and tone-adjusted) add value relative to the FRED-only dataset at different points in time rather than being restricted to short episodes.
As we include forecasting methods that feature non-linear predictive relationships, it is not obvious how to determine the marginal effect of a predictor on the variable of interest. Furthermore, even though we use heterogeneous forecasting methods, we want to ensure comparability of measures of predictor importance across methods. To accomplish this task, and to shed light on the predictors with the highest impact, for each forecasting model, we rely on a linear approximation to the prediction distribution for each prediction model, as proposed by woody2021 and used by clark2022f and Pruser2023. Concretely, we approximate the quantile predictions $ Q_{\tau ,t+h}$ with a linear regression model that uses regularization. For each quantile $\tau $, the following Lasso-type optimization problem is solved:
where $t_{0}$ denotes the beginning of the hold-out period, ${\beta } _{\tau }=\left( \beta _{\tau ,1},\ldots,\beta _{\tau ,J}\right) ^{\prime }$ denotes the vector of coefficients, $J$ denotes the number of potential predictors, and $\lambda \geq 0$ is a penalty parameter that controls the shrinkage intensity and is chosen via cross-validation. For the predictor importance analysis we use our most comprehensive set of predictor variables, including both FRED and all text-based predictors. We focus on the $10\%$ and 90% quantiles because our previous analysis has shown the most pronounced patterns in the tails.
Figure (ref) and (ref) show the number of nonzero coefficients for the nowcasts and one month ahead forecasts of the 10% and the 90% quantile, respectively. Across different target variables and quantiles, the structure of included FRED and text-based predictors is balanced. Within text-based predictors, roughly the same number of tone-adjusted and unadjusted topic proportions are included. In some cases, however, only text-based predictors survive as in the case for nowcasts of employment with Ridge or QR Forests.
Figure (ref) and (ref) in Appendix (ref) report the five most influential predictors for nowcasts and one month ahead forecasts at the $10\%$-quantile, respectively. Figure (ref) and (ref) in Appendix (ref) do the same for the $90\%$-quantile. Importance is measured by the absolute values of the coefficients associated with the (standardized) predictors.
Overall, the coefficients that are not set to zero, are of different types: lagged values of the respective target variable, FRED economic indicators, unadjusted topic proportions and tone-adjusted topic proportions. Regarding the text-based predictors, the chosen topic proportions make narrative sense for the respective target variable. For instance, both Topic 46 and SentTopic 46 were chosen for inflation nowcasts and one month ahead forecasts. The five most probable words in that topic are related to the housing market (home, house, property, estate and owner). Developments in the housing market can obviously be related to inflation dynamics. Topic 31, which is related to household finance, was also chosen for inflation nowcasting by QR Forests. For nowcasts and one month ahead forecasts of employment, Topic 47 and SentTopic 47 were chosen by Gaussian Processes which makes sense, given the topic's job market related words job, \textit{work}, \textit{force}, \textit{unemployment} and \textit{employment}. In addition, Topic 55 and SentTopics 55 were frequently selected for nowcasts and one month ahead forecasts of employment and production. The topic is related to nutrition and health issues, given the topic's most prominent words \textit{food}, \textit{disease}, \textit{meat}, \textit{animal} and \textit{product}.
\color{black}
We have analyzed the added value of text-based predictors for quantile predictions of macroeconomic time series. Our high-dimensional setup included forecasting methods with both linear and non-linear quantile-specific predictive relationships, and we considered different sets of predictors.
Although forecast performance varied across quantiles and target variables, altogether, combinations of FRED and text-based predictors produced the most accurate forecasts, especially in the tails. Tone-adjusted text-based predictors added further value as predictors compared to unadjusted ones alone. Overall, Gaussian Process Regressions and QR Forests prevailed in terms of tail forecast accuracy, suggesting that non-linear predictive relationships are a promising route to follow in tail forecasting.
In summary, our empirical results suggest that researchers using FRED data for tail risk forecasting should exploit the additional information contained in textual data. We find that the most precise tail risk predictions can be achieved when using non-linear models with economic indicators and (tone-adjusted) topic proportions. Our results can be of value to researchers, forecasters and policy makers alike.
\spacing{1.00}
\addcontentsline{toc}{section}{\refname}