Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
82,365 characters · 22 sections · 94 citation commands
Nowcasting Growth using Google Trends Data: A Bayesian Structural Time Series Model
\def\spacingset#1{ {#1}} \spacingset{1}
\def\spacingset#1{ {#1}} \spacingset{1}
\if00 \fi
\if10 {
} \fi
{\it Keywords:} Global-Local Priors, Non-Centred State Space, Shrinkage, Google Trends
\spacingset{1.45}
The primary object of nowcasting models is to produce `early' forecasts of target variables associated with long delays in data publication by exploiting the real time data publication schedule of the explanatory data set. While prediction is the primary goal here, the selected models can also sometimes provide structural interpretations ex post. Nowcasting is particularly relevant to central banks and other policymakers who are tasked with conducting forward looking policies on the basis of key economic variables such as GDP or inflation. Inflation data are, however, published with a lag of up to 7 weeks with respect to their reference period, and precise estimates of GDP can take years.\footnote{The exact lag in publications of GDP and inflation depends as well on which vintage of data the econometrician wishes to forecast. Since early vintages of aggregate quantities such as GDP can display substantial variation between vintages, this is not a trivial issue.} Since even monthly macroeconomic data arrive with considerable lag, it is now common to combine, next to traditional macroeconomic data, ever more information from Big Data sources such as internet search terms, satellite data, scanner data, etc. which have the advantage of being available in near real time bok2018macroeconomic. The recent Covid-19 pandemic has given further impetus to this trend, as faster indicators have proven especially useful in modelling the unprecedentedly sharp movements in the economy that traditional macroeconometric models fail to capture in a timely manner antolin2021advances,woloszko2020tracking.
In this paper we add to the burgeoning literature on using Google search data in the form of Google Trends (GT), which measure the relative search volume of certain search terms entered into the Google search engine, to nowcast aggregate economic time-series. In particular, we investigate the benefits of using monthly Google search information for nowcasting quarterly U.S. real GDP growth in real-time compared to traditional macro data and survey information. We contribute to this literature by being, to our best knowledge, the first paper to investigate the benefit of search information above and beyond macroeconomic data for the U.S. including performance during the Covid-19 pandemic period.
To deal with the specificities of the data, we propose robust nowcasting methods that are amenable to situations in which the policymaker needs to combine traditional and non-traditional data sources while providing tractable variable selection properties. For this purpose, we adapt current generation state space and regression priors to the widely popular Bayesian Structural Time Series model (BSTS) scott2014predicting. Results from our nowcasting application show that Google's search information improves nowcasts of GDP growth, particularly early on in the quarter before macroeconomic data are published. We show that our extensions allow for accuracy gains of up to 40% during certain nowcast periods in point as well as in density nowcasts compared to the original BSTS model of scott2014predicting while retaining its interpretability. These results are confirmed in a simulation study which checks robustness to a variety of data-generating processes.
In the following, we firstly discuss the state and regression priors as well as posteriors for our extended BSTS models and provide efficient sampling algorithms. In section 3, we elaborate further on the data used for nowcasting, including dealing with mixed frequency, the data publication calendar and the specificities of the Google Trends data set. In section 4, we present results based on our empirical application of nowcasting U.S. GDP growth, which is followed in section 5 by the results from our simulation study. Finally, Section 6 concludes with a discussion and avenues for future research.
The Bayesian Structural Time Series (BSTS) model, as proposed by scott2014predicting, provides a conceptually attractive model for nowcasting aggregate economic time-series with heterogeneous data sources, as it flexibly estimates latent time-trends, seasonality and deviations or `irregular' dynamics through variable selection using a high-dimensional shrinkage prior. Denote the target variable to be nowcast by $y_t = (y_1,\cdots,y_T)'$ and the K-dimensional explanatory data set as $x_t = (x_1',\cdots,x_T')'$ which for now are sampled at the common frequency, $t$. Then our model is as follows:
((ref)) is a linear state space model with Gaussian errors and states $\{\tau_t,\alpha_t,\delta_t\}_{t=1}^T$ which capture long-run trends and $S$ seasonal components $\delta_t$. The deviation from $\tau_t$ describes variation from a long-run trend which, when applied to the level of GDP can be interpreted as the output gap watson1986univariate,grant2017bayesian. $\alpha_t$ allows for a drift term in the trend which is often observed in stock variables such as in GDP, aggregate consumption and inflation grant2017bayesian,chan2017stochastic.
Variable selection on the possibly high-dimensional $K\times 1$ response vector $\beta$ in the BSTS model of scott2014predicting is done via a two component conjugate spike-and-slab prior. Estimation is standard george1993variable, and states $\tau_t$, $\alpha_t$ and $\delta_t$ are estimated jointly via the forward filtering backward sampling (FFBS) algorithm of durbin2002simple based on the Kalman filter. This implementation relies on Normal-Inverse Gamma (N-IG) priors for the states and state variances for conditional conjugacy. While the BSTS model is a natural model for many time-series applications, we bring 3 important methodological innovations which make it more robust to overfitting trend estimation and variable selection with heterogeneous high dimensional data.
In line with previous nowcasting studies, this paper focuses on nowcasting GDP growth rather than levels. However, two problems arise when applying model ((ref)) directly to growth variables. As growth variables are often approximately stationary, conceptually, the inclusion of $\alpha_t$ implies that GDP growth follows a boundless drift for which there is little structural justification or empirical evidence. A non-drifting stochastic trend, on the other hand, has been shown to markedly improve nowcasts of GDP growth as shown in antolin2017tracking, especially when the state variances are tightly controlled by priors such that the stochastic trend does not wander too wildly. This suggests that modeling time-variation is preferred over de-trending a priori.\footnote{Modelling an I(1) component in U.S. GDP growth is additionally consistent with Harvey's local-linear trend model harvey1985trends, the hodrick1997postwar filter and stock2012disentangling.} The underlying rationale for this improvement is the well known empirical finding of changes in long-run GDP growth kim1999has,mcconnell2000output,jurado2015measuring. Econometrically, the additional problem is that the IG priors with no prior mass on zero, as implemented for Bayesian linear state space methods, can bias posterior state variances away from zero, thereby potential leading to false support for state dynamics which can hurt forecast performance.
We extend model ((ref)) to flexibly let the data shut down state dynamics, and therefore broaden the applicability of model ((ref)), by adopting the non-centred parameterisation of the state space as suggested by fruhwirth2010stochastic. The non-centred parameterisation models state variances directly in the observation equation, which with normal priors, exerts much stronger shrinkage than IG priors.\footnote{Formally, a normal prior on the state standard deviation can be shown to imply a Gamma prior on the state variance.} This allows additionally for valid inference on testing for zero posterior variance via Savage-Dickey density ratios, as will be further discussed in section 5. Testing for zero posterior variance would be very challenging in a frequentist hypothesis testing approach because the null hypothesis of constant state in model ((ref)) lies on the boundary of the parameter space.
The non-centred model considered for the empirical application is equivalently written as:
and
with starting values $\tilde{\tau}_0 = \tilde{\alpha}_0 = 0$. Note that the seasonal component is left out for estimation due to the small sample length of the Google Trends data set and differing seasonal patterns between monthly and quarterly data.\footnote{For further discussion, please see section (ref)} To see that ((ref)) and ((ref)) is equivalent to ((ref)), let:
Hence, by setting $y_t = \tau_t + x_t'\beta + \epsilon_t$, it is clear that
which recovers ((ref)). Since $\sigma_{\tau,\alpha}$ are allowed to have support on the real line, they are not identified in multiplication with the states: the likelihood is invariant to signs of $\sigma_{\alpha}$ and $\sigma_{\tau}$. Consequently, mixing of the posterior state standard deviations can be poor and their distributions are likely to be bi-modal fruhwirth2010stochastic. This issue is addressed by randomly permuting signs in the Gibbs sampler as explained below. Similar to fruhwirth2010stochastic, we assume normal priors centred at 0 for $\sigma_i: \sigma_i \sim N(0,V_i)$ $\forall i \in \{\tau,\alpha\}$.
Collecting all state space parameters in $\theta = (\tau_0,\alpha_0,\sigma_{\tau},\sigma_{\alpha})$, we assume an independent multivariate normal prior with diagonal covariance matrix:
While the state processes $\{\tilde{\tau},\tilde{\alpha}\}^T_{t=1}$ can be estimated by any state space algorithm, we opt for the precision sampler method of chan2009efficient which is outlined in Appendix ((ref)) along with the state posteriors. In contrast to FFBS type algorithms, it samples the states without recursive estimation which speeds up computation significantly.
Our second enhancement concerns the SSVS prior. Variable selection in the BSTS model of scott2014predicting is done via a two component conjugate spike-and-slab prior which utilises a variant of Zellner’s g-prior and fixed expected model size. While computationally fast due to conjugacy, many high-dimensional problems benefit from prior independence moran2018variance and a fully hierarchical formulation to let the data decide on the most likely value of the parameters ishwaran2005spike.
Therefore, we follow ishwaran2005spike's extension to the SSVS prior, the Normal-Inverse-Gamma prior:
where $j \in (1,\cdots,K)$. The intuition remains the same as compared to the spike-and-slab prior of scott2014predicting in that the covariate's effect is modeled by a mixture of normals where it is either shrunk close to zero via a narrow distribution around zero (the spike component) or estimated freely though a relatively diffuse normal distribution (the slab component). Sorting into each component is handled through an indicator variable, $\gamma_j$, and the hyperparameter $c$ is chosen to be a very small number, thereby forcing shrinkage of noise variables to close to zero. While in the original BSTS model, the indicator variable, $\gamma_j$, depends on a fixed prior $\pi_0$ which governs the prior inclusion probability of a variable, ((ref)) allows for it to be estimated from the data through another level of hierarchy. We set $b_1=b_2=1$, which effectively assumes that any expected model size is a priori possible and thus allows for sparse but also dense model solutions as recommended by giannone2021economic. Finally, the prior variance $\delta^2_j$ is also allowed to be hierarchical. Posteriors are standard and described in the Appendix ((ref)). The posterior of $\gamma_j$ is of special interest to the analyst as it gives a data informed measure of importance of a variable. Specifically, $p(\gamma|y)$ can be interpreted as the posterior inclusion probability of a variable. \\
Our third and final enhancement of the BSTS models extends the employed shrinkage priors to the horseshoe prior. Like many recently popularised shrinkage priors, the horseshoe prior belongs to the broader class of global-local priors which take the following general form:
The idea of this family of priors is that the global scale $\nu$ controls the overall shrinkage applied to the regression, while the local scale $\lambda_j$ allows for the local possibility of regressors to escape shrinkage if they have large effects on the response. A variety of global-local shrinkage priors have been proposed polson2010shrink, but here we focus on arguably the most popular, the horseshoe prior of carvalho2010horseshoe which employs two half Cauchy distributions for $\nu$ and $\lambda_j$:
These two fat tailed scale distributions imply a shrinkage profile that has the spike-and-slab prior in its limit and therefore offers a continuous approximation to the SSVS piironen2017sparsity (see section (ref) in the appendix for further discussion). An additional attractive feature of the horseshoe prior is that it is completely automatic with respect to its hyperparameters and has been shown to be excellent at forecasting in several previous studies huber2019inducing,huber2020nowcasting,cross2020macroeconomic,follett2019achieving. Due to its special connection to frequentist shrinkage priors polson2010shrink, it offers not only good finite sample performance but also favourable asymptotic behaviour compared to competing global priors bhadra2019lasso. chakraborty2020bayesian in particular show that the fractional posterior mean as a point estimator is rate optimal in the minimax sense using ((ref)).
Nevertheless, fitting the horseshoe prior can be challenging when the scale parameters are not strongly identified by the data, which is particularly critical in cases where the likelihood is flat, for example, separable data in logistic regression piironen2017sparsity.\footnote{We thank an anonymous reviewer for having facilitated this discussion.} We provide in the appendix a ((ref)) robustness check based on piironen2017sparsity that are able to alleviate any identifiability concerns for the empirical study below.
Posteriors are described in the appendix ((ref)).
Although the horseshoe prior shrinks noise variables towards zero, the importance of a variable for nowcasts may not be immediately clear from posterior summary statistics of the coefficients, especially when the posterior is multi-modal. To aid interpretability and simultaneously preserve predictive ability, we employ the signal adaptive variable selection (SAVS) algorithm of ray2018signal to the posterior coefficients on a draw-by-draw basis. The algorithm uses a useful heuristic, inspired by frequentist lasso estimation, to threshold posterior regression coefficients to zero:
where $X_j = (x_{j1},\cdots,x_{jT})'$ is the $j^{th}$ column of the regressor matrix X, $sign(x)$ returns the sign of $x$ and $\hat{\beta}$ represents a draw from the regression posterior. The parameter $\kappa_j$ in ((ref)) acts as a threshold for each coefficient akin to the penalty parameter in lasso regression which can be selected via cross-validation. ray2018signal propose
which ranks the coefficients inverse-squared proportionally and provides good performance compared to alternate penalty levels ray2018signal,huber2019inducing. To see the similarity to lasso style regularisation, the solution to ((ref)) can be obtained by the following minimisation problem which is closely related to the adaptive lasso zou2006adaptive:
Here, $\overline{\phi}$ is the sparsified regression vector. Analogous to the SSVS posterior, the relative frequency of non-zero entries in the posterior coefficient vector can be interpreted as posterior inclusion probabilities. Integrating over the uncertainty of the parameters, we obtain the predictive distribution $p(\tilde{y}|y)$, which is similar to a Bayesian Model Averaged (BMA) posterior huber2019inducing.
With the conditional posteriors for the regression and state components at hand (see Appendix (ref)), we sample states as well as regression parameters with the following Gibbs sampler:
As mentioned in Section (ref), states are sampled in a non-recursive fashion which exploits sparse matrix computation and precision sampling. The exact sampling algorithm is given in Appendix (ref). As discussed in Section (ref), after sampling $\theta$ in step 2, we randomly permute signs of $(\tilde{\tau},\tilde{\alpha}),(\sigma_{\tau},\sigma_{\alpha})$ to aid mixing. Step 4 of the sampler will depend on the prior and its respective hyperpriors. While the posterior sampling scheme for the SSVS is standard, we use the efficient posterior sampler of bhattacharya2016fast to sample the regression coefficients of the horseshoe prior. Compared to Cholesky based sampling as used for the SSVS, computation speed is markedly improved; see Appendix (ref). Note that in step 4, we perform SAVS sparsification via ((ref)) on an iteration basis.
In this paper, we relate monthly macro data commonly used for nowcasting based on giannone2016exploiting and internet search information via U-MIDAS skip-sampling to real quarterly U.S. GDP growth. The U-MIDAS approach to mixed frequency belongs to the broader class of `partial system' models banbura2013now, which directly relate higher frequency information to the lower frequency target variable by vertically realigning the covariate vector. The benefit of this mixed frequency method compared to restricted MIDAS and full system state space methods is its simplicity in that existing models and priors can directly be applied to U-MIDAS sampled data as well as its competitive performance, especially when the frequency mismatch between the target and the regressors is small foroni2015unrestricted,foroni2014comparison, as is the case in our application. Switching notation from equation ((ref)) to make it explicit that $y_t$ is quarterly while $x_t$ is sampled at a higher, i.e., monthly frequency, denote $x_{t,M}=(x_{1,t,M}, \cdots, x_{K,t,M})$ and $\beta_m = (\beta_{1,M},\cdots,\beta_{K,M})'$ where $M=(1,2,3)$ denotes the monthly observation of the covariate within quarter, $t$. By concatenating each monthly column, we obtain a $T\times 3K$ regressor matrix $\boldsymbol{X}$ and a $3K \times 1$ regression coefficient vector $\boldsymbol{\beta}$. This vertical realignment is visualised for a single representative regressor below:
The macro data set pertains to an updated version of the database of giannone2016exploiting (henceforth, `macro data') which contains 13 time series which are closely watched by professional and institutional forecasters including real indicators (industrial production, house starts, total construction expenditure etc.), price data (CPI, PPI, PCE inflation), financial market data (BAA-AAA spread) and credit, labour and economic uncertainty measures (volume of commercial loans, civilian unemployment, economic uncertainty index etc.). We augment this data set with the composite Purchasing Managers Index (PMI) and the University of Michigan Consumer Confidence Index (UMCI). These are often used as leading indicators for producer and consumer sentiment, respectively. The target variable for this application is deseasonalised U.S. real GDP growth (GDP growth) data as reported in the FRED dataset.\footnote{Here, the deseasonalisation pertains to the X13-ARIMA method and was performed prior to download from the FRED-MD website.}$^{,}$\footnote{We thank an anonymous reviewer who brought to our attention that instead of mixing pre-deseasonalisation techniques between macroeconomic data and Google Trends discussed below, one could also deseasonalise with common techniques such as the Loess filter. In doing so, the results in this paper remain qualitatively identical. Details are available upon request.}
As early data vintages of macroeconomic data and GDP figures can exhibit substantial variation compared to final vintages croushore2006forecasting,sims2002role, there is no unambiguous choice of variable in evaluating nowcast models on historical data. Further complications can arise through changing definitions or methods of measurements carriero2015realtime. In order to judge the expected performance of the proposed models from a real-time perspective, we only make use of the latest vintages of the series available at the point in time of the nowcast. \footnote{The only exception is real GDP growth, for which, following previous nowcast studies carriero2015realtime,clark2011real, we use the second vintage for nowcast evalutation.} The stylised release calendar ((ref)) simulates the data availability within the months during which nowcasts are conducted. For instance, at the 24th nowcast period, all data which became available during periods 1-24 will be updated according to their latest available vintages dating prior to the release of PCE and PCEPI, which are published typically during the last week of a given month. Real-time vintages are downloaded from the FRED database using the `FredFetch' Matlab package.\footnote{The Matlab package is available from \url{https://github.com/MattCocci/FredFetch}. PMI data were downloaded from \url{quandl.com} using Quandl code `ISM/MANPMI'.}
Google Trends (GT) are indices produced by Google on the relative search volume popularity of a given search term, search topic or pre-specified search category, conditional on a given time frame and location. The difference between individual search terms and topics/categories is that the latter measures the search popularity for a basket of search terms which are content-wise related to the specified topic or category. In particular, categories are further split into a 5-level hierarchy of categories which are fixed a priori,\footnote{For an overview of categories and sub-categories, please see \url{https://github.com/pat310/google-trends-api/wiki/Google-Trends-Categories}} and topics can be assembled depending on the term one is interested in. For example, the user can specify the topic `Recession', whose related search queries contain, among others, `recession', `downturn', and `economic depression'. Likewise, the category `Welfare & Unemployment' relates to search queries about `unemployment, `food stamps' and `social security office'. A large literature on using individual Google Trends search terms\footnote{See, for example: guzman2011internet,mclaren2011using,askitas2009google,fondeur2013can,carriere2013nowcasting.} have shown that these data can improve predictions for economic time-series which have a clear connection to the specific search term used, such as using `unemployment benefits' to predict unemployment smith2016google. However, this approach has two potential limitations.
First, using broad search terms to capture general macroeconomic activity bears the risk of capturing spurious search behaviour. For example, the search term `jobs' might contain search volume for `Steve Jobs'. Second, since many search terms will be related to multiple topics, there may be lack of interpretability. To reduce search term ambiguity and interpretability in relationship to real GDP growth, we use Google topics and categories instead of individual search terms, and choose these based on their relationship with various aspects of the economy. As forcefully argued by woloszko2020tracking and fetzer2020coronavirus, these mostly alleviate spuriously correlated search terms as the user can confine the search purpose. This is benefited by the fact that Google refines this basket of search terms, by taking into account where users click after the search has been conducted woloszko2020tracking. Further, categories and topics can be conceptualised as factors based on search terms with the same meaning/purpose. Although the exact basket of search terms corresponding to a topic/category is not a priori accessible to the user, any topic or category with little predictive power will ultimately be shrunk to zero via the shrinkage priors employed in the proposed models.
Our sample comprises 37 Google Trends which were chosen based on capturing activity in various parts of the economy ranging from crisis/recession, labour market, personal finance, consumption to supply side activities.\footnote{This list was inspired by previous research such as woloszko2020tracking.} Our chosen list of topics and categories is as follows:
The relatively large proportion of search items related to consumption of goods and services reflects the large role of consumption in determining U.S. GDP. vosen2011forecasting and woo2019forecasting have shown that similar search items track and predict the UMCI index and private consumption very well, thereby capturing consumer sentiment. Labour topics and categories track the popularity of search terms related to job search and benefits demand which smith2016google, d2017predictive and fondeur2013can have shown to predict the unemployment rate in various countries. Topics related to personal finance and investment may signal wealth effects woloszko2020tracking, which tend to positively correlate with the business cycle (see figure (ref) in appendix). Topics around housing have been shown to be indicative of housing prices wu2015future,askitas2009google. The recession, business news and bankruptcy themed search items typically increase during economic downturns which therefore act as signals of economic distress and recessions castelnuovo2017google,chen2012forecasting.
While the selection of our search items is subjective, in general, there is no consensus on how to optimally select search terms for final estimation. Methods proposed in the previous literature can be summarised as: (i) pre-screening through correlation with the target variable as found via Google Correlate scott2014predicting,niesert2020can,choi2012predicting;\footnote{Unfortunately, Google Correlate has suspended updating their databases past 2017.} (ii) cross-validation ferrara2019google; (iii) use of prior economic intuition where search terms are selected through backward induction smith2016google,ettredge2005using,askitas2009google; and (iv) root terms, which similarly specify a list of search terms through backward induction, but additionally download “suggested” search terms from the Google interface. This serves to broaden the semantic variety of search terms in a semi-automatic way. As methodologies based on pure correlation do not preclude spurious relationships scott2014predicting,niesert2020can,ferrara2019google, we opt for our (somewhat subjective) selection to best guarantee economically relevant Google Trends.
Since search terms can display seasonality, we deseasonalise all Google Trends by the Loess filter, as recommended by scott2014predicting, which is implemented with the “stl” command in R.\footnote{To mitigate against inaccuracy stemming from sampling error, we downloaded the set of Google Trends seven times between 1 August 2021 to 8 August 2021 and took the cross-sectional average. Since we used the same IP address and google-mail account, there might still be some unaccounted measurement errors. However, using topics and categories instead of individual search terms, we observe much lower sampling variance.}
Although one of the main benefits of Google Trends is their timely availability, which can be as granular as displaying search popularity minutely for the past hour, the purpose of the empirical application is to showcase the flexibility of the proposed models in taking advantage of the heterogeneous information contained in adding new data sources to traditional macroeconomic data, even with little data processing efforts. Due to the simplicity of obtaining monthly Google Trends information, we sample the Google Trends information at the monthly frequency. Nevertheless, the proposed methodology can easily be extended to update monthly Google Trends with higher frequency search information via bridge methods as in ferrara2019google or could directly be included in the model via expansion of the covariate matrix.\footnote{Due to the already very high-dimensionality of the data set, we retain such extensions for future investigation. Constraining the parameter space via MIDAS sampling might make estimation more feasible.}
The indicative real-time calendar can be found in Table 1 and has been constructed after the data's real publication schedule. It comprises a total of 37 nowcast periods which make for an equal number of information sets $\Omega^v_t$ for $v = 1,\cdots,37$ which are used to construct nowcasts as explained in Section (ref). Google Trends are treated as released prior to any other macro information pertaining to a given month, since as argued, Google Trends information can essentially be continuously sampled.
To get a visual understanding of the Google search information, we plot in figure ((ref)) the first 3 principal components of the U-MIDAS transformed Google Trends information. Figure ((ref)) shows the factor loadings.
The first three principal components show very heterogeneous behaviour, but the dynamics conform to the economic intuition suggested from the loadings. The first component loads positively on supply side activity such as `Business services', `Construction, consulting & contracting', `Manufacturing' as well as `Investing' and negatively on recessionary themes and `Jobs' which spike during the crises and decrease during recoveries (see figure (ref) for indicative time-series plots of individual Google search items). Accordingly, the first component decreases strongly during the financial crisis (and to a smaller degree also during height of the Covid-19 recession) and picks up the rapid increase in economic activity after Q2 2020 very well.
The second component can be understood as a measure of consumer sentiment and financial health as it loads mostly on consumption items and negatively on topics such as `Bankruptcy' and `Foreclosure'; both terms were very popular during the financial crisis, but not so much the pandemic crisis.\footnote{This may reflect the different nature of the Covid-19 pandemic induced recession and the positive impact government policies had.} Accordingly, the second principal component shows a large dip during the financial but only a minor dip during the Covid-19 recession.
The third principal component loads very strongly on business news/recession/crisis items which increase in popularity during periods characterised by economic anxiety, thus spiking around the financial crisis and the pandemic period. Hence, it can be interpreted as an indicator of economic distress.
Understanding further how the skip-sampled Google Trend series correlate with the macro data set may help us anticipate which information will likely be picked up in the model. Figure ((ref)) shows a correlation heatmap between the macro data set and the first three Google Trends principal components. Please note that an increase in the UMCI and PMI indicate improved consumer and producer sentiments, respectively.
As expected, the first rather procyclical component correlates highly with the fed-funds rate which tends to also rise with the business cycle. The second component is strongly positively correlated with the UMCI index and negatively with the BAA spread (which increases during deteriorating financial conditions), indicating that it indeed captures something close to consumer sentiment and financial health. Since the third component acts as a recession signal, which abruptly spikes during crises, but is otherwise flat, it is not surprising that macroeconomic variables do not correlate very highly with it. This also indicates that search items in this group might add information that is not well captured by the other included macroeconomic information.
The predictive model used to generate in-sample and out-of-sample predictions is given in equations ((ref))-((ref)) where the first $T=45$ observations are used as training sample.\footnote{As a further alternative to the proposed BSTS models, we investigate in the appendix as well whether past GDP growth dynamics are more appropriately (in terms of nowcasting) modelled via ARMA components. The results show that the LLT components within the BSTS model are clearly preferred over ARMA type dynamics. We thank an anonymous reviewer for making this suggestion.} We estimate three variants of the model based on priors ((ref), (ref), and (ref)) and the original BSTS model of scott2014predicting, as well as an AR(4) model for comparison. In line with standard BSTS applications, we first compare the in-sample cumulative absolute one-step-ahead forecast errors, generated from the state space, as well as inclusion probabilities of the variables so as to shed light on which variables produce better fit and explain the outcome. Out-of-sample nowcasts are generated from the posterior predictive distribution $p(y_{T+1}|\Omega^v_{T})$ for growth observation $y_{T+1}$, conditional on the real-time information set $\Omega^v_{T}$, where $(v=1,\cdots,37)$ refers to nowcast periods within the real-time calendar (Table (ref)). This results in 37 different nowcasts which are generated on a rolling basis until the end of the forecast sample, $T_{end}$. As recommended by carriero2015realtime, variables that have not yet been published until nowcast period $v$ are zeroed out.
Point forecasts are computed as the mean of the posterior predictive distribution and are compared via real time root-mean-squared-forecast-error (RT-RMSFE) which are calculated for each nowcast period as:
where $\hat{y}^v_{T+j|\Omega^v_{T+j-1}}$ is the mean of the posterior prediction for nowcast period $v$ using information until $T+j-1$. \\
Forecast density fit is measured by the mean real-time log-predictive density score (RT-LPDS) and real-time continuous rank probability score (RT-CRPS):
where, for brevity of notation, $\zeta_{1:T+j-1}$ collects all model parameters as defined for each model, which are estimated with expanding in-sample information until $T+j-1$ and M stands for iterations of the Gibbs sampler after burn-in. Note that in ((ref)), $y^{v,A,B}_{T+j|\Omega^{v}_{T+j-1}}$ are independently drawn from the posterior predictive density $p(y^{v}_{T+1|\Omega^{v}_{T+j-1}}|y_T)$.
As shown by fruhwirth1995bayesian, in a setting where time-varying and fixed components for a structural state space model are chosen, the LPDS can be interpreted as a log-marginal likelihood based on the in-sample information and therefore provides a model founded scoring rule. The RT-CRPS can be thought of as the probabilistic generalisation of the mean-absolute-forecast-error. Similar to the log-score, it belongs to the broader class of strictly proper scoring rules gneiting2007strictly which allows for comparing density forecasts in a consistent manner.\footnote{We do not report calibration tests, as there are too few out-of-sample observations to meaningfully determine calibration.}$^,$\footnote{Although the CRPS is a symmetric scoring rule, it penalises outliers less aggressively than the log-score which is of advantage in small forecast samples such as ours.} To facilitate discussion, our objective is to maximise the RT-LPDS and minimise the RT-CRPS. For all forecast metrics, the predictive distribution used for ((ref), (ref), and (ref)) is traditionally generated in state space models via the prediction equations of the Kalman filter harvey1990forecasting. Instead, we use the simpler approximate method of cogley2005bayesian, which we found to make no practical difference in our sample.\footnote{The Kalman filter provides conditionally optimal forecast densities in terms of squared forecast error. However, if there is mispecification or if the forecast horizon is very short, then approximate methods can do just as well empirically. A similar logic holds when comparing direct and iterative forecasting methods such as in marcellino2006comparison. We thank an anonymous reviewer for bringing this to our attention.} The method is described in Appendix (ref).
Finally, to test whether a state variance is equal to zero, we make use of the Dickey-Savage density ratio evaluated at $\sigma^{\tau,\alpha}$ = 0:
It can be shown that for nested models, the DS statistic is equivalent to the Bayes factor between the prior and the posterior distribution of the parameter of interest at zero verdinelli1995computing. The intuition for the test is simple: if the prior probability-density-function (PDF) allocates more mass at 0 than the posterior at that point, there is evidence in favour of the unrestricted model, i.e., $\sigma^{\tau,\alpha} \neq 0$. While the priors for the state variances have well known forms and thus can be evaluated analytically, we estimate the denominator for all models through Monte Carlo integration.
Figure (ref) shows the in-sample cumulative-one-step ahead prediction errors using the proposed priors where the information set pertains to the entire estimation sample without ragged edges. From Figure (ref), it is clear that the horseshoe prior BSTS (HS-BSTS) provides the best in-sample predictions at all time periods. The HS-BSTS-SAVS and SSVS-BSTS initially provide similar fit, however diverge in performance around the financial and Covid-19 crises, especially the HS-SAVS-BSTS. It is striking that, compared to the former two, the HS-BSTS provides very stable performance as indicated by a nearly linear increase in errors even during the financial crisis and the Covid pandemic. It is also apparent that the SAVS algorithm is not able to retain the fit of the HS prior alone, which, as we show in the next subsection, is in contrast to the out-of-sample results. \\
To understand the driving variables behind the posterior predictive distributions, the posterior marginal inclusion probabilities for the SSVS and SAVS model are plotted for the top ten most drawn variables in Figures (ref) and (ref) respectively. The colors of the bars indicate the sign on a continuous scale of white (positive relationship) to black (negative relationship) of the variable when included in the model, and the prefix `GT' indicates Google Trend variables. The number [0,1,2] appended to a variable indicates the temporal position within a given quarter, with 0 being the latest month.
For all three models (for the BSTS model, see Figure (ref)), the posterior inclusion probabilities show that the most drawn Google search information pertains to the category `business news' and topic `investing' with clearly negative and positive impact respectively on GDP growth. The posteriors on these Google Trends variables (Figure (ref)) show that not only is the impact statistically significant but also economically so. As search intensity for business news goes up (down), GDP growth forecasts are adjusted downwards (upwards). Vice versa for the investing topic. Since the `business news' category spikes in recessions, but is otherwise flat, this suggests that during periods of heightened recessionary probability and economic fear, people engage and search more for business news which therefore acts as an indicator of expected economic distress. Similar reasoning has led to a large literature on using economic sentiment extracted from news media to model and forecast economic activity Kalamara2019,kalamara2022making,baker2016measuring,aprigliano2021power,gentzkow2019text,alexopoulos2015power,manela2017news,nyman2021news,shapiro2020measuring.\footnote{It would be interesting for future research to investigate whether Google Trends and news sentiment extracted from articles substitute or complement each other in modelling recessionary risks.}
Conversely, `investing' items are presumably searched more often when households and individuals are financially more stable which is when they engage in looking for investment opportunities. As seen in Figure (ref), this series tends to positively correlate with the business cycle. This interpretation of the investment topic corroborates findings of woloszko2020tracking who show that the investment topic has a positive impact in a panel data nowcasting exercise for GDP growth.
Figures (ref) and (ref) also reveal some interesting patterns about how macroeconomic data are employed in the models. The SSVS prior tends to select only the most dominant of the skip-sampled information, while the SAVS extended HS prior allocates significant inclusion probability to all months within a quarter. For example, while the SSVS prior selects the variable `unrate0', i.e., the unemployment rate for the last month in a given quarter, the HS-SAVS prior allocates nearly the same inclusion probability to data for all months on construction starts and M2. This result is likely driven by the fact that the SSVS prior discretises the model space and therefore, with correlated data, will tend to include only the variable with the highest marginal likelihood. Continuous shrinkage priors, on the other hand, make use of all covariates since they are always included in the model.
In line with this discussion, the posterior inclusion probability heatmaps in Figures (ref) and (ref) show generally that the SSVS and HS prior also display very different degrees of model uncertainty. Note that in the figures, inclusion probabilities to the left of the dashed line pertain to the macroeconomic data set, and to the right, the Google search data. The HS prior tends to display substantial uncertainty over inclusion, particularly for the Google Trends data, which makes sense given the similarity in signal within Google Trends categories and topics. By contrast, the SSVS tends to load only on a few Google search items and not explore posteriors of correlated GTs.
Nonetheless, the fact that the HS prior identifies individual macroeconomic series such as `pce2' as a clear signal indicates that mixed frequency information matter for predictive purposes, which would otherwise be lost in averaging information across quarters.
These results contribute to the recently popularised studies of sparsity within economic prediction problems giannone2021economic,cross2020macroeconomic in at least two ways. Firstly, they indicate that different sparsity patterns can emerge within data sources (here, macroeconomic and Google Trends) and within mixed frequency information. And secondly, different sparsity patterns can emerge depending on the prior used. Section (ref) further investigates robustness of the proposed priors to different sparsity settings, and provide recommendations.
Finally, our in-sample results clearly demonstrate (Figure (ref)) that there is support in the data for a local trend, but not a local linear trend: the posterior for $\sigma^{\tau}$ is clearly bi-modal with less mass on zero than the prior, while the posterior for $\sigma_{\alpha}$ has substantially more mass on zero than the prior. The Bayes factors are 28.69 and 0.42 for the state standard deviations respectively.
We now turn to out-of-sample nowcasting performance, where nowcasts are produced following the real time data publication calendar as explained in Section 2. Due to the extraordinary economic circumstances of the Covid-19 pandemic, we split the evaluation sample into: (a) pre-Covid (ending Q4 2019); and (b) during Covid (ending Q2 2021). We first evaluate point- and then density fit for the pre-Covid period.
RT-RMSFE are plotted in Figure (ref) for the competing non-centred BSTS estimators, as well as the AR(4) benchmark and the original BSTS model. Note that in all nowcast figures, we represent nowcast periods in which Google Trends are published by grey vertical bars. The following points emerge from Figure (ref). Firstly, it is clear that all proposed BSTS models based on the non-centred state space offer large performance gains (for certain nowcast periods up to 40%) over the original BSTS model. Secondly, all models nearly monotonically increase in precision as more data are released, where, as expected, the BSTS models outperform the AR benchmark\footnote{The HS-BSTS and HS-SAVS-BSTS outperform the AR model at all nowcast periods significantly as measured by the diebold1998vevaluating test at conventional significance levels. The SSVS-BSTS model does so only with the second nowcast period.} as soon as the first data becomes available. Thirdly, among non-centred BSTS models, the HS-SAVS-BSTS does the best; however, it is closely followed by the HS-BSTS. This indicates that the SAVS algorithm successfully shuts down contributions of noise variables and thus gives further validity to the variable selection results discussed above. With only a modest decrease of 2-5% in RT-MSFE relative to the plain HS-BSTS model, it is evident, however, that the horseshoe prior already provides aggressive shrinkage. Compared to the SSVS-BSTS, the horseshoe prior based BSTS models offer 7-25% improvements in terms of point forecast accuracy, especially so in the beginning nowcast periods.
Finally, we find that there is a large decrease in point-forecast error due to Google Trends releases prior to macroeconomic data becoming available which is consistent across all models considered. Improvements are in the range of 7-25% for the given models compared to the first period nowcasts.\footnote{Strictly speaking, these nowcasts are also forecasts due to the information set containing only information from the previous quarter.} The subsequent value of Google Trends for point forecasts is a function of how much a given model loads on the Google Trends search variables and how well the shrinkage prior can separate signal from noise. Hence, improvements for the HS-BSTS models are modest after the first GT release and the SSVS-BSTS experiences a noticeable improvement of 15% in the final GT release, whereas the original BSTS model of scott2014predicting becomes less precise with latter GT releases. The explanation is that the original BSTS model generally struggles with the dimensionality of data set which leads to ineffective variable selection and consequently poor nowcasting performance. \\
Similar to the real-time point forecasts, we plot real-time LPDS (RT-LPDS) and CRPS (RT-CRPS) in Figures (ref) and (ref), respectively. The RT-LPDS and RT-CRPS mostly confirm the main findings from the point nowcasts. In contrast to the point nowcasts, however, there is now a much more clear cut performance improvement in density fit when the Google Trends information are released in periods 15 and 27, especially so for the BSTS model of scott2014predicting. This divergence in performance hints at the fact that part of the value of including Google Trends information is to better characterise forecast uncertainty which in turn aids density calibration. Next, we explore this distinctive feature of Google Trends releases for nowcasts during the pandemic.
The pre-Covid results highlighted that the value of Google Trends are largest before any macroeconomic information are available for the given quarter. While the aim of the nowcasting application for the proposed models is not to provide very granular (weekly or higher) nowcasting models,\footnote{Higher frequencies would expand the covariate set within the U-MIDAS sampling framework even further. At higher frequencies, predictions could instead be based on single covariates and combined, for example via Bayesian model averaging, or alternatively the frequency weights could be constrained via lower parametric basis functions which is akin to conventional MIDAS estimation.}$^{,}$\footnote{See woloszko2020tracking for a panel data approach to a weekly GDP index based on Google Trends.} we now illustrate briefly how, even at the relatively coarse monthly level, Google search information improves predictions during the pandemic.
Figure (ref) plots mean forecasts with their credible 95% intervals for the HS-BSTS model based on only macroeconomic data (left column) and the full data set (right column) for the 15th (upper row) and 27th (lower row) nowcast periods respectively. The nowcast periods were chosen to showcase the best possible nowcasts based on information from the end of the second and third month within a quarter respectively. While neither model is able to capture the full extent of the trough during the Covid-19 recession, the model including search term information provides a clear sense of heightened downside risks through a large asymmetric dip of the lower part of credible forecast interval. This is in line with the findings from woloszko2020tracking that for many OECD countries, Google Trends information is able to provide timely downside risk indications. Uniquely, our nowcast exercise highlights how Google Trends can indicate large downward swings in GDP growth over and beyond contributions from macroeconomic data. The fact that Google search information has a greater impact on forecast uncertainty rather than point forecasts further indicates that future research should investigate the potential benefits of using alternative data sources for modelling conditional heteroskedasticity such as in GARCH or stochastic volatility type models.
It is also clear that both models struggle to nowcast the equally large upswing that follows the pandemic trough. Inspired by recent VAR forecasting literature during the pandemic lenza2020estimate,carriero2021addressing, we also explore a new BSTS model based on fatter tailed t-distributed errors. The logic behind models with fat tails (compared to those of the normal distribution) is to acknowledge that the large macroeconomic fluctuations, for example during the Covid pandemic, are hard to forecast and thus should be modelled through increased forecast uncertainty such that, importantly, large outliers do not adversely affect inference on model parameters.\footnote{This assumes that the outlier represents an `irregular' observation due to a shock rather than the co-evolution of macroeconomic variables.} Statistically, this is achieved in the posterior by down-weighting outliers through the error covariances.
The estimated model uses the same regression and state components as model ((ref))-((ref)). However, we assume $\epsilon_t \sim N(0,\sigma^2\psi_t)$, where $\psi_t$ is distributed as $\mathcal{G}^{-1}(\eta/2,\eta/2)$ where $\mathcal{G}^{-1}$ denotes the inverse-Gamma distribution and $\eta$ the degree of freedom parameter of the t-distribution. Smaller degrees of freedom indicate fatter tails. To estimate the BSTS-t model, we leverage a mixture representation of the t-distribution for which derivations and sampling steps are detailed in appendix ((ref)). Treating $\eta$ as a random variable, figure ((ref)), based on the whole estimation sample shows that there is clear evidence for fatter tails, which are mostly due to the large outliers during the Covid pandemic.
The nowcasts from this model (Figure (ref)) show that, in line with the finding of small posterior degrees of freedom for the error distribution, the forecast intervals are much wider compared to the normal BSTS models. The lower forecast interval now captures the trough during the pandemic already in the 15th nowcasting period. Surprisingly, we find that the t-model's mean prediction comes much closer to GDP growth realisation at the height of the recession in period 27. In fact, the additional nowcast period forecasts in Figure ((ref)) show that already with the third publication of the Google search information within Q2 2020, we see a large downward adjustment which had not materialised in period 24 before the GT release. The posterior inclusion probabilities ((ref)) in the appendix reveal that this is because the model loads less heavily on PCE inflation of the first month (`pce2') which had not reacted much until Q2 2020, as opposed to the Google Trends and construction starts data. Yet, this new model still struggles with the upswing and presents the trade-off that even pre-Covid the predictive uncertainty is very large compared to the normal BSTS models. We believe that future research may investigate whether and how different data and modelling techniques are able to accurately forecast not only the trough, but also the peak after sharp downturns.
The empirical application above showed that the proposed BSTS models perform better in point as well as density forecasts compared to the original model of scott2014predicting and that both the SAVS augmented horseshoe prior as well as the SSVS-BSTS exhibit a relatively sparse selection of macroeconomic data. This finding is in contrast to previous studies using macroeconomic data such as giannone2021economic and cross2020macroeconomic who find that priors yielding dense models generally outperform sparsity favouring priors. Since an innovation in this paper is the estimation of a latent local-linear trend which might filter out co-movement in the macroeconomic data, we compare the ability of the proposed priors to the original BSTS model scott2014predicting in capturing both sparse and dense environments. Further, to make the simulations closer to our empirical application, we additionally test the priors’ ability to detect zero state variances.
Specifically, we simulate local-linear-trend models as ((ref))-((ref)) having either the trend variance or the local trend variance set to zero, both equal to zero, or neither equal to zero. Accordingly, we generate 20 simulated samples for $(\sigma^{\tau},\sigma^{\alpha})=\{(0.5,0),(0,0.5), (0,0),(0.5,0.5)\}$ together with either a dense or a sparse DGP, where the sparse coefficient vector is set to
and the dense coefficient vector is
where $p_d$ is set to 2/3. For both coefficient vectors, the dimensionality, K, is set to 300 which is high dimensional compared to the number of observations $T=150$. We account explicitly for mixed frequencies by first generating the covariate matrix according to a multivariate normal distribution with mean 0 and a covariance matrix with its $(i,j)^{th}$ element defined as $0.5^{|i-j|}$ and then skip-sample each covariate individually after the U-MIDAS methodology as in ((ref)). In the simulations, the true regression coefficient values as well as state variances are known; hence, we compare the performance of the different priors via coefficient bias for the regression coefficients and Dickey-Savage density ratios evaluated at zero state variances. Bias is calculated as
where $\hat{\beta}$ refers to the mean of the posterior distribution. We estimate the original BSTS model with the expected model size, $\pi_0$, equal to the true number of non-zero coefficients.
As can be seen from Table (ref), both the non-centred BSTS models as well as the original BSTS model of scott2014predicting do better in sparse than in dense DGPs which is similar to the finding of cross2020macroeconomic. The largest gains of the proposed BSTS models over scott2014predicting can be found for dense DGPs where the proposed estimators offer gains in estimation accuracy well in excess of 50% . In sparse designs, however, the latter slightly outperforms the former. This is expected given that the spike-and-slab prior uses a point mass prior on zero and that the true expected model size is used. At the same time, it is encouraging that the differences in accuracy are very small. Among the proposed estimators in dense designs, the HS prior BSTS versions are 30-40% more accurate compared to the SSVS-BSTS which is in line with our findings from the empirical application. Hence, these results offer the conclusion that continuous shrinkage priors are clearly preferred over spike-and-slab models in dense DGPs with a latent local-linear trend component.
The Dickey-Savage density ratio tests confirm that the non-centred state space models are able to correctly identify which of the state variances are significant and which are not, even in high dimensional regression settings. However, the test is sensitive to correctly pinning down the regression coefficient vector: in dense designs, where the SSVS prior does worse than the horseshoe prior, the DS tests in cases (0,0.5) and (0,0) show false support for significant $\sigma^{\tau}$.\footnote{Note that we do not report DS tests for the original BSTS model. This is due to the fact that the prior on the state variance has no mass on zero and therefore is not testable.} \\
In this paper, we investigated the added benefit of including a collection of Google Trends (GT) topics and categories in nowcasts of U.S. real GDP growth through the lens of current-generation Bayesian structural time series (BSTS) models. We extended the BSTS of scott2014predicting to a non-centred formulation which allows shrinkage of state variances to zero in order to avoid overfitting states and therefore let the data speak about the latent structure. We further extended and compared priors used for the regression part which are agnostic about the underlying model dimensions to accommodate both sparse and dense solutions, as well as the widely successful horseshoe prior of carvalho2010horseshoe. To make the posterior of the horseshoe prior interpretable, we applied sparsifying algorithms borrowed from the machine learning literature, which improve upon the excellent fit of the horseshoe prior itself.
We find that Google Trends improve point as well as density nowcasts in real time within the sample under investigation, where largest improvements appear prior to publication of macroeconomic information. This finding is robust across all considered models. The highest posterior inclusion probability for prediction of GDP growth across all models is obtained with the Google topics/categories `business news' and `investing'. The time-series dynamics and model impact of these GTs suggest that they provide timely signals of economic anxiety and wealth effects, respectively. Structural implications of this finding may be investigated with larger Google Trend samples and for other countries. The superior performance of the proposed models over the original BSTS model is confirmed in a simulation study which shows that among the proposed models, the horseshoe prior BSTS performs best and the largest gains in estimation accuracy can be expected in dense DGPs. It is further confirmed that the non-centred state priors are able to correctly identify the latent structure, however they are sensitive to the efficacy of the regression prior to detect signals from noise.
Finally, we applied our models to the Covid-19 pandemic period and find that Google Trends information help characterise the uncertainty during the Covid recession and subsequent recovery period. An extension of the BSTS model to student-t errors is also shown to benefit the timeliness of the forecast revisions to the changes in the macroeconomic data.
Our work suggests some important avenues for future research. An aspect which remained unexplored in this study is that Google Trends might have time varying importance in relationship to the macroeconomic variable under investigation, as highlighted by koop2019macroeconomic. Search terms can be highly contextual and might therefore be able to predict turning points in some periods but not in others. While, given the limited quarterly observations of Google Trends, our current investigation of this research question is somewhat limited, this will improve in significance over time. Also, nowcasting in contexts where the design is partly dense and partly sparse is a challenging problem. Our work sheds some light on this question, but it also motivates further research in this direction.