Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
76,416 characters · 16 sections · 87 citation commands
Hierarchical Regularizers for Reverse Unrestricted Mixed Data Sampling Regressions
\affil{Department of Quantitative Economics, Maastricht University}
{\it Keywords}: mixed-frequency models, MIDAS models, forecasting, group lasso
\paragraph{Acknowledgements.} The third author was financially supported by the Dutch Research Council (NWO) under grant number VI.Vidi.211.032. In addition, we acknowledge support through the HiTEc Cost Action CA21163. We gratefully acknowledge the comments by participants at the 16th International Conference on Computational and Financial Econometrics (CFE 2022) and the 9th Annual Conference of the International Association for Applied Econometrics (IAAE 2023) and thank Claudia Foroni and Stephan Smeekes for helpful discussions.
In recent years, a growing body of literature on mixed-frequency (MF) models arose to exploit the information available in series recorded at different frequencies. So far, the literature has mostly focused on modeling low-frequency (LF) variables by means of more timely, high-frequency (HF) variables to improve the former's forecast or nowcast accuracy. One commonly used univariate MF model is the MIxed DAta Sampling (MIDAS) regression (ghysels2004midas) that links low-frequency to high-frequency data through tightly restricted distributed lag polynomials. Various extensions to the MIDAS framework have been proposed in the literature. foroni2015umidas, for instance, consider univariate unrestricted MIDAS (U-MIDAS) to enable ordinary least squares estimation instead of non-linear least squares (NLS). ghysels2016macroeconomics propose a multivariate extension to model relations between high- and low-frequency series in a stacked mixed-frequency VAR (MF-VAR) system; state-space alternatives of the latter are provided by, amongst others, mariano2003new, schorfheide2015real and koelbl2020new. Recently, high-dimensional MF models have been proposed for the univariate (e.g., han2020high; mogliani2021bayesian; babii2022machine), multivariate (for factor models see e.g., marcellino2010factor; foroni2014comparison; andreou2019inference; for Bayesian estimation e.g., schorfheide2015real; ghysels2016macroeconomics; gotz2016testing; mccracken2020real; cimadomo2021nowcasting; paccagnini2021identifying and for sparsity-inducing regularizers e.g., hecq2022hierarchical) and panel settings (e.g., babii2022panel).
A less explored research area concerns MF models suited for modeling HF variables by means of LF variables. This can be of interest when fundamental economic variables, usually available at a lower frequency, play an important role for forecasting higher frequency variables as they contain more valuable information for longer-term horizons. While the stacked MF-VAR of ghysels2016macroeconomics can be used to predict HF variables using LF ones (despite its main usage for LF variables)\footnote{Another example is dalBianco2012short who model the euro-dollar exchange rate at weekly frequency exploiting a set of monthly financial and macroeconomic variables using a state-space framework as in mariano2003new.}, dedicated MF models for HF components have been proposed in the literature. foroni2018rmidas extend the univariate MIDAS and U-MIDAS models to the Reverse (R-MIDAS) and Reverse Unrestricted MIDAS (RU-MIDAS) model where the HF dependent variable is regressed on the LF explanatory variables. Indeed, in RU-MIDAS models each HF component has its own model equation, thereby allowing the (lagged) dynamics between the high and low frequency variables to change for every HF component. We consider this RU-MIDAS model as it can be more flexible and parsimonious than the multivariate stacked MF-VAR and does not require the specification of a model for the LF variable. Moreover, the RU-MIDAS can automatically incorporate HF information released within an ongoing LF period (i.e., a conditional model), whereas the MF-VAR would require one to use a structural MF-VAR that describes the contemporaneous correlations among variables to update the forecast for every HF period. Thus, the RU-MIDAS model is more appropriate for direct forecasts of the HF variable than the MF-VAR.
Previous studies have, however, mostly considered RU-MIDAS models where the frequency mismatch $m$ between the HF and LF variable is small, for instance a setting with monthly HF components and a quarterly LF component (see e.g., foroni2018rmidas). Our interest lies in settings where this frequency mismatch $m$ is large, think for instance of a setting with daily HF components and a monthly LF component; such settings only recently attracted attention (see e.g., foroni2023low). In this case RU-MIDAS regressions are, however, severely affected by the “curse of dimensionality". There are two main sources of this curse. First, the larger the frequency mismatch the more equations, namely one for each HF component, need to be estimated. Since each equation comes with its own separate set of parameters, this in turn increases the total number of parameters in the RU-MIDAS model. Second, the number of observations available for the estimation of each HF equation is limited to the number of LF observations. Hence, the larger $m$, the fewer observations are available for estimation since the sample size of the high-frequency variable is divided by $m$. Without further adjustments, one would be limited to use RU-MIDAS with a small frequency mismatch. In this paper, we address this curse of dimensionality by proposing a pooled RU-MIDAS model with a hierarchical convex regularizer. First, to address the issue of sample size reduction, we propose to pool the regression coefficients regulating the lagged HF dynamics across the $m$ different HF equations. Pooling relies on the assumption that the lagged effect of a HF variable on itself does not change with the HF component that is modelled. For instance, by pooling we constrain the lagged effect of the previous day to be the same for each day within a month. Similarly for the lagged effect of order two and so on. We advocate to pool in RU-MIDAS models where the frequency mismatch $m$ is large as it allows one to use the full HF sample size instead of the limited LF sample size to estimate the HF coefficients, which is a considerable difference when the frequency mismatch is large. As a by-product, by pooling we decrease the total number of parameters that need to be estimated.
Second, to further reduce the curse of dimensionality in the pooled RU-MIDAS model one could, for instance, turn to distributed lag polynomials that restrict the coefficients similar to restricted R-MIDAS models. However, R-MIDAS (similar to the regular MIDAS) requires computationally expensive NLS estimation. Alternatively, foroni2023low consider RU-MIDAS in a Bayesian framework, whose good performance has been demonstrated for forecasting electricity prices foroni2023low and shipping freight-cost indices bouri2022role. We instead consider sparsity-inducing regularizers which form an appealing alternative (see hastie2015statistical for an introduction) and have successfully been applied to regular MIDAS regressions (babii2022machine) and MF-VARs (hecq2022hierarchical). However, to the best of our knowledge they have not been explored yet as tool for dimension reduction in RU-MIDAS regressions. In this paper, we adapt the mixed-frequency hierarchical regularizer of hecq2022hierarchical to the RU-MIDAS framework. The regularizer is build upon the group lasso with nested groups (see e.g., nicholson2020high) and encourages a hierarchical sparsity pattern that prioritizes the inclusion of coefficients according to the recency of the information the corresponding series contains about the state of the economy.
We consider two empirical applications. First, we consider HF volatility forecasting through LF macroeconomic variables. The relationship between financial volatility and macroeconomic or financial variables has been analyzed in the literature, including the mixed-frequency context (e.g., engle2013stock; asgharian2013importance; conrad2015anticipating; conrad2020two; fang2020predicting). Most studies find that macroeconomic variables only pay off when it comes to long-term forecasting but not so much for short-term horizons. We contribute to this literature by assessing the performance of the proposed regularized pooled RU-MIDAS model in forecasting the daily realized volatility of the S&P 500 through a diverse collection of monthly macroeconomic variables. We find that the pooled RU-MIDAS model leads to (slightly) better forecast performance compared to the original RU-MIDAS model. However, the addition of macroeconomic variables does not improve forecast accuracy of the daily S&P 500 realized variance compared to the popular HAR model of corsi2009simple.
In our second empirical application, we consider HF bicycle rental demand forecasting through LF ridership variables. To this end, we use publicly available hourly demand rental data from a bicycle-sharing system in New York City for the year 2023. We forecast the hourly demand for bicycles using publicly available daily ridership data on traffic estimates from other transportation types in New York City. To the best of our knowledge, such mobility application has not yet been explored in the mixed-frequency literature. We find that the pooled RU-MIDAS model leads to better forecast performance compared to natural univariate benchmarks that only use past information of the HF target variable. These forecast gains are noticeable for short intra-day forecast horizons, and then gradually disappear when considering day-ahead forecasting.
The remainder of this paper is structured as follows. Section (ref) introduces the pooled RU-MIDAS with hierarchically structured parameters and describes the regularized estimation procedure. Section (ref) presents the results on the empirical application for volatility forecasting, Section (ref) on the empirical application for bicycle rental forecasting. Section (ref) concludes. Additional results are available in the Appendix.
We start by revising the approximate RU-MIDAS regression of foroni2018rmidas based on linear lag polynomials in Section (ref).\footnote{We slightly deviate from foroni2018rmidas's (foroni2018rmidas) notation to facilitate the introduction of hierarchical structures in Section (ref). First, we make the LF variable observable at the end of LF period, together with the last HF observation within the LF period. In contrast, in foroni2018rmidas the LF variable is observable at the beginning of the LF period together with the first HF observation within the LF period. Second, we define the dummy variables in the dummy notation (see equation (ref)) differently. Third, the first lag of the LF variable always corresponds to the previous LF period.} Then, we introduce the pooled RU-MIDAS model in Section (ref), followed by the hierarchical sparsity structure on its parameters in Section (ref). Finally, the regularized estimator for the pooled RU-MIDAS model is presented in Section (ref).
Let $x_t$ denote the high-frequency (HF) variable which we observe at $t = \frac{1}{m}, \frac{2}{m}, \dots, \frac{m-1}{m}, 1, 1 + \frac{1}{m}, ... $. and $y_t$ the low-frequency (LF) variable which is only observed every $m$ periods at $t=1,2,\dots, T$, where $m$ denotes the frequency mismatch. For simplicity, we consider one lag for the LF variable and $m$ lags for the HF variable.\footnote{We can easily extend the methodology to include multiple high and low frequency variables with higher-order lags at the cost of more complicated notation.} The reverse unrestricted MIDAS (RU-MIDAS) model is given by
where $\alpha_i$, $\beta_{j,i}$ ($i,j = 1,\ldots,m$) are the unknown parameters, $\varepsilon_{t}$ the error term and the variables are assumed to be mean-centered such that no intercept is included. For illustrative purposes, consider equation (ref) for the monthly/quarterly setting with $m=3$. Then, $\alpha_i$ and $\beta_{j,i}$ ($j = 1,2,3$) correspond to the effects of, respectively, the lagged quarterly and monthly variables on the $i$th ($i = 1,2,3$) monthly HF variable. Hence, the subscript $i$ of the coefficients refers to the effect of the corresponding regressor on month $i$, whereas $j$ indicates the number of months the monthly regressor is lagged by. For $i=1$ (first month so for $t = \frac{1}{3}, 1+\frac{1}{3},...$), $\alpha_1$ contains the effect of the previous quarter onto the first month and $\beta_{1,1},\beta_{2,1}$ and $\beta_{3,1}$ the effect of respectively month three, two and one of the previous quarter on the first month. Similarly for $i=2$ (second month so for $t = \frac{2}{3}, 1+\frac{2}{3},...$), $\alpha_2$ contains the effect of the previous quarter, however $\beta_{1,2}$ now contains the effect of month one of the current quarter, i.e., the latest available monthly observation (first monthly lag), and $\beta_{2,2},\beta_{3,2}$ refer to respectively month three and two of the previous quarter. For $i=3$ (third month so for $t = 1, 2,...$), the latest available monthly observations are month two ($\beta_{1,3}$) and one ($\beta_{2,3}$) of the current quarter, followed by month three of the previous quarter ($\beta_{3,3}$). The RU-MIDAS model thus easily allows for the inclusion of HF information released within the current LF period. Figure (ref) provides a visual summary of these lagged dynamics for the RU-MIDAS model with the monthly/quarterly set-up. Finally, note that for ease of exposition, we include the same LF lag for each HF response in equation (ref), yet the model can easily accommodate different LF lags across HF responses to take, for instance, data releases into account.
foroni2018rmidas propose to group the system of equations in equation (ref) into a single equation
through the use of dummy variables $D_1$, $D_2$,\dots, $D_m$ taking respectively the value of one in the first ($t = \frac{1}{m}, 1+\frac{1}{m},...$), second ($t = \frac{2}{m}, 1+\frac{2}{m},...$), ... and $m$-th ($t = 1, 2,...$) HF observation within the LF period. We return to this dummy notation in Section (ref) due to its convenience for introducing the pooled RU-MIDAS model.
In a low-dimensional setting, the parameters of the RU-MIDAS model can be estimated by simple OLS. However, if the frequency mismatch $m$ is large, the model suffers from the curse of dimensionality since for one LF and $m$ HF lags $(1+m)m$ parameters need to be estimated with only $T$ low-frequency observations available, thereby making OLS unreliable. Indeed, first, the larger $m$ the more equations $i=1,\ldots,m$ need to be estimated which in turn increases the total number of parameters since each equation has its own set of parameters. Second, a large $m$ drastically decreases the sample size of the HF variable as the number of observations available for the estimation of each equation is limited to the number of LF observations $T$. The curse exacerbates if the number of LF and/or HF lags grows. Next, we propose to counter this sample size reduction by pooling corresponding lagged HF coefficients across equations.
To overcome the reduction in sample size, we propose to pool the coefficients $\beta_{j,i}$ of the lagged HF variable. For illustrative purposes, consider again equation (ref) for the monthly/quarterly setting. Assume that the lagged effect of the previous month onto the current month is the same regardless of the HF period $i=1,2,3$. Then we can restrict the coefficients by pooling, thereby imposing the restrictions $\beta_{1,1} = \beta_{1,2} = \beta_{1,3}$. Similarly, we assume that the lagged two months effect onto the current month is the same regardless of the HF period, i.e., $\beta_{2,1} = \beta_{2,2} = \beta_{2,3}$ and correspondingly for the the third month $\beta_{3,1} = \beta_{3,2} = \beta_{3,3}$. In sum, re-consider Figure (ref)(b) where the constraints imposed under pooling for the monthly/quarterly set-up are visualized through colour coding of the parameter labels, i.e., parameters with the same font color are pooled together.
Generally, let us introduce the “pooled RU-MIDAS" model in dummy notation as
Equation (ref) differs from (ref) in several aspects. In the RU-MIDAS model (ref), we consider a different equation for each HF variable which implies that the parameters capturing the lagged LF and HF effects vary for each HF component. In the pooled version (ref), in contrast, we only allow the effect of the lagged LF variable to vary for each HF component, whereas the lagged HF effects are constrained to be the same. Therefore, we now have $Tm$ observations available for the estimation of the lagged HF coefficients $\beta_1,\dots,\beta_m$ instead of just $T$, the number of LF observations. Furthermore, we greatly reduce the total number of estimated parameters from $(1+m)m$ to $2m$. Since even the later number can be considerable if the frequency mismatch $m$ is large and/or not all $m$ lagged HF or LF coefficients are necessarily of importance in modeling the HF variables, we further impose, in the next section, a hierarchical sparsity structure on the lagged parameters in equation (ref).
To further reduce the dimensionality in equation (ref), we resort to sparse estimation. While one could opt for “patternless" sparsity through the lasso (tibshirani1996regression), we propose to use the dynamic structure of the pooled RU-MIDAS in equation (ref) as an additional source of information to guide the estimation. We use a hierarchical sparse estimator to encourage structured sparsity patterns appropriate to the context of RU-MIDAS regressions, similarly to the one introduced in hecq2022hierarchical for stacked MF-VARs.
Consider the hierarchical sparsity pattern that naturally arises in the autoregressive (AR) coefficients of a pooled RU-MIDAS regression. Let $\bm{\beta}$ be the vector collecting the HF coefficients in equation (ref), namely $\bm{\beta} = [\beta_{1},..., \beta_{m}]^\prime$, and define $\bm{\alpha} = [\alpha_{1},..., \alpha_{m}]^\prime$ similarly for the LF coefficients. For each of the $2$ sub-vectors, we impose a hierarchical priority structure for parameter inclusion. We prioritize the inclusion of lagged coefficients according to the recency of the information the lagged time series contains relative to $x_t$, in line with the recency-based priority structure introduced in hecq2022hierarchical. The more recent the information contained in the lagged HF or LF predictor, the more informative, and thus the higher its priority of inclusion. For instance, for the HF lags in $\bm{\beta}$, the first lag $\beta_{1}$ contains the most recent information, followed by the second HF lag $\beta_{2}$, etc. Thus, for $j< j'$ we prioritize parameter $\beta_{j}$ over $\beta_{j'}$ in the model. Restrictions encouraging decay across lags are commonly used in the mixed-frequency literature, see for instance babii2022machine for univariate or ghysels2016macroeconomics for multivariate models. A similar recency structure arises in the coefficients of the LF variable. The first coefficient $\alpha_{1}$ contains the most recent information as it captures the effect of the previous LF period onto the first HF period which is followed by $\alpha_2$, the effect of the previous LF period onto the second HF period etc. Therefore, for $i< i'$ we prioritize parameter $\alpha_{i}$ over $\alpha_{i'}$. In the next section, we introduce the regularized estimator that encourages such hierarchical sparsity structures for the estimated parameters in the pooled RU-MIDAS model (ref). Through our empirical studies in Sections (ref) and (ref), we also demonstrate the flexibility of the regularized estimator in accommodating other hierarchical sparsity structures.
We use a nested group lasso (zhao2009composite) to attain the hierarchical structures detailed in Section (ref). Define $\bm{\beta}^{(j:m)} = [\beta_{j}, \dots, \beta_{m}]^\prime$ for $1 \leq j \leq m$. A nested structure then arises with $\bm{\beta}^{(1:m)} \supset \dots \supset \bm{\beta}^{(m:m)}$. Similarly, let $\bm{\alpha}^{(i:m)} = [\alpha_{i}, \dots, \alpha_{m}]^\prime$ for $1 \leq i \leq m$. We define the hierarchical group lasso estimator for the pooled RU-MIDAS model as
where $\lambda \geq 0$ is a tuning parameter and $\mathcal{P}^{\text{Hier}}(\bm{\alpha},\bm{\beta})$ denotes the hierarchical group penalty given by
The hierarchical structure for the high-frequency lags is built in through the condition that if $\bm{\beta}^{(j:m)} = \bf{0}$, then $\bm{\beta}^{(j':m)} = \bf{0}$ where $j < j'$, equivalently for $\bm{\alpha}$. We apply the proximal gradient algorithm to efficiently solve the optimization problem in equation (ref), as detailed in Appendix A of hecq2022hierarchical. For the selection of the tuning parameter $\lambda$, we rely on the Bayesian Information Criterion (BIC), but we also conduct an additional analysis to investigate the sensitivity of our results to the usage of time series cross-validation instead of the BIC (see Section (ref)). Finally, to reduce the bias introduced by lasso-type estimators, we use a “post-lasso" procedure where we apply the least squares on the variables retained by the hierarchical group lasso. This leads to improved empirical performance on our considered applications as discussed in Section (ref). We illustrate the performance of the pooled RU-MIDAS on two empirical studies as discussed in the next sections.
We consider the regularized estimator of the pooled RU-MIDAS to forecast the realized variance of the S&P 500 using several LF macroeconomic variables. We first review the literature on volatility forecasting with macroeconomic variables in Section (ref). We then discuss the data and our forecast set-up in Section (ref). Results are presented in Sections (ref) to (ref).
In finance, movements in volatility are closely tracked since they have a considerable impact on capital investment, consumption, and economic activities (see e.g., tsay2005analysis; bauwens2012handbook). Therefore, forecasting volatility and the understanding of its drivers is of interest to both academics and practitioners. In recent years, the relationship between financial volatility and macroeconomic variables has received increased attention in the literature, we contribute to this strand by investigating the predictive information contained in LF macroeconomic variables for forecasting HF financial volatility.
Two problems, however, that affect the construction of such volatility forecasts is that volatility is unobservable and that macroeconomic variables are usually only available at a lower frequency. To overcome these difficulties, engle2013stock propose a GARCH-MIDAS model to bridge the gap between daily stock returns and low-frequency (e.g., monthly, quarterly) explanatory variables. This model has been extensively used in the literature (e.g., asgharian2013importance; conrad2015anticipating; conrad2020two; fang2020predicting), but the evidence on which macroeconomic variables have predictive ability for stock market volatility is mixed. Nevertheless, coincident/lagging indicators such as inflation and industrial production (engle2013stock) and the expectations about them (conrad2015anticipating), as well as leading indicators such as term spread and housing starts (conrad2015anticipating, fang2020predicting) seem to often improve forecast performance. Additionally, the findings highlight that including macroeconomic variables pays off particularly when it comes to long-term forecasting and not so much for short-term horizons.
While GARCH models treat volatility as latent, accurate nonparametric estimates of volatility have become available since the emergence of high-frequency data (see e.g., andersen2001distribution; barndorff2002econometric). Realized variances, computed as the sum of the squared intraday returns for a particular day, make volatility “observable" and volatility forecast models relying on such data have become popular alternatives (see e.g., hua2013forecasting; haugom2014forecasting).\footnote{Apart from assessing forecast accuracy, one could explore Granger causality relations, see for instance gotz2016testing who test Granger-causality from a HF variable, namely the SP500 bipower variation to a LF macro variable, namely industrial production.} In this context, the heterogeneous autoregressive (HAR) model (corsi2009simple) has become particularly popular because of its simplicity and good performance to capture long memory features and for forecasting. The HAR model explains the log-realized variance as a linear function of the log-realized variances of yesterday and the average over the last five days (last week) and last 20 days (last month) which we will refer to in the remainder as the day-week-month (dwm) lag structure. The HAR model can be seen as a MIDAS regression model with step functions as a weighting scheme, since dimension reduction is achieved by imposing equality constraints on the coefficients in the autoregressive model of respectively lags 2 to 5 on the one hand and lags 6 to 20 on the other hand. The number of autoregressive parameters is thereby reduced from 20 (for the unrestricted AR(20)) to 3 (for the HAR).
Recently, it has become popular to compare the forecast performance of models with a more general lag structure selected by lasso-type estimators with the parsimonious dwm lag structure embedded in the HAR (e.g., audrino2016lassoing; audrino2017testing; ding2021forecasting; wilms2021multivariate). Though the results are mixed, there is some evidence that the dwm lag structure of the HAR model is hard to beat by more sophisticated lasso-based lag structures. So far, the addition of macroeconomic variables to such models has, however, not been explored yet. We build on this strand of the literature and investigate whether HF daily realized variances can be forecast more accurately by accounting for information in LF macroeconomic variables. Our proposed model pooled RU-MIDAS model, in contrast to the HAR, achieves dimension reduction by (i) pooling lagged HF effects across HF equations, and (ii) encouraging hierarchical sparsity on both the LF and the HF lagged effects. We thus combine imposed equality restrictions with data-driven sparsity restrictions.
We forecast the daily realized variance of the S&P 500 using a wide set of LF macroeconomic variables. Regarding the HF variable, we consider the daily log realized variance based on five-minute returns of the Standard & Poor’s 500 market index “SP500" taken from the Oxford–Man Institute of Quantitative Finance.\footnote{Alternatively, we could have taken a more robust estimator of the integrated variance such as the bipower variation or the median realized volatility but this is not the focus of our paper.} The sample spans from January 3rd 2000 to December 31st 2019 to avoid the influence of the pandemic. Our mixed-frequency models require a fixed frequency mismatch $m$, which we set equal to $m=20$ as in gotz2016testing. To ensure $m=20$ in each month, we therefore interpolate additional values for non-existing days for months with less than 20 trading days and disregard excessive days for months with more than 20 trading days at the beginning of the corresponding month. Figure (ref) plots the log-realized variances of the S&P500.
Regarding the LF macroeconomic variables, we consider monthly variables on various aspects of the economy: amongst others output, income, prices and employment, consumer sentiment and financial variables, see Table (ref) for an overview. The monthly macroeconomic series are directly taken from the FRED-MD dataset which is available at the Federal Reserve Bank of St. Louis FRED database (see mccracken2020real for details). We apply the transformation codes in column “T-code" to remove stochastic trends.\footnote{The transformation codes correspond to the codes in the FRED-MD dataset with the exception of “CPIAUCSL" where we found the first log difference to be sufficient to remove the stochastic trend during the period we consider.} Rather than using all macroeconomic series in the FRED-MD dataset, we use a diverse yet careful selection of macroeconomic series with potential forecasting power for forecasting daily realized volatility. Furthermore, we also consider the monthly news-based Economic Policy Uncertainty Index (“Epui") by baker2016measuring as additional “forward-looking" explanatory variable. Finally, while some of the macroeconomic time series described in Table (ref) are available at a higher frequency than monthly, we use monthly data for all these series to preserve the same daily-monthly mixed-frequency mismatch in our RU-MIDAS model.
In our analysis, we consider eleven different pooled RU-MIDAS models where each model includes a different LF macroeconomic predictor.\footnote{We also assessed the predictive information from the joint model that includes all low-frequency variables, but this did not result in improved forecast performance. Such a joint predictive model did, however, lead to a considerably higher computational burden. Results are available from the authors upon request.} We include $20$ HF lags (one month) and one LF lag in the model. To compare forecast performance of the different models and forecast methods, we perform a rolling window forecast exercise with a window size of 5 years ($T=1200$) and various forecast horizons ranging from daily ($h=1$), over weekly ($h=5$), and monthly ($h=20$), to several months ($h=40, 60, 120$). To forecast more than one day-ahead in time (i.e., forecast horizon $h > 1$), we use a direct forecast approach (marcellino2006forecast). For a given forecast horizon $h$, the direct approach re-estimates model (ref) but with $x_{t+h-1}$ as the dependent variable. The advantage of using such direct approach is that it does not require the forecast of the lagged low-frequency variable. However, the model specification changes for each forecast horizon. Finally, forecast performance is measured based on the out-of-sample mean absolute forecast error (MAFE) but results with root mean squared forecast error (RMSFE) are available upon request.
We have two main objectives in our empirical application that are subsequently investigated. First, we compare the forecast performance between the original RU-MIDAS and the proposed pooled RU-MIDAS model in Section (ref), to investigate how the degree of parameter flexibility in the mixed-frequency model influences the forecast accuracy. Second, we analyze in Section (ref) whether including LF macroeconomic variables for forecasting HF realized variance pays off by comparing the forecast performance between our pooled RU-MIDAS model and the HAR. We end with sensitivity analyses in Section (ref).
In this section, we compare the forecast performance of the original RU-MIDAS and the proposed pooled RU-MIDAS model, where the parameters of both models are estimated using the hierarchical estimator for comparability. The estimation procedure for the pooled RU-MIDAS model is detailed in Sections (ref) and (ref), for the original RU-MIDAS model, we set hierarchical sparsity structures for each HF equation $i=1,\dots,m$ individually. For instance, for the high-frequency AR lags ($\bm{\beta}_i$) the first high-frequency lag $\beta_{1,i}$ contains the most recent information, followed by the second HF lag $\beta_{2,i}$, etc. Thus, for $j< j'$ we prioritize parameter $\beta_{j,i}$ over $\beta_{j',i}$ in the model. As we only consider one LF lag and therefore $\alpha_i$ is a scalar, the group-lasso penalty for the LF lag boils down to a simple $\ell_1$-penalty.
Figure (ref) depicts the out-of-sample MAFEs for the pooled RU-MIDAS (x-axis) versus the original RU-MIDAS (y-axis) across the six forecast horizons (panels). Each panel includes eleven MAFEs (dots), one for each model with a particular macroeconomic variable as indicated in the coloured legend. The black line, being the 45 degree line, is added to facilitate comparison between the pooled RU-MIDAS model and the original one since dots above this line imply that the pooled model attains better average forecast accuracy (i.e. lower MAFEs). Across all forecast horizons (apart from $h=120$) and all macroeconomic variables, the pooled model delivers better forecast accuracy. This indicates that pooling corresponding lagged HF variables pays off for the considered data set. At short forecast horizons ($h=1, 5$), the MAFEs are nearly indistinguishable across macroeconomic variables. For horizons $h=20, 40, 60, 120$, the dot corresponding to “Houst" (red) stands out. The model including Housing obtains higher MAFEs for both the pooled as well as the original RU-MIDAS model, hence the red dot is positioned in the upper right corner relative to the other dots.
For each dot (i.e., each pairwise comparison between pooled RU-MIDAS and original RU-MIDAS) in Figure (ref), we perform a Diebold-Mariano (DM) test (diebold1995comparing) to test for significant difference in forecast accuracy. At horizon $h=1,5,20$, the pooled version is more accurate at 1% significance level across all eleven models. At horizon $h=40, 60$, the pooled version performs better at (at least) 5% significance level across all models with the exception of “Houst" for which the improvement is statistically insignificant. For $h=120$, there is no significant improvement. In sum, we observe that the addition of the pooling constraints to the RU-MIDAS model never leads to significantly worse forecast performance compared to the original RU-MIDAS model, often even better. We therefore continue with the pooled RU-MIDAS model and investigate whether the addition of LF macroeconomic variables can improve forecast performance over the standard HAR model.
In this section, we assess the predictive information in the LF macroeconomic variables when forecasting the S&P 500 daily realized variance. To this end, we compare the forecast performance between the pooled RU-MIDAS model, which includes macroeconomic variables, and the HAR that does not include macroeconomic variables. The parameters in the HAR model are estimated by OLS.
We first evaluate overall forecast performance, as computed via the MAFEs averaged across time. Figure (ref) depicts the MAFE for the pooled RU-MIDAS estimated by the hierarchical estimator (dots, one for each macroeconomic variable) and the HAR (black line) across the six horizons. We clearly notice that, for both models, the forecast performance deteriorates with the forecast horizon (as expected). Moreover, the HAR consistently outperforms the pooled RU-MIDAS model by delivering lower MAFEs. Particularly, the addition of “Houst" (red dot) deteriorates forecast accuracy in comparison to the other macroeconomic variables for which the difference is almost indistinguishable. Hence, overall, we find no evidence that the addition of macroeconomic variables improves forecast performance over the standard HAR. Similarly, fang2020predicting observed that the inclusion of past volatility information rather than the inclusion of macroeconomic variables improves forecasting accuracy. There are some exceptions at horizon $h=120$, namely for the models including the macroeconomic variables “Epui", “Fedfunds", “Indpro", “Termsp",“Unrate" and “Vix"\footnote{The models with macroeconomic variables “Cpi", “Csent", “M2", “Oil" also obtain a lower MAFE than the HAR, however, no macroeconomic lags are selected in these models (see Table (ref) in the Appendix) so they only differ from the HAR in terms of their HF lag structure.}, which display a somewhat lower MAFE than the one obtained with the HAR model. It should, however, be noted that differences are largely negligible.
To further analyze the differences in forecast performance, Table (ref) presents the MAFE of the HAR, and the pooled RU-MIDAS estimated by the hierarchical estimator (results as displayed in Figure (ref)), as well as by OLS to better assess the influence of the regularization. For each forecast horizon, we consider these three estimation procedures, the two RU-MIDAS in combination with the eleven models, each including a different macroeconomic predictor, leading to 23 models in total. For each horizon, we obtain the models included in the 75% Model Confidence Set (hansen2011model), whose MAFEs are displayed in bold.
We start by discussing the results for horizon $h=1$. Here, many combinations are included in the MCS with the HAR being the best performing one (on average). However, care is needed when evaluating the performance of the mixed-frequency models with macroeconomic variables included in the MCS vis-a-vis the HAR. Indeed, Table (ref) in the Appendix reveals that the pooled RU-MIDAS models do not include any macroeconomic variables, since all their corresponding coefficients are set to zero by the regularized estimator. Hence, the HAR models and the pooled RU-MIDAS models only differ from each other in terms of the HF lag structure. The HAR model includes the HF lags via the parsimonious dwm lag structure, which cannot be improved upon for short forecast horizon. While the pooled RU-MIDAS model estimated by OLS includes all 20 HF lags, the pooled RU-MIDAS model estimated by the hierarchical estimator includes on average around 4 HF lags (min. 3 lags, max. 6 lags) in the model. Our results confirm the previous findings by audrino2016lassoing, audrino2017testing, and wilms2021multivariate on the limited evidence that the more general lag structure allowed by lasso-based approaches can beat the parsimonious day-week-month lag structure embedded in the HAR in terms of out-of-sample forecast performance.
Across other forecast horizons, the HAR maintains its top performance. Hence, we do not find evidence that the considered LF macroeconomic variables contain predictive information for forecasting the daily S&P 500 daily realized variance. The only exception is $h=120$, where there is a small gain in terms of average MAFE over the HAR when using the pooled RU-MIDAS models with various macroeconomic variables. It should, however, be noted that only in a minority of rolling windows (see Table (ref) in the Appendix) this improved forecast accuracy over the HAR stems from the inclusion of macroeconomic variables.
Finally, let us discuss the performance of the pooled RU-MIDAS models estimated by OLS and the hierarchical estimator. We find that their performance is, overall, largely similar. The reason for that is that we employ a rolling window of five years to have sufficient low-frequency observations. However, with only one lag of a single low-frequency variable and twenty high-frequency lags included in the model, we need to estimate 40 coefficients while having a sample size of $T=1200$ high-frequency observations. Hence, the curse of dimensionality is rather limited in our set-up.
In sum, we find that macroeconomic variables do not help to improve the forecast performance of the daily S&P 500 realized variance over the simple HAR. This finding is robust across a wide range of macroeconomic variables and forecast horizons. The superiority of the HAR model over more sophisticated statistical learning/machine learning methods is also confirmed by the recent study of audrino2024hard including 1445 stocks from the US equity market. Indeed, the HAR model turns out to be exceptionally hard to beat especially so when it is re-estimated daily in a rolling window set-up with window size around two and a half to four years.
{\it Rolling Window Size.} We assess the sensitivity of our findings to the choice of rolling window size. We re-did the forecast exercise using a rolling window size of 2 and 10 years instead of 5 years; the results are reported in respectively Table (ref) and (ref) in the Appendix. While the overall findings remain unchanged, we do observe, as expected, a more pronounced advantage of the regularized pooled RU-MIDAS model over the OLS-based one if one were to use a smaller rolling window size of 2 years. For rolling window of 10 years, the advantage of using a regularized method in contrast to simple OLS largely disappears due to the sample size. Moreover, there is a minor benefit of including macroeconomic variables, as evidenced by the pooled RU-MIDAS estimated by OLS often being included in the MCS together with the HAR.
{\it Tuning Parameter Selection.} We re-did the forecast exercise using time series cross-validation instead of the BIC to select the tuning parameter $\lambda$ in equation (ref). We hereby use a rolling-window set-up, compute for each value of the tuning parameter the Mean Absolute Forecast Error (MAFE) across the last 20% (i.e.\ one year) observations as cross-validation score, and select the value of the tuning parameter with the lowest MAFE. Results are reported in Tables (ref) and (ref) of the Appendix. Table (ref) reveals that the usage of time series cross-validation instead of BIC to select the tuning parameter leads to a considerably higher percentage of selected LF macroeconomic coefficients. Nonetheless, the addition of the LF macroeconomic variables in the model does not lead to improved out-of-sample forecast performance, as can be seen from Table (ref). The MAFEs using either BIC or time series cross-validation are, overall, very similar.
{\it Post-lasso Estimation.} We re-did the forecast exercise without the post-lasso step to investigate its influence on the results. Table (ref) in the Appendix reports the results for the pooled RU-MIDAS model both with and without post-lasso estimation. We observe that the post-LASSO estimation step leads to improved forecast performance, and this across all considered models, macroeconomic variables and forecast horizons.
{\it High-frequency Lag Structure.} The pooled RU-MIDAS model is very flexible since it can easily incorporate other lag structures. For instance, one can easily incorporate the parsimonious dwm-lag structure used in the HAR into the penalized regression set-up of the pooled RU-MIDAS model. To this end, we re-did the forecast exercise using three different HAR-type models augmented with macroeconomic variables: The (i) “HAR+macro OLS" procedure is a HAR model (with dwm lag structure) augmented with a low-frequency macroeconomic variable in the pooled RU-MIDAS set-up and estimated through OLS. (ii) “HAR+macro HIER" is the same model as above but estimation is done with the hierarchical regularizer. The hierarchical sparsity penalties can be applied with minimal adaptation to prioritize the daily lag over the weekly lag, which in turn could be prioritized over the monthly lag. We also use a sparsity penalty on the coefficients corresponding to the macroeconomic variables. (iii) “HAR+macro HYBRID" is the same model as above but the HAR part of the regression is left unpenalized, and the low-frequency macroeconomic predictors are penalized. Table (ref) in the Appendix presents the MAFEs of these three procedures. We do not find improved forecast performance over the standard HAR model, so our findings remain robust to the usage of the high-frequency lag structure in the pooled RU-MIDAS model. Overall, “HAR+macro HYBRID" performs very similar to the HAR model since none of the low-frequency macroeconomic variables are selected (across most of the rolling windows). It outperforms “HAR+macro HIER" (overall apart from $h=120$) which in turn outperforms “HAR+macro OLS".
In our second empirical application, we use the pooled RU-MIDAS model to forecast demand for bicycle rentals in New York city using traffic estimates from several other transportation types. While data from this bicycle-sharing system has been previously consider in hu2024fast,rombouts2024cross, we are, to the best of our knowledge, the first to explore the predictive power of the ridership variables for bicycle rentals through mixed-frequency models. In Section (ref) we describe the data together with our forecast set-up. In Section (ref), we use the pooled RU-MIDAS model to forecast bicycle rental demand and present the results.
We forecast the hourly demand for bicycle rentals of the bicycle sharing system Citi Bike in New York City using a diverse set of LF ridership variables on various other transportation types.
Regarding the HF variable, we collect publicly available hourly data from Citi Bike which we process to obtain hourly time series measuring demand for bicycle rentals in New York City. The sample size spans January 1st 2023 to December 31st 2023, leading to a total of $T=8760$ hourly observations. Figure (ref), top panel, displays the hourly demand for bicycle rentals. Although not directly visible from Figure (ref), there is clear intra-day as well as day-of-the-week seasonality present in the data.
Regarding the LF variables, we collect publicly available daily ridership data on traffic estimates for seven different transportation types in New York City provided by the Metropolitan Transportation Authority, see Table (ref) for an overview. Figure (ref), bottom panel, displays the daily traffic totals over the same sample period and this for the seven different ridership types on a log-scale. The day-of-the-week seasonality is apparent from Figure (ref).
In our analysis, we consider seven different pooled RU-MIDAS models with mixed frequency mismatch $m=24$ where each model includes a different LF ridership predictor. We include $24\times7 = 168$ HF lags (one week) and 7 LF lags (one week) in the model. To accommodate the seasonal patterns present in the bicycle and ridership data, the recency-based hierarchy of the penalty terms in equation (ref) can be easily adjusted to a seasonal-based hierarchy structure. In particular, for the HF lags, we prioritize the inclusion of the lag corresponding to the same hour of the previous week before the inclusion of all other lags according to their recency. Similarly, regarding the LF lags, we prioritize the inclusion of the last week lag before the inclusion of the other lags according to their recency. We consider a similar rolling-window forecasting set-up as the one described in Section (ref). We use a rolling window size of 30 days and consider the following forecast horizons: $h=1,2,3,6,12$ hour(s)-ahead to cover intra-day forecast horizons and finally $h=24$ to cover day-ahead forecasting.
We compare the forecast performance of the pooled RU-MIDAS models to two natural univariate benchmarks tailored towards the seasonality in the bicycle rental data. First, we use a seasonal version of the random walk (RW) model where the forecast of a specific hourly time slot is the value observed in the same hour and day of the previous week. This is a popular forecasting model for platform applications since it works well in the case of strong seasonality patterns as observed in our data. As a second univariate benchmark model, we consider SARIMA models as implemented in the auto.arima function of the forecast package Rforecast-ref1, Rforecast-ref2 in R Rcoreteam. We search the optimal model for our hourly bike rental series in the space of seasonal ARIMA models with seasonal frequency $s=24$.
Table (ref) presents the MAFE of the seasonal RW, the SARIMA model, and the pooled RU-MIDAS model estimated by either OLS or the hierarchical estimator. For each horizon, we obtain the models included in the 75% MCS whose MAFEs are displayed in bold.
We first discuss the results for $h=1$. Both versions of the pooled RU-MIDAS model outperform the seasonal RW model and the SARIMA model, and this regardless of the ridership type. Hence, there is predictive information contained in the LF ridership data for the HF demand of bicycle rentals. Table (ref) also confirms that across all ridership types, the LF ridership variables are selected (i.e.\ their corresponding coefficients are estimated as non-zero) in a large majority of the rolling windows, namely more than 85% apart from the Bridges and Tunnels ridership type where they are still included in more than 65% of the rolling windows. The hierarchical estimator leads to the best forecast performance across all ridership types, and the pooled RU-MIDAS model with the ridership types Metro and Staten Island Railway are the best performing ones based on the MCS.
Turning to the other intra-day forecast horizons ($h=2,3,6,12$), we notice that the pooled RU-MIDAS model with hierarchical estimator still outperforms the univariate benchmarks, although the margin by which it outperform these benchmarks, especially so the seasonal RW, becomes smaller as the forecast horizon increases. Table (ref) reveals that the predictive power of the LF ridership variables gradually disappears as the forecast horizons increases, since fewer and fewer ridership variables are included in the pooled RU-MIDAS model. The selection percentages drop from more than 50% for horizon $h=2$ to less than 1% for horizon $h=12$. The MCS results reveal that the pooled RU-MIDAS models with hierarchical estimator still delivers the best performing models for horizons $h=2$ (LIRR, Metro and Staten Island Railway) and $h=3$ (all ridership types except for Bridges and Tunnels; the OLS pooled RU-MIDAS model with LIRR is also included in the MCS). For horizons $h=6$ and $h=12$, in contrast, the best performing pooled RU-MIDAS models perform equally good as the seasonal RW model.
Finally, for the day-ahead forecast horizon $h=24$, the pooled RU-MIDAS models still maintain their excellent performance when inspecting their MAFEs, yet the MAFEs are identical across ridership types, and Table (ref) further reveals that the good forecast performance stems from the sole inclusion of the HF lags selected in the model, since no LF lags are selected.
In sum, the pooled RU-MIDAS model with hierarchical estimator offers clear forecast gains at short intra-day horizons over standard univariate benchmarks, but the predictive power of the LF ridership variables dies out quickly across larger forecast horizons. For day-ahead forecast horizons (or further), there is no advantage of using a mixed-frequency model over a standard univariate model for the HF series of interest.
\color{black}
We consider mixed-frequency RU-MIDAS regressions with a large frequency mismatch $m$ and propose two extensions to the framework to tackle their curse of dimensionality. We (i) pool the regression coefficients regulating the lagged HF dynamics across the $m$ different HF equations to increase the sample size and (ii) we use a hierarchical regularizer to discriminate between recent and obsolete information. The RU-MIDAS model with hierarchical sparsity structures is quite flexible and can easily combine various high-frequency and low-frequency series in the model as well as incorporate other lag priority structures as demonstrated on two empirical applications.
We first use the proposed regularized pooled RU-MIDAS model to forecast the daily realized volatility of the S&P 500 through the information contained in a diverse set of monthly macroeconomic variables. While we find that the pooled RU-MIDAS model improves forecast performance upon the original (non-pooled) one, we do not find evidence that the monthly macroeconomic variables contain important information to improve forecast performance upon the popular benchmark HAR model that only exploits past volatility information. We also use the pooled RU-MIDAS model to forecast hourly demand for bicycle rentals in New York City through the information contained in a diverse set of daily ridership variables of other transportation types. Here, we do find evidence that the daily ridership variables contain important predictive power to improve forecast performance upon natural univariate benchmarks. This gain in forecast performance is visible for intra-day forecast horizons but not for day-ahead forecast horizons. It would be interesting to further explore the potential of regularized RU-MIDAS models for other application areas.
\paragraph{Data availability statement.} Regarding the volatility application, the FRED-MD dataset is publicly available at \url{https://research.stlouisfed.org/econ/mccracken/fred-databases/}. The Economic Policy Uncertainty Index is available at \url{https://www.policyuncertainty.com}. The realized variance was taken from the Oxford–Man Institute of Quantitative Finance, however the public access has been terminated. Regarding the ridership application, the Citi Bike dataset is publicly available at \url{https://citibikenyc.com/system-data}. The daily ridership data with traffic estimates is publicly available at \url{https://data.ny.gov}. Computer code is available from \url{https://github.com/MarieTernes/hierarchical-RUMIDAS}.