Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
67,865 characters · 11 sections · 80 citation commands
Machine Learning Time Series Regressions With an Application to Nowcasting
\def\spacingset#1{ {#1}} \spacingset{1}
\if00 {
\thispagestyle{empty}
} \fi
\if10 {
} \fi
{\it Keywords:} high-dimensional time series, heavy-tails, tau-mixing, sparse-group LASSO, mixed frequency data, textual news data. \\
\setcounter{page}{0}
\spacingset{1.45}
The statistical imprecision of quarterly gross domestic product (GDP) estimates, along with the fact that the first estimate is available with a delay of nearly a month, pose a significant challenge to policy makers, market participants, and other observers with an interest in monitoring the state of the economy in real time; see, e.g., ghysels2018forecasting for a recent discussion of macroeconomic data revision and publication delays. A term originated in meteorology, nowcasting pertains to the prediction of the present and very near future. Nowcasting is intrinsically a mixed frequency data problem as the object of interest is a low-frequency data series (e.g., quarterly GDP), whereas the real-time information (e.g., daily, weekly, or monthly) can be used to update the state, or to put it differently, to {\it nowcast} the low-frequency series of interest. Traditional methods used for nowcasting rely on dynamic factor models that treat the underlying low frequency series of interest as a latent process with high frequency data noisy observations. These models are naturally cast in a state-space form and inference can be performed using likelihood-based methods and Kalman filtering techniques; see banbura2013now for a recent survey.
So far, nowcasting has mostly relied on the so-called standard macroeconomic data releases, one of the most prominent examples being the Employment Situation report released on the first Friday of every month by the US Bureau of Labor Statistics. This report includes the data on the nonfarm payroll employment, average hourly earnings, and other summary statistics of the labor market activity. Since most sectors of the economy move together over the business cycle, good news for the labor market is usually good news for the aggregate economy. In addition to the labor market data, the nowcasting models typically also rely on construction spending, (non-)manufacturing report, retail trade, price indices, etc., which we will call the traditional macroeconomic data. One prominent example of nowcast is produced by the Federal Reserve Bank of New York relying on a dynamic factor model with thirty-six predictors of different frequencies; see bok2018macroeconomic for more details.
Thirty-six predictors of traditional macroeconomic series may be viewed as a small number compared to hundreds of other potentially available and useful nontraditional series. For instance, macroeconomists increasingly rely on nonstandard data such as textual analysis via machine learning, which means potentially hundreds of series. A textual analysis data set based on {\it Wall Street Journal} articles that has been recently made available features a taxonomy of 180 topics; see bybee2019structure. Which topics are relevant? How should they be selected? thorsrud2018words constructs a daily business cycle index based on quarterly GDP growth and textual information contained in the daily business newspapers relying on a dynamic factor model where time-varying sparsity is enforced upon the factor loadings using a latent threshold mechanism. His work shows the feasibility of traditional state space setting, yet the challenges grow when we also start thinking about adding other potentially high-dimensional data sets, such as payment systems information or GPS tracking data. Studies for Canada (galbraith2018nowcasting), Denmark (carlsen2010dankort), India (raju2019nowcasting), Italy (aprigliano2019using), Norway (aastveit2020nowcasting), Portugal (duarte2017mixed), and the United States (barnett2016nowcasting) find that payment transactions can help to nowcast and to forecast GDP and private consumption in the short term; see also moriwaki2019nowcasting for nowcasting unemployment rates with smartphone GPS data, among others. We could quickly reach numerical complexities involved with estimating high-dimensional state space models, making the dynamic factor model approach potentially computationally prohibitively complex and slow, although some alternatives to the Kalman filter exist for the large data environments; see e.g., chan2009efficient and delle2019efficient.
In this paper, we study nowcasting a low-frequency series -- focusing on the key example of US GDP growth -- in a data-rich environment, where our data not only includes conventional high-frequency series but also nonstandard data generated by textual analysis of financial press articles. We find that our nowcasts are either superior to or at par with those posted by the Federal Reserve Bank of New York (henceforth NY Fed). This is the case when (a) we compare our approach with the NY Fed using the same data, or (b) when we compare our approach using an expanded high-dimensional data set. The former is a comparison of methods, whereas the latter pertains to the value of the additional (nonstandard) big data. To deal with such massive nontraditional data sets, instead of using the likelihood-based dynamic factor models, we rely on a different approach that involves machine learning methods based on the regularized empirical risk minimization principle and data sampled at different frequencies. We adopt the MIDAS (Mixed Data Sampling) projection approach which is more amenable to high-dimensional data environments. Our general framework also includes the standard same frequency time series regressions.
Several novel contributions are required to achieve our goal. First, we argue that the high-dimensional mixed frequency time series regressions involve certain data structures that once taken into account should improve the performance of unrestricted estimators in small samples. These structures are represented by groups covering lagged dependent variables and groups of lags for a single (high-frequency) covariate. To that end, we leverage on the sparse-group LASSO (sg-LASSO) regularization that accommodates conveniently such structures; see simon2013sparse. The attractive feature of the sg-LASSO estimator is that it allows us to combine effectively the approximately sparse and dense signals; see e.g., carrasco2016sample for a comprehensive treatment of high-dimensional dense time series regressions as well as mogliani2020bayesian for a complementary to ours Bayesian view of penalized MIDAS regressions.
We recognize that the economic and financial time series data are persistent and often heavy-tailed, while the bulk of the machine learning methods assumes i.i.d.\ data and/or exponential tails for covariates and regression errors; see belloni2018high for a comprehensive review of high-dimensional econometrics with i.i.d.\ data. There have been several recent attempts to expand the asymptotic theory to settings involving time series dependent data, mostly for the LASSO estimator. For instance, kock2015oracle and uematsu2019high establish oracle inequalities for regressions with i.i.d.\ errors with sub-Gaussian tails; wong2017lasso consider $\beta$-mixing series with exponential tails; wu2016performance, han2017high, and chernozhukov2019lasso establish oracle inequalities for causal Bernoulli shifts with independent innovations and polynomial tails under the functional dependence measure of wu2005nonlinear; see also medeiros2016l1 and medeiros2017adaptive for results on the adaptive LASSO based on the triplex tail inequality for mixingales of jiang2009uniform.
Despite these efforts, there is no complete estimation and prediction theory for high-dimensional time series regressions under the assumptions comparable to the classical GMM and QML estimators. For instance, the best currently available results are too restrictive for the MIDAS projection model, which is typically an example of a causal Bernoulli shift with dependent innovations. Moreover, the mixing processes with polynomial tails that are especially relevant for the financial and macroeconomic time series have not been properly treated due to the fact that the sharp Fuk-Nagaev inequality was not available in the relevant literature until recently. The Fuk-Nagaev inequality, see fuk1971probability, describes the concentration of sums of random variables with a mixture of the sub-Gaussian and the polynomial tails. It provides sharp estimates of tail probabilities unlike Markov's bound in conjunction with the Marcinkiewicz-Zygmund or Rosenthal's moment inequalities.
This paper fills these gaps in the literature relying on the Fuk-Nagaev inequality for $\tau$-mixing processes of babiietalinference and establishes the nonasymptotic and asymptotic estimation and prediction properties of the sg-LASSO projections under weak tail conditions and potential misspecification. The class of $\tau$-mixing processes is fairly rich covering the $\alpha$-mixing processes, causal linear processes with infinitely many lags of $\beta$-mixing processes, and nonlinear Markov processes; see dedecker2004coupling,dedecker2005new for more details, as well as carrasco2002mixing and francq2019garch for mixing properties of various processes encountered in time series econometrics. Our weak tail conditions require at least $4+\epsilon$ finite moments for covariates, while the number of finite moments for the error process can be as low as $2+\nu$, provided that covariates are sufficiently integrable. From the theoretical point of view, we impose {\it approximate sparsity}, relaxing the assumption of exact sparsity of the projection coefficients and allowing for other forms of misspecification (see giannone2018economic for further discussion on the topic of sparsity). Lastly, we cover the LASSO and the group LASSO as special cases.
The rest of the paper is organized as follows. Section (ref) presents the setting of (potentially mixed frequency) high-dimensional time series regressions. Section (ref) characterizes nonasymptotic estimation and prediction accuracy of the sg-LASSO estimator for $\tau$-mixing processes with polynomial tails. We report on a Monte Carlo study in Section (ref) which provides further insights regarding the validity of our theoretical analysis in small sample settings typically encountered in empirical applications. Section (ref) covers the empirical application. Conclusions appear in Section (ref).
\paragraph{Notation:} For a random variable $X\in\ensuremath{\mathbf{R}}$, let $\|X\|_q=(\ensuremath{\mathbb{E}}|X|^q)^{1/q}$ be its $L_q$ norm with $q\geq 1$. For $p\in\ensuremath{\mathbf{N}}$, put $[p] = \{1,2,\dots,p\}$. For a vector $\Delta\in\ensuremath{\mathbf{R}}^p$ and a subset $J\subset [p]$, let $\Delta_J$ be a vector in $\ensuremath{\mathbf{R}}^p$ with the same coordinates as $\Delta$ on $J$ and zero coordinates on $J^c$. Let $\mathcal{G}$ be a partition of $[p]$ defining the group structure, which is assumed to be known to the econometrician. For a vector $\beta\in\ensuremath{\mathbf{R}}^p$, the sparse-group structure is described by a pair $(S_0,\mathcal{G}_0)$, where $S_0=\{j\in[p]:\;\beta_j\ne 0 \}$ and $\mathcal{G}_0 = \left\{G\in\mathcal{G}:\; \beta_{G} \ne 0\right\}$ are the support and respectively the group support of $\beta$. We also use $|S|$ to denote the cardinality of arbitrary set $S$. For $b\in\ensuremath{\mathbf{R}}^p$, its $\ell_q$ norm is denoted as $|b|_q = \left(\sum_{j\in[p]}|b_j|^q\right)^{1/q}$ for $q\in[1,\infty)$ and $|b|_\infty = \max_{j\in[p]}|b_j|$ for $q=\infty$. For $\mathbf{u},\mathbf{v}\in\ensuremath{\mathbf{R}}^T$, the empirical inner product is defined as $\langle \mathbf{u},\mathbf{v}\rangle_T = T^{-1}\sum_{t=1}^T u_tv_t$ with the induced empirical norm $\|.\|_T^2=\langle.,.\rangle_T=|.|_2^2/T$. For a symmetric $p\times p$ matrix $A$, let $\mathrm{vech}(A)\in\ensuremath{\mathbf{R}}^{p(p+1)/2}$ be its vectorization consisting of the lower triangular and the diagonal elements. For $a,b\in\ensuremath{\mathbf{R}}$, we put $a\vee b = \max\{a,b\}$ and $a\wedge b = \min\{a,b\}$. Lastly, we write $a_n\lesssim b_n$ if there exists a (sufficiently large) absolute constant $C$ such that $a_n\leq C b_n$ for all $n\geq 1$ and $a_n\sim b_n$ if $a_n\lesssim b_n$ and $b_n\lesssim a_n$.
Let $\{y_t:t\in[T]\}$ be the target low frequency series observed at integer time points $t\in[T]$. Predictions of $y_t$ can involve its lags as well as a large set of covariates and lags thereof. In the interest of generality, but more importantly because of the empirical relevance we allow the covariates to be sampled at higher frequencies - with same frequency being a special case. More specifically, let there be $K$ covariates $\{x_{t-(j-1)/m,k},j\in[m],t\in[T],k\in[K]\}$ possibly measured at some higher frequency with $m\geq 1$ observations for every $t$ and consider the following regression model
where $\phi(L) = I - \rho_1L - \rho_2L^2 - \dots - \rho_JL^J$ is a low-frequency lag polynomial and $\psi(L^{1/m};\beta_k)x_{t,k} = 1/m \sum_{j=1}^m\beta_{j,k}x_{t-(j-1)/m,k}$ is a high-frequency lag polynomial. For $m$ = 1, we have a standard autoregressive distributed lag (ARDL) model, which is the workhorse regression model of the time series econometrics literature. Note that the polynomial $\psi(L^{1/m}; \beta_k)x_{t,k}$ involves the same $m$ number of high-frequency lags for each covariate $k\in[K]$, which is done for the sake of simplicity and can easily be relaxed; see Section (ref).
The ARDL-MIDAS model (using the terminology of andreou2013should) features $J+1+m\times K$ parameters. In the big data setting with a large number of covariates sampled at high-frequency, the total number of parameters may be large compared to the effective sample size or even exceed it. This leads to poor estimation and out-of-sample prediction accuracy in finite samples. For instance, with $m$ = 3 (quarterly/monthly setting) and 35 covariates at 4 lagged quarters, we need to estimate $m\times K=420$ parameters. At the same time, say the post-WWII quarterly GDP growth series has less than 300 observations.
The LASSO estimator, see tibshirani1996regression, offers an appealing convex relaxation of a difficult nonconvex best subset selection problem. It allows increasing the precision of predictions via the selection of sparse and parsimonious models. In this paper, we focus on the structured sparsity with additional dimensionality reductions that aim to improve upon the unstructured LASSO estimator in the time series setting.
First, we parameterize the high-frequency lag polynomial following the MIDAS regression or the distributed lag econometric literature (see ghysels2006predicting) as
where $\beta_k$ is $L$-dimensional vector of coefficients with $L\leq m$ and $\omega:[0,1]\times \ensuremath{\mathbf{R}}^L\to \ensuremath{\mathbf{R}}$ is some weight function. Second, we approximate the weight function as
where $\{w_l:\; l=1,\dots,L \}$ is a collection of functions, called the dictionary. The simplest example of the dictionary consists of algebraic power polynomials, also known as almon1965distributed polynomials in the time series regression analysis literature. More generally, the dictionary may consist of arbitrary approximating functions, including the classical orthogonal bases of $L_2[0,1]$; see Online Appendix Section (ref) for more examples. Using orthogonal polynomials typically reduces the multicollinearity and leads to better finite sample performance. It is worth mentioning that the specification with dictionaries deviates from the standard MIDAS regressions and leads to a computationally attractive convex optimization problem, cf. FMO13.
The size of the dictionary $L$ and the number of covariates $K$ can still be large and the approximate sparsity is a key assumption imposed throughout the paper. With the approximate sparsity, we recognize that assuming that most of the estimated coefficients are zero is overly restrictive and that the approximation error should be taken into account. For instance, the weight function may have an infinite series expansion, nonetheless, most can be captured by a relatively small number of orthogonal basis functions. Similarly, there can be a large number of economically relevant predictors, nonetheless, it might be sufficient to select only a smaller number of the most relevant ones to achieve good out-of-sample forecasting performance. Both model selection goals can be achieved with the LASSO estimator. However, the LASSO does not recognize that covariates at different (high-frequency) lags are temporally related.
In the baseline model, all high-frequency lags (or approximating functions once we parameterize the lag polynomial) of a single covariate constitute a group. We can also assemble all lag dependent variables into a group. Other group structures could be considered, for instance combining various covariates into a single group, but we will work with the simplest group setting of the aforementioned baseline model. The sparse-group LASSO (sg-LASSO) allows us to incorporate such structure into the estimation procedure. In contrast to the group LASSO, see yuan2006model, the sg-LASSO promotes sparsity between and within groups, and allows us to capture the predictive information from each group, such as approximating functions from the dictionary or specific covariates from each group.
To describe the estimation procedure, let $\mathbf{y}$ = $(y_{1},\dots,y_T)^\top,$ be a vector of dependent variable and let $\mathbf{X}$ = $(\iota, \mathbf{y}_{1},\dots,\mathbf{y}_{J},Z_1W,\dots,Z_KW),$ be a design matrix, where $\iota = (1,1,\dots,1)^\top$ is a vector of ones, $\mathbf{y}_{j} = (y_{1-j},\dots, y_{T-j})^\top$, $Z_k = (x_{k,t-{(j-1)/m}})_{t\in[T],j\in[m]}$ is a $T\times m$ matrix of the covariate $k\in[K]$, and $W=\left(w_l\left((j-1)/m\right)/m\right)_{j\in[m],l\in[L]}$ is an $m\times L$ matrix of weights. In addition, put $\beta$ = $(\beta_0^\top,\beta_1^\top,\dots,\beta_K^\top)^\top$, where $\beta_0 = (\rho_0,\rho_1,\dots,\rho_J)^\top$ is a vector of parameters pertaining to the group consisting of the intercept and the autoregressive coefficients, and $\beta_k\in\ensuremath{\mathbf{R}}^L$ denotes parameters of the high-frequency lag polynomial pertaining to the covariate $k\geq 1$. Then, the sparse-group LASSO estimator, denoted $\hat\beta$, solves the penalized least-squares problem
with a penalty function that interpolates between the $\ell_1$ LASSO penalty and the group LASSO penalty
where $\|b\|_{2,1} = \sum_{G\in\mathcal{G}}|b_G|_2$ is the group LASSO norm and $\mathcal{G}$ is a group structure (partition of $[p]$) specified by the econometrician. Note that estimator in equation ((ref)) is defined as a solution to the convex optimization problem and can be computed efficiently, e.g., using an appropriate coordinate descent algorithm; see simon2013sparse.
The amount of penalization in equation ((ref)) is controlled by the regularization parameter $\lambda>0$ while $\alpha\in[0,1]$ is a weight parameter that determines the relative importance of the sparsity and the group structure. Setting $\alpha=1$, we obtain the LASSO estimator while setting $\alpha = 0$, leads to the group LASSO estimator, which is reminiscent of the elastic net. In Figure (ref) we illustrate the geometry of the penalty function for different groupings and different values of $\alpha$ covering (a) LASSO with $\alpha$ = 1, (b) group LASSO with one group, $\alpha$ = 0, and two sg-LASSO cases (c) one group and (d) two groups both with $\alpha$ = 0.5. In practice, groups are defined by a particular problem and are specified by the econometrician, while $\alpha$ can be fixed or selected jointly with $\lambda$ in a data-driven way such as using the cross-validation.
We focus on a generic high-dimensional linear projection model with a countable number of regressors
where $x_{t,0}=1$ and $m_t \triangleq \sum_{j=0}^\infty x_{t,j}\beta_j$ is a well-defined random variable. In particular, to ensure that $y_t$ is a well-defined economic quantity, we need $\beta_j\downarrow 0$ sufficiently fast, which is a form of the approximate sparsity condition, see belloni2018high. This setting nests the high-dimensional ARDL-MIDAS projections described in the previous section and more generally may allow for other high-dimensional time series models. In practice, given a (large) number of covariates, lags thereof, as well as lags of the dependent variable, denoted $x_t\in\ensuremath{\mathbf{R}}^p$, we would approximate $m_t$ with $x_t^\top\beta \triangleq \sum_{j=0}^px_{t,j}\beta_j$, where $p<\infty$ and the regression coefficient $\beta\in\ensuremath{\mathbf{R}}^p$ could be sparse. Importantly, our settings allows for the approximate sparsity as well as other forms of misspecification and the main result of the following section allows for $m_t\ne x_t^\top\beta$.
Using the setting of equation ((ref)), for a sample $(y_t,x_t)_{t=1}^T$, write
where $\mathbf{y}=(y_1,\dots,y_T)^\top$, $\mathbf{m} = (m_1,\dots, m_T)^\top$, and $\mathbf{u}=(u_1,\dots,u_T)^\top$. The approximation to $\mathbf{m}$ is denoted $\mathbf{X}\beta$, where $\mathbf{X} = (x_1,\dots,x_T)^\top$ is a $T\times p$ matrix of covariates and $\beta=(\beta_1,\dots,\beta_p)^\top$ is a vector of unknown regression coefficients.
We measure the time series dependence with $\tau$-mixing coefficients. For a $\sigma$-algebra $\mathcal{M}$ and a random vector $\xi\in\ensuremath{\mathbf{R}}^l$, put
where $\mathrm{Lip}_1=\left\{f:\ensuremath{\mathbf{R}}^l\to\ensuremath{\mathbf{R}}:\;|f(x)-f(y)|\leq |x-y|_1\right\}$ is a set of $1$-Lipschitz functions. Let $(\xi_t)_{t\in\ensuremath{\mathbf{Z}}}$ be a stochastic process and let $\mathcal{M}_t=\sigma(\xi_t,\xi_{t-1},\dots)$ be its canonical filtration. The $\tau$-mixing coefficient of $(\xi_t)_{t\in\ensuremath{\mathbf{Z}}}$ is defined as
If $\tau_k\downarrow0$ as $k\to\infty$, then the process $(\xi_t)_{t\in\ensuremath{\mathbf{Z}}}$ is called $\tau$-mixing. The $\tau$-mixing coefficients were introduced in dedecker2004coupling as dependence measures weaker than mixing. Note that the commonly used $\alpha$- and $\beta$-mixing conditions are too restrictive for the linear projection model with an ARDL-MIDAS process. Indeed, a causal linear process with dependent innovations is not necessary $\alpha$-mixing; see also andrews1984non for an example of AR(1) process which is not $\alpha$-mixing. Roughly speaking, $\tau$-mixing processes are somewhere between mixingales and $\alpha$-mixing processes and can accommodate such counterexamples. At the same time, sharp Fuk-Nagaev inequalities are available for $\tau$-mixing processes which to the best of our knowledge is not the case for the mixingales or near-epoch dependent processes; see babiietalinference.
dedecker2004coupling,dedecker2005new discuss how to verify the $\tau$-mixing property for causal Bernoulli shifts with dependent innovations and nonlinear Markov processes. It is also worth comparing the $\tau$-mixing coefficient to other weak dependence coefficients. Suppose that $(\xi_t)_{t\in\ensuremath{\mathbf{Z}}}$ is a real-valued stationary process and let $\gamma_k = \|\ensuremath{\mathbb{E}}(\xi_{k}|\mathcal{M}_0) - \ensuremath{\mathbb{E}}(\xi_{k})\|_1$ be its $L_1$ mixingale coefficient. Then we clearly have $\gamma_k\leq \tau_k$ and it is known that
where $Q$ is the generalized inverse of $x\mapsto \Pr(|\xi_0|>x)$ and $G$ is the generalized inverse of $x\mapsto\int_0^x Q(u)\ensuremath{\mathrm{d}} u$; see babiietalinference, Lemma A.1.1. Therefore, the $\tau$-mixing coefficient provides a sharp control of autocovariances similarly to the $L_1$ mixingale coefficients, which in turn can be used to ensure that the long-run variance of $(\xi_t)_{t\in\ensuremath{\mathbf{Z}}}$ exists. The $\tau$-mixing coefficient is also bounded by the $\alpha$-mixing coefficient, denoted $\alpha_k$, as follows
where the first inequality follows by dedecker2004coupling, Lemma 7 and the second by H\"{o}lder's inequality with $q,r\geq 1$ such that $q^{-1}+r^{-1}=1$. It is worth mentioning that the mixing properties for various time series models in econometrics, including GARCH, stochastic volatility, or autoregressive conditional duration are well-known; see, e.g., carrasco2002mixing, francq2019garch, babii2019commercial; see also dedecker2007weak for more examples and a comprehensive comparison of various weak dependence coefficients.
In this section, we introduce the main assumptions for the high-dimensional time series regressions and study the estimation and prediction properties of the sg-LASSO estimator covering the LASSO and the group LASSO estimators as special cases. The following assumption imposes some mild restrictions on the stochastic processes in the high-dimensional regression equation ((ref)).
It is worth mentioning that the stationarity condition is not essential and can be relaxed to the existence of the limiting variance of partial sums at costs of heavier notations and proofs. Condition (i) requires that covariates have at least $4$ finite moments, while the number of moments required for the error process can be as low as $2+\epsilon$, depending on the integrability of covariates. Therefore, (i) may allow for heavy-tailed distributions commonly encountered in financial and economic time series, e.g., asset returns and volatilities. Given the integrability in (i), (ii) requires that the $\tau$-mixing coefficients decrease to zero sufficiently fast; see Online Appendix, Section (ref) for moments and $\tau$-mixing coefficients of ARDL-MIDAS. It is known that the $\beta$-mixing coefficients decrease geometrically fast, e.g., for geometrically ergodic Markov chains, in which case (ii) holds for every $a,b>0$. Therefore, (ii) allows for relatively persistent processes.
For the support $S_0$ and the group support $\mathcal{G}_0$ of $\beta$, put
For some $c_0>0$, define $\mathcal{C}(c_0) \triangleq \left\{\Delta\in\ensuremath{\mathbf{R}}^p:\; \Omega_1(\Delta)\leq c_0\Omega_0(\Delta) \right\}$. The following assumption generalizes the restricted eigenvalue condition of bickel2009simultaneous to the sg-LASSO estimator and is imposed on the population covariance matrix $\Sigma = \ensuremath{\mathbb{E}}[\mathbf{X}^\top\mathbf{X}/T]$.
Recall that if $\Sigma$ is a positive definite matrix, then for all $\Delta\in\ensuremath{\mathbf{R}}^p$, we have $\Delta^\top\Sigma\Delta \geq \gamma|\Delta|_2^2$, where $\gamma$ is the smallest eigenvalue of $\Sigma$. Therefore, in this case Assumption (ref) is trivially satisfied because $|\Delta|_2^2\geq \sum_{G\in\mathcal{G}_0}|\Delta_{G}|_2^2$. The positive definiteness of $\Sigma$ is also known as a completeness condition and Assumption (ref) can be understood as its weak version; see babii2017completeness and references therein. It is worth emphasizing that $\gamma>0$ in Assumption (ref) is a universal constant independent of $p$, which is the case, e.g., when $\Sigma$ is a Toeplitz matrix or a spiked identity matrix. Alternatively, we could allow for $\gamma\downarrow 0$ as $p\to \infty$, in which case the term $\gamma^{-1}$ would appear in our nonasymptotic bounds slowing down the speed of convergence, and we may interpret $\gamma$ as a measure of ill-posedness in the spirit of econometrics literature on ill-posed inverse problems; see carrasco2007linear.
The value of the regularization parameter is determined by the Fuk-Nagaev concentration inequality, appearing in the Online Appendix, see Theorem (ref).
The regularization parameter in Assumption (ref) is determined by the persistence of the data, quantified by $a$, and the tails, quantified by $\varsigma = qr/(q+r)$. This dependence is reflected in the dependence-tails exponent $\kappa$. The following result describes the nonasymptotic prediction and estimation bounds for the sg-LASSO estimator, see Online Appendix (ref) for the proof.
Theorem (ref) provides nonasymptotic guarantees for the estimation and prediction with the sg-LASSO estimator reflecting potential misspecification. In the special case of the LASSO estimator ($\alpha=1$), we obtain the counterpart to the result of belloni2012sparse for the LASSO estimator with i.i.d.\ data taking into account that we may have $m_t\ne x_t^\top\beta$. At another extreme, when $\alpha=0$, we obtain the nonasymptotic bounds for the group LASSO allowing for misspecification which to the best of our knowledge are new, cf. negahban2012unified and van2016estimation. We call $s_\alpha$ the effective sparsity constant. This constant reflects the benefits of the sparse-group structure for the sg-LASSO estimator that can not be deduced from the results currently available for the LASSO or the group LASSO.
Next, we consider the asymptotic regime, in which the misspecification error vanishes when the sample size increases as described in the following assumption.
The following corollary is an immediate consequence of Theorem (ref).
If the effective sparsity constant $s_\alpha$ is fixed, then $p=o(T^{\kappa - 1})$ is a sufficient condition for the prediction and estimation errors to vanish, whenever $\mu\geq 2\kappa-1$. In this case Assumption (ref) (ii) is vacuous. More generally, $s_\alpha$ is allowed to increase slowly with the sample size. Convergence rates in Corollary (ref) quantify the effect of tails and persistence of the data on the prediction and estimation accuracies of the sg-LASSO estimator. In particular, lighter tails and less persistence allow us to handle a larger number of covariates $p$ compared to the sample size $T$. In particular $p$ can increase faster than $T$, provided that $\kappa>2$.
We assess via simulations the out-of-sample predictive performance (forecasting and nowcasting), and the MIDAS weights recovery of the sg-LASSO with dictionaries. We benchmark the performance of our novel sg-LASSO setup against two alternatives: (a) unstructured, meaning standard, LASSO with MIDAS, and (b) unstructured LASSO with the unrestricted lag polynomial. The former allows us to assess the benefits of exploiting group structures, whereas the latter focuses on the advantages of using dictionaries in a high-dimensional setting.
To assess the predictive performance and the MIDAS weight recovery, we simulate the data from the following DGP:
where $u_t\sim_{i.i.d.}N(0,\sigma_u^2)$ and the DGP for covariates $\{x_{k,t-(j-1)/m}:j\in[m],k\in[K]\}$ is specified below. This corresponds to a target of interest $y_t$ driven by two autoregressive lags augmented with high frequency series, hence, the DGP is an ARDL-MIDAS model. We set $\sigma^2_u=1$, $\rho_1=0.3$, $\rho_2=0.01$, and take the number of relevant high frequency regressors $K$ = 3. In some scenarios we also decrease the signal-to-noise ratio by setting $\sigma^2_u$ = 5. We are interested in quarterly/monthly data, and use four quarters of data for the high frequency regressors so that $m$ = 12. We rely on a commonly used weighting scheme in the MIDAS literature, namely $\omega(s;\beta_k)$ for $k$ = 1, 2 and 3 are determined by beta densities respectively equal to $\mathrm{Beta}(1,3),\mathrm{Beta}(2,3)$, and $\mathrm{Beta}(2,2)$; see ghysels2007midas or ghysels2019estimating, for further details.
The high frequency regressors are generated as either one of the following:
For the AR simulation design, we initiate the processes as $x_0\sim N\left(0,\sigma^2/(1-\rho^2)\right)$ and $y_0\sim N\left(0,\sigma^2(1-\rho_2)/((1+\rho_2)((1-\rho_2)^2-\rho_1^2))\right).$ For the VAR, the initial value of $(y_t)$ is the same, while $X_0 \sim N(0,I_K)$. In all cases, the first 200 observations are treated as burn-in. In the estimation procedure, we add 7 noisy covariates which are generated in the same way as the relevant covariates and use 5 low-frequency lags. The empirical models use a dictionary which consists of Legendre polynomials up to degree $L = 10$ shifted to the $[0,1]$ interval with the MIDAS weight function approximated as in equation ((ref)). The sample size is $T\in\{50, 100, 200\},$ and for all the experiments we use 5000 simulation replications.
We assess the performance of different methods by modifying the assumptions on the error terms of the high-frequency process $\mathbf{\varepsilon}_h$, considering multivariate high-frequency processes, changing the degree of Legendre polynomials $L$, increasing the noise level of the low-frequency process $\sigma^2_u$, using only half of the high-frequency lags in predictive regressions, and adding a larger number of noisy covariates. In the case of VAR high-frequency process, we set $\Phi$ to be block-diagonal with the first $5\times 5$ block having entries $0.15$ and the remaining $5\times 5$ block(s) having entries $0.075$.
We estimate three different LASSO-type regression models. In the first model, we keep the weighting function unconstrained, and therefore we estimate 12 coefficients per high-frequency covariate using the unstructured LASSO estimator. We denote this model LASSO-U-MIDAS (inspired by the U-MIDAS of foroni2015unrestricted). In the second model we use MIDAS weights together with the unstructured LASSO estimator; we call this model LASSO-MIDAS. In this case, we estimate $L+1$ number of coefficients per high-frequency covariate. The third model applies the sg-LASSO estimator together with MIDAS weights. Groups are defined as in Section (ref); each low-frequency lag and high-frequency covariate is a group, therefore, we have $K+5$ groups. We select the value of tuning parameters $\lambda$ and $\alpha$ using the 5-fold cross-validation, defining folds as adjacent blocks over the time dimension to take into account the time series dependence. This model is denoted sg-LASSO-MIDAS.
For regressions with aggregated data, we consider: (a) Flow aggregation (FLOW): $x_{k,t}^A$ = $1/m\sum_{j=1}^mx_{k,t-(j-1)/m}$, (b) Stock aggregation (STOCK): $x_{k,t}^A$ = $x_{k,t}$, and (c) Middle high-frequency lag (MIDDLE): single middle value of the high-frequency lag with ties solved in favor of the most recent observation (i.e., we take a single $6^{\rm th}$ lag if $m=12$). In these cases, the models are estimated using the OLS estimator, which is unfeasible when the number of covariates becomes equal to the sample size and we leave results blank in this case.
Detailed results are reported in the Online Appendix. Tables (ref)--(ref), cover the average mean squared forecast errors for one-step-ahead forecasts and nowcasts. The sg-LASSO with MIDAS weighting (sg-LASSO-MIDAS) outperforms all other methods in all simulation scenarios. Importantly, both sg-LASSO-MIDAS and unstructured LASSO-MIDAS with nonlinear weight function approximations perform much better than all other methods when the sample size is small ($T=50$). In this case, sg-LASSO-MIDAS yields the largest improvements over alternatives, in particular, with a large number of noisy covariates (bottom-right block). These findings are robust to increases in the persistence parameter of covariates $\rho$ from 0.2 to 0.7. The LASSO without MIDAS weighting has typically large forecast errors. Comparing across simulation scenarios, all methods seem to perform worse with heavy-tailed or persistent covariates. In these cases, however, the impact on the sg-LASSO-MIDAS method is lesser compared to the other methods. This simulation evidence supports our theoretical results and findings in the empirical application. Lastly, forecasts using flow-aggregated covariates seem to perform better than other simple aggregation methods in all simulation scenarios, but significantly worse than the sg-LASSO-MIDAS.
In Table (ref)--(ref) we report additional results for the estimation accuracy of the weight functions. In Figure (ref)--(ref), we plot the estimated weight functions from several methods. The results indicate that the LASSO without MIDAS weighting can not accurately recover the weights in small samples and/or low signal-to-noise ratio scenarios. Using Legendre polynomials improves the performance substantially and the sg-LASSO seems to improve even more over the unstructured LASSO.
We nowcast US GDP with macroeconomic, financial, and textual news data. Details regarding the data sources appear in the Online Appendix Section (ref). Regarding the macro data, we rely on 34 series used in the Federal Reserve Bank of New York nowcast model, discarding two series ("PPI: Final demand" and "Merchant wholesalers: Inventories") due to very short samples; see bok2018macroeconomic for more details regarding this data.
For all macro data, we use real-time vintages, which effectively means that we take all macro series with a delay. For example, if we nowcast the first quarter of GDP one month before the quarter ends, we use data up to the end of February, and therefore all macro series with a delay of one month that enter the model are available up to the end of January. We use Legendre polynomials of degree three for all macro covariates to aggregate twelve lags of monthly macro data. In particular, let $x_{t+(h+1-j)/m,k}$ be $k^{{\text{th}}}$ covariate at quarter $t$ with $m=3$ months per quarter and $h=2-1=1$ months into the quarter (2 months into the quarter minus 1 month due to publication delay), where $j = 1,2,\dots,12$ is the monthly lag. We then collect all lags in a vector
and aggregate $X_{t,k}$ using a dictionary $W$ consisting of Legendre polynomials, so that $X_{t,k}W$ defines as a single group for the sg-LASSO estimator.
In addition to macro and financial data, we also use the textual analysis data. We take 76 news attention series from bybee2019structure and use Legendre polynomials of degree two to aggregate three monthly lags of each news attention series. Note that the news attention series are used without a publication delay, that is, for the one-month horizon, we take the series up to the end of the second month. Moreover, the bybee2019structure news topic models involve rolling samples, avoiding look ahead biases when used in our nowcasts.
We compute the predictions using a rolling window scheme. The first nowcast is for 2002 Q1, for which we use fifteen years (sixty quarters) of data, and the prediction is computed using 2002 January (2-month horizon) February (1-month), and March (end of the quarter) data. We calculate predictions until the sample is exhausted, which is 2017 Q2, the last date for which news attention data is available. As indicated above, we report results for the 2-month, 1-month, and the end-of-quarter horizons. Our target variable is the first release, i.e., the advance estimate of real GDP growth. We tune sg-LASSO-MIDAS regularization parameters $\lambda$ and $\alpha$ using 5-fold cross-validation, defining folds as adjacent blocks over the time dimension to take into account the time series nature of the data. Finally, we follow the literature on nowcasting real GDP and define our target variable to be the annualized growth rate.
Let $x_{t,k}$ be the $k$-th high-frequency covariate at time \(t\). The general ARDL-MIDAS predictive regression is
where \(\phi(L)\) is the low-frequency lag polynomial, \(\mu\) is the regression intercept, and $\psi(L^{1/m};\beta_k)x_{tk},k=1,\dots,K$ are lags of high-frequency covariates. Following Section (ref), the high-frequency lag polynomial is defined as
where for $k^{\rm th}$ covariate, \(h_k\) indicates the number of leading months of available data in the quarter \(t\), $q_k$ is the number of quarters of covariate lags, and we approximate the weight function $\omega$ with the Legendre polynomial. For example, if $h_k=1$ and $q_k=4$, then we have $1$ month of data into a quarter and use $q_km=12$ monthly lags for a covariate $k$.
We benchmark our predictions against the simple AR(1) model, which is considered to be a reasonable starting point for short-term GDP growth predictions. We focus on predictions of our method, sg-LASSO-MIDAS, with and without financial data combined with series based on the textual analysis. One natural comparison is with the publicly available Federal Reserve Bank of New York, denoted NY Fed, model implied nowcasts.
We adopt the following strategy. First, we focus on the same series that are used to calculate the NY Fed nowcasts. The purpose here is to compare {\it models} since the data inputs are the same. This means that we compare the performance of dynamic factor models (NY Fed) with that of machine learning regularized regression methods (sg-LASSO-MIDAS). Next, we expand the data set to see whether additional financial and textual news series can improve the nowcast performance.
In Table (ref), we report results based on real-time macro data used for the NY Fed model, see bok2018macroeconomic. The results show that the sg-LASSO-MIDAS performs much better than the NY Fed nowcasts at the longer, i.e.\ 2-month, horizon. Our method significantly beats the benchmark AR(1) model for all the horizons, and the accuracy of the nowcasts improve with the horizon. Our end-of-quarter and 1-month horizon nowcasts are similar to the NY Fed ones, with the sg-LASSO-MIDAS being slightly better numerically but not statistically. We also report the average Superior Predictive Ability test of quaedvlieg2019multi over all three horizons and the result reveals that the improvement of the sg-LASSO-MIDAS model versus the NY Fed nowcasts is significant at the 5% significance level.
The comparison in Table (ref) does not fully exploit the potential of our methods, as it is easy to expand the data series beyond the small number used by the NY Fed nowcasting model. In Table (ref) we report results with additional sets of covariates which are financial series, advocated by andreou2013should, and textual analysis of news. In total, the models select from 118 series -- 34 macro, 8 financial, and 76 news attention series. For the moment we focus only on the first three columns of the table. At the longer horizon of 2 months, the method seems to produce slightly worse nowcasts compared to the results reported in Table (ref) using only macro data. However, we find significant improvements in prediction quality for the shorter 1-month and end-of-quarter horizons. In particular, a significant increase in accuracy relative to NY Fed nowcasts appears at the 1-month horizon. We report again the average Superior Predictive Ability test of quaedvlieg2019multi over all three horizons with the same result that the improvement of sg-LASSO-MIDAS versus the NY Fed nowcasts is significant at the 5% significance level. Lastly, we report results for several alternatives, namely, PCA-OLS, ridge, LASSO, and Elastic Net, using the unrestricted MIDAS scheme. Our approach produces more accurate nowcasts compared to these alternatives.
Table (ref) also features an entry called SPF (median), where we report results for the median survey of professional nowcasts for the 1-month horizon, and analyze how the model-based nowcasts compare with the predictions using the publicly available Survey of Professional Forecasters maintained by the Federal Reserve Bank of Philadelphia. We find that the sg-LASSO-MIDAS model-based nowcasts are similar to the SPF-implied nowcasts. We also find that the NY Fed nowcasts are significantly worse than the SPF.
In Figure (ref) we plot the cumulative sum of squared forecast error (CUMSFE) loss differential of sg-LASSO-MIDAS versus NY Fed nowcasts for the three horizons. The CUMSFE is computed as
for model $M1$ versus $M2.$ A positive value of $\text{CUMSFE}_{t,t+k}$ means that the model $M1$ has larger squared forecast errors compared to model $M2$ up to $t+k,$ and negative values imply the opposite. In our case, $M1$ is the New York Fed prediction error, while $M2$ is the sg-LASSO-MIDAS model. We observe persistent gains for the 2- and 1-month horizons throughout the out-of-sample period. When comparing the sg-LASSO-MIDAS results with additional financial and textual news series versus those based on macro data only, we see a notable improvement at the 1-month horizon and a more modest one at the end-of-quarter horizons. In Figure (ref), we plot the average CUMSFE for the 1-month and end-of-quarter horizons and observe that the largest gains of additional financial and textual news data are achieved during the financial crisis.
The result in Figure (ref) prompts the question whether our results are mostly driven by this unusual period in our out-of-sample data. To assess this, we turn our attention again to the last two columns of Table (ref) reporting diebold1995comparing test statistics which exclude the financial crisis period. Compared to the tests previously discussed, we find that the results largely remain the same, but some alternatives seem to slightly improve (e.g.\ LASSO or Elastic Net). Note that this also implies that our method performs better during periods with heavy-tailed observations, such as the financial crisis. It should also be noted that overall there is a slight deterioration of the average Superior Predictive Ability test over all three horizons when we remove the financial crisis.
In Figure (ref), we plot the fraction of selected covariates by the sg-LASSO-MIDAS model when we use the macro, financial, and textual analysis data. For each reference quarter, we compute the ratio of each group of variables relative to the total number of covariates. In each subfigure, we plot the three different horizons. For all horizons, the macro series are selected more often than financial and/or textual data. The number of selected series increases with the horizon, however, the pattern of denser macro series and sparser financial and textual series is visible for all three horizons. The results are in line with the literature -- macro series tend to co-move, hence we see a denser pattern in the selection of such series, see e.g. bok2018macroeconomic. On the other hand, the alternative textual analysis data appear to be very sparse, yet still important for nowcasting accuracy, see also thorsrud2018words.
This paper offers a new perspective on the high-dimensional time series regressions with data sampled at the same or mixed frequencies and contributes more broadly to the rapidly growing literature on the estimation, inference, forecasting, and nowcasting with regularized machine learning methods. The first contribution of the paper is to introduce the sparse-group LASSO estimator for high-dimensional time series regressions. An attractive feature of the estimator is that it recognizes time series data structures and allows us to perform the hierarchical model selection within and between groups. The classical LASSO and the group LASSO are covered as special cases.
To recognize that the economic and financial time series have typically heavier than Gaussian tails, we use a new Fuk-Nagaev concentration inequality, from babiietalinference, valid for a large class of $\tau$-mixing processes, including $\alpha$-mixing processes commonly used in econometrics. Building on this inequality, we establish the nonasymptotic and asymptotic properties of the sparse-group LASSO estimator.
Our empirical application provides new perspectives on applying machine learning methods to real-time forecasting, nowcasting, and monitoring with time series data, including unconventional data, sampled at different frequencies. To that end, we introduce a new class of MIDAS regressions with dictionaries linear in the parameters and based on orthogonal polynomials with lag selection performed by the sg-LASSO estimator. We find that the sg-LASSO outperforms the unstructured LASSO in small samples and conclude that incorporating specific data structures should be helpful in various applications.
Our empirical results also show that the sg-LASSO-MIDAS using only macro data performs statistically better than NY Fed nowcasts at 2-month horizons and overall for the 1- and 2-month and end-of-quarter horizons. This is a comparison involving the same data and, therefore, pertains to models. This implies that machine learning models are, for this particular case, better than the state space dynamic factor models. When we add the financial data and the textual news data, a total of 118 series, we find significant improvements in prediction quality for the shorter 1-month and end-of-quarter horizons.
We thank participants at the Financial Econometrics Conference at the TSE Toulouse, the JRC Big Data and Forecasting Conference, the Big Data and Machine Learning in Econometrics, Finance, and Statistics Conference at the University of Chicago, the Nontraditional Data, Machine Learning, and Natural Language Processing in Macroeconomics Conference at the Board of Governors, the AI Innovations Forum organized by SAS and the Kenan Institute, the 12th World Congress of the Econometric Society, and seminar participants at the Vanderbilt University, as well as Harold Chiang, Jianqing Fan, Jonathan Hill, Michele Lenza, and Dacheng Xiu for comments. We are also grateful to the referees and the editor whose comments helped us to improve our paper significantly. All remaining errors are ours.
\spacingset{1.45} \setcounter{page}{1} \setcounter{section}{0} \setcounter{equation}{0} \setcounter{table}{0} \setcounter{figure}{0}