EconBase
← Back to paper

Factor-augmented sparse MIDAS regressions with an application to nowcasting

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

63,939 characters · 11 sections · 78 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Factor-augmented sparse MIDAS regressions with an application to nowcasting

\def\spacingset#1{ {#1}} \spacingset{1}

\if11 \fi \if01 {

center[center omitted — 105 chars of source]

} \fi

abstractThis article investigates factor-augmented sparse MIDAS (Mixed Data Sampling) regressions for high-dimensional time series data, which may be observed at different frequencies. Our novel approach integrates sparse and dense dimensionality reduction techniques. We derive the convergence rate of our estimator under misspecification due to the MIDAS approximation error, $\tau$-mixing dependence, and polynomial tails. Our method's finite sample performance is assessed via Monte Carlo simulations. We apply the methodology to nowcasting U.S. GDP growth and demonstrate that it outperforms both sparse regression and standard factor-augmented regression during the COVID-19 pandemic. These findings indicate that the growth through this period was influenced by both idiosyncratic (sparse) and common (dense) shocks. \textcolor{black}{The approach is implemented in the midasml R package, available on CRAN.}

{\it Keywords:} factor models, high-dimensional data, mixed-frequency data, nowcasting, COVID-19 pandemic

\spacingset{1.5}

Introduction

In the current economic climate, where accurate nowcasting of macroeconomic variables such as U.S. GDP is critical for policymakers, addressing the challenges posed by high-dimensional mixed-frequency data has become a key econometric concern. Nowcasting involves predicting the current or near-future values of low-frequency outcome variables using high-frequency data—a task complicated by the need for effective variable selection and control over parameter proliferation. Motivated by the challenges of the data used in nowcasting, this paper introduces a novel estimation method specifically designed for high-dimensional mixed-frequency data.

High dimensionality—where the number of predictors is large and can even exceed the number of observations—poses challenges for traditional estimation methods. Two prominent strategies have emerged: factor-augmented regressions and sparse methods. Factor-augmented regressions, as introduced by stock1999forecasting, stock2002forecasting, reduce dimensionality by extracting a small set of latent factors via principal component analysis (PCA), under the assumption of a dense factor structure. These approaches are widely used for their efficiency and interpretability, especially when most predictors are assumed to contribute meaningfully. In contrast, sparse methods like LASSO tibshirani1996regression, bickel2003simultaneous focus on selecting a small subset of relevant variables, promoting parsimony. In mixed-frequency settings—where variables are observed at different intervals, such as weekly financial and monthly macroeconomic series—additional complexities arise. MIDAS regressions andreou2010regression address this by applying weighting schemes that directly relate high-frequency predictors to low-frequency outcomes. \textcolor{black}{Recent work adapts sparse and factor-augmented regressions to this context: siliverstovs2017short,uematsu2019high propose LASSO approaches using an UMIDAS scheme}, babii2022machine use a sparse-group LASSO on MIDAS-weighted variables to exploit natural group structures, while marcellino2010factor, andreou2013should, and koh2023inference examine factor MIDAS regressions, where factors estimated by PCA are aggregated via nonlinear MIDAS schemes.

In this paper, we propose the novel approach of factor-augmented sparse MIDAS regression for forecasting with mixed-frequency high-dimensional data, which harnesses the advantages of both LASSO and factor-augmented regression. Our method works as follows. First, we construct MIDAS-weighted variables using pre-set polynomials, enabling the flexible handling of mixed-frequency data. Second, we estimate factors through PCA applied on the MIDAS-weighted variables. Third, we estimate a high-dimensional linear regression model using the sparse-group LASSO estimator in which we include the MIDAS-weighted covariates and the extracted factors as predictors. This integrated approach captures both dense signals (through the extracted factors) and the sparse signals distributed across different frequencies. By applying the sparse-group LASSO estimator, we effectively model hierarchical relationships among predictors, as it accommodates both grouped variables and sparsity within groups babii2022machine.

We establish the theoretical properties of our estimator, deriving rates of convergence that account for potential misspecification due to the MIDAS approximation or approximate sparsity. Our theoretical framework accommodates $\tau$-mixing processes with polynomial tails, which are often encountered in financial and macroeconomic time series. This setting allows our method to remain robust even when facing heavy-tailed distributions or complex dependencies, as commonly observed in nowcasting applications. Note that it is important to use $\tau$-mixing—which was introduced by dedecker2004coupling—as opposed to more traditionally used $\beta$- or $\alpha$-mixing in our setting. The reason is that both $\beta$ and $\alpha$ mixing coefficients are too strong for the process we consider, namely, an autoregressive distributed lag (ARDL) model, see babii2022machine for a formal analysis of $\tau$-mixing processes within the context of the ARDL model. In the present paper, we extend the modeling approach babii2022machine to a factor-augmented model and show its validity under the $\tau$-mixing assumption. Relative to babii2022machine, the challenge is to obtain results when factors are estimated by PCA. \textcolor{black}{The theoretical results show that, when the number of variables used to estimate the factors is sufficiently large, factor estimation does not affect the convergence rate of the estimator. This reinforces the reliability of the method and strengthens the confidence with which researchers can apply it in practice.}

Numerical results underscore the practical effectiveness of the proposed factor-augmented sparse MIDAS approach. Monte Carlo simulations show its superior finite-sample performance relative to existing methods. Empirically, we apply the method to nowcast U.S. GDP growth using panels of weekly financial and monthly macroeconomic data. By modeling macroeconomic predictors through a factor structure and integrating both sparse and dense signals, our approach outperforms sparse-only and standard factor-augmented regressions—particularly during the COVID-19 pandemic. These results suggest that large shock during pandemic was driven by both idiosyncratic (sparse) and common (dense) components. Importantly, we find that factor augmentation offers limited gains in stable periods, where sparsity alone suffices, but becomes essential during episodes of heightened macroeconomic uncertainty. This distinction highlights the value of our method in adapting to changing economic conditions and improving forecast accuracy when it matters most.

\noindentLiterature review. Our theoretical results contribute to the literature on high-dimensional econometrics. babii2022machine study the properties of the sparse-group LASSO with mixed-frequency data (without estimated factors). We extend their work by proving an estimation error bound and deriving convergence rates while allowing for estimated factors and $\tau$-mixing dependence. To obtain these results, we establish convergence rates for factor estimation using PCA under $\tau$-mixing, thereby providing novel insights into the statistical properties of PCA in time series settings. Considering $\tau$-mixing processes is important, as highlighted by dedecker2004coupling, babii2022machine among others, as other commonly used dependence assumptions appear to be too strong for the models we consider. We also want to stress the differences between the factor MIDAS method of marcellino2010factor,andreou2013should,ferrara2019nowcasting, koh2023inference and our approach. Their approach assumes a factor model on a high-dimensional, high-frequency set of variables. They implement a two-step procedure: first, the latent factors are estimated via PCA, and then these factors are projected onto the target variable using a low-dimensional nonlinear MIDAS regression. Instead, our approach uses a high-dimensional linear MIDAS regression, which also allows the many original predictors to play a role as long as they exhibit a sparse coefficient pattern.

Several recent studies have investigated high-dimensional factor-augmented sparse regression models outside the mixed-frequency context. fan2023bridging propose a general framework that simultaneously accommodates sparse and dense signals. Building on this, beyhum2024testing introduce a fully data-driven procedure to test for the presence of sparse components in such models, showing that sparsity is statistically significant for most FRED-MD series. krampe2021factor propose a factor-augmented sparse VAR and demonstrate its superior forecasting performance. Sparse-plus-dense structures have also been explored in high-dimensional panel data settings by hansen2019factor and vogt2022cce, who incorporate factor structures alongside sparse idiosyncratic components. \textcolor{black}{Our work is the first to consider factor-augmented sparse regression in a mixed-frequency framework.} The previously mentioned papers do not consider MIDAS weighting, their theoretical results do not allow for approximation errors, and they rely on stronger dependence measures than $\tau$-mixing.

\textcolor{black}{Alternative approaches have been proposed to go beyond purely sparse or factor-augmented regressions. A related strand of the literature on targeted predictors—such as bai2008forecasting and bessec2013short—adopts a two-step procedure: strong predictors are first selected using soft or hard thresholding, and factors are then extracted from the selected subset. This strategy emphasizes dense modeling conditional on pre-selection, in contrast to the integrated sparse-plus-dense framework we develop. franjic2024nowcasting propose a sparse mixed-frequency dynamic factor model that directly combines sparsity with a dynamic factor structure, aiming to exploit both high-frequency signals and cross-sectional dependence. mogliani2021bayesian cast the sparse MIDAS framework in a Bayesian setting using structured shrinkage priors, i.e., group LASSO type, while kohns2025flexible consider time-varying coefficients with structured group-shrinkage priors.}

Finally, our work contributes to the growing literature on nowcasting, which faced significant challenges during the COVID-19 pandemic due to unprecedented economic shocks. foroni2022forecasting adjust nowcasts using prediction errors from the financial crisis, while huber2023nowcasting employ non-parametric Bayesian regression trees for robust GDP forecasts. carriero2024addressing address pandemic-related outliers using outlier-augmented stochastic volatility BVARs, and hauzenberger2024nowcasting propose Bayesian MIDAS regressions with Gaussian Processes to capture nonlinear effects. Additionally, diebold2020covid evaluates the effectiveness of the ADS index in tracking real-time economic activity during the pandemic. Our findings show that factor-augmented sparse methods, grounded in a rigorous statistical framework, offer a robust alternative for nowcasting in periods of heightened uncertainty. Empirically, we demonstrate that dense macroeconomic information becomes especially valuable during crises, yielding significantly improved predictions throughout the COVID-19 period.

\noindentNotation. For an integer $N\in \mathbb{N}$, let $[N]=\{1,\dots, N\}$. The transpose of a $n_1 \times n_2$ matrix $H$ is written $H^{\top}$. Its $k^{th}$ singular value is $\sigma_k(H)$. Let us also define the Euclidean norm $\left\| H\right\|_2^2=\sum_{i=1}^{n_1} \sum_{j=1}^{n_2}H_{i,j}^2$ and the sup-norm $\|H\|_\infty=\max\limits_{i\in[n_1], j\in[n_2]}|H_{i,j}|$. The quantity $n_1 \vee n_2$ is the maximum of $n_1$ and $n_2$. For a real-valued random variable $\zeta$ and $g>0$, we let ${\left\vert\kern-0.4ex\left\vert\kern-0.4ex\left\vert \zeta \right\vert\kern-0.4ex\right\vert\kern-0.4ex\right\vert}_g={\mathbb{E}}[|\zeta|^g]^\frac1g$. For a $d$-dimensional random vector $\zeta$, we define ${\left\vert\kern-0.4ex\left\vert\kern-0.4ex\left\vert \zeta \right\vert\kern-0.4ex\right\vert\kern-0.4ex\right\vert}_g=\sup\limits_{u\in{\mathbb{R}}^d: \ \|u\|_2\le 1}{\left\vert\kern-0.4ex\left\vert\kern-0.4ex\left\vert u^\top \zeta \right\vert\kern-0.4ex\right\vert\kern-0.4ex\right\vert}_g$. For a vector $v\in{\mathbb{R}}^N$, and $G\subset[N]$, we let $v_G$ be the vector in ${\mathbb{R}}^N$ such that $(v_G)_i=v_i$ if $i\in G$ and $(v_G)_i=0$ otherwise.

Factor-augmented sparse MIDAS regression

In this section, we describe our novel factor-augmented sparse MIDAS regression. Let $\bigl\{y_t,\ t \in[T]\bigr\}$ be the low-frequency outcome variable we want to predict where $T$ denotes the sample size. For example, in our application, we consider U.S. real GDP growth, which is measured quarterly. We use information in two sets of high-frequency regressors denoted $$\widetilde{z}=\left\{\widetilde{z}_{t-(j-1)/m_{z},k},\ t\in [T], k \in[K_{z}],j\in[m_{z}]\right\}$$ and $$\widetilde{x}=\left\{\widetilde{x}_{t-(j-1)/m_{x},k},\ t\in [T],k \in[K_{x}],j\in [m_{x}]\right\}.$$ Note that, throughout the paper, we use a “$\sim$" to denote high-frequency variables. In total, there are $K=K_{z}+K_{x}$ regressors which we observe more frequently than the target variable $y_t$, e.g., weekly and monthly. The quantities $m_{z}$ and $m_{x}$ determine the number of high-frequency lags we use of variables in $\widetilde{z}$ or $\widetilde{x}$, respectively. In our application, $\widetilde{z}$ is a high-dimensional weekly panel of financial variables, while $\widetilde{x}$ is a high-dimensional monthly panel of macroeconomic predictors. We distinguish these data types because they differ in frequency, and because we assume that $\widetilde{x}$ follows a factor model whereas $\widetilde{z}$ does not.

The nowcasting task of $y_t$ by $\widetilde{z}$ and $\widetilde{x}$ is a high-dimensional problem because (i) $\widetilde{z}$ and $\widetilde{x}$ contain many variables and (ii) the higher sampling frequency of the variables $\widetilde{z}$ and $\widetilde{x}$ creates parameter proliferation because many lags can be introduced in the model. This calls for the use of several complementary dimension reduction approaches.

The model is

equation[equation omitted — 76 chars of source]

where

equation[equation omitted — 437 chars of source]

is the true regression function. Note that our theory does not rely on the definition of $m_t$ in (ref) but only requires that the true $m_t$ is well approximated by our predictors. Hence, our theory implicitly allows the true model to contain nonlinearities. The definition of $m_t$ in (ref) serves to justify our approach heuristically. Here, $\rho=(\rho_0,\dots,\rho_J)^\top\in{\mathbb{R}}^{J+1}$ are the autoregressive coefficients, $J$ is the number of lags, $\omega_{z,k},\ k\in[K_z]$, $\omega_{x,k},\ k\in [K_x]$ and $ \omega_{f,r},\ r\in[K_f]$ are some functions from $[0,1]$ to ${\mathbb{R}}$ determining the linear regression coefficients of the different high-frequency lags, and $$\widetilde{f}=\left\{\widetilde{f}_{t-(j-1)/m_{x},r},\ t\in [T],r \in[K_{f}],j\in[m_{x}]\right\}$$ are factors linked to $\widetilde{x}$ through the factor model:

equation[equation omitted — 170 chars of source]

for $t\in [T],k\in[K_{x}],j\in[m_{x}],$ where $\left\{\widetilde{b}_{k ,r},\ k\in[K_{x}],r \in[K_{f}]\right\}$ are nonrandom loadings and $\left\{\widetilde{u}_{t-(j-1)/m_{x},k},\ t\in [T],k \in[K_{x}],j\in[m_{x}] \right\}$ are error terms. Note that we only assume that $\widetilde{x}$ (and not $\widetilde{z}$) follows a factor model. This is because, following the literature, it is uncommon to assume a factor structure among financial variables.

The model in (ref)-(ref) contains both the original regressors and the factors. Estimating all the coefficients of the high-frequency lags in (ref) is a complex econometric task. To overcome this difficulty, we approximate the high-frequency lag polynomials through a MIDAS approach. Specifically, we consider a dictionary $(w_d)_{d\geq 0}$ of functions from $[0,1]$ to ${\mathbb{R}}$ for which there exist linear approximation coefficients $\left\{ \widetilde\alpha_{k,d},\ k\in[K_z], d\in[D]\right\}$, $\left\{ \widetilde\beta_{k,d},\ k\in[K_x], d\in[D]\right\}$ and $\left\{ \widetilde\gamma_{r,d},\ r\in[K_f], d\in[D]\right\}$ such that, for all $u\in[0,1]$,

equation[equation omitted — 363 chars of source]

where $D$ is the number of polynomials from the dictionary used to approximate high-frequency lag coefficients linearly. In practice, we use Legendre polynomials up to degree $3$ as approximating functions, which implies that $D=4$.

For all $t\in[T]$, let us define the MIDAS-weighted variables

equation[equation omitted — 553 chars of source]

Let $p_z=DK_z$, $p_x=DK_x$, $R=DK_f$, $z_{t}=(z_{t,1},\dots,z_{t,p_z})^\top$, $x_t=(x_{t,1},\dots, x_{t,p_x})^\top$, $u_t= (u_{t,1},\dots,u_{t,p_x})^\top$, $f_t=(f_{t,1},\dots,f_{t,R})^\top$. Due to (ref), we have

equation[equation omitted — 167 chars of source]

where $$ a_t =m_t - \left(\rho_0 + \sum_{j=1}^{\text{J}}\rho_j y_{t-j} + z_{t}^\top\alpha +x_t^\top \beta + f_{t}^\top\gamma\right)\approx 0,$$ is the approximation error and

equation*[equation* omitted — 256 chars of source]

are the linear regression coefficients. Note that no multicollinearity issue between $x_t$ and $f_t$ arises in model (ref). This is because by (ref), equation (ref) can be rewritten as $y_t= \rho_0 + \sum_{j=1}^{\text{J}}\rho_j y_{t-j} + z_{t}^\top\alpha +u_t^\top \beta + f_{t}^\top(\gamma +B^\top\beta) +a_t +\varepsilon_t,$ and $u_t$ and $f_t$ are not multicollinear. Remark also that that $x_t$ follows the factor model

equation[equation omitted — 83 chars of source]

where $B=(b_1,\dots, b_{p_x})^\top$ is a $p_x\times R$ matrix, with $b_{D(k-1)+d}= (\widetilde b_{k,1},\dots,\widetilde{b}_{k,R})^\top,\ k\in[K_z],d\in[D]$.

To present the estimation procedure, we rewrite (ref) in a matrix form. We introduce further notation. Let $w_t=(1,y_{t-1},\dots,y_{t-J}, z_t^\top,x_t^\top)^\top $, $W=(w_1,\dots,w_T)^\top$, $F=(f_1,\dots,f_T)^\top$, $A=(a_1,\dots,a_T)^\top$ and $\mathcal{E} = (\varepsilon_{1}, \dots, \varepsilon_{T})^\top$. $W$ is a $T\times p$ matrix where $p=1+J+p_z+p_x$, $F$ is a $T\times R$ matrix, while $A$ and $\mathcal{E}$ are $T\times 1$ vectors. The regression model (ref), therefore, can be written in a matrix form

equation[equation omitted — 77 chars of source]

where $\delta=(\rho_0,\rho_1,\dots,\rho_J,\alpha^\top,\beta^\top)^\top.$ To proceed with the estimation of (ref), we estimate the factors $F$ by PCA. We let the columns of $\widehat{F}/\sqrt{T}$ be the eigenvectors corresponding to the leading $R$ eigenvalues of $XX^\top$. In practice, the number of factors $R$ is unknown and we replace it by an estimator $\widehat{R}$. In our empirical application to nowcasting, we use the growth ratio estimator of ahn2013eigenvalue, but there exist various alternatives in the literature bai2002determining,onatski2010determining,bai2019rank,fan2022estimating. Finally, we use the estimator

equation[equation omitted — 244 chars of source]

where $\lambda\in{\mathbb{R}}$ is the tuning parameter, and the sparse-group LASSO (sg-LASSO) penalty function is

equation*[equation* omitted — 65 chars of source]

which interpolates between the $\ell_1$ LASSO and group LASSO norms. Here, $\mu\in[0,1]$ is the interpolation parameter between the two norms. The latter is defined as $\|d\|_{2,1}= \sum_{G\in\mathcal{G}}\|d_G\|_2$. Such a penalty is beneficial in MIDAS regression settings because it encourages within-group sparsity, which allows for accurate weight function estimation, as well as across-group sparsity, which controls the dimensionality of regressors babii2022machine. In practice, $\lambda$ and $\mu$ are selected via cross-validation. Specifically, we use 5-fold cross-validation, defining folds as adjacent blocks of sub-samples over the time dimension to take into account the time series dependence. We experimented with the number of folds and the overall conclusions remain the same. We stick with 5-fold cross-validation to match the sg-LASSO-MIDAS approach of babii2022machine, which we benchmark against. For the nowcasting application, following babii2022machine, we choose a group structure corresponding to the original high-frequency covariates, that is $\mathcal{G}=\{G_1,\dots,G_{1+K_z+K_x}\}$, where $G_1 =[J+1]$ and $G_\ell =\{J+1+(\ell-1)d:\ d\in[D]\} $ for $\ell\in\{2,\dots, 1+K_z+K_x\}$. This structure promotes sparsity across the original predictors.

This is a sparse plus dense dimension reduction approach. Indeed, sparsity is induced through the penalty $\Omega(\cdot)$ and the presence of the factors, which are linear combinations of the variables in $X$, introduces a dense pattern in the estimation procedure. By combining three different dimension reduction approaches (MIDAS, factor models/PCA and LASSO), we propose a method appropriate for the challenges at hand.

Lastly, we note that in our empirical application, we also explore factor-augmented and sparse MIDAS regressions as alternatives to our approach. Specifically, the factor-augmented MIDAS regression assumes that coefficients related to the sparse component are zero and employs OLS for estimation. Formally, this approach solves

equation[equation omitted — 144 chars of source]

On the other hand, the sparse MIDAS regression imposes zero restrictions on $\gamma$ coefficients, for which we apply the sg-LASSO estimator:

equation[equation omitted — 150 chars of source]

These methods are referred to as FAMIDAS (equation (ref)) and sg-LASSO-MIDAS (equation (ref)), respectively. Our main approach in equation (ref), which uses both sparse and dense signals, is denoted as sg-LASSO-FAMIDAS.

Theory

Let us now investigate the theoretical properties of our approach. Our theoretical analysis allows for the approximation error $a_t$, which can be due to the MIDAS approximation, approximate sparsity or nonlinearites. Notably, we do not need to assume that the true model $m_t$ is linear or sparse (or takes the form (ref)), but only that it is well approximated by a sparse linear combination of the predictors that we use. We consider an asymptotic regime where $T\to \infty$, and $p_x$ and $p_z$ go to infinity as a function of $T$. The number of factors is fixed with $T$. It would be possible to let it grow with $T$, but this would greatly complicate the theoretical analysis and, more importantly, the exposition of results.\footnote{See, e.g., beyhum2022factor,freeman2023linear for an analysis with a growing number of factors.} The group structure $\mathcal{G}$ and the parameters $\delta$ and $\gamma$ are not random but can vary with $T$. For this theoretical analysis, $\mu$ is fixed with $T$. We make the following assumptions.

AssumptionIt holds that \begin{enumerate}[(i)] • For all $t\in[T]$, ${\mathbb{E}}[f_tf_t^\top]=I_R$ and $B^\top B$ is diagonal; • All the eigenvalues of the $R\times R$ matrix $p_x^{-1} B^\top B$ are bounded away from $0$ and $\infty$ as $p_x\to \infty$; • $\|B\|_\infty=O(1)$. \end{enumerate}
AssumptionThe following holds: \begin{enumerate}[(i)] • For all $t\in[T],k\in[p_x], \ell\in [p], r\in[R]$, it holds that \begin{align*}&{\mathbb{E}}[u_{t,k}]={\mathbb{E}}[u_{t,k}f_{t,r}]={\mathbb{E}}[u_{t,k}\varepsilon_t]={\mathbb{E}}[f_{t,r}\varepsilon_t]={\mathbb{E}}[w_{t,\ell}\varepsilon_t]=0;\end{align*} • There exist $q > 2$ and $C_1>0$, such that, for all $t\in[T]$, we have $${\left\vert\kern-0.4ex\left\vert\kern-0.4ex\left\vert u_{t} \right\vert\kern-0.4ex\right\vert\kern-0.4ex\right\vert}_{2q}+ {\left\vert\kern-0.4ex\left\vert\kern-0.4ex\left\vert \varepsilon_{t} \right\vert\kern-0.4ex\right\vert\kern-0.4ex\right\vert}_{2q}+ {\left\vert\kern-0.4ex\left\vert\kern-0.4ex\left\vert f_{t} \right\vert\kern-0.4ex\right\vert\kern-0.4ex\right\vert}_{2q}+ {\left\vert\kern-0.4ex\left\vert\kern-0.4ex\left\vert w_{t} \right\vert\kern-0.4ex\right\vert\kern-0.4ex\right\vert}_{2q}+{\left\vert\kern-0.4ex\left\vert\kern-0.4ex\left\vert p_x^{-1/2}\sum_{k\in[p_x]}u_{t,k}b_k \right\vert\kern-0.4ex\right\vert\kern-0.4ex\right\vert}_{2q} \le C_1;$$ • There exists $C_2>0$ such that, for all $t\in[T]$, $\max\limits_{k\in [p]}\|E[w_{t,k}u_t]\|_2 \le C_2$ and $$\textcolor{black}{\max_{s,t\in[T]}{\mathbb{E}}\left[\left\{p_x^{-1/2}\left(u_s^\top u_t-{\mathbb{E}}[u_s^\top u_t]\right)\right\}^2\right]\le C_2.}$$ \end{enumerate}

Assumption (ref) is the same as Assumption 3 in fan2023bridging. Its conditions (ref) and (ref) form a strong factor assumption as in bai2003inferential. Assumption (ref) contains restrictions on the moments of random variables. Its condition (ref) is a no-correlation condition. It would hold if (a) $u_t$ is mean zero and uncorrelated with $f_t,\varepsilon_t$ and (b) $\varepsilon_t$ is mean zero and uncorrelated with $f_t,w_t$. Condition (ref) assumes that variables have strictly more than $4$ finite moments, having, therefore, polynomial tails. Finally, condition (ref) bounds some moments. \textcolor{black}{By the inequality of Cauchy-Schwarz, ${\mathbb{E}}\left[\left\{p_x^{-1/2}\left(u_s^\top u_t-{\mathbb{E}}[u_s^\top u_t]\right)\right\}^2\right]$ is bounded if $\max_{s,t\in[T]}{\mathbb{E}}\left[\left\{p_x^{-1/2}\left(u_s^\top u_t-{\mathbb{E}}[u_s^\top u_t]\right)\right\}^4\right]$ is bounded. A bound on the latter quantity is standard in the literature on factor models, see fan2013large,fan2023bridging.} The bound on $\max\limits_{k\in [p]}\|E[w_{t,k}u_t]\|_2$ is new to the literature. It essentially assumes that $w_{t,k}$ can only be correlated with a few entries of $u_t$.

Next, to control the time series dependence of the variables, we use the concept of $\tau$-mixing. The definition of $\tau$-mixing coefficients is as follows. For a $\sigma$-algebra $\mathcal{M}$ and a random vector $\xi\in{\mathbb{R}}^l$, let

equation*[equation* omitted — 154 chars of source]

where $\mathrm{Lip}_1=\left\{f:{\mathbb{R}}^l\to{\mathbb{R}}:\;|f(x)-f(y)|\leq |x-y|_1\right\}$ is a set of $1$-Lipschitz functions. Let $\{\xi_t\}_t$ be a stochastic process and let $\mathcal{M}_t=\sigma(\xi_t,\xi_{t-1},\dots)$ be its filtration. The $\tau$-mixing coefficients of $\{\xi_t\}_t$ are defined as

equation*[equation* omitted — 147 chars of source]

The process $\{\xi_t\}_t$ is called $\tau$-mixing when $\tau_s\downarrow0$ as $s\to\infty$ . As noted in the introduction, using $\tau$-mixing conditions is less restrictive than $\alpha$ or $\beta$-mixing. The latter are too strong for the process we consider, namely, an autoregressive distributed lag (ARDL) model, see dedecker2004coupling and babii2022machine.

AssumptionThe process $$\left\{\left(f_{t}^\top, u_t^\top,\varepsilon_t, w_t^\top, \left(p_x^{-1/2}\sum_{k\in[p_x]}u_{t,k}b_k\right)^\top\right)^\top\right\}_t $$ is stationary, and, for some constants $C_3>0$ and $a>(q-1)/(q-2)$, its $\tau$-mixing coefficients, denoted by $\tau_s$ satisfy $\tau_s\le C_3s^{-a}$ for all $s\in\mathbb{N}$.

We are ready to state some results on the estimation of the factors. To do so, we introduce further notation. Let $V$ be the matrix with the $R$ largest eigenvalues of $T^{-1}XX^\top$. Next, we define $H=T^{-1}V^{-1}\widehat{F}^\top FB^\top B$, which is a rotation matrix such that the factors rotated by this matrix are well estimated.\footnote{In the Online Appendix, we prove that $V$ is invertible with probability going to $1$, see Lemma (ref).} This rotation matrix is the same as the one used in the literature to show consistency of estimated factors, see bai2003inferential and bai2006confidence. Moreover, we let $$h_T= \left(\frac{p}{T^{\kappa-1}}\right)^{1/\kappa}\vee \sqrt{\frac{\log(2p)}{T}},$$ where $\kappa= \frac{(a+1)q-1}{a+q-1} $. We have the following lemma.

LemmaUnder Assumptions (ref), (ref) and (ref), we have \begin{enumerate}[(i)] • $\left\|\widehat{F}-FH^\top \right\|_2=O_P\left(\sqrt{\frac{T}{p_x}} +1\right)$; • $\left\|\left(\widehat{F}-FH^\top\right)^\top \mathcal{E}\right\|_2=O_P\left(\sqrt{\frac{T}{p_x}} +1 \right);$$\left\|\left(\widehat{F}-FH^\top\right)^\top W\right\|_\infty=O_P\left(\left(\frac{T}{p_x}+\sqrt{\frac{T}{p_x}}\right) \left(\sqrt{p_x}h_T+1\right) \right) .$ \end{enumerate}

Lemma (ref) shows consistency of factor estimates. To the best of our knowledge, this is the first result of this type under $\tau$-mixing. The rates in statements (ref) and (ref) are the same as in the standard literature, see bai2003inferential, bai2006confidence and fan2023bridging. Hence the presence of $\tau$-mixing does not change the rate of convergence in $\ell_2$-norm. It only appears in the rate of convergence in $\ell_\infty$-norm such as statement (ref).

Now, let $\mathcal{S}_0=\{k\in[p]:\ \delta_k\ne 0\}$ and $\mathcal{G}_0=\{G\in\mathcal{G}:\ \delta_G\ne 0\}$ be respectively be the support and the group support of $\delta$. We define the effective sparsity $\sqrt{s_\mu}=\mu \sqrt{|\mathcal{S}_0|} +(1-\mu)\sqrt{|\mathcal{G}_0|}$ and the maximum group size $G^*=\max_{G\in\mathcal{G}}|G|$. To avoid notational complexities arising when $s_\mu=0$, we assume that $s_\mu \ge 1$ (otherwise, it suffices to replace $s_\mu$ by $s_\mu\vee 1$ in all statements). We also define $$\Sigma = {\mathbb{E}}[w_tw_t^\top]- {\mathbb{E}}[w_tf_t^\top] {\mathbb{E}}[w_tf_t^\top]^\top,$$ which is the population Gram matrix of the $w_t$ once they have been projected on the vector space orthogonal to the factors $f_t$. We make the following assumption.

AssumptionThe following holds: \begin{enumerate}[(i)] • $\|\gamma\|_2=O(1)$, $p=o(T^{(\kappa-1)/2})$ and $$G^*s_\mu \left[\left(\left(\frac{p^2}{T^{\kappa-1}}\right)^{1/\kappa}\vee \sqrt{\frac{\log(2p)}{T}}\right)+\frac{1}{\sqrt{p_x}}\right]=o(1);$$ • There exists $\nu>0$ independent of $T$ such that $\sigma_{p}(\Sigma)\ge \nu.$ \end{enumerate}

Assumption (ref) (ref) is a rate condition. It assumes that $p$ cannot grow too quickly with $T$, this is needed because of the presence of time series dependence and polynomial tails. It also restricts the rate at which the sparsity level $s_\mu$ can grow. In the literature on the LASSO estimator, it is typically assumed that $s_*\sqrt{\log(2p)/T}=o(1)$, where $s_*$ is the sparsity level bickel2003simultaneous. Our condition is stronger because of time series dependence, polynomial tails and the factors are estimated. Condition (ref) of Assumption (ref) assumes that the smallest eigenvalue of $\Sigma$ is bounded from below. This allows to show that the classical restricted eigenvalue condition of the LASSO literature holds bickel2003simultaneous.

Then, we state a bound on the prediction error of our estimator, that is the $\ell_2$-norm of the prediction $W\widehat{\delta}+\widehat{F}\widehat{\gamma}$ minus the target prediction $ W\delta+F\gamma$. To state the bound, we need further notation. Let $P_{\widehat{F}}=\frac1T\widehat{F}\widehat{F}^\top$ be the orthogonal projector on the vector space spanned by the columns of $\widehat{F}$ and $M_{\widehat{F}}=I_T-P_{\widehat{F}}$. We also introduce $\widetilde{W} = M_{\widehat{F}}W $ which corresponds to $W$ projected on the orthogonal of the vector space spanned by the columns of $\widehat{F}$. Finally, let $\Omega^*(\cdot)$ be the dual norm of $\Omega(\cdot)$, that is, for any $d\in{\mathbb{R}}^p$, $\Omega^*(d)=\sup\limits_{v\in{\mathbb{R}}^p}\frac{v^\top d}{\Omega(d)}.$

TheoremUnder Assumptions (ref), (ref), (ref) and (ref), and letting $\lambda\ge 2\Omega^*\left(\frac1T\widetilde{W}^\top \mathcal{E}\right)$, with probability going to $1$, we have \begin{align*}&\frac{1}{\sqrt{T}}\left\|W\widehat{\delta}+\widehat{F}\widehat{\gamma}-W\delta-F\gamma\right\|_2\\ &\le \left(\frac{288s_\mu\lambda^2}{\nu }+ \frac{4}{T}\left(\|A\|_2+\left\|M_{\widehat{F}}F\gamma\right\|_2\right)^2\right)^{1/2} +\frac{1}{\sqrt{T}}\left( \left\|M_{\widehat{F}}F\gamma\right\|_2+ \left\| P_{\widehat{F}} \mathcal{E}\right\|_2\right). \end{align*}

As standard with estimation error bounds on the LASSO bickel2003simultaneous, we need that the penalty term is large enough, that is $\lambda\ge 2\Omega^*\left(\frac1T\widetilde{W}^\top \mathcal{E}\right)$. The quantity $\Omega^*\left(\frac1T\widetilde{W}^\top \mathcal{E}\right)$ is the effective noise of the problem lederer2021estimating. In practice, $\lambda$ is chosen by cross-validation. As in the literature, the bound here depends on the approximation error $\|A\|_2$, see bickel2003simultaneous and babii2022machine. Compared to the traditional LASSO literature, our bound contains two additional terms: $\left\|M_{\widehat{F}}F\gamma\right\|_2$ and $\left\| P_{\widehat{F}} \mathcal{E}\right\|_2$. They are present because we do not use the true $F$ for estimation but rather its estimate $\widehat{F}$. The term $\left\|M_{\widehat{F}}F\gamma\right\|_2$ is the $\ell_2$-norm of the projection of $F\gamma$ on the orthogonal of the vector space generated by the columns of $\widehat{F}$, while the quantity $\left\| P_{\widehat{F}} \mathcal{E}\right\|_2$ is the $\ell_2$-norm of the projection of $\mathcal{E}$ on the vector space generated by the columns of $\widehat{F}$. Using in particular Lemma (ref), we can bound $\left\|M_{\widehat{F}}F\gamma\right\|_2$, $\left\| P_{\widehat{F}} \mathcal{E}\right\|_2$ and $\Omega^*\left(\frac1T\widetilde{W}^\top \mathcal{E}\right)$ in probability. This allows to obtain the following corollary, which states the rate of convergence of the prediction error of our estimator.

CorollaryUnder Assumptions (ref), (ref), (ref) and (ref), and letting $\lambda= 2\Omega^*\left(\frac1T\widetilde{W}^\top \mathcal{E}\right)$, we have \begin{align*}\frac{1}{T}\left\|W\widehat{\delta}+\widehat{F}\widehat{\gamma}- W\delta-F\gamma\right\|_2^2&= O_P\left(s_\mu h_T^2+ \|A\|_2^2 +\frac{1}{p_x} \right). \end{align*}

When $s_\mu h_T^2p_x\to \infty$, the term $\frac{1}{p_x}$ in the rate of convergence of Corollary (ref) becomes negligible and our rate of convergence becomes the same as that of babii2022machine. Hence, the condition $s_\mu h_T^2p_x\to \infty$ guarantees that the error in estimating the factors does not influence the rate of convergence of the prediction errors. This condition requires that $p_x$ grows sufficiently quickly with respect to $T$. The standard rate of convergence of the LASSO estimator in prediction error would replace $h_T^2$ by $\log(p)/T$. We have $h_T^2$ in our rate because of the presence of $\tau$-mixing and polynomial tails.

Monte Carlo simulations

In this section, we present a Monte Carlo study to examine the finite sample performance of our proposed method. We simulate data using the following data generating process. We use model (ref)-(ref), where $\rho_0 = 0.5$, $J = 2$, $\rho_1 = 0.2$, $\rho_2 = 0.1$ are fixed across all scenarios. We consider the sample sizes $T \in \{50, 100, 200\}$, which are representative of typical settings in nowcasting applications. We study several scenarios, with different degrees of cross-sectional dependence and fatness of the tails of the variables.

The error term $\varepsilon_t$ follows a mean-zero AR(1) process with an autocorrelation coefficient of $0.4$ and variance $0.1$. The error in this AR(1) process follows either a Gaussian or a student-$t(5)$ distribution.

The panel $\widetilde{z}$ represents high-frequency regressors that enter the model only through a sparse pattern. These variables are simulated at a weekly frequency assuming $m_z= 13$. We set $p_z = 30$ and the weekly variables are generated as AR(1) processes with an autocorrelation coefficient of $\rho_z=0.4$ and mean-zero error terms that have either a Gaussian or a student-$t(5)$ distribution. The error terms have a cross-sectional covariance matrix specified as $(1-\rho_z^2)\Sigma$, where $\Sigma^z_{ij} = c_z^{|i-j|}$ for $i, j \in [p_z]$, where $c_z\in\{0.1,0.4\}$ controls the degree of cross-sectional dependence among the regressors. We set $\omega_{z,1}$ and $\omega_{z,2}$ equal to the Beta$(1, 2)$ and Beta$(2, 2)$ densities, respectively. The other variables in $\widetilde{z}$ do not enter the model, that is $\omega_{z,k}\equiv 0$ for all $k\ge 3 $.

Next, the factors in $\widetilde{f}$ are monthly variables, that is $m_x=3$. We fix the number of factors to $K_f=1$ and simulate the factor as a zero-mean AR(1) process with an autoregressive coefficient of $0.4$ and a mean-zero error with unit variance which follows either a Gaussian or a student-$t(5)$ distribution. Given these factors, we generate the $p_x=130$ monthly regressors in $\widetilde{x}$ according to the factor model (ref), where the loadings $\left\{\widetilde{b}_{k ,r},\ k\in[K_{x}],r \in[K_{f}]\right\}$ are i.i.d. uniform random variables on the interval $[-1,1]$. The idiosyncratic components $\left\{\widetilde{u}_{t-(j-1)/m_{x},k},\ t\in [T],k \in[K_{x}],j\in[m_{x}] \right\}$ are modeled as zero-mean autoregressive processes with an autoregressive coefficient of $\rho_u=0.4$, with error terms following a Gaussian or a student-$t(5)$ distribution with cross-sectional covariance matrix $\Sigma\left(1-\rho_u^2\right)$. The matrix $\Sigma$ is defined as $\Sigma_{ij} = c_u^{|i-j|}$ for $i, j \in [p_x]$, where $c_u\in\{0.1,0.4\}$ controls the degree of cross-sectional dependence among the monthly regressors. For $\omega_{x,1}$ and $\omega_{x,2}$ we use Beta$(1, 2)$ and Beta$(2, 2)$ densities, respectively. For all $k\ge 3 $, $\omega_{x,k}\equiv 0$, that is only the first two variables in $\widetilde{x}$ enter the model.

For the model estimation, we use Legendre polynomials of degree 3 to approximate the Beta density functions used in the weighting process ($D=4$). This choice ensures a flexible yet accurate representation of the underlying relationships. It is important to note that the model is simulated based on Beta density functions which we aim to approximate, hence the regressors that enter the sg-LASSO estimator are estimated based on Legendre polynomials. Lastly, the the number of factors is estimated using the growth ratio estimator of ahn2013eigenvalue. The methods evaluated in our study correspond to those described in Section (ref).

Table (ref) in the Online Appendix presents the mean squared error (MSE) results of three different methods—M1 (FAMIDAS), M2 (sg-LASSO-MIDAS), and M3 (sg-LASSO-FAMIDAS)—under various simulation settings. Panel A reports results for the baseline scenario with weak cross-sectional dependence ($c_z = c_u = 0.1$), while Panel B shows results for a scenario with stronger cross-sectional dependence ($c_z = c_u = 0.4$). In both panels, the performance is evaluated under two types of error distributions: Gaussian and student-$t(5)$. The reported MSE values are based on out-of-sample nowcasts.

Across all sample sizes $T\in \{50, 100, 200\}$, M3 (sg-LASSO-FAMIDAS) consistently achieves the lowest MSE compared to M1 and M2, indicating superior predictive accuracy. As the sample size increases, all methods demonstrate improved performance (i.e., lower MSE), but the gains are most pronounced for M3. For example, in the baseline scenario (Panel A), M3’s MSE decreases from 1.1760 to 0.6133 under the Gaussian errors when $T$ increases from 50 to 200. In the stronger cross-sectional dependence scenario (Panel B), the relative performance of M3 remains robust, maintaining the lowest MSE across all settings. Notably, under the student-$t(5)$ distribution, M3 shows a significant relative improvement over M1 and M2, particularly for smaller sample sizes ($T = 50$), suggesting that the approach is more resilient to heavy-tailed errors when the true data generating process is sparse plus dense.

\textcolor{black}{In addition, we implemented designs where we set $\omega_{z,k} = \omega_{x,k} = 0$ (called the “dense DGP”) and $\omega_{f,k} = 0$ (called the “sparse DGP”). The results are reported in the Online Appendix, Tables (ref) and (ref), respectively. These results show that our proposed sparse plus dense approach outperforms sparse-only and dense-only approaches at almost no cost. Specifically, under the dense DGP, our method performs similarly to FAMIDAS, while under the sparse DGP, it achieves performance comparable to or better than sg-LASSO-MIDAS, which explicitly imposes sparsity without dense component. Note that the performance gains in the sparse DGP arise only when cross-sectional dependence is strong, likely because the factor component absorbs this dependence, thereby improving the estimation of the sparse signals—see, e.g., fan2020factor.}

Overall, the results indicate that sg-LASSO-FAMIDAS (M3) outperforms the other methods in terms of MSE, especially as the sample size increases or when errors exhibit heavy tails, highlighting its effectiveness in handling different data-generating scenarios.

Nowcasting GDP growth

For our empirical application, we focus on nowcasting current-quarter U.S. GDP growth, evaluating performance at three monthly horizons: two months before the end of the quarter ($h=2$), one month ahead ($h=1$), and at quarter-end ($h=0$). Since GDP data is typically released with a delay—often a month or more after the quarter ends—we mimick realistic nowcasting scenarios. For example, at the end of January (Q1), the model is trained on data until December (Q4), using available high-frequency data from January to nowcast Q1 GDP. The parameter $h$ indicates the number of months remaining in the target quarter. We use $m_z$ and $m_x$ to denote the number of recent high-frequency observations used from financial and macroeconomic panels, respectively. For monthly macro predictors, we include data from the current and previous quarters: $m_x = 6$ when $h=0$, $m_x = 5$ when $h=1$, and $m_x = 4$ when $h=2$. For weekly financial data, this translates to $m_z = 26$, $21$, and $17$ weeks, respectively, aligning with the number of weeks available from both the current and previous quarters.

Data

We apply our methods to nowcast the initial, or “advance," release of real U.S. GDP growth, which is typically published near the end of the first month of the following quarter. (An exception is the 2018 Q4 release, which was delayed due to a government shutdown.) Further details on this event are available from the Bureau of Economic Analysis. For our analysis, we use real-time GDP vintages from the ALFRED database maintained by the St. Louis FED. The in-sample period spans 1984 Q1 to 2007 Q4, with nowcasting beginning in 2008 Q1 and proceeding quarter by quarter using an expanding window approach.

\noindentMonthly macroeconomic regressors $\boldsymbol{\tilde{x}}$. The monthly covariates and factors are drawn from the FRED-MD real-time dataset; see mccracken2016fred for details. All variables are transformed according to the authors’ recommended procedures. We retain only those series that are consistently available across all vintages used for out-of-sample forecasting and exclude financial variables from the monthly panel. The final dataset includes 76 monthly macroeconomic series.

\noindentWeekly financial regressors $\boldsymbol{\tilde{z}}$. We use weekly financial variables primarily drawn from those used to construct the Chicago Fed’s National Financial Conditions Index (NFCI). While the NFCI is a latent factor built from weekly, monthly, and quarterly data, we focus solely on the majority of its weekly components, excluding lower-frequency series. Our final panel includes 36 weekly financial series. Since financial data is available in real time, publication delays are not a concern. However, some series have shorter histories. To address this, we apply matrix completion with nuclear-norm regularization to impute missing values, ensuring a balanced panel. This imputation is performed at the original weekly frequency; see Section D of the Online Appendix for details.

We provide the full list of series for monthly macro and weekly financial data with additional details in Section B of the Online Appendix.

Empirical results

Our candidate model set includes the AR(4) model, which we regard as the simplest and designate as our benchmark model. Consequently, our reported root mean squared forecast errors are presented relative to the AR(4). Notably, our conclusions remain consistent even when employing alternative benchmarks such as the random walk or AR(1). Our main results are based on the assumption that monthly macro covariates follow a factor model, hence entering the nowcasting equation with sparse plus dense signal. The weekly financial series are additional regressors that influence the target variables in a sparse way. We have four autoregressive terms ($J=4$). As in the simulations, we use Legendre polynomials of degree 3 as the dictionary ($D=4$). In practice, we choose all tuning parameters by cross-validation adjusting for time series dependence, see, e.g., Section 4 in babii2022machine. We estimate the number of factors using the growth ratio estimator of ahn2013eigenvalue.

Table (ref) presents results for two sub-samples: Panel A covers the full out-of-sample period from 2008 Q1 to 2022 Q2, while Panel B focuses on the pre-COVID period ending in 2019 Q4. RMSEs for the autoregressive benchmark model are reported in absolute terms, with all other results expressed relative to this benchmark. As expected, the benchmark model exhibits a notable increase in prediction errors during the COVID-19 period, reflecting its inability to incorporate high-frequency information. Our results show that the factor-augmented sparse MIDAS regression substantially outperforms both FAMIDAS (dense) and sg-LASSO-MIDAS (sparse) methods over the full sample, with the largest gains observed during the pandemic. This underscores the effectiveness of combining sparse and dense components under heightened economic uncertainty. In contrast, prior to COVID-19, the performance of the sparse and factor-augmented sparse models is comparable, suggesting that the benefits of factor augmentation are most pronounced in volatile environments. Lastly, we note that the performance differences are statistically significant, as confirmed by the average superior predictive ability (aSPA) test of quaedvlieg2021multi; see Section (ref) of the Online Appendix for details.

table[table omitted — 1,037 chars of source]

In Figure (ref), we present the square-root cumulative sum of squared forecast errors (CUMSUM) for three competing methods across two sub-samples. These graphs are designed to visualize the performance differences between the models throughout the out-of-sample period. The forecast errors are calculated for the end-of-quarter horizon, and the CUMSUM is computed as follows:

equation*[equation* omitted — 90 chars of source]

where $\hat \epsilon_{q,j}$, $j\in\{\text{FAMIDAS, sg-LASSO-MIDAS, sg-LASSO-FAMIDAS}\}$ are the out-of-sample nowcast errors. We plot the CUMSUM for the full sample period, corresponding to Panel A in Table (ref), and up to the COVID pandemic which corresponds to Panel B in the same Table.

The plots highlight the differences between the methods, particularly during the onset of the COVID-19 outbreak. Prior to the pandemic, the performances of the sg-LASSO-FAMIDAS and sg-LASSO-MIDAS models are comparable, with the former providing more accurate nowcasts during the financial crisis and the latter performing better during stable periods between crises. This suggests that factor augmentation may not significantly impact performance during stable periods, despite the need to estimate additional unpenalized regression coefficients and the factors themselves. As the COVID-19 period unfolds, both the FAMIDAS and sg-LASSO-MIDAS methods show a marked decline in prediction accuracy, whereas the sg-LASSO-FAMIDAS method demonstrates notable resilience in handling the substantial shock introduced by the pandemic.

figure[figure omitted — 683 chars of source]

Figure (ref) in the online appendix presents the nowcasts from the three forecasting methods across the three horizons, alongside the advance release of real GDP, which serves as the target. These plots, corresponding to the RMSEs in Table (ref), illustrate how each model tracks the advance estimate and how additional information improves accuracy over the quarter. We focus on the COVID-19 period, specifically Q2 and Q3 of 2020. In both quarters, the factor-augmented sparse regression model delivers highly accurate nowcasts by quarter-end. Since the advance estimate becomes available only a month into the following quarter, this highlights the effectiveness of timely data use. As expected, nowcast accuracy improves as more information becomes available. However, sg-LASSO-MIDAS, which uses only sparse signals, tends to understate the magnitude of GDP growth, while FAMIDAS, which relies solely on dense signals, produces more erratic forecasts. Overall, the figure shows that factor-augmented sparse MIDAS regression yields more balanced and reliable nowcasts during periods of high volatility.

\textcolor{black}{Economic rationale}

\textcolor{black}{The results suggest that prior to the COVID pandemic, the data-generating process was predominantly sparse, whereas the COVID shock introduced a combination of idiosyncratic and common shocks in macroeconomic variables, but only idiosyncratic shocks in financial markets, which help nowcasting GDP (see the discussion and results backing the claim in a paragraph below). This pattern can be explained economically. The stringent physical restrictions during the pandemic had widespread effects on the real economy, generating common shocks across macro variables—see diebold2020covid. In contrast, financial markets, although initially disrupted in early 2020, rebounded quickly, leading to a disconnect from real economic activity—a phenomenon widely discussed in policy circles (e.g., igan2020disconnect). This disconnect and the absence of broad-based market co-movement help explain the lack of a dense financial component in our model. In the next two paragraphs, we discuss complementary analyses that confirm the aforementioned interpretation.}

\textcolor{black}{First, we formally test for the presence of sparse and dense components. Using the data-driven bootstrap test of beyhum2024testing, we test the null that the coefficients on the sparse component are zero. For the full real-time sample at the end-of-quarter horizon, the sparse part is significant (p = 0.002), and remains so pre-COVID (p = 0.037), indicating that idiosyncratic shocks contribute to explaining GDP. To test the significance of the macro factor (dense) component, we use the debiased LASSO with Gaussian multiplier bootstrap as in dezeure2017high, where factors are not penalized in the initial LASSO regression. We find that PCA factors are significant in the full sample (p = 0.003) but not pre-COVID (p = 0.668). These test results mirror our out-of-sample findings: the sparse model performs best pre-COVID, while the combined sparse-plus-dense models improve substantially when including the COVID period. (We leave the theoretical validation of the tests by beyhum2024testing and dezeure2017high in our factor-augmented sparse regression setting with $\tau$-mixing data to further research.)}

\textcolor{black}{We further assess sg-LASSO-FAMIDAS by replacing PCA-based macro factors with widely used observed indicators: the Aruoba-Diebold-Scotti index (ADS; aruoba2009real), the Chicago Fed National Activity Index (CFNAI; brave2019new), and the National Financial Conditions Index (NFCI; brave2017introducing, amburgey2023real). Using real-time vintages from the respective Federal Reserve Banks, we treat ADS and NFCI as weekly and CFNAI as monthly, preserving the original MIDAS structure and lag lengths. As shown in Table (ref) in the online appendix, ADS and CFNAI yield more accurate nowcasts than NFCI, with ADS performing best—likely due to its higher frequency and more stable inputs, consistent with diebold2020covid. Since ADS and CFNAI track real economic activity, while NFCI captures financial conditions, these findings support our conclusion that macro variables carry both sparse and dense signals, whereas financial variables contribute primarily sparse signals. Notably, our PCA-based factor approach still outperforms these observed alternatives in predictive accuracy. Lastly, Table (ref) shows that adding PCA-based financial factors to the regression yields similar results, reinforcing our conclusion that financial variables mainly contribute sparse signals.}

\textcolor{black}{Robustness}

\textcolor{black}{We conduct several robustness checks. First, we vary the factor model specification by imposing weak factors via sparse loadings following uematsu2022estimation as well as the sparse PCA approach of zou2006sparse, and assess alternative estimators for the number of factors, including those of bai2002determining and bai2019rank. We also evaluate a methodology where PCA factors are first computed from high-frequency macro predictors, then included in a MIDAS regression to nowcast GDP, as in, e.g., andreou2020mixed. To assess sensitivity to structural breaks, we implement a rolling window scheme. }

\textcolor{black}{Results appear in Tables (ref), (ref) and (ref) in the online appendix. Estimating factors from high-frequency data does not improve performance; the baseline sg-LASSO-FAMIDAS—where factors are extracted from MIDAS-weighted regressors—consistently outperforms the alternatives. For factor selection, the eigenvalue growth ratio of ahn2013eigenvalue delivers the best results, with the eigenvalue ratio estimator yielding similar performance. Notably, regardless of the selection method, the sparse-plus-dense specification consistently outperforms sparse-only and dense-only models, suggesting it captures additional—though slightly noisier—predictive signals that improve the accuracy. Rolling window results (Table (ref)) align with expanding window results, indicating little impact of structural breaks, consistent with, e.g., boot2020does.}

Conclusion

In this paper, we propose a factor-augmented sparse MIDAS regression for high-dimensional mixed-frequency data. By combining sparse and dense components, the method captures complex cross-frequency dependencies. We establish its theoretical properties under $\tau$-mixing and allow for model misspecification due to approximate sparsity or imperfect MIDAS lag structures, extending the literature with new convergence results for factor-augmented sparse regression.

Monte Carlo and empirical results on nowcasting U.S. GDP growth show that the method outperforms purely sparse or dense models, especially during volatile periods such as the COVID-19 pandemic. Sparse signals dominate in stable times, while dense factor information is crucial under stress. The approach offers a robust framework for forecasting in complex environments and could be extended to nonlinear models or alternative high-frequency indicators.

\if11 { {

Funding sources

Jad Beyhum was financially supported by the Research Fund KU Leuven through the grant STG/23/014. Jonas Striaukas gratefully acknowledges the financial support from the European Commission, MSCA-2022-PF Individual Fellowship, Project 101103508, MACROML. } } \fi \if01 \fi