Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
61,199 characters · 13 sections · 73 citation commands
Testing for sparse idiosyncratic components in factor-augmented regression models
\if11 \fi
\if01 {
} \fi
{\it Keywords:} sparse plus dense, high-dimensional inference, LASSO, factor models
\spacingset{1.5}
In this paper, we investigate a factor-augmented sparse regression model. Our analysis involves an observed sample of $T$ real-valued outcomes $y_1,\dots,y_T$, and high-dimensional regressors $x_1,\dots,x_T\in{\mathbb{R}}^p$, which are interconnected as follows:
Here, $\varepsilon_t\in {\mathbb{R}}$ represents a random error, $u_t$ is a $p$-dimensional random vector of idiosyncratic shocks, $f_t$ is a $K$-dimensional random vector of factors, and $B$ is a $p\times K$ (nonrandom) matrix of loadings. The parameters of interest are $\gamma^*\in {\mathbb{R}}^K$ and $\beta^*\in {\mathbb{R}}^p$ and the right-hand side of (ref) is unobserved. We consider the case where the number $p$ of regressors is large with respect to the sample size $T$ and a sparsity condition on the high-dimensional parameter vector $\beta^*$ is imposed. In the asymptotic regimes we study, $T$ goes to infinity, $p$ is allowed to grow with $T$ while $K$ remains fixed. The model formulation in equation (ref) effectively merges two popular approaches in handling high-dimensional datasets: factor regression (stock2002forecasting,bai2006confidence) and sparse high-dimensional regression (tibshirani1996regression,bickel2009simultaneous). Such a model allows the outcome to be related to the regressors through both common and idiosyncratic shocks and may better explain the data than factor regression or sparse regression alone (see fan2023latent,fan2021bridging, which introduce and study model (ref)). As noted in fan2021bridging, this type of structure has many applications in forecasting, causal inference and to describe correlation networks. Note that, as in stock2002forecasting,bai2006confidence,fan2021bridging, we could augment the model (ref) with additional regressors $w_t$ entering the first equation of (ref) but not the second one. This case is discussed in the Appendix Section (ref).
We develop a test for the hypothesis:
where our theory outlines the set of sparse alternatives against which our test has power. Our specification test sheds light on the data-generating process by allowing us to determine if the underlying model is dense (as is the factor regression model) or sparse plus dense, as is the factor-augmented sparse regression model. This determination will then tell us if the relation between the regressors and the outcome is only driven by common shocks (factor regression) or if idiosyncratic shocks also play a role (factor-augmented sparse regression). The question of the adequacy of sparse or dense representations has recently garnered significant attention (see, e.g., giannone2021economic,kolesar2023fragility). However, existing studies mostly focus on the differences between sparse and dense models and have found that dense models are often more adequate. In contrast, we compare a dense model with a sparse plus dense alternative.
In this paper, we propose a new bootstrap test for (ref). Our test's principle is to compare two estimators of $\sum_{t=1}^T u_t\varepsilon_t$. The first estimator is only consistent under the null, while the second estimator relies on the LASSO, and is, therefore, consistent under sparse alternatives. Our proposed test does not require estimating covariance matrices and is easy to implement. Following lederer2021estimating, we outline a data-driven rule to select the tuning parameter of the LASSO estimator and prove its theoretical validity. We establish the validity of the test within a theoretical framework that accommodates scenarios where the number of variables, denoted by $p$, can significantly exceed $T$, and the explanatory variables exhibit strong mixing and possess polynomial tails. We use simulations to evaluate the finite sample properties of our procedure. Our test controls size and exhibits good power against sparse alternatives even when $p$ greatly exceeds $T$ or the data are heavy-tailed and serially correlated. A potential limitation of our approach might be that our approach also rejects when $\beta^*$ is nonzero but has a dense structure. To assess this issue, we conduct simulations with dense $\beta^*$. We find that our test exhibits very low power against such alternatives (in absolute terms and also relatively to sparse alternatives with the same signal-to-noise ratio). Hence, our test exhibits some robustness against dense alternatives. This result is intuitive: the LASSO estimator on which our test relies sets to zero very small coefficients pertaining to dense alternatives. Finally, we apply our test to several commonly studied datasets in macroeconomics and finance and often reject the null. This suggests that sparsity can help describe economic data once a dense component (here, modeled through the factors) is included in the model. This result complements the recent studies giannone2021economic,kolesar2023fragility, which concluded that dense representations were often more appropriate for economic data. The R package \textquotesingle FAS\textquotesingle implements our approach.\\
Related literature. Our paper relates to the growing literature on factor-augmented sparse models, see hansen2019factor,fan2023do, fan2023latent, fan2021bridging, vogt2022cce, beyhum2023factor,barigozzi2023fnets among others. In particular, fan2021bridging proposes a general framework for factor-augmented sparse models. As noted by a reviewer, our test can be used to check that the sparse idiosyncratic components are jointly significant after the third step of the methodology by fan2021bridging. Note that fan2021bridging's model includes lags of idiosyncratic terms and additional regressors. In the Appendix, we explain how to extend our test to the case with additional regressors. Section (ref) of the Online Appendix outlines an adapted test with a lag. We do not formally prove that our test works in these cases, but we conjecture that the asymptotic properties extend to these more general models. Simulations reported in Section (ref) of the Online Appendix corroborate this presumption. We also note that the factor-augmented sparse model studied in the present paper is a generalization of an earlier model studied in fosten2017model,fosten2017confidence which augments the factor regression of stock2002forecasting with a low-dimensional set of idiosyncratic terms.
fan2023latent recently introduced the Factor-Adjusted deBiased Test (FabTest) for evaluating (ref). However, the FabTest exhibits several limitations. The test relies on a desparsified LASSO estimator based on model (ref). To achieve desparsification, fan2023latent utilized the nodewise LASSO method proposed by zhang2014confidence and van2014asymptotically for estimating the precision matrix of the idiosyncratic shocks. However, this approach introduces $p$ additional tuning parameters, in addition to the one used in the original LASSO regression. Although the tuning parameters are selected through cross-validation in practice, fan2023latent did not provide a theoretical justification for this selection procedure. Besides, inferential theory for LASSO-type regressions is not well understood when the tuning parameter is selected by cross-validation. Moreover, the test's performance may deteriorate due to errors associated with the nodewise LASSO estimates, and it incurs a heavy computational cost. Another limitation of the FabTest is its reliance on estimating the variance of $\varepsilon_t$, which can lead to imprecise results where variance estimation is challenging. Additionally, fan2023latent only established the validity of the FabTest for i.i.d. sub-Gaussian data (see Section 2 in fan2023latent). An alternative is to use the partial covariance test developed by fan2021bridging and applied in fan2023do. This test's principle is to estimate the covariance matrix of the idiosyncratic terms (including the idiosyncratic term of $y_t$, that is, $u_t^\top\beta^*+\varepsilon_t$) and then use a Gaussian bootstrap under the null to obtain the critical value of the test statistic. In contrast to our test, this partial covariance test does not make use of the LASSO estimator and requires estimating a high-dimensional covariance matrix, which is challenging.
Finally, we would like to note that this paper contributes to various other strands of literature. First, it connects to recent literature considering testing for high-dimensional parameters. There exists several approaches, see fan2015power,zhu2018linear,chernozhukov2019inference,lederer2021estimating,he2023most and references therein. Our strategy draws inspiration from lederer2021estimating, a recent paper that introduces a bootstrap procedure for selecting the penalty parameter of LASSO in a standard sparse linear regression. They employ this procedure to test the null hypothesis that a specific high-dimensional parameter equals zero. We adapt their approach to the case with unobserved factors, time series dependence and polynomial tails, which poses a challenge beyond the scope of the results in lederer2021estimating. Second, our work is related to the literature on inference on parameters of additional low-dimensional regressors in the factor regression model of stock2002forecasting, see bai2006confidence,gonccalves2014bootstrapping,gonccalves2020bootstrapping. Third, our work connects with the literature on specification tests for models involving unobserved factors. Many papers test for the validity of the assumption that loadings are time-independent in the approximate factor model itself — the second equation in (ref) — (breitung2011testing, chen2014detecting, han2015tests, yamamoto2015testing, su2017time, su2020testing, baltagi2021estimating, xu2022testing, fu2023testing), while corradi2014testing tests for time-independence of all coefficients in the factor regression model of stock2002forecasting. Our approach complements this literature by proposing a specification test of the factor regression model under a different alternative, namely the factor-augmented sparse regression model. \\
\noindentOutline. The paper is organized as follows. In Section (ref), we outline our testing procedure. Its asymptotic properties are studied in Section (ref). Then, Section (ref) contains Monte Carlo simulations. The empirical applications can be found in Section (ref). Section (ref) concludes. The Appendix contains an adapted testing procedure for a model with additional regressors and lags. The Online Appendix includes all proofs and additional discussions and simulation and empirical results.\\
\noindentNotation. For an integer $N\in \mathbb{N}$, let $[N]=\{1,\dots, N\}$. The transpose of a $n_1 \times n_2$ matrix $A$ is written $A^{\top}$. Its $k^{th}$ singular value is $\sigma_k(A)$. Let us also define the Euclidean norm $\left\| A\right\|_2^2=\sum_{i=1}^{n_1} \sum_{j=1}^{n_2}A_{ij}^2$ and the sup-norm $\|A\|_\infty=\max\limits_{i\in[n_1], j\in[n_2]}|A_{ij}|$. The quantity $n_1 \vee n_2$ is the maximum of $n_1$ and $n_2$, $n_1 \wedge n_2$ is the minimum of $n_1$ and $n_2$. For $N\in \mathbb{N}$, $I_N$ is the identity matrix of size $N\times N$. For a real-valued random variable $Z$ and $g>0$, we let ${\left\vert\kern-0.4ex\left\vert\kern-0.4ex\left\vert Z \right\vert\kern-0.4ex\right\vert\kern-0.4ex\right\vert}={\mathbb{E}}[|Z|^g]^\frac1g$. For a $d$-dimensional random vector $Z$, we define ${\left\vert\kern-0.4ex\left\vert\kern-0.4ex\left\vert Z \right\vert\kern-0.4ex\right\vert\kern-0.4ex\right\vert}_g=\sup\limits_{u\in{\mathbb{R}}^d: \ \|u\|_2\le 1}{\left\vert\kern-0.4ex\left\vert\kern-0.4ex\left\vert u^\top Z \right\vert\kern-0.4ex\right\vert\kern-0.4ex\right\vert}_g$.
In this subsection, we explain our testing procedure, which is then summarized in algorithmic form in subsection (ref). To facilitate understanding, we rewrite the model in matrix form as follows:
where $Y=(y_1,\dots,y_T)^\top$, $F=(f_1,\dots, f_T)^\top$ is a $T\times K$ matrix, $U=(u_1,\dots,u_T)^\top$ and $X=(x_1,\dots,x_T)^\top$ are $T\times p$ matrices and $\mathcal{E}=(\varepsilon_1,\dots,\varepsilon_T)^\top.$
It is important to note that, under the null hypothesis $H_0$, we have $U^\top (Y-F\gamma^*)=U^\top\mathcal{E}$. This observation suggests a testing procedure that involves computing an estimate $2T^{-1}\left\|U^\top (Y-F\gamma^*)\right\|_{\infty}$ and comparing it with the (estimated) quantiles of $2T^{-1}\left\|U^\top \mathcal{E}\right\|_{\infty}$.\footnote{We have a factor 2 in front of $T^{-1}\left\|U^\top (Y-F\gamma^*)\right\|_{\infty}$ and $T^{-1}\left\|U^\top \mathcal{E}\right\|_{\infty}$ because $2T^{-1}\left\|U^\top \mathcal{E}\right\|_{\infty}$ is the effective noise of the problem, a natural concept in the literature on the LASSO, see lederer2021estimating.}
We can estimate $U^\top (Y-F\gamma^*)$ by principal components analysis. First, let $\widehat{K}$ be one of the many estimators of the number of factors $K$ available in the literature (see for instance bai2002determining,onatski2010determining,ahn2013eigenvalue,bai2019rank,fan2022estimating). As in fan2023latent, we let the columns of $\widehat{F}/\sqrt{T}$ be the eigenvectors corresponding to the leading $\widehat{K}$ eigenvalues of $XX^\top$ and $\widehat{B}= X^\top \widehat{F}(\widehat{F}^\top \widehat{F})^{-1}=T^{-1} X^\top \widehat{F} .$ Then, we project the data on the orthogonal of the vector space generated by the estimated factors. Let $\widehat{P}=T^{-1} \widehat{F}\left(\widehat{F}^\top \widehat{F}/T\right)^{-1}\widehat{F}^\top=T^{-1} \widehat{F}\widehat{F}^\top$ be the projector on the vector space generated by the columns of $\widehat{F}$. A natural estimate for $U$ is $\widehat{U}=X-\widehat{F}\widehat{B}^\top =\left(I_T-\widehat{P}\right) X$. Similarly, we let $\widetilde{Y}=\left(I_T-\widehat{P}\right)Y$ be an estimate of $Y-F\gamma^*$. The final estimate of $2T^{-1}\left\|U^\top (Y-F\gamma^*)\right\|_{\infty}$ is our test statistic
Next, to estimate the quantiles of the distribution of $2T^{-1}\left\|U^\top \mathcal{E}\right\|_{\infty},$ we need an estimate of $\mathcal{E}$. We obtain it through the following LASSO estimator:
where $\lambda>0$ is a penalty parameter, the choice of which will be fully data-driven in both theory and practice. For $t\in [T]$, we denote by $\widetilde{y}_t$ the $t^{th}$ element of $\widetilde{Y}$ and $\widehat{u}_t$ as the $T\times 1$ vector corresponding to the $t^{th}$ row of $\widehat{U}$. For a given $\lambda>0$, let $\widehat{\varepsilon}_{\lambda,t}=\widetilde{y}_t -\widehat{u}_t^\top\widehat{\beta}_{\lambda},\ t\in[T]$ be the estimate of $\varepsilon_t$. For a fixed $\alpha \in(0,1)$, we can then estimate $q_\alpha$, the $(1-\alpha)$ quantile of the distribution of $2T^{-1}\|U^\top \mathcal{E}\|_{\infty}$, by the Gaussian multiplier bootstrap. Let $e=(e_1,\dots,e_T)$ be a standard normal random vector independent of the data $(X,Y)$ and define the criterion $$\widehat{Q}(\lambda,e)=\left\|\frac2T \sum_{t=1}^T\widehat{u}_t\widehat{\varepsilon}_{\lambda,t} e_t\right\|_\infty.$$ The estimate $\widehat{q}_\alpha(\lambda)$ of $q_\alpha$ is then the $(1-\alpha)$-quantile of the distribution of $\widehat{Q}(\lambda,e)$ given $X$ and $Y$. Formally, $\widehat{q}_\alpha(\lambda)=\inf\left\{q:{\mathbb{P}}_e(\widehat{Q}(\lambda,e)\le q)\ge 1-\alpha\right\}$, where ${\mathbb{P}}_e(\cdot)={\mathbb{P}}(\cdot|X,Y)$.
The only remaining element is the procedure to select $\lambda$. We adapt the approach of lederer2021estimating to our setting. Our choice of $\lambda$ is
We explain in Section (ref) how to compute $\widehat{\lambda}_\alpha$ in practice. The infimum in (ref) exists because for all $\lambda \ge \bar{\lambda}=2T^{-1}\left\|\widehat{U}^\top\widetilde{Y}\right\|_\infty$, it holds that $\widehat{\beta}_\lambda =\widehat{\beta}_{\bar \lambda}=0.$ Moreover, since $\widehat{U}\widehat{\beta}_\lambda$ is a continuous function of $\lambda$, $\widehat{q}_\alpha(\lambda)$ is also continuous in $\lambda$ and the infimum is attained at a point $\widehat{\lambda}_\alpha>0$ such that $q_\alpha\left(\widehat{\lambda}_\alpha\right)= \widehat{\lambda}_\alpha$. Let us recall briefly the heuristics behind the choice of $\lambda$ and refer the reader to lederer2021estimating for more details. First, note that when $\lambda$ is close to $q_{\alpha}$, standard convergence bounds for the LASSO suggest that $\widehat{\beta}_\lambda$ is a precise estimate of $\beta^*$, so that $\widehat{\varepsilon}_{\lambda,t}$ is a good estimate of $\varepsilon_t$ and, in turn, $\widehat{q}_\alpha(\lambda)$ is close to $q_\alpha$. Second, when $\lambda$ becomes (much) larger than $q_\alpha$, the error $\widehat{\varepsilon}_{\lambda,t}-\varepsilon_t$ becomes large and dependent of $\widehat{u}_t$, which in turn increases $\widehat{q}_\alpha(\lambda)$ and leads it to be larger than $q_{\alpha}$. We then let our estimator of $q_\alpha$ be $\widehat{\lambda}_\alpha=\widehat{q}_\alpha\left(\widehat{\lambda}_\alpha\right)$.
The test rejects $H_0$ at the level $\alpha$ when our test statistic is given in (ref) is larger than the estimate $\widehat{\lambda}_\alpha$ of $q_\alpha$. Therefore, our testing procedure is free of tuning parameters stemming from the LASSO regression in equation (ref).
Algorithm (ref) below explains how to conduct the test in practice. Let us discuss Step 4 of Algorithm (ref) in detail. It approximates $\widehat{\lambda}_\alpha$ as defined in (ref). It is advisable to set the grid size $M$ and the number of bootstrap samples $L$ to be as large as possible. As mentioned in lederer2021estimating, one can speed up Step 4b. by computing the LASSO with a warm start along the penalty parameter path (see, e.g., friedman2007pathwise), i.e., for each decreasing $\lambda$, the new coefficient estimate is computed by using the previous (i.e., the one that was computed with a larger value of $\lambda$) as a starting value. Furthermore, Step 4c. can be accelerated through parallelization techniques. In our implementation, we use both suggestions, which greatly speed up the computations. We also note that to compute the $p$-value of the test, it suffices to conduct it on a grid of values of $\alpha$ and let the $p$-value be equal to the largest value of $\alpha$ in this grid such that the test of level $\alpha$ rejects $H_0$.
In this section, we provide the asymptotic properties of the test in a theoretical framework allowing for time series dependence in the factors and the idiosyncratic shocks and polynomial tails. We place ourselves in an asymptotic regime where $T$ goes to infinity and $p$ goes to infinity as a function of $T$. The number of factors $K$ is fixed with $T$. It would be possible to let it grow, see, for instance, beyhum2022factor. The distributions of the factors $f_{t}$ and the error terms $\varepsilon_t$ do not depend on $T$, while the distribution of the other variables are allowed to vary with $T$. All the constants we introduce are universal in the sense that they do not vary with the sample size. Our assumptions are similar to that of fan2021bridging but significantly weaker than that of fan2023latent, which imposes that the variables are i.i.d. sub-Gaussian.
We introduce further notation. The loading $b_{jk}$ corresponds to the $j^{th}$ element of the $k^{th}$ column of $B$. Let also $b_j=(b_{j1},\dots,b_{jK})^\top$. For $t\in[T]$, let $$z_t= \left(u_t^\top,f_t^\top,\varepsilon_t,\frac{1}{\sqrt{p}}\sum_{\ell=1}^pu_{t\ell}b_{\ell}^\top, \frac{1}{\sqrt{p}}\sum_{\ell=1}^pu_{t\ell}f_{t}^\top, \frac{1}{\sqrt{p}}\sum_{\ell=1}^pu_{t\ell}\varepsilon_t\right)^\top.$$ Finally, define $\Sigma={\mathbb{E}}[u_tu_t^\top]$.
We make the following assumptions.
Assumption (ref) means that the estimator $\widehat{K}$ of the number of factors $K$ is consistent. Examples of $\widehat{K}$ and sufficient conditions for its consistency can be found in bai2002determining,onatski2010determining,ahn2013eigenvalue,bai2019rank,fan2022estimating. Assumption (ref) is the same as Assumption 3 in fan2021bridging. Its conditions (ref) and (ref) constitute a strong factor assumption (bai2003inferential). Assumption (ref) restricts the moments and the tail behavior of the variables. Assumption (ref) (ref),(ref),(ref) contain conditions on the moments of the different variables similar to that of the literature (bai2006confidence,fan2013large). We assume that the variables in Assumption (ref) (ref) have polynomial tails with common parameter $q+\zeta$. It would be possible to have different tail parameters for each variable but we avoid doing so to simplify our presentation. As in fan2021bridging the number of finite moments is at least $8$. In the similar context of inference on factor regression models, Assumption E.2. in bai2006confidence and Assumption 7 in gonccalves2014bootstrapping impose conditions analogous to the restriction that $\{u_t\varepsilon_t\}_t$ is uncorrelated across $t$ in Assumption (ref) (ref). A sufficient condition for the latter restriction is that $\{\varepsilon_t\}_t$ is uncorrelated across $t$ and independent of $\{u_t\}_t$. Note that this assumption could be avoided by using a block bootstrap method, but this would complicate the test. It may also not be justified because (most of) the serial correlation in the data may be picked up by the factors $f_t$ and not by the error term $\varepsilon_t$. Next, Assumption (ref) (ref) is a restriction on the cross-section correlation of the idiosyncratic shocks. This condition holds, for instance, if $\{u_t\}_t$ and $\{b_\ell\}_\ell$ are independent and $\max_{j\in[p]}\sum_{\ell=1}^p|\Sigma_{j\ell}| =O(1)$. Assumption (ref) means that the process $\{z_t\}_t$ has strong mixing coefficients decaying polynomially, which is a restriction on the time-series dependence of the variables. A similar assumption is made in fan2021bridging. Note that Assumption (ref) restricts the full distribution of the process $\{z_t\}_t$, while Assumption (ref) (ref) just imposes a condition on a particular serial correlation.
Let us introduce $\varphi^*= \gamma^*-B^\top \beta^*$ . To interpret $\varphi^*$, note that the first equation of (ref) can be rewritten $y_t=f_t^\top\varphi^*+x_t^\top\beta^*+\varepsilon_t$, which becomes a usual high-dimensional sparse regression model when $\varphi^*=0$. The next assumption concerns the relative growth rates of $T,p,\|\varphi^*\|_2$ and $\|\beta^*\|_1$.
Condition (i) restricts the rate at which $p$ can grow with respect to $T$. We allow it to grow at a polynomial rate, which is more restrictive than the standard restriction for the LASSO (exponential rate). We need to be more restrictive because of the presence of polynomial tails and time series dependence. Assumption (ref) (ref) contains sparsity restrictions on the alternative hypotheses. When $\|\beta^*\|_\infty=O(1)$, condition (ref) corresponds, up to logarithmic factors, to the standard consistency condition for the LASSO with bounded regressors and errors with sub-Gaussian tails that is $\sqrt{\log(p)/T}(s_0\vee 1)=o(1)$, where $s_0$ is the number of nonzero coefficients of $\beta^*$. We can impose a similar sparsity condition as in the standard LASSO literature with sub-Gaussian errors, despite dealing with polynomial tails and serial correlation, because we utilize the high-dimensional central limit theorem for polynomial-tailed time series from fan2021bridging. Under the rate condition specified in Assumption (ref) (ref), this theorem allows us to essentially revert to the scenario with sub-Gaussian errors. Condition (ref) is a slightly more restrictive version of the standard condition that $\sqrt{T}/(T\wedge p)=o(1)$ for inference in the factor regression model.\footnote{This condition is equivalently stated as $\sqrt{T}/p=o(1)$ in bai2006confidence,corradi2014testing and many others.} Indeed, since $\varphi^*$ is of size $K$, it is reasonable to assume that $\|\varphi^*\|_2=O(1)$. Under this condition, (ref) corresponds to $\sqrt{T}/(T\wedge p)=o(1)$ up to logarithmic factors. As noted by a referee, the role of (iii) is to ensure that the error coming from the estimation of the factors is asymptotically negligible (see for instance bai2006confidence). It may be possible to relax it using a more elaborate bootstrap scheme, see gonccalves2014bootstrapping. Additionally, it is worth noting that our proofs reveal that Assumption (ref) is stronger than necessary, and the validity of the test could be established under more complex but weaker rate conditions. However, for the sake of clarity, we present Assumption (ref) instead of a more intricate condition.
We have the following theorem.
The proof of Theorem (ref) can be found in Online Appendix (ref). Statement (ref) means that the empirical size of the test tends to the nominal size. Statement (ref) shows that the test has asymptotic power equal to $1$ against sequences of alternatives such that $\sqrt{\frac{\log(T\vee p)}{T\wedge p}}=o_P\left(T^{-1}\left\|U^\top U\beta^*\right\|_\infty\right) $. As noted in lederer2021estimating, such a condition is inevitable because the presence of the error $\varepsilon_t$ prevents us from distinguishing true $U\beta^*$ and $\varepsilon_t$ when $U\beta^*$ is too small. We discuss further this condition in Section (ref) of the Online Appendix.
In this section, we provide a Monte Carlo study shedding light on the finite sample performance of our proposed testing procedure. We use the following model as our data-generating process (DGP):
We generate samples with $T=\{200, 400\}$ observations, $p=\{T, 5T\}$ variables and $K=2$ factors. The loadings $B=\{b_{jk}\}_{j\in[p],k\in[K]}$ are such that $b_{jk}\sim\mathcal{U}[-1,1]$. The factors are generated as $f_t=\rho_f f_{t-1}+ \tilde f_t$ for $t=2,\dots, T$, where $\tilde f_t$ are i.i.d. $\mathcal{N}\left(0,I_K\left(1-\rho_f^2\right)\right)$. The idiosyncratic components $\{u_t\}$ are such that $u_t=\rho_{u}u_{t-1}+\tilde u_t $ for $t=2,\dots, T$, where $\tilde u_t$ are i.i.d. $\mathcal{N}\left(0,\Sigma\left(1-\rho_{u}^2\right)\right)$, with $\Sigma_{ij}=c_u^{|i-j|},i,j\in[p]$., where $c_u$ reflects the amount of cross-sectional dependence. We also let $\varepsilon_t=\rho_{e} \varepsilon_{t-1}+\tilde \varepsilon_t $ for $t=2,\dots, T$, where $\tilde \varepsilon_t$ are i.i.d. $\mathcal{N}\left(0,\left(1-\rho_{e}^2\right)\right)$.
The parameters $\rho_f$, $\rho_{u}$, and $\rho_e$ control the level of time series dependence. The stationary distributions of $f_t$, $u_t$, $\varepsilon_t$ are, respectively, $\mathcal{N}(0,I_K)$, $\mathcal{N}(0,\Sigma)$ and $\mathcal{N}(0,1)$. We initialize $f_0$, $u_0$ and $\varepsilon_0$ as such. We consider three dependency designs, where we vary the cross-sectional dependence via parameter $s$ and the time series dependence via parameters $(\rho_f, \rho_u, \rho_e)$ as follows:
Our theory does not formally allow the third design, but we want to show that our test performs well even under weak serial correlation of $\{\varepsilon_t\}_t$.
Finally, we consider two cases for the target parameter $\beta^*$. For the first case we set $\beta^*=(1,0,\dots,0)^\top\times m$, where $m\in\{0,0.1,0.2,0.3,0.4\}$ controls the signal strength. This choice of $\beta^*$ corresponds to a sparse design. For the second case, we consider $\beta^*=(1/\sqrt{p}, \dots, 1/\sqrt{p})^\top\times m$, which corresponds to dense design. In both cases, we set $\gamma^*=(0.5,0.5)^\top$. Note that the choice of $\beta^*$ ensures that for the dependence Design 1, the signal-to-noise ratio is the same for both sparse and dense cases.
We compute the rejection probabilities of our test at the significance levels $\alpha\in\{0.1,0.05, 0.01\}$ over $2000$ replications. To implement our test, we set $M=100$ and choose an equidistant grid of values for $\lambda$, using $L=1000$ bootstrap replications. The results are insensitive to the choice of $L$ and $M$ as long as they are sufficiently large, which is expected since their primary role is in the approximation of theoretical quantities. In our experience, $L=M=100$ already yields very precise results. The number of factors $K$ is estimated through the eigenvalue ratio estimator of ahn2013eigenvalue.
The results for $T=200$ and sparse and dense alternatives are reported in Tables (ref)-(ref). In the Online Appendix, we present simulations under the same data-generating processes, but with the larger sample size $T=400$ (Tables (ref) and (ref)). First, we observe that the empirical size is close to the nominal levels for all designs, indicating that both cross-sectional and serial dependence have little effect on the empirical size of our testing procedure. Notably, the results show consistent power when the alternative hypothesis is sparse (Table (ref)) across all data-generating processes (DGPs), suggesting that the dependency design does not significantly impact the power of our test. However, the power tends to slightly decrease in scenarios with large $p$ cases for both $T=200$ and $T=400$. Interestingly, in the case of dense alternatives (Table (ref)), we see a significant drop (relative to sparse alternatives) in power across all designs. This highlights that our testing procedure has low power when $\beta^*\neq0$ is dense. An intuition for this result is as follows. Our test uses the LASSO estimator, which enforces sparsity. When $\beta^*$ is dense, the LASSO estimator often sets $\widehat{\beta}_{\widehat{\lambda}_\alpha}$ to $0$, leading to non-rejection of the null.
In this subsection, we consider the same data-generating process as in the main design Table (ref), but we change elements of it to obtain results for heavy-tailed data and in cases when the number of factors is over- and under-estimated. For the heavy-tailed data scenarios, we generate factors, idiosyncratic shocks and regression errors from student-$t(5)$ distribution. This allows us to investigate the impact of the heavy tails on the proposed testing procedure. Note that our theoretical analysis imposes a condition for the existence of first $8$ finite moments, so that generating student-$t(5)$ errors goes beyond what our assumptions allow. We generate heavy-tailed factors $f_t$, idiosyncratic components $u_t$ and the error terms $\epsilon_t$ by generating $\tilde{f}_t$, $\tilde{u}_t$ and $\tilde{\epsilon}_t$ from a student-$t(5)$ rather than a Gaussian distribution. Besides changing the distribution of the data, we also compute the performance of our test when the number of factors is either over-estimated or under-estimated. In this case, our data generating process is the same as for Table (ref), but rather than estimating the number of factors $K$ using the eigenvalue ratio estimator, we set it to $K=5$ (over-estimated case) and $K=1$ (under-estimated case).
Results are reported in Tables (ref)-(ref) and (ref)-(ref) for heavy-tailed data and different number of factors, respectively. In the former case, we see a slight deterioration of our method compared to the Gaussian DGPs. This is not surprising since the residuals of the LASSO regression, as well as the factors, may be less accurately estimated in finite samples due to heavy tails, see, e.g., babii2020inference. Next, we analyze the results of the case of the over-estimated number of factors. The performance is slightly affected compared to the case where we use the eigenvalue ratio estimator to determine the number of factors (see Table (ref)), but the differences are small. When comparing with the under-estimated case, we see that our testing procedure is over-sized. These results are in line with the literature on inference with factors models, see, e.g., moon2015linear.
It is also interesting to investigate the performance of our testing procedure with lagged idiosyncratic shocks. In this case, we generate the data from the following model
where all the elements are as in the main DGPs except that we add lagged idiosyncratic shock, set $\beta_1^* = \beta^*=(1,0,\dots,0)^\top\times m$ and all elements of $\beta_2^*=0$. The algorithm to compute the test for this model appears in the Online Appendix (ref). We report results for sparse and Gaussian DGPs which appear in Online Appendix Tables (ref)-(ref). Results show a slight deterioration of performance in terms of power, while the empirical size appears to be similar as in the main scenario Table (ref) (see also Table (ref) in Online Appendix for a large $T$ comparison). The small decrease in power is due to an increase in the dimensionality of the regressors.
In this section, we present three empirical examples using well-established macroeconomic and financial datasets.
First, we examine the FRED-MD dataset, which includes 121 monthly macroeconomic series covering various sectors of the US economy. For further details, see mccracken2016fred. We also study a quarterly variant of FRED-MD with 202 variables, named FRED-QD, for which results are reported in the Online Appendix. Second, we analyze a financial dataset comprising 100 representative anomalies from the literature, which we use to explain aggregate market returns and industry portfolio returns. For more information, we refer to dong2022anomalies. Lastly, we investigate the network structure in asset returns using a finance dataset. Following the approach of fan2021bridging, we consider a cross-section of monthly stock returns and a set of observable factors to study the relationships among financial firms.
We study these datasets for several reasons. First, FRED-MD and its quarterly variant FRED-QD are among the most widely studied high-dimensional macroeconomic datasets. Providing further insights into these datasets could be valuable for the extensive empirical macroeconomic literature. Second, the empirical asset pricing literature has proposed and examined many factors, known as anomalies, that explain asset prices. Our research offers new insights into applying long-short anomalies to explain the market and industry portfolio returns rather than individual asset returns. Third, applying our method to firm-level stock returns illustrates using a model with observed regressors, as detailed in Appendix Section (ref). Overall, these diverse applications, using macroeconomic and financial datasets, demonstrate the versatility of our approach across different econometric settings, varying in sample sizes relative to the number of regressors. Lastly, we note that in all empirical applications, we select the number of factors using the eigenvalue ratio estimator of ahn2013eigenvalue.
In the macroeconomic application, we investigate the FRED-MD dataset over the sample period 1980 January to 2019 December containing $T=480$ observations. This is a commonly used dataset in various macroeconomic studies. In our analysis, we regress each of the 121 variables available in the dataset on common and idiosyncratic shocks. Specifically, we estimate the following model:
where $y_{t+1}$ represents one of the variables in the dataset at time $t+1$, and $x_t$ includes all the remaining regressors at time $t$. Note that we transform the original series using commonly applied transformations as suggested by mccracken2016fred. To estimate $f_t$, we apply PCA using the eigenvalue ratio estimator to determine the number of common factors. We apply our procedure to test $H_0$. Results for each category of variable appear in Table (ref).
First, results show that in most categories, we frequently reject the null hypothesis at the 10%, 5%, and even 1% significance levels. This provides evidence of sparsity in the idiosyncratic shocks of the macroeconomic data. The categories where we reject the null most often are Output and Income, Consumption, Orders, and Inventories, Labor Market, Interest and Exchange Rates, and Prices. For the remaining categories, the null hypothesis is not frequently rejected, suggesting that the sparse component is not important for those series. Interestingly, for the housing category, we never reject the null, while the other two categories for which we reject less frequently, i.e., Money and Credit and \textit{Stock Market}, could be classified as financial rather than macroeconomic data. In addition to the results for the FRED-MD dataset, we also obtain similar rejection ratios for the FRED-QD dataset, which is a quarterly macroeconomic dataset similar to FRED-MD, but containing more variables and less observations. Results appear in Table (ref). The results for the FRED-QD dataset also show a similar pattern, providing evidence of sparsity in the idiosyncratic shocks.
In our second application, we test the sparse component of a financial dataset comprising a set of regressor variables shown to predict market excess returns (dong2022anomalies). Specifically, we examine whether the idiosyncratic sparse component is important in explaining market excess returns as well as returns of 49 industry portfolios when regressing on a representative sample of 100 long-short anomaly portfolio returns. The data spans from January 1970 to December 2017, therefore, the effective sample size is $T=575$. We estimate the same model as in (ref).
dong2022anomalies have shown that these regressors lead to accurate forecasts of stock market excess return by employing various machine learning and forecast combination methods. For more details on the data and a full list of target variables, see Section (ref) of the Online Appendix. Results are reported in Table (ref).
Our results indicate that for more than 50% of 49 industry returns, the idiosyncratic shocks are significant at the 10% level, meaning we reject the null hypothesis at the 10% level. For the aggregate market returns, the p-value is 0.042, providing evidence in favor of the sparse component. Additionally, we reject the null at the 1% significance level for nine industry portfolios: Apparel (Clths), Automobiles and Trucks (Autos), Aircraft (Aero), Rubber and Plastic Products (Rubbr), Construction Materials (BldMt), Real Estate (RlEst), \textit{Printing and Publishing} (Books), \textit{Non-Metallic and Industrial Metal Mining} (Mines), and \textit{Business Supplies} (Paper). Furthermore, we analyze returns of 10 industry portfolios which are based on a broader classification compared to 49 industries. Results are presented in Table (ref) in the Online Appendix, which confirm a similar pattern. Therefore, we find strong evidence supporting the presence of sparse idiosyncratic components in equity returns for a variety of industries and the aggregate market.
In the third empirical application, we examine the sparse idiosyncratic components of individual stock returns using the dataset from jensen2023there. Specifically, we use monthly stock return data for a sample of 728 firms with no missing data from January 1991 to December 2022, resulting in $T=384$ and $p=727$ (the regressors $x_{it}$ are the returns of other firms). We test a model similar to fan2021bridging where our test can be seen as a diagnostic check after the third step in the approach put forward by fan2021bridging. However, our analysis diverges from fan2021bridging by focusing on a balanced panel of firms. We select firms for which we have a complete time series of returns and observed regressors. Furthermore, in our industry or sector analyses, we group firms into 49 industries — the same as in finance I application —, whereas fan2021bridging use a different firm classification.
Specifically, denote $y^{(i)}_{t}$ the stock excess return of firm $i$ at time $t$. We regress the firm's excess return on a set of observable factors — see Table (ref) for the list of factors — as well as common and idiosyncratic components stemming from all other returns. The model is thus an example of the extension of the main model with additional covariates, see Online Appendix (ref). Denoting the returns of all firms except the $i^{th}$ firm as $x_{it} = (y_{jt})_{j\in[n]/i}$, where $n$ is the total number of firms, we report the rejection ratios $\beta^*$ by testing $\beta^*=0$ in the following model for each firm $i$:
where $w_t$ denotes the observable factors.
Results are reported in Table (ref). First, we see that the rejection ratios (the proportion of firms $i$ for which we reject the null) are relatively low and only slightly above the nominal significance levels. This suggests that the sparse component is much less significant in individual stock returns compared to other applications we considered, i.e., we find weak evidence of the presence of sparse idiosyncratic shocks in firm-level returns. To gain further insight, we report results by grouping firms into 49 industries, as reported in Table (ref) in the Online Appendix. We find that for most industries, the rejection ratios are small. However, a few industries exhibit a larger proportion of rejections, notably Precious Metals (Gold) and Communications (Telcm).
This paper proposes a new bootstrap test for the adequacy of the factor regression model against factor-augmented sparse alternatives. We establish the asymptotic validity of our test under time series dependence and polynomial tails. In a Monte Carlo study, we show that our procedure has excellent finite sample properties against sparse alternatives but low power against dense alternatives. We often reject the null when we apply our testing procedure to standard datasets in macroeconomics and finance. This suggests that sparsity is present - on top of a dense model - in several economic environments.
This message complements previous studies. Indeed, based on different approaches, giannone2021economic,kolesar2023fragility found evidence against the presence of sparsity in several economic datasets, comparing only sparse and only dense models. Our analysis instead suggests that sparsity may still have a role to play on top of a dense component. These findings also constitute arguments in favor of using sparse plus dense models, which have recently gathered a lot of interest in the econometrics literature.
A potential limitation of our analysis is that we may reject $H_0$ because $\beta^*$ is nonzero but dense. In our Monte-Carlo simulations, we compare DGPs with sparse and dense $\beta^*$ with the same signal-to-noise ratio and find that our test has high power against the sparse alternative but low power against dense deviations from the null. This finding should mitigate the previous concern that rejection might be due to dense $ \beta^*\ne 0$. A complementary approach free of this concern is testing sparsity directly. kolesar2023fragility takes this road by proposing a test for the null hypothesis of sparsity in a high-dimensional regression model (without factors). The limitations of their approach are that it is only valid when the number of variables is smaller than the sample size and that it does not reject the null when the regression coefficient is equal to $0$. This is a problem because, in our context, we would not conclude in favor of the existence of sparsity when $\beta^*=0$. For future research, it would be interesting to adapt the test of kolesar2023fragility to the factor-augmented regression model and apply it to the datasets we consider in the present paper.