Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
59,873 characters · 11 sections · 70 citation commands
Composite Quantile Factor Model
\doublespace
JEL Classification: C18, C21.
Keywords: Composite quantiles, factor analysis, panel data.
\doublespace
Factor model is a useful statistical tool to describe data with unobserved and systematic components. Detailed textbook discussions on classical factor analysis for data with a fixed or small number of variables can be found in lawleymaxwell1971factoranalysis,anderson2003introduction. Following the important work in stockandwatson1998nber,stockandwatson2002JASA,stockandwatson2002jbes,baiandng2002ecma,bai2003ecma on high-dimensional panel data, the research on factor analysis in panel data has been extended to many directions, and there now exists a large body of literature on factor analysis for high-dimensional panel data. See baiandwang2016are for a review on recent developments in this burgeoning field.
At the heart of factor analysis for high-dimensional panel data is the use of principal component analysis (PCA) method. The estimates for factors are chosen to be the normalized eigenvectors of the sample covariance matrix of the data, based on which we can estimate the factor loadings using the least-squares (LS) method. These eigenvectors coincide with the solutions in PCA, and the procedure's simplicity greatly contributes to its wide popularity as a research tool in empirical macroeconomics and finance.
Despite its simplicity, the PCA-based procedure typically imposes some higher-order moment conditions on the factor and error terms (see, e.g., the assumptions in stockandwatson2002JASA,bai2003ecma) in order to obtain desirable asymptotic results. However, the sample covariance matrix on which PCA operates may be irrelevant for data with infinite (or very large) variances and PCA becomes invalid. Weakening or even removing these conditions will be appealing since many data in applications such as finance are either heavy tailed or of unknown nature. Two approaches for robust factor analysis have emerged in recent literature. chenetal2021qfm introduce the quantile factor models (QFM), where both the factors and factor loadings are quantile-dependent. At the quantile position $ 0.5 $, QFM estimator can be interpreted as the least absolute deviation (LAD) estimator that may be robust to certain error distributions. andoandbai2020jasa provide a more general framework that adds a regression component with heterogeneous coefficients. By assuming the factors and errors follow a joint elliptical distribution, heetal2022JBESnomoment propose a second approach to replacing the sample covariance matrix in PCA with the spatial Kendall's tau matrix that can handle non-normal distributions with large or infinite variances, a situation in which the standard PCA fails.
This paper studies another approach to robust factor analysis and we term it composite quantile factor models (CQFM). Our approach is inspired by the interesting work in zouandyuan2008cqr. zouandyuan2008cqr notice that, in a linear regression with infinite error variance, the parameter estimator will no longer have root-$ n $ consistency or asymptotic normality; a robust procedure such as LAD can be used but its relative efficiency to the LS estimator can be very small. The authors propose to estimate the regression coefficients by simultaneously minimizing the standard quantile regression objective function at multiple quantile positions and call this procedure composite quantile regression (CQR). zouandyuan2008cqr demonstrate the good finite sample properties of the CQR estimator for several non-normal error distributions.
Although factor analysis is different from linear regression, they are intrinsically connected (see, e.g., stockandwatson1998nber for the use of LS method in deriving the solution to the factor model). If CQR works in linear regression, we conjecture a variant of it will also work in factor model. This paper studies the extension of CQR to the estimation of factor model by simultaneously minimizing the objective function at multiple quantiles. The resulting estimates are shown to have good finite sample properties under various non-normal error distributions. Because the CQFM visits different quantiles of data during estimation, the estimated factors can usually pick up more skewness (and kurtosis) information about the data, often resulting a better fit of the model. It is important to point out that, although CQFM uses the method of quantile regression, its estimates are for the mean factor model. This sets our paper apart from the work in andoandbai2020jasa,chenetal2021qfm, where the goal is to estimate parameters at a specific quantile position.
We make the following contributions to the growing literature on panel factor analysis. First, we introduce CQFM as a new method to perform factor analysis on the mean factor model that can capture features of data at different quantiles. Second, we develop the asymptotic distribution results for the estimated factors and factor loadings; an information criterion is also developed to consistently select the factor number. Third, we provide extensive simulation evidence to show that CQFM works well for several non-normal error distributions. A special case of CQFM is when one chooses to optimize the objective function at a single quantile position, say $ 0.5 $. This reduces CQFM to the QFM in chenetal2021qfm. In several simulation examples, we demonstrate the advantage of CQFM as a result of using information at multiple quantiles. We also develop an R package cqrfactor that implements the CQFM method in this paper. The cqrfactor package can be downloaded from \url{https://github.com/xhuang20/cqrfactor}.
The rest of the paper is organized as follows. Section 2 sets up the objective function for CQFM and discusses the estimation procedure and the asymptotic results. Section 3 discusses the information criterion for the selection of factor numbers. Section 4 presents all simulation results. Section 5 applies CQFM method to the modeling of the quarterly macroeconomic data in mccrackenandng2020FREDQD. Section 6 concludes. The online supplement contains all proofs, additional figures and tables.
Let $ Y_{it} $ be the observation at time $ t $ for the $ i $th cross-section unit. Consider the following factor model:
where $F_{0t}$ is an $r \times 1$ vector of factors, $\lambda_{0i}$ is an $r \times 1$ vector of factor loading, and $\varepsilon_{it}$ is the error term. Both $F_{0t}$ and $\lambda_{0i}$ are unobserved, and the goal is to estimate them jointly. We assume the number of factors ($r$) is known. An information criterion will be developed to estimate $r$ consistently in (ref). Rewrite (ref) in matrix form to have
where $Y_{it}$ and $\varepsilon_{it}$ are elements of $Y$ and $\varepsilon$, respectively, and
Let $ \tau$ be a quantile position with $0 < \tau < 1$ and $b_{0\tau}$ be the $ 100\tau\% $ quantile of $\varepsilon_{it}$. For the quantile factor model, we seek the estimates for $b_{0\tau}$, $F_{0t}$, and $\lambda_{0i}$ that minimize the following QFM objective function:
where $\rho_{\tau}(u) = u(\tau - \mathbf{I}(u \leq 0))$ is the check function in quantile regression. The estimates from (ref) depend on $\tau$ and can be written as $\hat{F}_t(\tau)$ and $\hat{\lambda}_i(\tau)$, an approach adopted in andoandbai2020jasa,chenetal2021qfm.
In CQFM, instead of estimating the model at a single quantile position $\tau$, we estimate the model simultaneously at multiple quantiles by choosing a sequence of $K$ quantiles, $0 < \tau_1<\tau_2<\cdots<\tau_K<1$, and minimizing the following objection function:
where $b_{\tau_{k}}$ estimates $b_{0\tau_{k}}$, the $100\tau_{k}\%$ quantile of $\varepsilon_{it}$. Let $\hat{\lambda}_i$ and $\hat{F}_t$ be the estimators for $\lambda_{0i}$ and $F_{0t}$ in (ref). Unlike the solutions to (ref), $\hat{\lambda}_i$ and $\hat{F}_t$ are not dependent on any specific quantile position, and they estimate parameters in the mean factor model. By minimizing the objective function across multiple quantiles, the estimators can adapt to data features at different quantiles while still giving estimates for the mean of the process. We usually select equally spaced quantiles with $\tau_{k} = \frac{k}{K+1}$ for $k = 1,2,\cdots,K$ and $K$ is an odd number such as $5$ or $7$. This will always include the $50\%$ quantile in estimation, but an even number of quantiles also works for (ref). The number $K$ can be viewed as a tuning parameter of CQFM. In the special case of $K=1$, i.e., when a single quantile position is used in (ref), CQFM reduces to the quantile factor model. Our R package can estimate both CQFM and QFM.
There is no closed-form solution to the minimization exercise in (ref), $\hat{b}_{\tau_{k}}$, $\hat{\lambda}_i$, and $\hat{F}_t$ need to be obtained through an iterative algorithm. Since $\lambda_i$ and $F_t$ appear as a product in (ref), they are not separately identifiable. We use the following normalization for identification purposes:
We describe the steps of the algorithm below. Let $s$ denote the iteration step and the ranges of the subscripts $i,t,k$ are the same as those appear in (ref).
A few remarks follow.
We make the following assumptions to derive the asymptotic results.
(ref) is almost identical to stockandwatson2002JASA and chenetal2021qfm and can help identify both $F_0$ and $\Lambda_0$. See baing2013pcrfactor for a more detailed discussion on the identification in factor models. (ref) is a standard one in quantile regression. This assumption is made for the unconditional distribution of $\varepsilon_{it}$. If we consider the conditional distribution of $\varepsilon_{it}$ given $F_{0t}$ in (ref), all expectations in the proof will be conditional. The i.i.d. requirement in (ref) is strong. However, this assumption simplifies the presentation of the asymptotic results and allows us to make direct comparison of our asymptotic results to the cross-section regression result in zouandyuan2008cqr; in addition, the simple form of our asymptotic results facilitates the efficiency comparison between the CQFM-based factors and the PCA-based factors in bai2003ecma (see a remark following (ref) for a discussion). We can modify (ref) so that it is conditional on $F_{0t}$, and the asymptotic covariance in (ref) will have the standard sandwich form. We discuss this in a remark following (ref). Our simulation section includes results for errors with heteroskedasticity and AR(1) structure, and CQFM continues to give good results especially when the sample size is large.
The following theorem gives the asymptotic distribution of the estimated factors and factor loadings.
The format of the limiting distributions in (ref) resembles the result for linear regression coefficients in zouandyuan2008cqr.
The number of factors is assumed to be known in (ref). We discuss the selection of factor number in this section. Since the important work in baiandng2002ecma on consistent factor number selection, there has been continued development of new methods in the literature. See baing2007jbesprimitive,amengualandwatson2007JBES,hallinandliska2007jasa,onatski2009ecma,lamandyao2012AOS for panel mean regression models and andoandbai2020jasa,chenetal2021qfm for panel quantile regression models.
Denote $r$ the estimated number of factors. To work with (ref), we propose the following information criterion (IC):
where
and we use $\hat{b}_{\tau_{k}}(r), \hat{\lambda}_i(r)$ and $\hat{F}_t(r)$ to denote estimates based on $r$ number of factors. This information criterion is similar to the one used in andoandbai2020jasa and $\textit{IC}_{p1}$ in baiandng2002ecma. (ref) shows the consistency of $IC(r)$. Let $C_{NT} = \min(N,T)$.
See the online supplement for the proof.
In this section, we use Monte Carlo simulation to study the finite sample properties of the CQFM method. To compare CQFM to other methods, we use the R code in heetal2022JBESnomoment to compute the robust two-step (RTS) factors and the matlab code in chenetal2021qfm to compute the QFM factors at quantile position $0.5$ (QFM(0.5)) and also the estimated factor numbers. When space permits, we also add the PCA results.
The number of quantiles in CQFM is an additional tuning parameter, and we choose $K = 5$ for demonstration purposes, which corresponds to the quantiles of $0.17, 0.33, 0.5, 0.67$ and $0.83$. A convergence criterion of $ 10^{-3}$ is used in the MM algorithm.
Consider the following 3-factor data generating process (DGP):
where $F_{0t,1} = 0.8 F_{0t-1,1} + e_{1t}$, $F_{0t,2} = 0.5 F_{0t-1,2} + e_{2t}$, $F_{0t,3} = 0.2 F_{0t-1,3} + e_{3t}$, and both $e$ and $\lambda_{0i,j}$ are i.i.d. $N(0,1)$. This is identical to the DGP in chenetal2021qfm except that we consider several asymmetrical i.i.d. errors. They are summarized in (ref). Let $\gamma_1$ and $\gamma_2$ be the skewness and excess kurtosis coefficient, respectively.
The R package sn is used to simulate the skewed normal and skewed t distributions in (ref). If one specifies the skewness parameter directly, the sn package restricts $|\gamma_1| < 0.99527$; we set $\gamma_1 = 0.99$ for the first two error distributions. The asymmetric Laplace error term is generated using the rlaplace function in the R package LaplacesDemon. The three numbers, $0,0.5,4$ correspond to the location, scale, and kappa parameter in the rlaplace function in R. For the asymmetric Laplace distribution, a kappa value of $4$ implies a skewness of about $-1.99$. Next, we consider a more skewed log-normal distribution with mean and s.d. equal to $0$ and $1.5$, respectively. These are the parameter values for the log-normal density, which implies the error term $\varepsilon_{it}$ has its mean, s.d., and skewness equal to $ 3.08, 8.97$ and $33.47$. For both the asymmetric Laplace distribution and the log-normal distribution, we subtract the theoretical mean from the simulated errors so that all error terms have zero mean. Finally, we consider a mixture of skewed normal distribution, where $sn(0,1,0.99)$ and $sn(0,9,0.99)$ denote the skewed normal distribution with $\mu_{\varepsilon} = 0, \sigma_{\varepsilon}=1, \gamma_1=0.99$ and $\mu_{\varepsilon} = 0, \sigma_{\varepsilon}=3, \gamma_1=0.99$, respectively. The weights for the mixture normal are $0.9$ and $0.1$.
We consider five different sample sizes: $\left(N,T\right) = (50,100), (100,50), (100,200), (200,100)$ and $(300,300)$. For each error distribution and sample size, we report the value of an evaluation metric based on $100$ replications for every estimation method.
(ref) in the online supplement report the results for $5$ symmetric error distributions, including $N(0,1)$, t distribution with $1$ degree of freedom, etc. (ref) report the results for the following heteroskedasticity and AR(1) asymmetric errors:
where $\lambda_{0i,4}$ and $F_{0t,4}$ are i.i.d. $N(0,1)$ and $\varepsilon_{it}$ and $u_{it}$ are asymmetric errors defined in (ref). CQFM is found to have good finite properties in these cases too.
Similar to chenetal2021qfm, (ref) reports the average adjusted $R^2$ from regressing $F_{0t,1}, F_{0t,2}$, and $F_{0t,3}$ on the $3$ estimated factors from the RTS, QFM(0.5), and CQFM methods. Results in (ref) assess how well the estimated factors span the space spanned by the true factors. (ref) in the supplement expands (ref) to include the PCA results.
For the first two error distributions with small skewness in (ref), all three methods perform well. Their differences in the adjusted $R^2$ mostly appear in the third digit. Still, we see CQFM performs slightly better than RTS and QFM. In the case of asymmetric Laplace error, the difference between CQFM and the other two methods start to grow larger. For example, for the sample size $(50,100)$, the adj. $R^2$ associated with $F_{0t,3}$ is $0.8682$ for QFM, while it is $0.9518$ for CQFM. The log-normal error distribution poses the greatest challenge to the other methods, as the regression yields much lower adj. $R^2$s compared to those of CQFM, and increasing sample size from $(50,100)$ to $(300,300)$ does not seem to help.
To further investigate the accuracy of the estimates, we compute the mean squared error (MSE) of the estimated components and report them in (ref). The MSE is defined as
(ref) clearly indicates that CQFM always yields the smallest MSE for the $5$ error distributions. Depending on the error distribution, CQFM's reduction in MSE can be huge -- its MSE can be a fraction of that of PCA.
We can draw several conclusions based on (ref). First, compared to QFM at a single quantile position $\tau = 0.5$, the higher adj. $R^2$ and smaller MSE of CQFM suggests that there is some benefit in performing the estimation at multiple quantile positions simultaneously. Second, CQFM continues to work well in cases such as the asymmetric Laplace and log-normal errors, implying that CQFM can be a useful alternative to PCA in certain cases.
The online supplement also includes the adj. $R^2$ and MSE results ((ref)) for $5$ symmetric error distributions. Overall, CQFM continues to provide robust estimates.
Next, we study the performance of the information criterion in (ref). (ref) reports the average estimated factor number and the frequency of correct factor number estimation. Both CQFM and PCA perform well in majority of the cases. For log-normal error, all methods fail in small sample. However, CQFM yields better results when sample size is large with $\text{Prob}(\hat{r} = 3) = 82\%$ in the case of $(300,300)$.
For the $5$ symmetric error distributions in (ref), our proposed IC with (ref) continues to work well except for the $t_1$ error. In this case, the rank-based approach in chenetal2021qfm gives good results when sample size is large. After standardizing the data and using (ref) in (ref), CQFM also gives satisfactory results when the sample size is large.
In this section, we use the quarterly macroeconomic data set, FRED-QD, in mccrackenandng2020FREDQD to study the properties of CQFM factors. We use the version “2023-06.csv", which contains $258$ quarterly observations from 1959/3/1 to 2023/3/1 for $246$ macroeconomic variables. The data link is: \url{https://research.stlouisfed.org/econ/mccracken/fred-databases/}. We use the matlab code in mccrackenandng2020FREDQD to prepare the data, including transforming all variables to stationary time series based on the tcode in mccrackenandng2020FREDQD, removing outliers, and using the EM algorithm to fill in missing values. The final data set has $255$ quarterly observations and $246$ variables ($T = 255, N = 246$).
The number of estimated factors varies across different methods. For example, the CQFM estimate is $1$ ($3$ if (ref) is used); the rank minimization method in chenetal2021qfm reports $4$ factors at $\tau = 0.5$; the $IC_{p1}$ and $IC_{p2}$ in baiandng2002ecma give $12$ and $8$ factors, respectively. Since our focus is on the property of the estimated factor, we follow stockandwaston2012nberrecession and choose the number $6$ across different methods. The scree plot in (ref) reveals why the proposed IC with (ref) selects only one factor: the first eigenvalue is $54$ and explains about $22\%$ of the variation in the (standardized) data while the second eigenvalue is $19$ and explains about $7.8\%$ of the variance in the data.
(ref) plots the first three CQFM and PCA factors (factors 4 to 6 are plotted in (ref)). The first three factors from CQFM and PCA are very similar to each other. Despite this visual similarity, the estimated factors exhibit different moment properties. (ref) summarizes the skewness and kurtosis of the six estimated factors from the four methods. We make a few observations. First, the CQFM-based factors tend to have larger skewness and kurtosis in the first few factors. This means, if the data have large skewness and/or kurtosis, the CQFM-based factors will likely give a better fit for the component ($\lambda_{0i}' F_{0t}$). Second, even if other methods such as PCA-based factors exhibit larger skewness and/or kurtosis in later factors -- for example, the 5th PCA factor exhibits larger skewness than CQFM, these larger value will unlikely be helpful in capturing the skewness and kurtosis in the data since it is typically the first few factors that determines the overall variability of the data. Third, compared to CQFM-based factors, the QFM-based factors exhibit less skewness and kurtosis, suggesting that estimation done at a single quantile position such as $\tau = 0.5$ may not be effective in capturing certain features of the data; the composite quantile approach is more effective in this regard. The small MSEs for CQFM in (ref) attest to the above arguments.
Next, following stockandwatson2002jbes, we use the diffusion indexes to forecast one-quarter-ahead macroeconomic variables. The forecasting function is given by
where $y_{it}$ is the original data transformed according to the tcode in FRED-QD “2023-06.csv" and $\hat{F}_t$ is the estimated $6$ factors at time $t$. The forecast period starts from $2000 Q1$ to $2023Q1$, a total of $93$ forecasts for each of the $246$ macroeconomic variables, and this forecast period covers three NBER-determined recessions, including the one induced by the recent pandemic. For each rolling forecast, we use a rolling window of $120$ quarters to estimate factors and the coefficients $\beta_j$ and $\beta_F$. We forecast the data that are transformed using the tcode in mccrackenandng2020FREDQD and convert the forecast back to data in their original levels.
(ref) reports the forecast root-MSE (RMSE) for the three most common macroeconomic variables, real gross domestic product (GDPC1), civilian unemployment rate (UNRATE), and consumer price index for all urban consumers (CPIAUCSL) in the FRED-QD data set. CQFM gives good results, but its performance is not the best for the CPI data. It's also somewhat surprising that the AR(4) model can sometimes do better than factor-augmented methods. In the last row, we compute the average of RMSE over the $246$ macroeconomic variables for each of the $5$ models, and CQFM gives the smallest average RMSE. Notice that the results for the $4$ factor-based models are obtained by simply choosing $6$ factors without any additional tuning of the model. mccrackenandng2020FREDQD consider $7$ factors, and, for the regression in (ref), they try $2^7-1 = 127$ different combinations of the $7$ factors. Similar approach can also be used here to possibly improve the performance of the factor-based models. In addition, many other aspects of diffusion index modeling can be tuned to yield a favorable model, which includes, but not limited to, the number lags of the factors (we consider only $1$ in our regression), the forecast horizon (3-month, 6-month, one-year, etc.), the inclusion of lag variables in the $Y$ matrix in (ref) in factor analysis, the use of balanced panel data vs. unbalance panel data with EM-algorithm-generated data, the types of data transformation used, whether to split the data before and after a recession, among others. In the case of CQFM, we can also tune the parameter $K$ to possibly improve its performance. (ref) is a simple demonstration of the use of CQFM-based factors. A more comprehensive study is needed to further study the properties of different factor-based models.
In this paper, we develop the method of composite quantile factor model. We demonstrate in both simulations and an empirical study that, compared to PCA and several other methods, CQFM can be more effective in modeling asymmetric data due to its capability of adapting to data at multiple quantile positions. Asymptotic distributional theory and an information criterion for consistent factor number selection are also discussed. PCA-based method is popular for factor analysis, and CQFM will be a useful addition to a researcher's toolkit when handling non-normal data.
Many extensions of the current research are possible, and we give two examples that are highly relevant to data modeling. One is the creation of sparsity in CQFM. Adding penalty functions to (ref) gives
zouandyuan2008cqr use the adaptive lasso in zou2006jasaadplasso to induce sparsity in linear regression, and many other penalty functions are available for $F$ and $\Lambda$. The other example, following the work in bai2009ecma, is to add a regression component to (ref) so that it becomes the panel data model with interactive fixed effects
where $X_{it}$ is a vector of regressors. The model in (ref) is a hybrid of CQR in zouandyuan2008cqr and CQFM in the current paper. Given the good finite sample properties of CQR and CQFM under certain non-normal data, we expect estimators from (ref) will also show some robustness to non-normal data. We leave these topics for future research.
\spacing{1.45}
\setcounter{page}{1} \spacing{1.42}