Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
109,382 characters · 21 sections · 81 citation commands
When can weak latent factors be statistically inferred?
\theoremstyle{plain} \newtheorem{lemma}{Lemma} \newtheorem{prop}{Proposition}\newtheorem{theorem}{Theorem}\setcounter{theorem}{0} \newtheorem{corollary}{Corollary} \newtheorem{assumption}{Assumption} \newtheorem{example}{Example} \newtheorem{definition}{\textbf{Definition}} \newtheorem{fact}{\textbf{Fact}} \newtheorem{condition}{\textbf{Condition}}\theoremstyle{definition}\theoremstyle{remark}\newtheorem{remark}{\textbf{Remark}}\newtheorem{claim}{\textbf{Claim}}\newtheorem{conjecture}{\textbf{Conjecture}}
Keywords: factor model, principal component analysis, weak factors, cross-sectional correlation, inference, signal-to-noise ratio.
\setcounter{tocdepth}{2}
The factor model, a pivotal tool for analyzing large panel data, has become a significant topic in finance and economics research cham1983factor,FamaFrench1993factors,StockWatson2002PCA,BaiNg2002ECTA,GigXiu2021JPE,fan2021recent. The estimation and inference for factor models are crucial in economic studies, particularly in areas like asset pricing and return forecasting. In the era of big data, the factor model has gained increased prominence in capturing the latent common structure for large panel data, where both the cross-sectional and temporal dimensions are ultra-high BaiPeng2016ARFEfactor,FanLiLiao2021ARFEfactor. Principal component analysis (PCA), known for its simplicity and effectiveness, is closely connected with the factor model and has long been a key research topic of interest in the econometric community StockWatson2002PCA,BaiNg2002ECTA,Bai2003ECTA,Onatski2012PCAweak,POET2013,BaiNg2013identifiPCA,BaiNg2023PCA.
As pointed out by GiglioXiu2023weakFactorPredict, most theoretical guarantees for the PCA approach to factor analysis rely on the pervasiveness assumption BaiNg2002ECTA,Bai2003ECTA. This assumption requires the signal-to-noise ratio (SNR), which measures the factor strength relative to the noise level, to grow with the rate of $\sqrt{N}$ -- the square root of the cross-sectional dimension. However, many real datasets in economics do not exhibit sufficiently strong factors to meet this pervasiveness assumption. When the SNR grows slower than $\sqrt{N}$, the resulting model is often called the weak factor model Onatski2009number,Onatski2010number. Extensive research has been dedicated to the weak factor model Onatski2012PCAweak,BKP2021weakFactor,frey2022weakFactorNumber,YoTa2022sWF_estimate,YoTa2022sWF_infer,BaiNg2023PCA,Jiang2023PCA,ChoiMing2024PCA, among which the PC estimators (the estimators via the PCA approach) have been a primary subject. Recently, BaiNg2023PCA,Jiang2023PCA,ChoiMing2024PCA studied the consistency and asymptotic normality of the PC estimators for factors and factor loadings in the weak factor model. In the extreme case (also called the super-weak factor model) where the SNR is $O(1)$, Onatski2012PCAweak showed that the PC estimators are inconsistent.
This paper establishes a novel and comprehensive theory for PCA in the weak factor model. Our theory is non-asymptotic and can be easily translated to asymptotic results. The conditions we propose for the asymptotic normality of PC estimators are optimal in the sense that, surprisingly different from the existing literature, the required growth rate of the SNR for consistency aligns with that for asymptotic normality, differing only by a logarithmic factor. In particular, in the regime $N\asymp T$, where $T$ is the temporal dimension, we prove that asymptotic normality holds as long as the SNR grows faster than a polynomial rate of $\log N$, and this result is a substantial advance compared with the existing results that require the SNR to grow with a polynomial rate of $N$ BaiNg2023PCA,Jiang2023PCA,ChoiMing2024PCA.
The most innovative part of our theory lies in establishing a closed-form, first-order approximation of the PC estimator, that is, we decompose the PC estimator into three components --- the ground truth, a first-order term, and the higher-order negligible term. We express the first-order term explicitly using the parameters in the factor model. This closed-form characterization paves the way for us to establish the asymptotic normality and design various test statistics for practical applications. Our theory is based on novel applications of the leave-one-out analysis and matrix concentration inequalities. Moreover, our findings provide valuable insights and practical implications for a range of econometric problems related with PCA, e.g., macroeconomic forecasting based on factor-augmented regressions BaiNg2006factorAugmented.
We demonstrate the practical applications of our theories using both synthetic and real datasets. First, we design an easy-to-implement statistic for factor specification test BaiNg2006factor --- whether an observed factor is in the linear space spanned by the latent common factors or not. A key innovation of our approach is the ability to conduct this test in any flexible subset of the whole period, owing to the row-wise error bound in our theory. We utilize the monthly return data of the S&P 500 constituents from 1995 to 2024, along with the data of the Fama-French three factors, then run the factor test in a rolling window manner FanLiaoYao2015power. Our results uncover a notable decline in the importance and explanatory power of the size factor during the 2008 financial crisis, and of both the size and value factors during the COVID-19 pandemic around 2019. These findings are supported by our test results, which reject the null hypothesis that these factors are in the linear space of latent factors.
Then, we design a test statistic for the structural break of betas StockWatson2009beta_break,breitung2011beta_break, and apply it to the aforementioned S&P 500 constituents data. Our novelty is that our test statistic works in the weak factor model without the pervasiveness assumption, which is required in prior work. We test for each stock to determine if the beta has changed before and after the three recessions covered by our data --- the Early 2000s Recession, the 2008 Great Recession, and the COVID-19 Recession. We find that in each recession, the sectors most affected by the shocks, where many stocks exhibited structural breaks in betas, correspond reasonably to the causes of these economic recessions. For example, during the 2008 Great Recession, marked by the subprime mortgage crisis, the financial sector experienced a strong impact, which aligns with its exposure to mortgage-backed securities and other related financial instruments. During the COVID-19 Recession, the Health Care and Real Estate sectors were significantly impacted, reflecting the uncertainties brought about by lockdowns and health crises due to the pandemic. Additionally, we develop statistical tests for the betas to evaluate the similarity in risk exposure between two stocks and construct valid confidence interval for the systematic risk of each stock.
The rest of the paper is organized as follows. Section (ref) introduces the model setup, basic assumptions, and notation. In Section (ref), we show our main results on the first-order approximations for the PC estimators, and present the asymptotic normality results as corollaries. Section (ref) provides a detailed comparison of our results with related work. In Section (ref), we showcase four applications in econometrics based on our main results. Section (ref) collects the numerical results in both simulated data and real data. More related works are discussed in Section (ref), and the paper concludes with a discussion on future directions in Section (ref).
Let $N$ be the number of cross-sectional units and $T$ be the number of observations. Consider the factor model for a panel of data $\{x_{i,t}\}:$
where $\bm{f}_{t}=(f_{1,t},f_{2,t},\ldots,f_{r,t})^{\top}$ is the latent factor, $r$ is the number of factors, $\bm{b}_{i}=(b_{i,1},b_{i,2},\ldots,b_{i,r})^{\top}$ is a vector of factor loadings, and $\varepsilon_{i,t}$ represents the idiosyncratic noise. Viewing $x_{i,t}$ as the excess return of the $i$-th asset at time $t$, the model ((ref)) is intimately linked with the multi-factor pricing model. This model originates from the Arbitrage Pricing Theory (APT) developed by Ross1976APT and finds extensive applications in finance.
To compact the notation, we denote by $\bm{B}=(\bm{b}_{1},\bm{b}_{2},\ldots,\bm{b}_{N})^{\top}$ the $N\times r$ factor loading matrix. Let $\bm{x}_{t}=(x_{1,t},x_{2,t},\ldots,x_{N,t})^{\top}$ and $\bm{e}_{t}=(\varepsilon_{1,t},\varepsilon_{2,t},\ldots,\varepsilon_{N,t})^{\top}$. Then, the factor model ((ref)) can be expressed as $\bm{x}_{t}=\bm{Bf}_{t}+\bm{e}_{t}$, or written in a matrix form as follows
where $\bm{X}=(\bm{x}_{1},\bm{x}_{2},\ldots,\bm{x}_{T})$, $\bm{F}=(\bm{f}_{1},\bm{f}_{2},\ldots,\bm{f}_{T})^{\top}$, and $\bm{E}=(\bm{e}_{1},\bm{e}_{2},\ldots,\bm{e}_{T})$ are the $N\times T$ panel data, the $T\times r$ factor realizations, and the $N\times T$ idiosyncratic noise matrix, respectively.
In our setup, the only observable part is the panel data $\bm{X}$. We are interested in the estimation and inference for both the latent factors and factor loadings via the PCA approach.
We present some basic assumptions as follows. Note that the factor model has the rotation (indeed affine transform) ambiguity BaiNg2013identifiPCA, that is, the factor loadings and latent factors are not identifiable since $\bm{BF}^{\top}=(\bm{B}\bm{H}^{-1})(\bm{F\bm{H}^{\top}})^{\top}$ holds for any invertible matrix $\bm{H}$. Without loss of generality, we assume that the columns of $\bm{B}$ are orthogonal and the covariance of $\bm{f}_{t}$ is the identity matrix, as stated in Assumption (ref) below.
Assumption (ref) is a standard identifiability condition for the factor model POET2013. The singular values $\{\sigma_{i}\}$ of the factor loading matrix $\bm{B}$ characterize the strengths of the latent factors.
Next, to accommodate the cross-sectional correlation in the noise and to facilitate leave-one-out analysis in our technical proof, we propose Assumption (ref) for the noise matrix $\bm{E}$, specifying its structure and the distribution of its entries.
The nonzero off-diagonal entries of $\bm{\Sigma}_{\varepsilon}$ characterize the cross-sectional correlations in the idiosyncratic noise $\bm{e}_{t}=(\varepsilon_{1,t},\varepsilon_{2,t},\ldots,\varepsilon_{N,t})^{\top}$. Though the noise terms $\bm{e}_{1},\bm{e}_{2},\ldots,\bm{e}_{T}$ are independent under Assumption (ref), our theory could potentially be generalized to the cases where the temporal correlations are present in the noise matrix $\bm{E}$, and this generalization is an interesting future direction. The formulation ((ref)) imposes additional structural constraints on the noise matrix. However, such assumptions are commonplace in the study of the weak factor model; see similar assumptions in Onatski2010number,Onatski2012PCAweak.
Then, to establish the non-asymptotic results via the concentration inequalities, we propose the following assumption on the distribution of the factors.
The conditions in Assumption (ref) are slightly stronger than those in the literature, e.g., BaiNg2023PCA, but remain standard and simplify our technical proofs.
We introduce some notation that will be used throughout the paper. For two sequences $a_{n}$ and $b_{n}$, we write $a_{n}\lesssim b_{n}$ (or equivalently, $b_{n}\gtrsim a_{n}$) if $a_{n}=O(b_{n})$, i.e., there exist a constant $C>0$ and an integer $N>0$ such that, $a_{n}\leq Cb_{n}$ holds for any $n>N$; write $a_{n}\asymp b_{n}$ if both $a_{n}\lesssim b_{n}$ and $a_{n}\gtrsim b_{n}$ hold and $a_{n}\ll b_{n}$ (or equivalently, $b_{n}\gg a_{n}$) if $a_{n}=o(b_{n})$.
For a symmetric matrix $\bm{M}$, we denote by $\lambda_{\min}(\bm{M})$ and $\lambda_{\max}(\bm{M})$ its minimum and maximum eigenvalues. For a vector $\bm{v}$, we denote $\Vert\bm{v}\Vert_{2}$, $\Vert\bm{v}\Vert_{1}$, and $\Vert\bm{v}\Vert_{\infty}$ as the $\ell_{2}$-norm, $\ell_{1}$-norm, and supremum norm, respectively. Consider any matrix $\bm{A}\in\mathbb{R}^{m\times n}$. We denote by $\bm{A}_{i,\cdot}$ and $\bm{A}_{\cdot,j}$ the $i$-th row and the $j$-th column of $\bm{A}$. We let $\Vert\bm{A}\Vert_{2},$ $\Vert\bm{A}\Vert_{\mathrm{F}},$ and $\Vert\bm{A}\Vert_{2,\infty}$ denote the spectral norm, the Frobenius norm, and the $\ell_{2,\infty}$-norm (i.e., $\Vert\bm{A}\Vert_{2,\infty}:=\sup_{1\leq i\leq m}\Vert\bm{A}_{i,\cdot}\Vert_{2}$), respectively. For an index set $S\subseteq\{1,2,\ldots,m\}$ (resp. $S\subseteq\{1,2,\ldots,n\}$), we use $|S|$ to denote its cardinality, and use $\bm{A}_{S,\cdot}$ (resp. $\bm{A}_{\cdot,S}$) to denote a submatrix of $\bm{A}$ whose rows (resp. columns) are indexed by $S$. We let $col(\bm{A})$, $\bm{A}^{+}$, and $\bm{P}_{\bm{A}}:=\bm{A}\bm{A}^{+}$ denote the column subspace of $\bm{A}$, the generalized inverse of $\bm{A}$, and the projection matrix onto the column space of $\bm{A}$.
We denote by $\mathsf{diag}(a_{1},a_{2},\ldots,a_{r})$ the $r\times r$ diagonal matrix whose diagonal entries are given by $a_{1},a_{2},\ldots,a_{r}$. Let $\bm{I}_{r}$ be the $r\times r$ identity matrix. We denote by $\mathcal{O}^{r\times r}$ the set of all $r\times r$ orthonormal (or rotation) matrices. For a non-singular $n\times n$ matrix $\bm{H}$ with SVD $\bm{U}_{H}\bm{\Sigma}_{H}\bm{V}_{H}^{\top}$, we denote by $\mathsf{sgn}(\bm{H})$ the following orthogonal matrix $\mathsf{sgn}(\bm{H})\coloneqq\bm{U}_{H}\bm{V}_{H}^{\top}$. Then we have that, for any two matrices $\widehat{\bm{U}},\bm{U}\in\mathbb{R}^{n\times r}$ with $r\leq n$, among all rotation matrices, the one that best aligns $\widehat{\bm{U}}$ and $\bm{U}$ is precisely $\mathsf{sgn}(\widehat{\bm{U}}^{\top}\bm{U})$ ma2017implicit, namely, \[ \mathsf{sgn}(\widehat{\bm{U}}^{\top}\bm{U})=\arg\min_{\bm{O}\in\mathcal{O}^{r\times r}}\Vert\widehat{\bm{U}}\bm{O}-\bm{U}\Vert_{\mathrm{F}}. \]
In this section, we demonstrate the desirable statistical performance of PCA in weak factor models. We will first present a master result (cf. Theorem (ref)) on subspace error decomposition. Based on this key result, we derive non-asymptotic distributional characterization for our estimators for the factors and factor loadings, paving the way to data-driven statistical inference.
First, we formally introduce the PC estimators of the factor and factor loadings in Algorithm (ref).
Our master result focuses on the subspace estimates $\widehat{\bm{U}}$ and $\widehat{\bm{V}}$. We define some relevant quantities first. Denote by $\bm{U}\bm{\Lambda}\bm{V}^{\top}$ the SVD of $T^{-1/2}\bm{BF}^{\top}$, where both $\bm{U}\in\mathbb{R}^{N\times r}$ and $\bm{V}\in\mathbb{R}^{T\times r}$ have orthonormal columns, and $\bm{\Lambda}=\mathsf{diag}(\lambda_{1},\lambda_{2},\ldots,\lambda_{r})\in\mathbb{R}^{r\times r}$ satisfies that $\lambda_{1}\geq\lambda_{2}\geq\cdots\geq\lambda_{r}\geq0$. We define $n:=\max(N,T)$.
Before presenting our master results, we list two key assumptions on the SNR. To characterize the SNR, we define \[ \theta:=\sigma_{r}/\sqrt{\Vert\bm{\Sigma}_{\varepsilon}\Vert_{2}},\qquad\vartheta_{k}:=\sigma_{r}/\Vert(\bm{\Sigma}_{\varepsilon}^{1/2})_{k,\cdot}\Vert_{1},\text{\qquad and\qquad}\vartheta:=\sigma_{r}/\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{1}. \] Here, the signal strength, which is the factor strength in our case, is effectively represented by $\sigma_{r}$, the smallest singular value of the factor loading matrix $\bm{B}$. The noise level is captured by the norm of the covariance matrix $\bm{\Sigma}_{\varepsilon}$ for the idiosyncratic noise $\bm{E}=\bm{\Sigma}_{\varepsilon}^{1/2}\bm{Z}$ (cf. ((ref))). The SNRs $\theta$, $\vartheta_{k}$, and $\vartheta$ are similar, though their denominators adopt different norms of $\bm{\Sigma}_{\varepsilon}$. The inclusion of $\ell_{1}$-norms $\Vert(\bm{\Sigma}_{\varepsilon}^{1/2})_{k,\cdot}\Vert_{1}$ and $\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{1}=\max_{1\leq k\leq N}\Vert(\bm{\Sigma}_{\varepsilon}^{1/2})_{k,\cdot}\Vert_{1}$, though less common than the spectral norm $\sqrt{\Vert\bm{\Sigma}_{\varepsilon}\Vert_{2}}=\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{2}$, is that our main results are on the row-wise error bound and we need these norms in the matrix inequalities for technical reasons. Then, two crucial assumptions on SNR are as follows.
Assumption (ref) implies Assumption (ref) due to the elementary inequality $\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{2}\leq\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{1}$. On the other hand, $\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{1}\leq s\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{2}$, where $s$ is the maximum number of nonzero elements of $\bm{\Sigma}_{\varepsilon}^{1/2}$ in each row, which is small when $\bm{\Sigma}_{\varepsilon}$ is sparse. In particular, when the noise covariance matrix $\bm{\Sigma}_{\varepsilon}$ is diagonal (i.e., no cross-sectional correlations, also known as the strict factor model Ross1976APT,fanfanlv2008covFactor,baishi2011Cov), we have that $\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{2}=\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{1}$. In this case, Assumption (ref) is equivalent with Assumption (ref). Throughout the paper, we denote by $C_{0}$, $C_{1}$, $C_{2}$, $C_{U}$, $C_{V}$, $c_{0}$, etc. the generic constants that may vary from place to place. We are now ready to present our main results as follows.
The perturbation bounds in Theorem (ref) pave the way for the statistical inference of the factor loadings and factors. The reasons why the quantities presented in Theorem (ref) relate to the factor loadings $\bm{B}$ and factors $\bm{F}$ are that, as demonstrated in Lemma (ref), the column subspaces of $\bm{U}$ and $\bm{B}$ are identical, as are those of $\bm{V}$ and $\bm{F}$.
To show how strong our SNR condition is, we examine scenarios where there is no correlation in the noise matrix $\bm{E}$, i.e., $\bm{\Sigma}_{\varepsilon}$ is diagonal as in the strict factor model. In this case, the SNRs satisfy $\vartheta=\theta=\sigma_{r}/\sigma_{\varepsilon}$, where $\sigma_{\varepsilon}=\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{2}=\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{1}$. Then Assumptions (ref) and (ref) match the minimal condition $\theta\gg\sqrt{n/T}$ required for the consistent estimations of the left and right singular subspaces yan2021inference, up to log factors.
To explain that why $\bm{\Psi}_{U}$ and $\bm{\Psi}_{V}$ are higher-order negligible terms, we compare their bounds with the magnitude of the first-order terms $\bm{G}_{U}$ and $\bm{G}_{V}$. For simplicity, we consider the setting where $r\asymp1$, $N\asymp T$, $\bm{\Sigma}_{\varepsilon}$ is well-conditioned (i.e., $\kappa_{\varepsilon}:=\lambda_{\text{max}}(\bm{\Sigma}_{\varepsilon})/\lambda_{\text{min}}(\bm{\Sigma}_{\varepsilon})\asymp1$), and $\bm{U}$ is incoherent chen2021Monograph; the matrix $\bm{V}$ can be proven to be incoherent since $\Vert\bm{V}\Vert_{2,\infty}\lesssim T^{-1/2}\log^{1/2}n$. Then, the typical size of the $k$-th row of the first-order term $\bm{G}_{U}$ can be measured by $\mathbb{E}^{1/2}[\Vert(\bm{G}_{U})_{k,\cdot}\Vert_{2}^{2}]$. Standard computations yield $\mathbb{E}^{1/2}[\Vert(\bm{G}_{U})_{k,\cdot}\Vert_{2}^{2}]\geq T^{-1/2}\sigma_{r}^{-1}\sqrt{\lambda_{\text{min}}(\bm{\Sigma}_{\varepsilon})}$. Similarly, for the first-order term $\bm{G}_{V}$, we have that $\mathbb{E}^{1/2}[\Vert(\bm{G}_{V})_{l,\cdot}\Vert_{2}^{2}]\geq T^{-1/2}\sigma_{r}^{-1}\sqrt{\lambda_{\text{min}}(\bm{\Sigma}_{\varepsilon})}$. Then, using the row-wise bounds for $\bm{\Psi}_{U}$ and $\bm{\Psi}_{V}$ in Theorem (ref), we obtain that
The last inequality in ((ref)) (resp. ((ref))) implies that the first-order term $(\bm{G}_{U})_{k,\cdot}$ (resp. $(\bm{G}_{V})_{l,\cdot}$) dominates the higher-order term $(\bm{\Psi}_{U})_{k,\cdot}$ (resp. $(\bm{\Psi}_{V})_{l,\cdot}$), provided that the SNRs grow at polynomial rate of $\log n$: $\theta^{-1}\vartheta_{k}\vartheta\gg\log^{3/2}n$ and $\theta\gg\log^{3/2}n$ (resp. $\theta\gg\log^{3/2}n$). These conditions are less stringent than the assumptions in the prior work BaiNg2023PCA,Jiang2023PCA,ChoiMing2024PCA that required the SNR to grow with a polynomial rate of $n$.
As we will show later, Theorem (ref) is the foundation to design test statistics and conduct statistical inference for the factors, factor loadings, and other parameters of interest. To the best of our knowledge, the non-asymptotic first-order approximations ((ref))--((ref)) are new in the statistics and econometrics literature. While similar findings exist in the fields of low-rank matrix completion and spectral methods, the closest result to ours is Theorem 9 of yan2021inference. However, they assumed that the entries of the noise matrix $\bm{E}$ are independent, while we allow the presence of cross-sectional correlations in the noise matrix $\bm{E}$.
Our results in Theorem (ref) are nontrivial generalizations of those in yan2021inference. We note that their proof of row-wise error bounds via leave-one-out (LOO) technique requires constructing an auxiliary matrix that is independent with the target row. This construction is obtained by zeroing the row with the same position in noise matrix. However, this approach does not work in our case because it is still correlated with the target row due to the cross-sectional correlation among all rows. We overcome this challenge through applications of matrix concentration inequalities, and our results do not need any additional structural assumption (e.g., sparsity) on the noise covariance matrix $\bm{\Sigma}_{\varepsilon}$. An additional advance of our theory compared with yan2021inference lies in that, our bounds for $\Vert(\bm{\Psi}_{U})_{k,\cdot}\Vert_{2}$ and $\Vert(\bm{\Psi}_{V})_{l,\cdot}\Vert_{2}$ are free of condition number $\kappa:=\sigma_{1}/\sigma_{r}$, making our theory work even when the factor loading matrix $\bm{B}$ is near-singular and ill-conditioned, i.e., $\kappa$ is very large. In particular, our result accomodates the heterogeneous case of BaiNg2023PCA.
In this section, we present the immediate consequences of Theorem (ref) --- estimation error bounds and distributional characterization for the PC estimators. Later in Section (ref), we will compare these results with the recent work (BaiNg2023PCA,Jiang2023PCA,ChoiMing2024PCA), and highlight the advantages of our theory.
Recall that our estimators for the factors $\bm{F}$ and factor loadings $\bm{B}$ in Algorithm (ref) are given by $\widehat{\bm{F}}=\sqrt{T}\widehat{\bm{V}}$ and $\widehat{\bm{B}}=\widehat{\bm{U}}\widehat{\bm{\bm{\Sigma}}}$ respectively. Let us take a moment to see how to quantify the estimation error for $\widehat{\bm{F}}$ in the face of rotational ambiguity. In view of Theorem (ref), we know that $\widehat{\bm{F}}\bm{R}_{V}$ should be close to $\sqrt{T}\bm{V}$ in Euclidean distance. In addition, recall that $\bm{U}\bm{\Lambda}\bm{V}^{\top}$ is the SVD of the rank-$r$ matrix $T^{-1/2}\bm{BF}^{\top}$, which suggests the existence of an invertible, $\sigma(\bm{F})$-measurable matrix $\bm{J}\in\mathbb{R}^{r\times r}$ such that $\bm{V}=T^{-1/2}\bm{FJ}$. Hence, by defining $\bm{R}_{F}\coloneqq\bm{J}\bm{R}_{V}^{\top}$, we may use $\widehat{\bm{F}}-\bm{F}\bm{R}_{F}$ to evaluate the estimation error of $\widehat{\bm{F}}$. Similarly, we define $\bm{R}_{B}\coloneqq(\bm{R}_{F}^{-1})^{\top}$ and use $\widehat{\bm{B}}-\bm{B}\bm{R}_{B}$ to measure the estimation error of $\widehat{\bm{B}}$.
We will show in Lemma (ref) that $\bm{R}_{F}$ and $\bm{R}_{B}$ are close to rotation matrices. Indeed, in the study of PC estimators for factor model (e.g., BaiNg2023PCA,Jiang2023PCA,ChoiMing2024PCA), it is quite common to take $\bm{F}\bm{R}_{F}$ and $\bm{B}\bm{R}_{B}$, rather than $\bm{F}$ and $\bm{B}$, as the groundtruth, though the specific choices of $\bm{R}_{F}$ and $\bm{R}_{B}$ vary across different works. Both the averaged and the row-wise estimation error bound can be deduced from Theorem (ref). We record the results in the following corollary.
Let us interpret the above bounds under some specific settings. For simplicity, we ignore the log factors $\log n$ and consider the setting discussed after Theorem (ref), with the assumption that $\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{1}\asymp\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{2}\asymp1$. Under this setting, to make the assumptions on $\Vert\bar{\bm{U}}\Vert_{2,\infty}$ hold, it suffices to assume that $\theta\gtrsim1$, where $\theta\equiv\sigma_{r}/\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{2}$ is the SNR. For factors, the upper bounds for the averaged and row-wise estimation error rates in ((ref)) are given by $\theta^{-2}$ and $\theta^{-1}$, respectively; for factor loadings, the two bounds in ((ref))--((ref)) are given by $T^{-1}+N^{-1}\theta^{-2}$ and $T^{-1/2}+N^{-1/2}\theta^{-1}$, respectively. All these bounds go to zero under the condition that $\theta\gg1$, which match the minimal condition required for the consistent estimations of the left and right singular subspaces if there is no correlation in the noise matrix yan2021inference.
Next, we present our results on the inference for factors and factor loadings in Corollaries (ref) and (ref), respectively.
Let us look at the assumptions about the SNRs $\theta$ and $\vartheta_{i}$ in Corollaries (ref) and (ref). Consider the setting where $r\asymp1$, $N\asymp T$, $c_{\varepsilon}\asymp1$, $\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{1}\asymp\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{2}\asymp1$, $\Vert(\bm{\Sigma}_{\varepsilon}^{1/2})_{i,\cdot}\Vert_{1}\asymp1$, and $\bm{\Sigma}_{\varepsilon}$ is well-conditioned. Then, the SNR condition in Corollary (ref) (resp. Corollary (ref)) is equivalent to $\theta\gg\log^{3/2}N$ (resp. $\theta\gg\log N$ and $\vartheta_{i}\gg\log^{3/2}N$), indicating that inference of factors and factor loadings is achievable as long as the SNR grows faster than a polynomial rate of $\log N$. Our SNR condition is less restrictive than the prior work BaiNg2023PCA,Jiang2023PCA,ChoiMing2024PCA that required the SNR to grow with a polynomial rate of $N$. Regardless of log factors, our SNR condition for inference is equivalent to $\theta\gg1$, which is optimal since it matches the same condition required for consistency as commented after Corollary (ref). Also, in the special case where $\bm{\Sigma}_{\varepsilon}$ is diagonal, i.e., there is no correlation in the noise matrix, our SNR condition for inference matches the minimal condition required for the subspace inference yan2021inference.
Both Corollaries (ref) and (ref) are stated in a non-asymptotic sense, and can be easily translated to asymptotic normality results. In practice, the confidence regions for the factors and the factor loadings can be constructed by replacing the asymptotic covariance matrices $\bm{\Sigma}_{F,t}$ and $\bm{\Sigma}_{B,i}$ with their consistent estimators: \[ \widehat{\bm{\Sigma}}_{F,t}=\widehat{\bm{\Sigma}}^{-1}\widehat{\bm{U}}^{\top}\widehat{\bm{\bm{\Sigma}}}_{\varepsilon}^{\tau}\widehat{\bm{U}}\widehat{\bm{\Sigma}}^{-1}\qquad\text{and}\qquad\widehat{\bm{\Sigma}}_{B,i}=\frac{1}{T}(\widehat{\bm{\bm{\Sigma}}}_{\varepsilon}^{\tau})_{i,i}\bm{I}_{r}, \] respectively. Here, $\widehat{\bm{\bm{\Sigma}}}_{\varepsilon}^{\tau}$ is the estimator of the noise covariance matrix which we will introduce in detail in Section (ref). The intuition behind the above consistent estimators is that, $\widehat{\bm{U}}$ and $\widehat{\bm{\Sigma}}$ are the consistent estimators of $\bm{U}\bm{R}_{U}$ and $\bm{R}_{U}^{\top}\bm{\Lambda}\bm{R}_{V}$, respectively, and we will show this fact in the proof of Theorem (ref).
In this section, we compare our main results with related ones established in prior works BaiNg2023PCA,Jiang2023PCA,ChoiMing2024PCA. All of them studied the PC estimators in weak factor models, and established estimation error bounds and asymptotic normality under different assumptions. Distinct from these three papers, all our results are entirely non-asymptotic, providing finite-sample characterizations for both the estimation error and the uncertainty level of statistical inference. From a technical viewpoint, in the regime $N\asymp T$, a notable advancement of our theory is that, our assumptions for inference require the growth rate of SNR to be faster than a polynomial rate of $\log N$, unlike the polynomial rate of $N$ required in these three papers. Also, all these three papers assume that the $r$ singular values $\{\sigma_{i}\}_{i=1}^{r}$ of factor loading $\bm{B}$ are distinct, while our theory does not require such eigengap condition.
Having provided an overview of the advantages of our theory, we now make the comparisons in detail.
To enhance the clarity of the comparison of the SNR assumption for inference, we detailed the results in Table (ref) under a specific setup: Assumptions (ref), (ref), (ref) hold; The cross-sectional and temporal dimensions satisfy that $N\asymp T\asymp n$; the number of factors satisfies $r\asymp1$; the noise covariance matrix satisfies $\kappa_{\varepsilon}\asymp1$, $\left\Vert \bm{\Sigma}_{\varepsilon}\right\Vert _{2}\asymp1$, and $\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{1}\asymp1$. Since the scale of the noise is reflected by $\bm{\Sigma}_{\varepsilon}$ and in our current setup we have that $\left\Vert \bm{\Sigma}_{\varepsilon}\right\Vert _{2}\asymp1$ and $\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{1}\asymp1$, the assumptions on the SNRs $\theta=\sigma_{r}/\sqrt{\left\Vert \bm{\Sigma}_{\varepsilon}\right\Vert _{2}}$ and $\vartheta=\sigma_{r}/\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{1}$ defined in Section (ref) are fully represented by the growth rate of the smallest singular value $\sigma_{r}$ of factor loading matrix $\bm{B}$.
Table (ref) demonstrates the advancement of our theory. For the conditions of ChoiMing2024PCA in Table (ref), the parameter $\omega>0$ is defined in Assumption B” (iv) therein on the noise structure and they need two additional growth rate assumptions on $\sigma_{r}$ (see Assumption D” (i) therein). For other regimes where $T\ll N$ or $N\ll T$, similar comparisons can be made, so we omit the details for the limit of space.
Our first-order approximations do not only advance the existing theories on the PCA under the weak factor model, but also pave the path for various statistical tests that are useful in economics and finance. In this section, we show four applications of our results.
Recall that the factor model is given by $\bm{X}=\bm{BF}^{\top}+\bm{E}$, where the matrix $\bm{F}$ is the realization of latent factors. Suppose we have time series data of some observed factors, e.g., the Fama-French factors. Our focus is on testing if the observed factor is in the linear space spanned by the latent factors $\bm{F}$. In particular, we examine this linear dependence in a flexible range of the whole period $[1,T]$. Formally, we consider an index set $S=\{t_{1},t_{2},\ldots,t_{|S|}\}\subseteq\{1,2,\ldots,T\}$ of interest, and we have the data $\bm{v}=(v_{t_{1}},v_{t_{2}},\ldots,v_{t_{|S|}})^{\top}\in\mathbb{R}^{|S|}$ of an observed factor recorded at the time index set $S$. We test the hypothesis as follows,
Under the null hypothesis $H_{0}$, we have that $v_{t}=\bm{F}_{t,\cdot}\bm{w}=\bm{f}_{t}^{\top}\bm{w}$ for any $t\in S$.
Under the strong factor model where the pervasiveness assumption holds, BaiNg2006factor studied this problem and designed test statistics for the whole set, i.e., $S=\{1,2,\ldots,T\}$. Our study extends to the case when $S$ is a subset of the whole time span $\{1,2,\ldots,T\}$ under the weak factor model. The subset $S$ can be any specific time window of interest. This scenario is economically meaningful because the relationship $v_{t}=\bm{f}_{t}^{\top}\bm{w}$ may only be valid for relatively short periods, not necessarily across the entire span; see BaiNg2006factor for the connections between the CAPM analysis and this problem. In Section (ref), we show that our factor specification test results strikingly reconcile with the economic cycles and financial crisis.
Our test statistic relies on the estimation of the noise covariance matrix $\bm{\Sigma}_{\varepsilon}$. We denote by $\widehat{\bm{\bm{\Sigma}}}_{\varepsilon}^{\tau}$ the estimator of $\bm{\Sigma}_{\varepsilon}$ and we will elaborate its construction in the next section. In Theorem (ref) below, we keep the error bound $\Vert\widehat{\bm{\bm{\Sigma}}}_{\varepsilon}^{\tau}-\bm{\Sigma}_{\varepsilon}\Vert_{2}$ in the final result.
The idea of the test statistic $\mathfrak{\widehat{\mathcal{T}}}(S,\bm{v})$ is to utilize the fact that, under the null hypothesis $H_{0}$, the residual vector of $\bm{v}$ is zero after projection onto the column space of $\bm{F}_{S,\cdot}$. Note that $\widehat{\bm{V}}_{S,\cdot}$ estimates the column space of $\bm{F}_{S,\cdot}$. So, we construct $\mathfrak{\widehat{\mathcal{T}}}(S,\bm{v})$ by computing the $\ell_{2}$-norm of the projection residual vector $\Vert(\bm{I}_{|S|}-\bm{P}_{\widehat{\bm{V}}_{S,\cdot}})\bm{v}\Vert_{2}^{2}=\bm{v}^{\top}(\bm{I}_{|S|}-\bm{P}_{\widehat{\bm{V}}_{S,\cdot}})\bm{v}$. The component $\widehat{\phi}$ estimates the variance of $(\bm{I}_{|S|}-\bm{P}_{\widehat{\bm{V}}_{S,\cdot}})\bm{v}$, and needs the estimator $\widehat{\bm{\bm{\Sigma}}}_{\varepsilon}^{\tau}$ of noise covariance matrix as input.
Our focus is on the regime where the size of subset $S$ is small. This is economically meaningful because the linear relationship between observed factors and latent factors usually holds only for a short horizon. Based on our distributional results in (ref), the case when $S$ is the whole period, i.e., $|S|=T$, can be handled straightforwardly by applying the Gaussian approximation to the Chi-square distribution. Our assumption on the size $|S|$ can be summarized as $r+\log n\ll|S|\ll T/(\kappa_{\varepsilon}^{2}r^{2}\log^{4}n)$. On the one hand, we require that $|S|$ cannot be too small so that $\bm{F}_{S,\cdot}$ is of full column rank and its singular value is close to $\sqrt{|S|}$, which is crucial for our proof. On other hand, we require that $|S|$ cannot be too large, otherwise the null distribution $\chi^{2}(|S|-r)$ with mean $|S|-r$ would diverge in such a case, complicating the proof of proximity between $\chi^{2}(|S|-r)$ and the law of test statistic $\mathfrak{\widehat{\mathcal{T}}}(S,\bm{v})$.
The assumption on SNR is that $\theta\gg\frac{n}{T}\sqrt{|S|}\kappa_{\varepsilon}^{4}r^{2}\log^{4}n$. Consider the setting where $r\asymp1$, $N\asymp T$, and $\bm{\Sigma}_{\varepsilon}$ is well-conditioned (i.e., $\kappa_{\varepsilon}\asymp1$). Then, the SNR condition becomes $\theta\gg\sqrt{|S|}\log^{4}n$, illustrating the effect of the subset size $|S|$ on the growth rate required for SNR $\theta$. The $\sqrt{|S|}$ factor appears in the SNR because we use the Chi-square distribution as the null distribution to accommodate the general case when $|S|$ is small. When $|S|\asymp1$, regardless of the log factors, the SNR condition is equivalent to $\theta\gg1$, which match the minimal condition required for the consistent estimations of the left and right singular subspaces if there is no correlation in the noise matrix yan2021inference.
BaiNg2006factor studied the factor specification test under the whole period case, i.e., $S=\{1,2,\ldots,T\}$. They derived the approximate cumulative distribution function (CDF) for their test statistic, but they did not provide a finite sample error bound for it and their results were under the strong factor model. In comparison, our results give a precise characterization of the error and adapt to the weak factor model. In our case, both the low SNR in the weak factor model and the subset $S$ in the hypothesis $H_{0}$ make it challenging to design the test statistics. The subspace perturbation bounds in Theorem (ref) enable us to conquer these difficulties and establish non-asymptotic analysis of our test statistic for any arbitrary subset $S$.
We accommodate the case where the cross-sectional dimension $N$ is much larger than the temporal dimension $T$. In such case, estimating the high-dimensional covariance matrix $\bm{\Sigma}_{\varepsilon}$ becomes challenging yet is vital for numerous statistical tests, including the factor specification test and the two-sample test for betas discussed in the next section. Following the approaches in, e.g., bickel2008covariance,fan2011covFsee,POET2013, we assume that the error covariance matrix $\bm{\Sigma}_{\varepsilon}$ is sparse in a suitable sense and use the adaptive thresholding method to estimate $\bm{\Sigma}_{\varepsilon}$. It's important to note that the assumed sparsity of $\bm{\Sigma}_{\varepsilon}$ is solely for our statistical tests and is not a required condition to establish our main theories on the subspace perturbation bounds in Theorem (ref).
The assumption is slightly weaker than those in, e.g., bickel2008covariance,POET2013, where they assumed that $\lambda_{\min}(\bm{\Sigma}_{\varepsilon})\asymp1$, $\lambda_{\max}(\bm{\Sigma}_{\varepsilon})\asymp1$, and $\max_{1\leq i\leq N}\sum_{j=1}^{N}|(\bm{\Sigma}_{\varepsilon})_{i,j}|^{q}\leq s_{0}$ for some sparsity parameter $s_{0}$. In particular, for $q=0$, it constrains on the maximum number of nonzero elements of $\bm{\Sigma}_{\varepsilon}$. This sparsity assumption for the noise covariance $\bm{\Sigma}_{\varepsilon}$ is also natural in economics and finance. As the latent common factors largely explain the co-movements in the panel data, the correlation among individual asset's idiosyncratic noises should be close to zero. A specific example of the sparse structure of $\bm{\Sigma}_{\varepsilon}$ arises from the remaining sector effects gaglia2016panelData.
We estimate the idiosyncratic noise matrix using $\widehat{\bm{E}}=\bm{X}-\widehat{\bm{B}}\widehat{\bm{F}}^{\top}$, where $\widehat{\bm{B}}$ and $\widehat{\bm{F}}$ are the PCA estimators of factor loadings and factors defined in Section (ref). Then, the pilot estimator of the covariance matrix $\bm{\Sigma}_{\varepsilon}$ is given by $\widehat{\bm{\bm{\Sigma}}}_{\varepsilon}=T^{-1}\widehat{\bm{E}}\widehat{\bm{E}}^{\top}$. We will show in Lemma (ref) later that $\max_{1\leq i,j\leq N}\bigl|(\widehat{\bm{\bm{\Sigma}}}_{\varepsilon}-\bm{\Sigma}_{\varepsilon})_{i,j}\bigl|\lesssim\epsilon_{N,T}\sqrt{(\bm{\Sigma}_{\varepsilon})_{i,i}(\bm{\Sigma}_{\varepsilon})_{j,j}}$, where $\epsilon_{N,T}>0$ is negligible as $N,T\rightarrow\infty$. Then we apply an adaptive thresholding method bickel2008covariance to obtain the sparse covariance matrix estimator $\widehat{\bm{\bm{\Sigma}}}_{\varepsilon}^{\tau}$ as follows, \[ (\widehat{\bm{\bm{\Sigma}}}_{\varepsilon}^{\tau})_{i,j}:=\Biggl\{
\] where $h(z,\tau)$ is a thresholding function with the threshold value $\tau$. Here, the threshold value $\tau_{i,j}$ is set adaptively to $\tau_{i,j}=C\epsilon_{N,T}\sqrt{(\widehat{\bm{\bm{\Sigma}}}_{\varepsilon})_{i,i}(\widehat{\bm{\bm{\Sigma}}}_{\varepsilon})_{j,j}}$ with some large constant $C>0$. The idea behind the adaptive thresholding is examining the sample correlation matrix and retaining entries exceeding $C\epsilon_{N,T}$ in magnitude. The specific examples of the function $h$ include the hard-thresholding function $h(z,\tau)=z1_{\{|z|<\tau\}}$, among other common thresholding functions like the soft-thresholding function and SCAD FanLi2001SCAD. In general, we require that the thresholding function $h(z,\tau)$ to satisfy: (i) $h(z,\tau)=0\text{ if }|z|<\tau$; (ii) $|h(z,\tau)-z|<\tau$. The estimation error is given as follows.
The above error bounds for the estimators $\widehat{\bm{\bm{\Sigma}}}_{\varepsilon}$ and $\widehat{\bm{\bm{\Sigma}}}_{\varepsilon}^{\tau}$ differ from the existing literature owing to our row-wise subspace perturbation bounds in Theorem (ref). Let us interpret the condition ((ref)) under some specific settings. Consider the setting where $r\asymp1$, $N\asymp T$, $\bar{\bm{U}}$ is $\mu$-incoherent, $\Vert\bm{B}\Vert_{2,\infty}\lesssim1$ and $\bm{\Sigma}_{\varepsilon}$ is well-behaved such that $\gamma_{\varepsilon,i}\asymp1$, $(\bm{\Sigma}_{\varepsilon})_{i,i}\asymp1$, and $\Vert\bm{\Sigma}_{\varepsilon}\Vert_{2}\asymp1$. Then, the condition ((ref)) is equivalent to $T\geq(\sigma_{r}\log^{2}n)/\epsilon_{N,T}$, which requires the sample size $T$ is sufficiently large. According to Lemma (ref), fulfilling our assumption on the estimation error $\Vert\widehat{\bm{\Sigma}}_{\varepsilon}^{\tau}-\bm{\Sigma}_{\varepsilon}\Vert_{2}$ required for Theorem (ref) merely requires that the estimation error $\epsilon_{N,T}$ satisfies \[ \epsilon_{N,T}\ll\bigl((\left\Vert \bm{\Sigma}_{\varepsilon}\right\Vert _{2}/s(\bm{\Sigma}_{\varepsilon}))|S|^{-1/2}\kappa_{\varepsilon}^{-4}r^{-1}\log^{-2}n\bigr)^{1/(1-q)}. \] In subsequent applications, we will directly use $\epsilon_{N,T}$, instead of $\widehat{\bm{\bm{\Sigma}}}_{\varepsilon}^{\tau}$, to express assumptions on the estimation error of the noise covariance matrix.
The discussion so far has studied the factor models with constant betas. However, in many cases, there is a possibility of time-variant betas, and it is important to test for structural changes in betas StockWatson2009beta_break,breitung2011beta_break,chen2014beta_break,han2015beta_break. We test for two time periods $\Gamma_{1}=\{t_{1}+1,t_{1}+2,\ldots,t_{1}+T_{1}\}$ and $\Gamma_{2}=\{t_{2}+1,t_{2}+2,\ldots,t_{2}+T_{2}\}$ that are possibly not consecutive, i.e., $t_{1}+T_{1}\leq t_{2}+1$. The factor models in these two periods are subject to a structural break in beta: for each given individual unit $i$,
where $i$ is the index of the cross-sectional unit of interest. Here, we assume that the factors are the same, but the betas can be different. For example, we might ask if a external shock such as the 2008 financial crises has caused the changes in factor loadings before and after the shock with a transition period $(T_1, T_2)$.
The data consist of two panels, $\bm{X}^{1}$ and $\bm{X}^{2}$, during $\Gamma_{1}$ and $\Gamma_{2}$, respectively, where $\bm{X}^{1}\in\mathbb{R}^{N\times T_{1}}$ and $\bm{X}^{2}\in\mathbb{R}^{N\times T_{2}}$. We test the hypothesis as follows: \[ H_{0}:\ \bm{b}_{i}^{1}=\bm{b}_{i}^{2}\ \leftrightarrow\ H_{1}:\ \bm{b}_{i}^{1}\neq\bm{b}_{i}^{2}. \] Different from our approach, chen2014beta_break,han2015beta_break tested for the entire factor loading, with their null hypothesis being $H_{0}:\ \bm{b}_{i}^{1}=\bm{b}_{i}^{2}\text{ for }\forall i$. StockWatson2009beta_break,breitung2011beta_break also studied the test for a given $i$ like ours, while StockWatson2009beta_break focused on empirical studies and breitung2011beta_break developed and established theories for the test statistics. Our work differs in that, as we will show below, the effectiveness of our test statistic does not require the pervasiveness assumption needed in breitung2011beta_break. In addition, our formulation allows the gap between the two periods $\Gamma_{1}$ and $\Gamma_{2}$, which can model the special transition periods we purposely, e.g., financial crisis, enabling us to test for structural breaks before and after these special periods.
To construct the test statistic, we first merge the two panels $\bm{X}^{1}$ and $\bm{X}^{2}$ into a $N\times(T_{1}+T_{2})$-dimensional data matrix $\bm{X}=(\bm{X}^{1},\bm{X}^{2})$. Next, we apply SVD as described in Algorithm (ref) to obtain an estimator of the factors: $\widehat{\bm{F}}=\sqrt{T}\widehat{\bm{V}}$, where $T=T_{1}+T_{2}$ and $\widehat{\bm{U}}\widehat{\bm{\Sigma}}\widehat{\bm{V}}^{\top}$ is the truncated rank-$r$ SVD of $T^{-1/2}\bm{X}$ with $\widehat{\bm{U}}\in\mathbb{R}^{N\times r}$ and $\widehat{\bm{V}}\in\mathbb{R}^{T\times r}$. Then, we split the estimated factors into $\widehat{\bm{F}}=((\widehat{\bm{F}}^{1})^{\top},(\widehat{\bm{F}}^{2})^{\top})^{\top}$ according to two time periods $\Gamma_{1}$ and $\Gamma_{2}$, where $\widehat{\bm{F}}^{j}\in\mathbb{R}^{T_{j}\times r}$. Subsequently, we obtain the estimator for $\bm{b}_{i}^{j}$ by regressing $\bm{X}_{i,\cdot}^{j}$ on $\widehat{\bm{F}}^{j}$: \[ \widehat{\bm{b}}_{i}^{j}=((\widehat{\bm{F}}^{j})^{\top}\widehat{\bm{F}}^{j})^{-1}(\widehat{\bm{F}}^{j})^{\top}(\bm{X}_{i,\cdot}^{j})^{\top}\qquad\text{for}\qquad j=1,2. \] Finally, under null hypothesis $H_{0}$, we show that $\widehat{\bm{b}}_{i}^{1}-\widehat{\bm{b}}_{i}^{2}$ is approximately Gaussian via the first-order approximation results in Theorem (ref), and construct the test statistic using the plug-in estimator of the covariance matrix.
To illustrate our assumptions, we consider the following setting similar to that discussed after Theorem (ref): $r\asymp1$, $T_{1}\asymp T\asymp N$, $T_{1}\gg T_{2}$, $\bm{\Sigma}_{\varepsilon}$ is well-conditioned and its sparsity parameter satisfies $s(\bm{\Sigma}_{\varepsilon})\lesssim\Vert\bm{\Sigma}_{\varepsilon}\Vert_{2}$, and the column subspace of $\bm{B}$ is incoherent. Then, fulfilling our assumptions merely requires that, $\theta\gg\log^{3/2}n$ and $\vartheta_{i}\gg T^{-1/2}\log n$ for the SNR, $T\gg\log^{3}n$ for the sample size, and $\epsilon_{N,T}\ll\min(\log^{-1}n,(N^{1/2}/\log n)^{1/(1-q)})$ for the estimation error of $\bm{\Sigma}_{\varepsilon}$. In particular, our SNR conditions are less restrictive than prior work breitung2011beta_break that needed the pervasiveness assumption, which required the SNR to grow as $\theta\gg N^{1/4}$. Our results adapt to the weak factor model with the cross-sectional correlations, and to the best of our knowledge, no prior work has developed the test statistics for the structural break test under this case.
In finance and economics, besides the latent common factors that drive the co-movements of asset returns, the factor loadings, also known as betas, are important as well, which measure the sensitivity of asset return to the movements of the factors. Consider the example that the panel data $\bm{X}$ is the stock return, where the $i$-th row of $\bm{X}$ represents the time series data of the $i$-th stock. Then, the factor loadings $\bm{B}$ assess the risk exposure of these stocks to the latent common factors $\bm{f}_{t}$, with the $i$-th row $\bm{b}_{i}$ being the $i$-th stock's beta.
For any distinct $i$ and $j$, we aim to test the hypothesis that $\bm{b}_{i}=\bm{b}_{j}$, i.e., if the $i$-th and the $j$-th stocks have the same risk exposure on the common factors. It is a statistical approach to evaluate the similarity in risk structure between two stocks. Our test statistic is similar to the idea of two-sample test: we show that $\widehat{\bm{B}}_{i,\cdot}-\widehat{\bm{B}}_{j,\cdot}$ is approximately Gaussian and then derive a Chi-square test statistic. In Theorem (ref) below, we construct the test statistic and show its validity.
To illustrate the SNR conditions ((ref))--((ref)), we consider the setting discussed after Theorem (ref), i.e., $r\asymp1$, $N\asymp T$, $\bm{\Sigma}_{\varepsilon}$ is well-conditioned, and the column subspace of $\bm{B}$ is incoherent. Then, fulfilling our assumptions ((ref))--((ref)) merely requires that $\theta\gg\log n$ and $\theta^{-1}\vartheta\min(\vartheta_{i},\vartheta_{j})\gg\log^{3/2}n$. These SNR conditions are less restrictive than the prior work BaiNg2023PCA,Jiang2023PCA,ChoiMing2024PCA that required the SNR to grow with a polynomial rate of $n$.
To the best of our knowledge, no prior work has developed the test statistics for this two-sample test of betas. Our results are applicable to the weak factor model with the cross-sectional correlations. The subspace perturbation bounds in Theorem (ref) pave the way for us to statistically assess the similarity for arbitrary two rows of factor loadings $\bm{B}$.
In the factor model $x_{i,t}=\bm{b}_{i}^{\top}\bm{f}_{t}+\varepsilon_{i,t}$, the risk associated with stock return $x_{i,t}$ is decomposed into two parts Bai2003ECTA --- systematic risk from the common component $\bm{b}_{i}^{\top}\bm{f}_{t}$ and idiosyncratic risk from the noise $\varepsilon_{i,t}$. Systematic risk, often referred to as market risk, is integral in financial economics as it represents the inherent risk affecting the entire market or market segment. This risk, driven by broader economic forces such as inflation, political events, and changes in interest rates, comes from the risk factors $\bm{f}_{t}$ that explain the systematic co-movements and impacts all the stocks.
A standard metric of the systematic risk is the variance of $\bm{b}_{i}^{\top}\bm{f}_{t}$, which is given by $\operatorname{{\rm Var}}(\bm{b}_{i}^{\top}\bm{f}_{t})= \Vert\bm{b}_{i}\Vert_{2}^{2}$. Our focus is on constructing the confidence interval (CI) for systematic risk $\Vert\bm{b}_{i}\Vert_{2}^{2}$. Similar to the idea we conduct the inference for beta $\bm{b}_{i}$ in Corollary (ref) where the estimator of $\bm{b}_{i}$ is $\widehat{\bm{B}}_{i,\cdot}$, we show that $\Vert\widehat{\bm{B}}_{i,\cdot}\Vert_{2}^{2}-\Vert\bm{b}_{i}\Vert_{2}^{2}$ is approximately Gaussian and then construct the CI for $\Vert\bm{b}_{i}\Vert_{2}^{2}$. We present the CI and its validity in Theorem (ref).
We now explain the assumptions under the setting discussed after Theorem (ref). In this case, to make ((ref)) hold, it suffices to assume that $\theta^{-1}\vartheta_{i}\vartheta\gg\log^{2}n$ and $\theta^{-1}\vartheta_{i}^{2}\gg\log^{2}n$; to make ((ref)) hold, it suffices to assume that $\theta\gg1$, $\vartheta_{i}\theta\gg\log n$, and $\vartheta_{i}\gg T^{-1/2}\log n$. All these conditions require only polynomial growth rates of $\log n$ for the SNRs $\vartheta_{i}$, $\vartheta$, and $\theta$. The assumption on $\Vert\bar{\bm{U}}\Vert_{2,\infty}$ is the same with that we required in Corollary (ref) to prove the estimation error for factor loading.
On one hand, as shown in Theorem (ref), the CI width $\widehat{\sigma}_{B,i}$ is proportional to the square root of the product of noise level estimator $(\widehat{\bm{\bm{\Sigma}}}_{\varepsilon}^{\tau})_{i,i}$ and systematic risk estimator $\Vert\widehat{\bm{B}}_{i,\cdot}\Vert_{2}^{2}$. On the other hand, in the proof of Theorem (ref), we show that the bias of systematic risk estimator $\Vert\widehat{\bm{B}}_{i,\cdot}\Vert_{2}^{2}$ is proportional to $\Vert\bm{b}_{i}\Vert_{2}^{2}$. Thus, the assumption that $\Vert\bm{b}_{i}\Vert_{2}^{2}/(\bm{\Sigma}_{\varepsilon})_{i,i}\leq c_{0}\delta^{2}r^{-1}\log^{-1}n$ is essential because the validity of inference hinges on the dominance of the CI width over the bias. Consequently, we assume that the ratio between systematic risk $\Vert\bm{b}_{i}\Vert_{2}^{2}$ and noise level $(\bm{\Sigma}_{\varepsilon})_{i,i}$ cannot be too large. The row-wise subspace perturbation bounds in Theorem (ref) facilitate us to conduct statistical inference for the systematic risk of any given stock in the panel data.
In this section, we conduct Monte Carlo simulations to demonstrate our inferential theories for the PC estimators in the weak factor models. Additionally, our empirical studies reveal that the testing results based on our test statistics surprisingly align with the economic cycles and financial crisis periods.
To make our simulations similar to the real applications, we use the standard Fama-French three-factor model: \[ x_{i,t}=\bm{b}_{i}^{\top}\bm{f}_{t}+\varepsilon_{i,t},\qquad1\leq i\leq N,1\leq t\leq T, \] where the dimension of the latent factor $\bm{f}_{t}$ is set to $r=3$. The idiosyncratic noise $\varepsilon_{i,t}$ exhibits cross-sectional correlations, and the noise covariance matrix $\bm{\Sigma}_{\varepsilon}$ is sparse.
We generate the factor loadings $\{\bm{b}_{i}\}_{i=1}^{N}$, the factors $\{\bm{f}_{t}\}_{t=1}^{T}$, and the noise terms $\{\bm{\varepsilon}_{t}\}_{t=1}^{T}$ independently from $\mathcal{N}(0,\bm{\Sigma}_{b})$, $\mathcal{N}(0,s_{f}\bm{\Sigma}_{f})$, and $\mathcal{N}(0,\bm{\Sigma}_{\varepsilon})$, respectively, where $\bm{\varepsilon}_{t}=(\varepsilon_{1,t},\ldots,\varepsilon_{N,t})^{\top}$. Both $\bm{\Sigma}_{b}$ and $\bm{\Sigma}_{f}$ are set to $r\times r$ identity matrices. We generate $\bm{\Sigma}_{\varepsilon}$ as a block-diagonal matrix $\bm{\Sigma}_{\varepsilon}=\text{\ensuremath{\mathsf{diag}}}(\bm{A}_{1},\bm{A}_{2},\ldots,\bm{A}_{J})$, where the number of blocks is set to $J=20$. We set $N=300$ and $T=200$. For each $\bm{A}_{i}$, we construct it as an equi-correlation matrix $\bm{A}_{i}=(1-\rho_{i})\bm{I}_{m}+\rho_{i}\bm{1}_{m}\bm{1}_{m}^{\top}$, where the block size $m$ is set to $m=N/J=15$, $\bm{1}_{m}$ is the $m$-dimensional vector of ones, and $\rho_{i}$ is drawn from a uniform distribution on $[0,0.5]$. To validate our test statistic for two-sample test of betas, we set $\bm{b}_{2}$ equal to $\bm{b}_{1}$ after generating the loading matrix $\bm{B}$. This slight modification allows us to examine our test statistics under the null hypothesis $\bm{b}_{1}=\bm{b}_{2}$. In our simulation results, we report the values of $\theta=\sigma_{r}/\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{2}$ to reflect different levels of SNR.
First, we demonstrate the practical validity of the confidence regions constructed using Corollaries (ref)--(ref) and Theorem (ref), for the factors, betas (i.e., factor loadings), and systematic risks, respectively.
Finally, we compute the mean and standard deviation for $\{\widehat{\mathsf{Cov}}_{F}(t)\}_{t=1}^{T}$, $\{\widehat{\mathsf{Cov}}_{B}(i)\}_{i=1}^{N}$, and $\{\widehat{\mathsf{Cov}}_{\ell_{2}}(i)\}_{i=1}^{N}$, then present these as $\mathsf{Mean}(\widehat{\mathsf{Cov}})$ and $\mathsf{Std}(\widehat{\mathsf{Cov}})$, with the results reported in Table (ref). As indicated in Table (ref), the coverage probabilities are close to 0.95 according to the mean value, and exhibit stability across different rows of $\bm{F}$ and $\bm{B}$ as evidenced by low standard deviation. As SNR $\theta$ goes down, the slight slippage of the coverage probabilities of the systematic risks reconciles with our comments after Theorem (ref) that, $\Vert\bm{b}_{i}\Vert_{2}^{2}/(\bm{\Sigma}_{\varepsilon})_{i,i}$, which is close to SNR, reflects the ratio between the bias and the CI width and cannot be too large. These favorable numerical results persist even under low SNR, supporting our inferential theory under the weak factor model.
Next, we demonstrate the effectiveness of our test statistics in Theorems (ref) and (ref) by showing their satisfactory size and power.
Tables (ref)--(ref) report the empirical rejections rates at 5% significance level over 200 Monte Carlo trials. Table (ref) (resp. (ref)) shows that for the test statistics in Theorem (ref) (resp. Theorems (ref) and (ref)), the results are favorable, exhibiting appropriate size and power, even under a weak signal setup where the SNR $\theta$ is small.
We analyze the monthly returns data of the S&P 500 constituents from the CRSP database for the period from January 1995 to March 2024. We apply the factor specification test in Theorem (ref) and the structural break test for betas in Theorem (ref) to the stock returns.
First, we consider the factor specification test. The observed factors we study are the Fama-French three factors: market (MKT), size (SMB), and value (HML), denoted as $\bm{v}^{(1)}$, $\bm{v}^{(2)}$, and $\bm{v}^{(3)}$, respectively. We obtain the time series data for these factors from Kenneth French's website. We conduct our tests via a rolling window approach, moving a 60-month window $[t+1,t+T]$ ($T=60$) through the dataset. To mitigate the survival bias, we keep the time series that have no more than 50% missing values in each window, i.e., the number of missing values is less than $T/2$, and fill the missing values by the median of each time series. For each window, the formed data matrix $\bm{X}^{t}$ is an $N\times T$ matrix, where the cross-sectional dimension $N$ varies with $t$. We assume that $\bm{X}^{t}$ satisfies the factor model as in ((ref)), i.e., \[ \bm{X}^{t}=\bm{B}^{t}(\bm{F}^{t})^{\top}+\bm{E}^{t}. \] Here, the superscript $t$ is added to each matrix to emphasize that we apply the PCA method across different time windows $[t+1,t+T]$. The number of factors is fixed as $r=3$. We conduct the factor specification test in Theorem (ref) to test the null hypothesis \[ \bm{v}_{S^{t},\cdot}^{(i)}=\bm{F}_{S^{t},\cdot}^{t}\bm{w}^{t,(i)}, \] for the three factors, corresponding to $i=1,2,3$, respectively. The time index subset $S^{t}$ is set as $S^{t}=[t+T-L+1,t+T]$ with $|S^{t}|=L=12$.
The above procedure implies that, for each 12-month period $S^{t}=[t+T-L+1,t+T]$, to test if the observed factors are in the column space of the latent factors, we look back and utilize a broader historical data window $[t+1,t+T]$ ($T=60$) to estimate the latent common factors $\bm{F}^{t}$. Then we test the null hypothesis for this specific 12-month period $S^{t}$, a subset at the end of the whole window $[t+1,t+T]$. Finally, we plot in Figure (ref) the test statistics against the time index $t$ for each factor, underscoring the 95% critical value. As highlighted by FanLiaoYao2015power, this rolling window manner not only utilizes the up-to-date information in the equity universe, but also alleviates the impacts of time-varying betas and sampling biases.
In Figure (ref), our findings indicate that during the financial crisis of 2007-2009, the null hypothesis that the size factor SMB lies in the latent factors' column space, is rejected at 95% confidence level. This suggests a diminished importance and reduced explanatory power of the size factor SMB for stock return data in this period. Note that the sizes of stocks can still be important or even more important than that in the normal period, but not necessarily captured by the variable SMB. Similar interpretations hold for the value factor HML around 2009. During the COVID period around 2019, both the size factor SMB and the value factor HML exhibit a loss in explanatory power. The spike for the size factor SMB during 2022 is probably due to the war between Russia and Ukraine that started in February 2022, which is an unexpected economic shock to the stock market. Notably, the market portfolio maintains its explanatory strength throughout, indicating its stability and resilience as an explanatory variable, even during distinct economic cycles.
Next, we consider the structural break test for betas for individual stocks. We test whether the betas have changed before and after the three economic recessions covered by the time horizon of our data -- the Early 2000s Recession (Mar. 2001--Nov. 2001) due to the dot com bubble, the 2008 Great Recession due to housing bubble and financial crisis (Dec. 2007--Jun. 2009), and the COVID-19 Recession (Feb. 2020--Apr. 2020). Here, the start and the end of each recession are according to the NBER's Business Cycle Dating Committee.
For each recession period $[t_{\text{start}},t_{\text{end}}]$, we first take the data $\bm{X}^{1}$ and $\bm{X}^{2}$ lying in the time window $[t_{\text{start}}-T_{1},t_{\text{start}}-1]$ and $[t_{\text{end}}+1,t_{\text{end}}+T_{2}]$, respectively, where $T_{1}=T_{2}=60$. Next, we merge the two panels to get $\bm{X}=(\bm{X}^{1},\bm{X}^{2})$, and fill the missing values in the same manner as in the factor specification test. We assume that $\bm{X}$ satisfies the factor model as follows
The number of factors is chosen by a scree plot of the eigenvalues of $\bm{X}$, which is $4$, $3$, and $3$ for the three recessions, respectively. We test the hypothesis $H_{0}:\ \bm{b}_{i}^{1}=\bm{b}_{i}^{2}\ \leftrightarrow\ H_{1}:\ \bm{b}_{i}^{1}\neq\bm{b}_{i}^{2}$ for each cross-sectional unit $i$ using the test statistic in Theorem (ref). To report the test results, we group the stocks into 11 sectors by Global Industrial Classification Standard (GICS), and then count the numbers of stocks in each sector that reject the null at the 95% confidence level. The test results for the three recessions are illustrated in Figures (ref), (ref), and (ref), respectively.
In Figure (ref), we observe that the sectors with the highest number of stocks experiencing structural breaks in betas are Financials, Information Technology, Industrials, and Utilities. The impact on the Information Technology sector is likely due to the dot-com bubble burst, which was one of the triggers of the Early 2000s Recession. Another significant cause of this recession was the 9/11 attacks. Such severe economic shocks are possible reasons why typically stable sectors like Industrials and Utilities were also affected.
In Figure (ref), the 2007-2009 financial crisis, marked by the subprime mortgage crisis and the collapse of the United States housing bubble, affected many sectors. The financial sector experienced a strong impact due to direct exposure to mortgage-backed securities and other related financial instruments. The crisis led to the failure or collapse of many of the United States' largest financial institutions. Even though our analysis is inevitably influenced by survival bias, as we can only analyze stocks that existed before and after the crisis, we still observe that the financial sector had the most affected stocks. The impact in Real Estate reflects the direct consequences of the housing bubble burst. For sectors related to consumer spending, such as Consumer Discretionary, Consumer Staples, and Health Care, the shocks can be attributed to reduced consumer spending due to increased unemployment and economic uncertainty.
In Figure (ref), the economic effects of the pandemic can be seen in many affected sectors, such as Real Estate, Industrials, Health Care, and Financials. The significant changes in Real Estate likely reflect the uncertainties brought about by lockdowns and health crises, resulting in severe fluctuations in property values and rent payments. The impact on the Health Care sector may be tied to the heightened demand for medical services and supplies, alongside volatility in biotechnology investments. Notably, Utilities and Energy showed smaller changes, suggesting relative stability in these sectors despite overall market volatility. During the 2008 Great Recession, these sectors also demonstrated similar stability. This resilience could be attributed to the essential nature of services provided by these sectors, making them less susceptible to economic disruptions.
The factor model is an important topic in finance and economics. The early econometric studies on factor model can be dated back to forni2000factor,StockWatson2002PCA. Most previous works on factor model assumes that all factors are strong or pervasive, that is, the SNR grows at rate $\sqrt{N}$ BaiNg2002ECTA,Bai2003ECTA,POET2013. When the SNR grows at a rate slower than $\sqrt{N}$, the model is often called the weak factor model, which has been a popular research topic in recent years. The method for determining the number of factors under the weak factor model has been studied by a few papers Onatski2009number,Onatski2010number,frey2022weakFactorNumber. Several recent works have pursued the estimation and inference in the weak factor model. To name a few examples, BaiNg2023PCA,Jiang2023PCA,ChoiMing2024PCA studied the consistency and asymptotic normality of PC estimators; YoTa2022sWF_estimate,YoTa2022sWF_infer,wei2023sWF studied the sparsity-induced weak factor models where the low SNR is due to the sparsity of the factor loading matrix; BKP2021weakFactor proposed an estimator of factor strength and established its theoretical guarantee; Onatski2012PCAweak showed that the PC estimators are inconsistent in the extreme case (also known as the super-weak factor model) where the SNR is $O(1)$.
The PCA is one of the most popular methods for the factor model, and has been studied in many papers mentioned previously StockWatson2002PCA,BaiNg2002ECTA,Bai2003ECTA,Onatski2012PCAweak,POET2013,BaiNg2023PCA,Jiang2023PCA,ChoiMing2024PCA. Among the enormous literature, BaiNg2023PCA,Jiang2023PCA,ChoiMing2024PCA are the most recent and closest to our paper, since we all focus on the PC estimators, especially on the inference side, under the weak factor model. Besides the theoretical analysis, the variants of PCA under the weak factor model have also been applied to empirical asset pricing, macroeconomic forecasting, and many other important problems in finance gigXiu2021testWeak,GiglioXiu2023weakFactorPredict. The estimation of factor model is closely related to the low-rank matrix denoising in the statistical machine learning community. In recent years, studying the factor model from the view of low-rank matrix denoising has provided lots of exciting findings and understandings (see FanLiLiao2021ARFEfactor,yan2021inference for comprehensive reviews), and our theory is partly inspired by this low-rank structure of factor model.
From a technical perspective, our estimation procedure is a spectral method, and is related to previous studies of spectral methods on PCA and subspace estimation abbe2020ell_p,cai2019subspace,chen2020bridging,yan2021inference,zhou2023deflated. Our analysis relies on the leave-one-out techniques that have found wide applications in analyzing spectral estimators and nonconvex optimization algorithms el2015impact,abbe2017entrywise,ma2017implicit,chen2019noisy,chen2020convex; see the recent monograph chen2021Monograph for more details. In addition, the problem and analysis in this paper is also related to past works on inference for other low-rank models chen2019inference,xia2021statistical,chernozhukov2021inference,choi2023inference,choi2024inference,yan2024entrywise.
In this paper, we establish a novel theory for PCA under the weak factor model, offering significant advancements on the inference of PCA over the existing literature. The weak factor model removes the pervasiveness assumption that requires of SNR growing at a $\sqrt{N}$ rate, where $N$ is the cross-sectional dimension. Our theory covers both estimation and inference of factors and factor loadings. Notably, in the regime $N\asymp T$, where $T$ is the temporal dimension, we show that the asymptotic normality of PC estimators holds as long as the SNR grows faster than a polynomial rate of $\log N$, while the previous work required a polynomial rate of $N$. The optimality of our theory is in the sense that, the required growth rate of the SNR for consistency aligns with that for asymptotic normality, differing only by some logarithmic factors. Our theory paves the way to design easy-to-implement test statistics for practical applications, e.g., factor specification test, structural break test for betas, and build confidence regions for crucial model parameters, e.g., betas and systematic risks. We validate our statistical methods through extensive Monte Carlo simulations, and conduct empirical studies to find noteworthy correlations between our test results and specific economic cycles.
While the current scope of our theory is considerably wide-ranging, it is possible to further widen it and there are lots of topics that are worth pursuing. For instance, how to extend the theory to the setting where each observed variables have only fouth moment via robustification of the covariance input fan2021robust? How to generalize our inferential theory to the case where the time-serial correlation also appears in the idiosyncratic noise? If the panel data is missing at random, are the PC estimators still valid under the weak factor model, and how to conduct statistical inference in this case? What benefit could our new theory bring to forecasting methods based on the factor-augmented regression? Among many directions, these topics can be investigated in future research.
J. Fan is supported in part by the NSF grants DMS-2053832 and DMS-2210833, and ONR grant N00014-22-1-2340. Y. Yan is supported in part by the Norbert Wiener Postdoctoral Fellowship from MIT.