The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
109,382 characters
When can weak latent factors be statistically inferred?
\theoremstyle{plain} \newtheorem{lemma}{\textbf{Lemma}} \newtheorem{prop}{\textbf{Proposition}}\newtheorem{theorem}{\textbf{Theorem}}\setcounter{theorem}{0}
\newtheorem{corollary}{\textbf{Corollary}} \newtheorem{assumption}{\textbf{Assumption}}
\newtheorem{example}{\textbf{Example}} \newtheorem{definition}{\textbf{Definition}}
\newtheorem{fact}{\textbf{Fact}} \newtheorem{condition}{\textbf{Condition}}\theoremstyle{definition}\theoremstyle{remark}\newtheorem{remark}{\textbf{Remark}}\newtheorem{claim}{\textbf{Claim}}\newtheorem{conjecture}{\textbf{Conjecture}}
\title{\textbf{When can weak latent factors be statistically inferred?}}
\author{Jianqing Fan\thanks{Department of Operations Research and Financial Engineering, Princeton
University, Princeton, NJ 08544, USA; Email: \texttt{\{jqfan,yuheng\}@princeton.edu}.} \and Yuling Yan\thanks{Institute for Data, Systems, and Society, Massachusetts Institute
of Technology, Cambridge, MA 02142, USA; Email: \texttt{[email removed]}.} \and Yuheng Zheng\footnotemark[1]}
\maketitle
\begin{abstract}
This article establishes a new and comprehensive estimation and inference theory for principal component analysis (PCA) under the weak factor model that allow for cross-sectional dependent idiosyncratic components under the nearly minimal factor strength relative to the noise level or signal-to-noise ratio. Our theory is applicable regardless of the relative growth rate between the cross-sectional dimension $N$ and temporal dimension $T$. This more realistic assumption and noticeable result require completely new technical device, as the commonly-used leave-one-out trick is no longer applicable to the case with cross-sectional dependence. Another notable advancement of our theory is on PCA inference --- for example, under the regime where $N\asymp T$, we show that the asymptotic normality for the PCA-based estimator holds as long as the signal-to-noise ratio (SNR) grows faster than a polynomial rate of $\log N$. This finding significantly surpasses prior work that required a polynomial rate of $N$. Our theory is entirely non-asymptotic, offering finite-sample characterizations for both the estimation error and the uncertainty level of statistical inference. A notable technical innovation is our closed-form first-order approximation of PCA-based estimator, which paves the way for various statistical tests. Furthermore, we apply our theories to design easy-to-implement statistics for validating whether given factors fall in the linear spans of unknown latent factors, testing structural breaks in the factor loadings for an individual unit, checking whether two units have the same risk exposures, and constructing confidence intervals for systematic risks. Our empirical studies uncover insightful correlations between our test results and economic cycles.
\end{abstract}
\noindent \textbf{Keywords: }factor model, principal component analysis,
weak factors, cross-sectional correlation, inference, signal-to-noise
ratio.
\setcounter{tocdepth}{2}
\tableofcontents{}
\section{Introduction }
The factor model, a pivotal tool for analyzing large panel data, has
become a significant topic in finance and economics research \citep[e.g.,][]{cham1983factor,FamaFrench1993factors,StockWatson2002PCA,BaiNg2002ECTA,GigXiu2021JPE,fan2021recent}.
The estimation and inference for factor models are crucial in economic
studies, particularly in areas like asset pricing and return forecasting.
In the era of big data, the factor model has gained increased prominence
in capturing the latent common structure for large panel data, where
both the cross-sectional and temporal dimensions are ultra-high \citep[see, e.g., recent surveys][]{BaiPeng2016ARFEfactor,FanLiLiao2021ARFEfactor}.
Principal component analysis (PCA), known for its simplicity and effectiveness,
is closely connected with the factor model and has long been a key
research topic of interest in the econometric community \citep[e.g.,][]{StockWatson2002PCA,BaiNg2002ECTA,Bai2003ECTA,Onatski2012PCAweak,POET2013,BaiNg2013identifiPCA,BaiNg2023PCA}.
As pointed out by \citet{GiglioXiu2023weakFactorPredict}, most theoretical
guarantees for the PCA approach to factor analysis rely on the pervasiveness
assumption \citep[e.g.,][]{BaiNg2002ECTA,Bai2003ECTA}. This assumption
requires the signal-to-noise ratio (SNR), which measures the factor
strength relative to the noise level, to grow with the rate of $\sqrt{N}$
-- the square root of the cross-sectional dimension. However, many
real datasets in economics do not exhibit sufficiently strong factors
to meet this pervasiveness assumption. When the SNR grows slower than
$\sqrt{N}$, the resulting model is often called the weak factor model
\citep[e.g.,][]{Onatski2009number,Onatski2010number}. Extensive research
has been dedicated to the weak factor model \citep[e.g.,][]{Onatski2012PCAweak,BKP2021weakFactor,frey2022weakFactorNumber,YoTa2022sWF_estimate,YoTa2022sWF_infer,BaiNg2023PCA,Jiang2023PCA,ChoiMing2024PCA},
among which the PC estimators (the estimators via the PCA approach)
have been a primary subject. Recently, \citet{BaiNg2023PCA,Jiang2023PCA,ChoiMing2024PCA}
studied the consistency and asymptotic normality of the PC estimators
for factors and factor loadings in the weak factor model. In the extreme
case (also called the super-weak factor model) where the SNR is $O(1)$,
\citet{Onatski2012PCAweak} showed that the PC estimators are inconsistent.
This paper establishes a novel and comprehensive theory for PCA in
the weak factor model. Our theory is non-asymptotic and can be easily
translated to asymptotic results. The conditions we propose for the
asymptotic normality of PC estimators are optimal in the sense that,
surprisingly different from the existing literature, the required
growth rate of the SNR for consistency aligns with that for asymptotic
normality, differing only by a logarithmic factor. In particular,
in the regime $N\asymp T$, where $T$ is the temporal dimension,
we prove that asymptotic normality holds as long as the SNR grows
faster than a polynomial rate of $\log N$, and this result is a substantial
advance compared with the existing results that require the SNR to
grow with a polynomial rate of $N$ \citep[e.g.,][]{BaiNg2023PCA,Jiang2023PCA,ChoiMing2024PCA}.
The most innovative part of our theory lies in establishing a closed-form,
first-order approximation of the PC estimator, that is, we decompose
the PC estimator into three components --- the ground truth, a first-order
term, and the higher-order negligible term. We express the first-order
term explicitly using the parameters in the factor model. This closed-form
characterization paves the way for us to establish the asymptotic
normality and design various test statistics for practical applications.
Our theory is based on novel applications of the leave-one-out analysis
and matrix concentration inequalities. Moreover, our findings provide
valuable insights and practical implications for a range of econometric
problems related with PCA, e.g., macroeconomic forecasting based on
factor-augmented regressions \citep[e.g.,][]{BaiNg2006factorAugmented}.
We demonstrate the practical applications of our theories using both
synthetic and real datasets. First, we design an easy-to-implement
statistic for factor specification test \citep[e.g.,][]{BaiNg2006factor}
--- whether an observed factor is in the linear space spanned by
the latent common factors or not. A key innovation of our approach
is the ability to conduct this test in any flexible subset of the
whole period, owing to the row-wise error bound in our theory. We
utilize the monthly return data of the S\&P 500 constituents from
1995 to 2024, along with the data of the Fama-French three factors,
then run the factor test in a rolling window manner \citep[e.g.,][]{FanLiaoYao2015power}.
Our results uncover a notable decline in the importance and explanatory
power of the size factor during the 2008 financial crisis, and of
both the size and value factors during the COVID-19 pandemic around
2019. These findings are supported by our test results, which reject
the null hypothesis that these factors are in the linear space of
latent factors.
Then, we design a test statistic for the structural break of betas
\citep[e.g.,][]{StockWatson2009beta_break,breitung2011beta_break},
and apply it to the aforementioned S\&P 500 constituents data. Our
novelty is that our test statistic works in the weak factor model
without the pervasiveness assumption, which is required in prior work.
We test for each stock to determine if the beta has changed before
and after the three recessions covered by our data --- the Early
2000s Recession, the 2008 Great Recession, and the COVID-19 Recession.
We find that in each recession, the sectors most affected by the shocks,
where many stocks exhibited structural breaks in betas, correspond
reasonably to the causes of these economic recessions. For example,
during the 2008 Great Recession, marked by the subprime mortgage crisis,
the financial sector experienced a strong impact, which aligns with
its exposure to mortgage-backed securities and other related financial
instruments. During the COVID-19 Recession, the Health Care and Real
Estate sectors were significantly impacted, reflecting the uncertainties
brought about by lockdowns and health crises due to the pandemic.
Additionally, we develop statistical tests for the betas to evaluate
the similarity in risk exposure between two stocks and construct valid
confidence interval for the systematic risk of each stock.
The rest of the paper is organized as follows. Section \ref{sec:model_assump_notation}
introduces the model setup, basic assumptions, and notation. In Section
\ref{sec:main_1st_approx}, we show our main results on the first-order
approximations for the PC estimators, and present the asymptotic normality
results as corollaries. Section \ref{sec:compare_SNR} provides a
detailed comparison of our results with related work. In Section \ref{sec:3_econ_applications},
we showcase four applications in econometrics based on our main results.
Section \ref{sec:Numerical-experiments} collects the numerical results
in both simulated data and real data. More related works are discussed
in Section \ref{sec:related work}, and the paper concludes with a
discussion on future directions in Section \ref{sec:Discussion}.
\section{Model, assumptions, and notation\label{sec:model_assump_notation}}
\subsection{Model setup}
Let $N$ be the number of cross-sectional units and $T$ be the number
of observations. Consider the factor model for a panel of data $\{x_{i,t}\}:$
\begin{equation}
x_{i,t}=\bm{b}_{i}^{\top}\bm{f}_{t}+\varepsilon_{i,t},\qquad1\leq i\leq N,1\leq t\leq T,\label{factor model entrywise}
\end{equation}
where $\bm{f}_{t}=(f_{1,t},f_{2,t},\ldots,f_{r,t})^{\top}$ is the
latent factor, $r$ is the number of factors, $\bm{b}_{i}=(b_{i,1},b_{i,2},\ldots,b_{i,r})^{\top}$
is a vector of factor loadings, and $\varepsilon_{i,t}$ represents
the idiosyncratic noise. Viewing $x_{i,t}$ as the excess return of
the $i$-th asset at time $t$, the model (\ref{factor model entrywise})
is intimately linked with the multi-factor pricing model. This model
originates from the Arbitrage Pricing Theory (APT) developed by \citet{Ross1976APT}
and finds extensive applications in finance.
To compact the notation, we denote by $\bm{B}=(\bm{b}_{1},\bm{b}_{2},\ldots,\bm{b}_{N})^{\top}$
the $N\times r$ factor loading matrix. Let $\bm{x}_{t}=(x_{1,t},x_{2,t},\ldots,x_{N,t})^{\top}$
and $\bm{e}_{t}=(\varepsilon_{1,t},\varepsilon_{2,t},\ldots,\varepsilon_{N,t})^{\top}$.
Then, the factor model (\ref{factor model entrywise}) can be expressed
as $\bm{x}_{t}=\bm{Bf}_{t}+\bm{e}_{t}$, or written in a matrix form
as follows
\begin{equation}
\bm{X}=\bm{BF}^{\top}+\bm{E},\label{factor model matrix form}
\end{equation}
where $\bm{X}=(\bm{x}_{1},\bm{x}_{2},\ldots,\bm{x}_{T})$, $\bm{F}=(\bm{f}_{1},\bm{f}_{2},\ldots,\bm{f}_{T})^{\top}$,
and $\bm{E}=(\bm{e}_{1},\bm{e}_{2},\ldots,\bm{e}_{T})$ are the $N\times T$
panel data, the $T\times r$ factor realizations, and the $N\times T$
idiosyncratic noise matrix, respectively.
In our setup, the only observable part is the panel data $\bm{X}$.
We are interested in the estimation and inference for both the latent
factors and factor loadings via the PCA approach.
\subsection{Basic assumptions}
We present some basic assumptions as follows. Note that the factor
model has the rotation (indeed affine transform) ambiguity \citep[e.g.,][]{BaiNg2013identifiPCA},
that is, the factor loadings and latent factors are not identifiable
since $\bm{BF}^{\top}=(\bm{B}\bm{H}^{-1})(\bm{F\bm{H}^{\top}})^{\top}$
holds for any invertible matrix $\bm{H}$.
Without loss of generality, we assume that the columns of $\bm{B}$
are orthogonal and the covariance of $\bm{f}_{t}$ is the identity
matrix, as stated in Assumption \ref{Assump_Bf_identification} below.
\begin{assumption} \label{Assump_Bf_identification}For $t=1,2,\ldots,T$,
the factor $\bm{f}_{t}$ has mean zero and the identity covariance
matrix. The factor loading matrix $\bm{B}$ has orthogonal columns:
\begin{equation}
\bm{B}^{\top}\bm{B}=\bm{\Sigma}^{2}\text{\qquad with\qquad}\bm{\Sigma}=\text{\ensuremath{\mathsf{diag}}}(\sigma_{1},\sigma_{2},\bm{\ldots},\sigma_{r}),\label{factor loading diagonal}
\end{equation}
where $\sigma_{1}\geq\sigma_{2}\geq\cdots\geq\sigma_{r}>0$. \end{assumption}
Assumption \ref{Assump_Bf_identification} is a standard identifiability
condition for the factor model \citep[e.g.,][]{POET2013}. The singular
values $\{\sigma_{i}\}$ of the factor loading matrix $\bm{B}$ characterize
the strengths of the latent factors.
Next, to accommodate the cross-sectional correlation in the noise
and to facilitate leave-one-out analysis in our technical proof, we
propose Assumption \ref{Assump_noise_Z_entries} for the noise matrix
$\bm{E}$, specifying its structure and the distribution of its entries.
\begin{assumption} \label{Assump_noise_Z_entries}The idiosyncratic
noise matrix $\bm{E}$ is given by
\begin{equation}
\bm{E}=\bm{\Sigma}_{\varepsilon}^{1/2}\bm{Z}.\label{noise matrix formula}
\end{equation}
Here, $\bm{\Sigma}_{\varepsilon}$ is a $N\times N$ positive definite
matrix and $\bm{\Sigma}_{\varepsilon}^{1/2}$ is the symmetric square
root of $\bm{\Sigma}_{\varepsilon}$; $\bm{Z}=(Z_{i,t})_{i=1}^{N}{}_{t=1}^{T}$
is a $N\times T$ matrix; The entries of $\bm{Z}$ are independent
sub-Gaussian random variables that satisfy
\[
\mathbb{E}[Z_{i,t}]=0,\text{ }\mathbb{E}[Z_{i,t}^{2}]=1,\text{ }\left\Vert Z_{i,t}\right\Vert _{\psi_{2}}=O(1),
\]
for $1\leq i\leq N,1\leq t\leq T$, where $\left\Vert \cdot\right\Vert _{\psi_{2}}$
is the sub-Gaussian norm \citep[see Definition 2.5.6 in][]{vershynin2016high}.\end{assumption}
The nonzero off-diagonal entries of $\bm{\Sigma}_{\varepsilon}$ characterize
the cross-sectional correlations in the idiosyncratic noise $\bm{e}_{t}=(\varepsilon_{1,t},\varepsilon_{2,t},\ldots,\varepsilon_{N,t})^{\top}$.
Though the noise terms $\bm{e}_{1},\bm{e}_{2},\ldots,\bm{e}_{T}$
are independent under Assumption \ref{Assump_noise_Z_entries}, our
theory could potentially be generalized to the cases where the temporal
correlations are present in the noise matrix $\bm{E}$, and this generalization
is an interesting future direction. The formulation (\ref{noise matrix formula})
imposes additional structural constraints on the noise matrix. However,
such assumptions are commonplace in the study of the weak factor model;
see similar assumptions in \citet{Onatski2010number,Onatski2012PCAweak}.
Then, to establish the non-asymptotic results via the concentration
inequalities, we propose the following assumption on the distribution
of the factors.
\begin{assumption}\label{Assump_factor_f}The factor
$\bm{F}=(\bm{f}_{1},\bm{f}_{2},\ldots,\bm{f}_{T})^{\top}$ is independent
with the noise matrix $\bm{E}$; The factors $\bm{f}_{1},\bm{f}_{2},\ldots,\bm{f}_{T}$
are independent and sub-Gaussian random vectors in $\mathbb{R}^{r}$
satisfying that
\[
\left\Vert \bm{f}_{t}\right\Vert _{\psi_{2}}=O(1)\text{\qquad for\qquad}1\leq t\leq T.
\]
\end{assumption}
The conditions in Assumption~\ref{Assump_factor_f} are slightly
stronger than those in the literature, e.g., \citet{BaiNg2023PCA},
but remain standard and simplify our technical proofs.
\subsection{Notation\label{subsec:Paper-organization-notation}}
We introduce some notation that will be used throughout the paper.
For two sequences $a_{n}$ and $b_{n}$, we write $a_{n}\lesssim b_{n}$
(or equivalently, $b_{n}\gtrsim a_{n}$) if $a_{n}=O(b_{n})$, i.e.,
there exist a constant $C>0$ and an integer $N>0$ such that, $a_{n}\leq Cb_{n}$
holds for any $n>N$; write $a_{n}\asymp b_{n}$ if both $a_{n}\lesssim b_{n}$
and $a_{n}\gtrsim b_{n}$ hold and $a_{n}\ll b_{n}$ (or equivalently,
$b_{n}\gg a_{n}$) if $a_{n}=o(b_{n})$.
For a symmetric matrix $\bm{M}$, we denote by $\lambda_{\min}(\bm{M})$
and $\lambda_{\max}(\bm{M})$ its minimum and maximum eigenvalues.
For a vector $\bm{v}$, we denote $\Vert\bm{v}\Vert_{2}$, $\Vert\bm{v}\Vert_{1}$,
and $\Vert\bm{v}\Vert_{\infty}$ as the $\ell_{2}$-norm, $\ell_{1}$-norm,
and supremum norm, respectively. Consider any matrix $\bm{A}\in\mathbb{R}^{m\times n}$.
We denote by $\bm{A}_{i,\cdot}$ and $\bm{A}_{\cdot,j}$ the $i$-th
row and the $j$-th column of $\bm{A}$. We let $\Vert\bm{A}\Vert_{2},$
$\Vert\bm{A}\Vert_{\mathrm{F}},$ and $\Vert\bm{A}\Vert_{2,\infty}$
denote the spectral norm, the Frobenius norm, and the $\ell_{2,\infty}$-norm
(i.e., $\Vert\bm{A}\Vert_{2,\infty}:=\sup_{1\leq i\leq m}\Vert\bm{A}_{i,\cdot}\Vert_{2}$),
respectively. For an index set $S\subseteq\{1,2,\ldots,m\}$ (resp.
$S\subseteq\{1,2,\ldots,n\}$), we use $|S|$ to denote its cardinality,
and use $\bm{A}_{S,\cdot}$ (resp. $\bm{A}_{\cdot,S}$) to denote
a submatrix of $\bm{A}$ whose rows (resp. columns) are indexed by
$S$. We let $col(\bm{A})$, $\bm{A}^{+}$, and $\bm{P}_{\bm{A}}:=\bm{A}\bm{A}^{+}$
denote the column subspace of $\bm{A}$, the generalized inverse of
$\bm{A}$,
and the projection matrix onto the column space of $\bm{A}$.
We denote by $\mathsf{diag}(a_{1},a_{2},\ldots,a_{r})$ the $r\times r$
diagonal matrix whose diagonal entries are given by $a_{1},a_{2},\ldots,a_{r}$.
Let $\bm{I}_{r}$ be the $r\times r$ identity matrix. We denote by
$\mathcal{O}^{r\times r}$ the set of all $r\times r$ orthonormal
(or rotation) matrices. For a non-singular $n\times n$ matrix $\bm{H}$
with SVD $\bm{U}_{H}\bm{\Sigma}_{H}\bm{V}_{H}^{\top}$, we denote
by $\mathsf{sgn}(\bm{H})$ the following orthogonal matrix $\mathsf{sgn}(\bm{H})\coloneqq\bm{U}_{H}\bm{V}_{H}^{\top}$.
Then we have that, for any two matrices $\widehat{\bm{U}},\bm{U}\in\mathbb{R}^{n\times r}$
with $r\leq n$, among all rotation matrices, the one that best aligns
$\widehat{\bm{U}}$ and $\bm{U}$ is precisely $\mathsf{sgn}(\widehat{\bm{U}}^{\top}\bm{U})$
\citep[see, e.g., Appendix D.2.1 in][]{ma2017implicit}, namely,
\[
\mathsf{sgn}(\widehat{\bm{U}}^{\top}\bm{U})=\arg\min_{\bm{O}\in\mathcal{O}^{r\times r}}\Vert\widehat{\bm{U}}\bm{O}-\bm{U}\Vert_{\mathrm{F}}.
\]
\section{Main results\label{sec:main_1st_approx}}
In this section, we demonstrate the desirable statistical performance
of PCA in weak factor models. We will first present a master result
(cf.~Theorem~\ref{Thm UV 1st approx row-wise error}) on subspace
error decomposition. Based on this key result, we derive non-asymptotic
distributional characterization for our estimators for the factors
and factor loadings, paving the way to data-driven statistical inference.
First, we formally introduce the PC estimators of the factor and factor
loadings in Algorithm \ref{alg:PCA-SVD}.
\begin{algorithm}[h]
\caption{The PCA-based method.}
\label{alg:PCA-SVD}\begin{algorithmic}
\STATE \textbf{{Input}}: panel data $\bm{X}$, rank $r$.
\STATE \textbf{{Compute} }the truncated rank-$r$ SVD $\widehat{\bm{U}}\widehat{\bm{\Sigma}}\widehat{\bm{V}}^{\top}$
of $T^{-1/2}\bm{X}$, where $\widehat{\bm{U}}\in\mathbb{R}^{N\times r}$
and $\widehat{\bm{V}}\in\mathbb{R}^{T\times r}$ have orthonormal
columns, and $\widehat{\bm{\Sigma}}=\mathsf{diag}(\widehat{\sigma}_{1},\widehat{\sigma}_{2},\ldots,\widehat{\sigma}_{r})\in\mathbb{R}^{r\times r}$
satisfies that $\widehat{\sigma}_{1}\geq\widehat{\sigma}_{2}\geq\cdots\geq\widehat{\sigma}_{r}$.
\STATE \textbf{{Output}} $\widehat{\bm{F}}:=T^{1/2}\widehat{\bm{V}}$
as the estimator of factors $\bm{F}$, and $\widehat{\bm{B}}:=\widehat{\bm{U}}\widehat{\bm{\bm{\Sigma}}}$
as the estimator of factor loadings $\bm{B}$.
\end{algorithmic}
\end{algorithm}
Our master result focuses on the subspace estimates $\widehat{\bm{U}}$
and $\widehat{\bm{V}}$. We define some relevant quantities first.
Denote by $\bm{U}\bm{\Lambda}\bm{V}^{\top}$ the SVD of $T^{-1/2}\bm{BF}^{\top}$,
where both $\bm{U}\in\mathbb{R}^{N\times r}$ and $\bm{V}\in\mathbb{R}^{T\times r}$
have orthonormal columns, and $\bm{\Lambda}=\mathsf{diag}(\lambda_{1},\lambda_{2},\ldots,\lambda_{r})\in\mathbb{R}^{r\times r}$
satisfies that $\lambda_{1}\geq\lambda_{2}\geq\cdots\geq\lambda_{r}\geq0$.
We define $n:=\max(N,T)$.
\subsection{A first-order characterization of subspace perturbation errors}
Before presenting our master results, we list two key assumptions
on the SNR. To characterize the SNR, we define
\[
\theta:=\sigma_{r}/\sqrt{\Vert\bm{\Sigma}_{\varepsilon}\Vert_{2}},\qquad\vartheta_{k}:=\sigma_{r}/\Vert(\bm{\Sigma}_{\varepsilon}^{1/2})_{k,\cdot}\Vert_{1},\text{\qquad and\qquad}\vartheta:=\sigma_{r}/\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{1}.
\]
Here, the signal strength, which is the factor strength in our case,
is effectively represented by $\sigma_{r}$, the smallest singular
value of the factor loading matrix $\bm{B}$. The noise level is captured
by the norm of the covariance matrix $\bm{\Sigma}_{\varepsilon}$
for the idiosyncratic noise $\bm{E}=\bm{\Sigma}_{\varepsilon}^{1/2}\bm{Z}$
(cf.~(\ref{noise matrix formula})). The SNRs $\theta$, $\vartheta_{k}$,
and $\vartheta$ are similar, though their denominators adopt different
norms of $\bm{\Sigma}_{\varepsilon}$. The inclusion of $\ell_{1}$-norms
$\Vert(\bm{\Sigma}_{\varepsilon}^{1/2})_{k,\cdot}\Vert_{1}$ and $\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{1}=\max_{1\leq k\leq N}\Vert(\bm{\Sigma}_{\varepsilon}^{1/2})_{k,\cdot}\Vert_{1}$,
though less common than the spectral norm $\sqrt{\Vert\bm{\Sigma}_{\varepsilon}\Vert_{2}}=\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{2}$,
is that our main results are on the row-wise error bound and we need
these norms in the matrix inequalities for technical reasons. Then,
two crucial assumptions on SNR are as follows.
\begin{assumption} \label{Assump_SNR 1 norm}There exists a sufficiently
large constant $C_{1}>0$ such that
\begin{equation}
\vartheta\equiv\frac{\sigma_{r}}{\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{1}}\geq C_{1}\sqrt{\frac{n}{T}\log n}.\label{SNR 1-norm logn}
\end{equation}
\end{assumption}
\begin{assumption} \label{Assump_SNR 2 norm}There exists a sufficiently
large constant $C_{2}>0$ such that
\begin{equation}
\theta\equiv\frac{\sigma_{r}}{\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{2}}\geq C_{2}\sqrt{\frac{n}{T}\log n}.\label{SNR 2-norm logn}
\end{equation}
\end{assumption}
Assumption \ref{Assump_SNR 1 norm} implies Assumption \ref{Assump_SNR 2 norm}
due to the elementary inequality $\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{2}\leq\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{1}$.
On the other hand, $\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{1}\leq s\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{2}$,
where $s$ is the maximum number of nonzero elements of $\bm{\Sigma}_{\varepsilon}^{1/2}$
in each row, which is small when $\bm{\Sigma}_{\varepsilon}$ is sparse.
In particular, when the noise covariance matrix $\bm{\Sigma}_{\varepsilon}$
is diagonal (i.e.,~no cross-sectional correlations, also known as
the strict factor model \citep{Ross1976APT,fanfanlv2008covFactor,baishi2011Cov}),
we have that $\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{2}=\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{1}$.
In this case, Assumption \ref{Assump_SNR 1 norm} is equivalent with
Assumption \ref{Assump_SNR 2 norm}. Throughout the paper, we denote
by $C_{0}$, $C_{1}$, $C_{2}$, $C_{U}$, $C_{V}$, $c_{0}$, etc.~the
generic constants that may vary from place to place. We are now ready
to present our main results as follows.
\begin{theorem}\label{Thm UV 1st approx row-wise error} Assume that
$T\geq C_{0}(r+\log n)$ and $n\geq C_{0}r\log n$ for some sufficiently
large constant $C_{0}>0$. Consider the first-order expansions
\begin{subequations}
\begin{align}
\widehat{\bm{U}}\bm{R}_{U}-\bm{U} & =\bm{G}_{U}+\bm{\Psi}_{U}\text{\qquad with\qquad}\bm{G}_{U}:=T^{-1/2}\bm{E}\bm{V}\bm{\Lambda}^{-1},\label{U 1st-order approx hat til}\\
\widehat{\bm{V}}\bm{R}_{V}-\bm{V} & =\bm{G}_{V}+\bm{\Psi}_{V}\text{\qquad with\qquad}\bm{G}_{V}:=T^{-1/2}\bm{E}^{\top}\bm{U}\bm{\Lambda}^{-1},\label{V 1st-order approx hat til}
\end{align}
\end{subequations}
where $\bm{R}_{U}:=\mathsf{sgn}(\widehat{\bm{U}}^{\top}\bm{U})$
and $\bm{R}_{V}:=\mathsf{sgn}(\widehat{\bm{V}}^{\top}\bm{V})$ are
two global rotation matrices. Under Assumptions \ref{Assump_Bf_identification},
\ref{Assump_noise_Z_entries}, and \ref{Assump_factor_f}, we have,
with probability at least $1-O(n^{-2})$, the remainder terms $\bm{\Psi}_{U}$
and $\bm{\Psi}_{V}$ are higher-order negligible terms that satisfy
the following bounds:
(i) Under Assumption \ref{Assump_SNR 1 norm}, there exists some universal
constant $C_{U}>0$ such that, uniformly for all $k=1,2,\ldots,N$,
\[
\Vert(\bm{\Psi}_{U})_{k,\cdot}\Vert_{2}\leq C_{U}\frac{\sqrt{n}}{\vartheta_{k}\vartheta T}\sqrt{r}\log^{3/2}n+C_{U}(\frac{n}{\theta^{2}T}+\frac{1}{\theta\sqrt{T}}\sqrt{r}\log n)\Vert\bm{U}_{k,\cdot}\Vert_{2}+C_{U}\frac{n}{\vartheta_{k}\vartheta T}\log n\Vert\bm{U}\Vert_{2,\infty}.
\]
(ii) Under Assumption \ref{Assump_SNR 2 norm}, there exists some
universal constant $C_{V}>0$ such that, uniformly for all $l=1,2,\ldots,T$,
\[
\Vert(\bm{\Psi}_{V})_{l,\cdot}\Vert_{2}\leq C_{V}\frac{\sqrt{n}}{\theta^{2}T}\sqrt{r}\log^{3/2}n+C_{V}(\frac{n}{\theta^{2}T}+\frac{1}{\theta\sqrt{T}})\sqrt{r}\log n\Vert\bm{V}_{l,\cdot}\Vert_{2}.
\]
\end{theorem}
The perturbation bounds in Theorem \ref{Thm UV 1st approx row-wise error}
pave the way for the statistical inference of the factor loadings
and factors. The reasons why the quantities presented in Theorem \ref{Thm UV 1st approx row-wise error}
relate to the factor loadings $\bm{B}$ and factors $\bm{F}$ are
that, as demonstrated in Lemma \ref{Lemma SVD BF good event}, the
column subspaces of $\bm{U}$ and $\bm{B}$ are identical, as are
those of $\bm{V}$ and $\bm{F}$.
To show how strong our SNR condition is, we examine scenarios where
there is no correlation in the noise matrix $\bm{E}$, i.e., $\bm{\Sigma}_{\varepsilon}$
is diagonal as in the strict factor model. In this case, the SNRs
satisfy $\vartheta=\theta=\sigma_{r}/\sigma_{\varepsilon}$, where
$\sigma_{\varepsilon}=\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{2}=\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{1}$.
Then Assumptions \ref{Assump_SNR 1 norm} and \ref{Assump_SNR 2 norm}
match the minimal condition $\theta\gg\sqrt{n/T}$ required for the
consistent estimations of the left and right singular subspaces \citep[e.g.,][]{yan2021inference},
up to log factors.
To explain that why $\bm{\Psi}_{U}$ and $\bm{\Psi}_{V}$ are higher-order
negligible terms, we compare their bounds with the magnitude of the
first-order terms $\bm{G}_{U}$ and $\bm{G}_{V}$. For simplicity,
we consider the setting where $r\asymp1$, $N\asymp T$, $\bm{\Sigma}_{\varepsilon}$
is well-conditioned (i.e., $\kappa_{\varepsilon}:=\lambda_{\text{max}}(\bm{\Sigma}_{\varepsilon})/\lambda_{\text{min}}(\bm{\Sigma}_{\varepsilon})\asymp1$),
and $\bm{U}$ is incoherent \citep[see, e.g., Definition 3.1 in][]{chen2021Monograph};
the matrix $\bm{V}$ can be proven to be incoherent since $\Vert\bm{V}\Vert_{2,\infty}\lesssim T^{-1/2}\log^{1/2}n$.
Then, the typical size of the $k$-th row of the first-order term
$\bm{G}_{U}$ can be measured by $\mathbb{E}^{1/2}[\Vert(\bm{G}_{U})_{k,\cdot}\Vert_{2}^{2}]$.
Standard computations yield $\mathbb{E}^{1/2}[\Vert(\bm{G}_{U})_{k,\cdot}\Vert_{2}^{2}]\geq T^{-1/2}\sigma_{r}^{-1}\sqrt{\lambda_{\text{min}}(\bm{\Sigma}_{\varepsilon})}$.
Similarly, for the first-order term $\bm{G}_{V}$, we have that $\mathbb{E}^{1/2}[\Vert(\bm{G}_{V})_{l,\cdot}\Vert_{2}^{2}]\geq T^{-1/2}\sigma_{r}^{-1}\sqrt{\lambda_{\text{min}}(\bm{\Sigma}_{\varepsilon})}$.
Then, using the row-wise bounds for $\bm{\Psi}_{U}$ and $\bm{\Psi}_{V}$
in Theorem \ref{Thm UV 1st approx row-wise error}, we obtain that
\begin{subequations}
\begin{equation}
(\mathbb{E}^{1/2}[\left\Vert (\bm{G}_{U})_{k,\cdot}\right\Vert _{2}^{2}])^{-1}\left\Vert (\bm{\Psi}_{U})_{k,\cdot}\right\Vert _{2}\lesssim(\frac{\theta}{\vartheta_{k}\vartheta}+\frac{1}{\theta})\log^{3/2}n+\frac{1}{\sqrt{N}}\log n\ll1,\label{U first-order term size ratio}
\end{equation}
and
\begin{equation}
(\mathbb{E}^{1/2}[\left\Vert (\bm{G}_{V})_{l,\cdot}\right\Vert _{2}^{2}])^{-1}\left\Vert (\bm{\Psi}_{V})_{l,\cdot}\right\Vert _{2}\lesssim(\frac{1}{\theta}+\frac{1}{\sqrt{T}})\log^{3/2}n\ll1.\label{V first-order term size ratio}
\end{equation}
\end{subequations}
The last inequality in (\ref{U first-order term size ratio}) (resp.
(\ref{V first-order term size ratio})) implies that the first-order
term $(\bm{G}_{U})_{k,\cdot}$ (resp. $(\bm{G}_{V})_{l,\cdot}$) dominates
the higher-order term $(\bm{\Psi}_{U})_{k,\cdot}$ (resp. $(\bm{\Psi}_{V})_{l,\cdot}$),
provided that the SNRs grow at polynomial rate of $\log n$: $\theta^{-1}\vartheta_{k}\vartheta\gg\log^{3/2}n$
and $\theta\gg\log^{3/2}n$ (resp. $\theta\gg\log^{3/2}n$). These
conditions are less stringent than the assumptions in the prior work
\citep[e.g.,][]{BaiNg2023PCA,Jiang2023PCA,ChoiMing2024PCA} that required
the SNR to grow with a polynomial rate of $n$.
As we will show later, Theorem \ref{Thm UV 1st approx row-wise error}
is the foundation to design test statistics and conduct statistical
inference for the factors, factor loadings, and other parameters of
interest. To the best of our knowledge, the non-asymptotic first-order
approximations (\ref{U 1st-order approx hat til})--(\ref{V 1st-order approx hat til})
are new in the statistics and econometrics literature. While similar
findings exist in the fields of low-rank matrix completion and spectral
methods, the closest result to ours is Theorem 9 of \citet{yan2021inference}.
However, they assumed that the entries of the noise matrix $\bm{E}$
are independent, while we allow the presence of cross-sectional correlations
in the noise matrix $\bm{E}$.
Our results in Theorem \ref{Thm UV 1st approx row-wise error} are
nontrivial generalizations of those in \citet{yan2021inference}.
We note that their proof of row-wise error bounds via leave-one-out
(LOO) technique requires constructing an auxiliary matrix that is
independent with the target row. This construction is obtained by zeroing
the row with the same position in noise matrix. However, this approach
does not work in our case because it is still correlated with the
target row due to the cross-sectional correlation among all rows.
We overcome this challenge through applications of matrix concentration
inequalities, and our results do not need any additional structural
assumption (e.g., sparsity) on the noise covariance matrix $\bm{\Sigma}_{\varepsilon}$.
An additional advance of our theory compared with \citet{yan2021inference}
lies in that, our bounds for $\Vert(\bm{\Psi}_{U})_{k,\cdot}\Vert_{2}$
and $\Vert(\bm{\Psi}_{V})_{l,\cdot}\Vert_{2}$ are free of condition
number $\kappa:=\sigma_{1}/\sigma_{r}$, making our theory work even
when the factor loading matrix $\bm{B}$ is near-singular and ill-conditioned,
i.e., $\kappa$ is very large. In particular, our result accomodates
the heterogeneous case of \citet{BaiNg2023PCA}.
\subsection{Implications: estimation guarantees and distributional characterizations
\label{sec:F B inference estimate}}
In this section, we present the immediate consequences of Theorem
\ref{Thm UV 1st approx row-wise error} --- estimation error bounds
and distributional characterization for the PC estimators. Later in
Section \ref{sec:compare_SNR}, we will compare these results with
the recent work (\citet{BaiNg2023PCA,Jiang2023PCA,ChoiMing2024PCA}),
and highlight the advantages of our theory.
Recall that our estimators for the factors $\bm{F}$ and factor loadings
$\bm{B}$ in Algorithm \ref{alg:PCA-SVD} are given by $\widehat{\bm{F}}=\sqrt{T}\widehat{\bm{V}}$
and $\widehat{\bm{B}}=\widehat{\bm{U}}\widehat{\bm{\bm{\Sigma}}}$
respectively. Let us take a moment to see how to quantify the estimation
error for $\widehat{\bm{F}}$ in the face of rotational ambiguity.
In view of Theorem~\ref{Thm UV 1st approx row-wise error}, we know
that $\widehat{\bm{F}}\bm{R}_{V}$ should be close to $\sqrt{T}\bm{V}$
in Euclidean distance. In addition, recall that $\bm{U}\bm{\Lambda}\bm{V}^{\top}$
is the SVD of the rank-$r$ matrix $T^{-1/2}\bm{BF}^{\top}$, which
suggests the existence of an invertible, $\sigma(\bm{F})$-measurable
matrix $\bm{J}\in\mathbb{R}^{r\times r}$ such that $\bm{V}=T^{-1/2}\bm{FJ}$.
Hence, by defining $\bm{R}_{F}\coloneqq\bm{J}\bm{R}_{V}^{\top}$,
we may use $\widehat{\bm{F}}-\bm{F}\bm{R}_{F}$ to evaluate the estimation
error of $\widehat{\bm{F}}$. Similarly, we define $\bm{R}_{B}\coloneqq(\bm{R}_{F}^{-1})^{\top}$
and use $\widehat{\bm{B}}-\bm{B}\bm{R}_{B}$ to measure the estimation
error of $\widehat{\bm{B}}$.
We will show in Lemma \ref{Lemma SVD BF good event} that $\bm{R}_{F}$
and $\bm{R}_{B}$ are close to rotation matrices. Indeed, in the study
of PC estimators for factor model (e.g., \citet{BaiNg2023PCA,Jiang2023PCA,ChoiMing2024PCA}),
it is quite common to take $\bm{F}\bm{R}_{F}$ and $\bm{B}\bm{R}_{B}$,
rather than $\bm{F}$ and $\bm{B}$, as the groundtruth, though the
specific choices of $\bm{R}_{F}$ and $\bm{R}_{B}$ vary across different
works. Both the averaged and the row-wise estimation error bound can
be deduced from Theorem \ref{Thm UV 1st approx row-wise error}. We
record the results in the following corollary.
\begin{corollary}[\textsf{Estimation guarantees for factors and
factor loadings}]\label{corollary:error bound B F}Suppose that the
assumptions in Theorem \ref{Thm UV 1st approx row-wise error} hold.
Then there exists a universal constant $C_{0}>0$ such that: with
probability at least $1-O(n^{-2})$, for factors, the averaged and
the row-wise estimation error are given by
\begin{equation}
\frac{1}{T}\big\Vert\widehat{\bm{F}}-\bm{F}\bm{R}_{F}\big\Vert_{\mathrm{F}}^{2}\leq C_{0}\frac{1}{\theta^{2}}\frac{n}{T}r\text{\qquad and\qquad}\big\Vert(\widehat{\bm{F}}-\bm{F}\bm{R}_{F})_{t,\cdot}\big\Vert_{2}\leq C_{0}\frac{1}{\theta}\sqrt{r}\log n,\forall t,\label{averaged and rowwise error F}
\end{equation}
respectively; for factor loadings, we define $\bar{\bm{U}}:=\bm{B\Sigma}^{-1}\in\mathbb{R}^{N\times r}$,
if there exists some universal constant $C_{U}>0$ such that $\Vert\bar{\bm{U}}\Vert_{2,\infty}\leq C_{U}\theta/\vartheta$,
$\Vert\bar{\bm{U}}\Vert_{2,\infty}\leq C_{U}\sqrt{T}\vartheta/n$
and $\Vert\bar{\bm{U}}\Vert_{2,\infty}\leq C_{U}\sqrt{T}\theta^{2}/(\vartheta n)$,
then the averaged estimation error bound is given by
\begin{subequations}
\begin{equation}
\frac{1}{N}\big\Vert\widehat{\bm{B}}-\bm{B}\bm{R}_{B}\big\Vert_{\mathrm{F}}^{2}\leq C_{0}\bigl(\frac{1}{T}\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{1}^{2}+(\frac{n^{2}}{\theta^{4}T^{2}}+\frac{1}{\theta^{2}T})\sigma_{r}^{2}\bigr)r^{3}\log^{4}n,\label{averaged error B}
\end{equation}
and the row-wise estimation error bound is given by
\begin{equation}
\big\Vert(\widehat{\bm{B}}-\bm{B}\bm{R}_{B})_{i,\cdot}\big\Vert_{2}\leq C_{0}\bigl(\frac{1}{\sqrt{T}}\Vert(\bm{\Sigma}_{\varepsilon}^{1/2})_{i,\cdot}\Vert_{1}+(\frac{n}{\theta^{2}T}+\frac{1}{\theta\sqrt{T}})\sigma_{r}\Vert\bar{\bm{U}}_{i,\cdot}\Vert_{2}\bigr)r\log^{2}n.\label{rowwise error B}
\end{equation}
\end{subequations}
\end{corollary}
Let us interpret the above bounds under some specific settings. For
simplicity, we ignore the log factors $\log n$ and consider the setting
discussed after Theorem \ref{Thm UV 1st approx row-wise error}, with
the assumption that $\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{1}\asymp\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{2}\asymp1$.
Under this setting, to make the assumptions on $\Vert\bar{\bm{U}}\Vert_{2,\infty}$
hold, it suffices to assume that $\theta\gtrsim1$, where $\theta\equiv\sigma_{r}/\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{2}$
is the SNR. For factors, the upper bounds for the averaged and row-wise
estimation error rates in (\ref{averaged and rowwise error F}) are
given by $\theta^{-2}$ and $\theta^{-1}$, respectively; for factor
loadings, the two bounds in (\ref{averaged error B})--(\ref{rowwise error B})
are given by $T^{-1}+N^{-1}\theta^{-2}$ and $T^{-1/2}+N^{-1/2}\theta^{-1}$,
respectively. All these bounds go to zero under the condition that
$\theta\gg1$, which match the minimal condition required for the
consistent estimations of the left and right singular subspaces if
there is no correlation in the noise matrix \citep[e.g.,][]{yan2021inference}.
Next, we present our results on the inference for factors and factor
loadings in Corollaries \ref{Thm F inference} and \ref{Thm B inference},
respectively.
\begin{corollary}[\textsf{Distributional theory for factors}]\label{Thm F inference}Suppose
that Assumptions \ref{Assump_Bf_identification}, \ref{Assump_noise_Z_entries},
and \ref{Assump_factor_f} hold. For any given target error level
$\delta>0$, assume that $T\geq C_{1}(r+\log n)$, $n\geq C_{1}r\log n$,
$T\geq C_{0}\kappa_{\varepsilon}\delta^{-2}r^{2}\log^{4}n$, $n\geq C_{0}\delta^{-1/2}$,
\[
\theta\geq C_{0}\delta^{-1}\sqrt{\kappa_{\varepsilon}}\frac{n}{T}r\log^{3/2}n,\qquad\text{and}\qquad\frac{\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\bar{\bm{U}}\Vert_{2,\infty}}{\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{2}}\leq c_{0}\delta\kappa_{\varepsilon}^{-1/2}r^{-5/4},
\]
hold for some universal constant $c_{0},C_{0}>0$ and sufficiently
large constant $C_{1}>0$. Then, for any $t=1,2,\ldots,T$, it holds
that
\begin{equation}
\sup_{\mathcal{C}\in\mathscr{C}^{r}}\left\vert \mathbb{P}((\widehat{\bm{F}}-\bm{F}\bm{R}_{F})_{t,\cdot}\in\mathcal{C})-\mathbb{P}(\mathcal{N}(0,\bm{\Sigma}_{F,t})\in\mathcal{C})\right\vert \leq\delta,\label{F Gaussian law result}
\end{equation}
where $\mathscr{C}^{r}$ is the collection of all convex sets in $\mathbb{R}^{r}$,
and the covariance matrix $\bm{\Sigma}_{F,t}$ is given by $\bm{\Sigma}_{F,t}=\bm{R}_{V}\bm{\Lambda}^{-1}\bm{U}^{\top}\bm{\Sigma}_{\varepsilon}\bm{U}\bm{\Lambda}^{-1}\bm{R}_{V}^{\top}$.
\end{corollary}
\begin{remark}If all the entries of the matrix $\bm{Z}$ are Gaussian,
i.e., the noise is Gaussian, then the result (\ref{F Gaussian law result})
holds without the assumption that $\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\bar{\bm{U}}\Vert_{2,\infty}/\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{2}\leq c_{0}\kappa_{\varepsilon}^{-1/2}r^{-5/4}\delta$.
Indeed, in the scenario where $\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{1}/\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{2}\asymp1$
and $\bar{\bm{U}}$ is $\mu$-incoherent,
fulfilling our assumption on $\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\bar{\bm{U}}\Vert_{2,\infty}/\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{2}$
merely requires that $N\geq\mu\kappa_{\varepsilon}r^{5/2}c_{0}^{-2}\delta^{-2}$
on the growth rate of cross-sectional dimension $N$.\end{remark}
\begin{corollary}[\textsf{Distributional theory for factor loadings}]\label{Thm B inference}Suppose
that Assumptions \ref{Assump_Bf_identification}, \ref{Assump_noise_Z_entries},
and \ref{Assump_factor_f} hold. Assume that there exists a constant
$c_{\varepsilon}>0$ such that $\lambda_{\min}(\bm{\Sigma}_{\varepsilon})>c_{\varepsilon}$.
For any given target error level $\delta>0$, assume that $T\geq C_{1}(r+\log n)$,
$n\geq C_{1}r\log n$, $T\geq C_{0}\delta^{-2}(r^{2}\log n+\kappa_{\varepsilon}r\log^{4}n)$,
\[
\vartheta_{i}\geq C_{0}\delta^{-1}\frac{1}{\sqrt{c_{\varepsilon}}}\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{1}\frac{n}{T}\sqrt{r}\log^{3/2}n,\qquad\text{and}\qquad\theta\geq C_{0}\delta^{-1}\sqrt{\kappa_{\varepsilon}}\frac{n}{T}\sqrt{r}\log n,
\]
for some universal constant $C_{0}>0$ and sufficiently large constant
$C_{1}>0$. Then, for any $i=1,2,\ldots,N$, it holds that
\[
\sup_{\mathcal{C}\in\mathscr{C}^{r}}\left\vert \mathbb{P}((\widehat{\bm{B}}-\bm{B}\bm{R}_{B})_{i,\cdot}\in\mathcal{C})-\mathbb{P}(\mathcal{N}(0,\bm{\Sigma}_{B,i})\in\mathcal{C})\right\vert \leq\delta,
\]
where $\mathscr{C}^{r}$ is the collection of all convex sets in $\mathbb{R}^{r}$,
and the covariance matrix $\bm{\Sigma}_{B,i}$ is given by $\bm{\Sigma}_{B,i}=T^{-1}(\bm{\Sigma}_{\varepsilon})_{i,i}\bm{I}_{r}$.
\end{corollary}
Let us look at the assumptions about the SNRs $\theta$ and $\vartheta_{i}$
in Corollaries \ref{Thm F inference} and \ref{Thm B inference}.
Consider the setting where $r\asymp1$, $N\asymp T$, $c_{\varepsilon}\asymp1$,
$\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{1}\asymp\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{2}\asymp1$,
$\Vert(\bm{\Sigma}_{\varepsilon}^{1/2})_{i,\cdot}\Vert_{1}\asymp1$,
and $\bm{\Sigma}_{\varepsilon}$ is well-conditioned. Then, the SNR
condition in Corollary \ref{Thm F inference} (resp. Corollary \ref{Thm B inference})
is equivalent to $\theta\gg\log^{3/2}N$ (resp. $\theta\gg\log N$
and $\vartheta_{i}\gg\log^{3/2}N$), indicating that inference of
factors and factor loadings is achievable as long as the SNR grows
faster than a polynomial rate of $\log N$. Our SNR condition is less
restrictive than the prior work \citep[e.g.,][]{BaiNg2023PCA,Jiang2023PCA,ChoiMing2024PCA}
that required the SNR to grow with a polynomial rate of $N$. Regardless
of log factors, our SNR condition for inference is equivalent to $\theta\gg1$,
which is optimal since it matches the same condition required for
consistency as commented after Corollary \ref{corollary:error bound B F}.
Also, in the special case where $\bm{\Sigma}_{\varepsilon}$ is diagonal,
i.e., there is no correlation in the noise matrix, our SNR condition
for inference matches the minimal condition required for the subspace
inference \citep[e.g.,][]{yan2021inference}.
Both Corollaries \ref{Thm F inference} and \ref{Thm B inference}
are stated in a non-asymptotic sense, and can be easily translated
to asymptotic normality results. In practice, the confidence regions
for the factors and the factor loadings can be constructed by replacing
the asymptotic covariance matrices $\bm{\Sigma}_{F,t}$ and $\bm{\Sigma}_{B,i}$
with their consistent estimators:
\[
\widehat{\bm{\Sigma}}_{F,t}=\widehat{\bm{\Sigma}}^{-1}\widehat{\bm{U}}^{\top}\widehat{\bm{\bm{\Sigma}}}_{\varepsilon}^{\tau}\widehat{\bm{U}}\widehat{\bm{\Sigma}}^{-1}\qquad\text{and}\qquad\widehat{\bm{\Sigma}}_{B,i}=\frac{1}{T}(\widehat{\bm{\bm{\Sigma}}}_{\varepsilon}^{\tau})_{i,i}\bm{I}_{r},
\]
respectively. Here, $\widehat{\bm{\bm{\Sigma}}}_{\varepsilon}^{\tau}$
is the estimator of the noise covariance matrix which we will introduce
in detail in Section \ref{sec:noise cov estimation}. The intuition
behind the above consistent estimators is that, $\widehat{\bm{U}}$
and $\widehat{\bm{\Sigma}}$ are the consistent estimators of $\bm{U}\bm{R}_{U}$
and $\bm{R}_{U}^{\top}\bm{\Lambda}\bm{R}_{V}$, respectively, and
we will show this fact in the proof of Theorem \ref{Thm UV 1st approx row-wise error}.
\section{Comparison with previous work\label{sec:compare_SNR}}
In this section, we compare our main results with related ones established
in prior works \citet{BaiNg2023PCA,Jiang2023PCA,ChoiMing2024PCA}.
All of them studied the PC estimators in weak factor models, and established
estimation error bounds and asymptotic normality under different assumptions.
Distinct from these three papers, all our results are entirely non-asymptotic,
providing finite-sample characterizations for both the estimation
error and the uncertainty level of statistical inference. From a technical
viewpoint, in the regime $N\asymp T$, a notable advancement of our
theory is that, our assumptions for inference require the growth rate
of SNR to be faster than a polynomial rate of $\log N$, unlike the
polynomial rate of $N$ required in these three papers. Also, all
these three papers assume that the $r$ singular values $\{\sigma_{i}\}_{i=1}^{r}$
of factor loading $\bm{B}$ are distinct, while our theory does not
require such eigengap condition.
Having provided an overview of the advantages of our theory, we now
make the comparisons in detail.
\begin{itemize}
\item \citet{BaiNg2023PCA} studied both the homogeneous case where $\sigma_{i}^{2}\asymp N^{\alpha}$
for $i=1,2,\ldots,r$, and the heterogeneous case where $\sigma_{i}^{2}\asymp N^{\alpha_{i}}$
with $1\geq\alpha_{1}\geq\alpha_{2}\geq\ldots\geq\alpha_{r}>0$. They
required $\alpha_{r}>1/2$ to prove the asymptotic normality for inference.
However, the case when $0<\alpha_{r}\leq1/2$ has not been covered
by their inferential theory. Our theory fills this gap and establishes
the inference results even when $\sigma_{r}$ do not grow with a polynomial
rate of $N$. Besides the inference, our theory provides row-wise
estimation error bounds for factors and factor loadings as detailed
in Section \ref{sec:F B inference estimate}, while the bounds in
\citet{BaiNg2023PCA} are Frobenius norm bounds measuring the average
error over all rows. The consistency and asymptotic normality results
in \citet{Jiang2023PCA} are similar to those in \citet{BaiNg2023PCA},
while the focus of \citet{Jiang2023PCA} is to identify the so-called
pseudo-true parameter that is consistently estimated by the PCA estimator.
We note that both their frameworks accommodate temporal correlation,
an aspect not covered by our theory.
\item \citet{ChoiMing2024PCA} adopted the leave-one-out
analysis similar in spirit to those used in matrix completion to investigate
the PCA estimators. The setup they studied is the homogeneous case
in \citet{BaiNg2023PCA} because they assumed that $N^{-\alpha}\bm{B}^{\top}\bm{B}\rightarrow\bm{\Sigma}_{\bm{B}}$,
which requires $\sigma_{i}^{2}\asymp N^{\alpha}$ for all the singular
values $\{\sigma_{i}\}_{i=1,2,\ldots,r}$ of factor loading $\bm{B}$.
When temporal correlation does not exist, similar to our results,
they also filled the gap to establish the inference results for $0<\alpha\leq1/2$.
In comparison, our theory is fully non-asymptotic and does not assume
any asymptotic growth assumptions on the singular values. Also, the
homogeneous case they studied assumes that the condition number $\kappa=\sigma_{1}/\sigma_{r}$
satisfies that $\kappa\asymp1$, while our theory does not need this
assumption and allows any growth rate of the condition number $\kappa$.
\end{itemize}
To enhance the clarity of the comparison of the SNR assumption for
inference, we detailed the results in Table \ref{table:SNR normality}
under a specific setup: Assumptions \ref{Assump_Bf_identification},
\ref{Assump_noise_Z_entries}, \ref{Assump_factor_f} hold; The cross-sectional
and temporal dimensions satisfy that $N\asymp T\asymp n$; the number
of factors satisfies $r\asymp1$; the noise covariance matrix satisfies
$\kappa_{\varepsilon}\asymp1$, $\left\Vert \bm{\Sigma}_{\varepsilon}\right\Vert _{2}\asymp1$,
and $\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{1}\asymp1$. Since
the scale of the noise is reflected by $\bm{\Sigma}_{\varepsilon}$
and in our current setup we have that $\left\Vert \bm{\Sigma}_{\varepsilon}\right\Vert _{2}\asymp1$
and $\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{1}\asymp1$, the assumptions
on the SNRs $\theta=\sigma_{r}/\sqrt{\left\Vert \bm{\Sigma}_{\varepsilon}\right\Vert _{2}}$
and $\vartheta=\sigma_{r}/\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{1}$
defined in Section \ref{sec:main_1st_approx} are fully represented
by the growth rate of the smallest singular value $\sigma_{r}$ of
factor loading matrix $\bm{B}$.
\begin{table}[h]
\caption{SNR assumptions for inference of factor and factor loading\label{table:SNR normality}}
\centering
\begin{tabular}{c|c|c}
\hline
& \multirow{2}{*}{Factor} & \multirow{2}{*}{Factor loading}\tabularnewline
& & \tabularnewline
\hline
\multirow{2}{*}{\citet{BaiNg2023PCA}} & \multirow{2}{*}{$\sigma_{r}\gg N^{1/4}$} & \multirow{2}{*}{$\sigma_{r}\gg N^{1/4}$}\tabularnewline
& & \tabularnewline
\hline
\multirow{2}{*}{\citet{Jiang2023PCA}} & \multirow{2}{*}{$\sigma_{r}\gg\max(N^{1/4},\kappa^{2})$} & \multirow{2}{*}{$\sigma_{r}\gg\kappa^{1/2}N^{1/4}$}\tabularnewline
& & \tabularnewline
\hline
\multirow{2}{*}{\citet{ChoiMing2024PCA}} & \multirow{2}{*}{$\sigma_{r}\asymp N^{\alpha/2}\text{ with }\alpha>0$} & \multirow{2}{*}{$\sigma_{r}\gg N^{1/6}\log^{\omega/6}N$}\tabularnewline
& & \tabularnewline
\hline
\multirow{2}{*}{Our Theory} & \multirow{2}{*}{$\sigma_{r}\gg\log^{3/2}N$} & \multirow{2}{*}{$\sigma_{r}\gg\log^{3/2}N$}\tabularnewline
& & \tabularnewline
\hline
\end{tabular}
\end{table}
Table \ref{table:SNR normality} demonstrates the advancement of our
theory. For the conditions of \citet{ChoiMing2024PCA} in Table \ref{table:SNR normality},
the parameter $\omega>0$ is defined in Assumption B'' (iv) therein
on the noise structure and they need two additional growth rate assumptions
on $\sigma_{r}$ (see Assumption D'' (i) therein). For other regimes
where $T\ll N$ or $N\ll T$, similar comparisons can be made, so
we omit the details for the limit of space.
\section{Applications in econometrics\label{sec:3_econ_applications}}
Our first-order approximations do not only advance the existing theories
on the PCA under the weak factor model, but also pave the path for
various statistical tests that are useful in economics and finance.
In this section, we show four applications of our results.
\subsection{The factor specification tests}
Recall that the factor model is given by $\bm{X}=\bm{BF}^{\top}+\bm{E}$,
where the matrix $\bm{F}$ is the realization of latent factors. Suppose
we have time series data of some observed factors, e.g., the Fama-French
factors. Our focus is on testing if the observed factor is in the
linear space spanned by the latent factors $\bm{F}$. In particular,
we examine this linear dependence in a flexible range of the whole
period $[1,T]$. Formally, we consider an index set $S=\{t_{1},t_{2},\ldots,t_{|S|}\}\subseteq\{1,2,\ldots,T\}$
of interest, and we have the data $\bm{v}=(v_{t_{1}},v_{t_{2}},\ldots,v_{t_{|S|}})^{\top}\in\mathbb{R}^{|S|}$
of an observed factor recorded at the time index set $S$. We test
the hypothesis as follows,
\begin{equation}
H_{0}:\qquad\text{There exists }\bm{w}\in\mathbb{R}^{r}\text{ such that }\bm{v}=\bm{F}_{S,\cdot}\bm{w}.\label{Null H0 factor test subset}
\end{equation}
Under the null hypothesis $H_{0}$, we have that $v_{t}=\bm{F}_{t,\cdot}\bm{w}=\bm{f}_{t}^{\top}\bm{w}$
for any $t\in S$.
Under the strong factor model where the pervasiveness assumption holds,
\citet{BaiNg2006factor} studied this problem and designed test statistics
for the whole set, i.e., $S=\{1,2,\ldots,T\}$. Our study extends
to the case when $S$ is a subset of the whole time span $\{1,2,\ldots,T\}$
under the weak factor model. The subset $S$ can be any specific time
window of interest. This scenario is economically meaningful because
the relationship $v_{t}=\bm{f}_{t}^{\top}\bm{w}$ may only be valid
for relatively short periods, not necessarily across the entire span;
see \citet{BaiNg2006factor} for the connections between the CAPM
analysis and this problem. In Section \ref{sec:Numerical-experiments},
we show that our factor specification test results strikingly reconcile
with the economic cycles and financial crisis.
Our test statistic relies on the estimation of the noise covariance
matrix $\bm{\Sigma}_{\varepsilon}$. We denote by $\widehat{\bm{\bm{\Sigma}}}_{\varepsilon}^{\tau}$
the estimator of $\bm{\Sigma}_{\varepsilon}$ and we will elaborate
its construction in the next section. In Theorem \ref{Thm factor test plug-in Chi-sq}
below, we keep the error bound $\Vert\widehat{\bm{\bm{\Sigma}}}_{\varepsilon}^{\tau}-\bm{\Sigma}_{\varepsilon}\Vert_{2}$
in the final result.
\begin{theorem}\label{Thm factor test plug-in Chi-sq} For any given
target error level $\delta>0$, assume there exist some universal
constants $c_{0},C_{0}>0$ and sufficiently large constant $C_{1}>0$
such that, $n\geq C_{1}r\log n$, $|S|\geq C_{1}(r+\log n)$, $T\geq C_{0}|S|\delta^{-2}\kappa_{\varepsilon}^{2}r^{2}\log^{4}n$,
\begin{equation}
\theta\geq C_{0}\delta^{-1}\frac{n}{T}\sqrt{|S|}\kappa_{\varepsilon}^{4}r^{2}\log^{4}n,\qquad\text{and}\qquad\frac{\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\bar{\bm{U}}\Vert_{2,\infty}}{\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{2}}\leq c_{0}\delta\kappa_{\varepsilon}^{-1}\sqrt{\frac{T}{n}}|S|^{-3/2};\label{assemble condition SNR and data size}
\end{equation}
and that for the estimator $\widehat{\bm{\bm{\Sigma}}}_{\varepsilon}^{\tau}$
of the noise covariance matrix $\bm{\Sigma}_{\varepsilon}$, it holds
\[
\frac{1}{\left\Vert \bm{\Sigma}_{\varepsilon}\right\Vert _{2}}\big\Vert\widehat{\bm{\Sigma}}_{\varepsilon}^{\tau}-\bm{\Sigma}_{\varepsilon}\big\Vert_{2}\leq c_{0}\delta|S|^{-1/2}\kappa_{\varepsilon}^{-4}r^{-1}\log^{-2}n.
\]
Then, under Assumptions \ref{Assump_Bf_identification}, \ref{Assump_noise_Z_entries},
\ref{Assump_factor_f}, and the null hypothesis $H_{0}$ in (\ref{Null H0 factor test subset}),
we have that
\begin{equation}
\left|\mathbb{P}\left(\mathfrak{\widehat{\mathcal{T}}}(S,\bm{v})\leq\chi_{1-\alpha}^{2}(|S|-r)\right)-(1-\alpha)\right|\leq\delta,\label{Chi-square stat final inequality}
\end{equation}
where the test statistic $\mathfrak{\widehat{\mathcal{T}}}(S,\bm{v})$
is defined by
\[
\mathfrak{\widehat{\mathcal{T}}}(S,\bm{v}):=\frac{1}{\widehat{\phi}}\bm{v}^{\top}(\bm{I}_{|S|}-\bm{P}_{\widehat{\bm{V}}_{S,\cdot}})\bm{v}\qquad\text{with}\qquad\widehat{\phi}:=\frac{1}{T}((\widehat{\bm{V}}_{S,\cdot})^{+}\bm{v})^{\top}\widehat{\bm{\Sigma}}^{-1}\widehat{\bm{U}}^{\top}\widehat{\bm{\bm{\Sigma}}}_{\varepsilon}^{\tau}\widehat{\bm{U}}\widehat{\bm{\Sigma}}^{-1}(\widehat{\bm{V}}_{S,\cdot})^{+}\bm{v}.
\]
\end{theorem}
\begin{remark}If all the entries of the noise matrix $\bm{Z}$ in
(\ref{noise matrix formula}) are Gaussian, then we do not need the
second assumption in \eqref{assemble condition SNR and data size}.
\end{remark}
The idea of the test statistic $\mathfrak{\widehat{\mathcal{T}}}(S,\bm{v})$
is to utilize the fact that, under the null hypothesis $H_{0}$, the
residual vector of $\bm{v}$ is zero after projection onto the column
space of $\bm{F}_{S,\cdot}$. Note that $\widehat{\bm{V}}_{S,\cdot}$
estimates the column space of $\bm{F}_{S,\cdot}$. So, we construct
$\mathfrak{\widehat{\mathcal{T}}}(S,\bm{v})$ by computing the $\ell_{2}$-norm
of the projection residual vector $\Vert(\bm{I}_{|S|}-\bm{P}_{\widehat{\bm{V}}_{S,\cdot}})\bm{v}\Vert_{2}^{2}=\bm{v}^{\top}(\bm{I}_{|S|}-\bm{P}_{\widehat{\bm{V}}_{S,\cdot}})\bm{v}$.
The component $\widehat{\phi}$ estimates the variance of $(\bm{I}_{|S|}-\bm{P}_{\widehat{\bm{V}}_{S,\cdot}})\bm{v}$,
and needs the estimator $\widehat{\bm{\bm{\Sigma}}}_{\varepsilon}^{\tau}$
of noise covariance matrix as input.
Our focus is on the regime where the size of subset $S$ is small. This is economically meaningful because the linear relationship between observed factors and latent factors usually holds only for a short horizon. Based on our distributional results in \eqref{Chi-square stat final inequality}, the case when $S$ is the whole period, i.e., $|S|=T$, can be handled straightforwardly by applying the Gaussian approximation to the Chi-square distribution. Our assumption on the size $|S|$ can be summarized as
$r+\log n\ll|S|\ll T/(\kappa_{\varepsilon}^{2}r^{2}\log^{4}n)$. On
the one hand, we require that $|S|$ cannot be too small so that $\bm{F}_{S,\cdot}$
is of full column rank and its singular value is close to $\sqrt{|S|}$,
which is crucial for our proof. On other hand, we require that $|S|$
cannot be too large, otherwise the null distribution $\chi^{2}(|S|-r)$
with mean $|S|-r$ would diverge in such a case, complicating the
proof of proximity between $\chi^{2}(|S|-r)$ and the law of test
statistic $\mathfrak{\widehat{\mathcal{T}}}(S,\bm{v})$.
The assumption on SNR is that $\theta\gg\frac{n}{T}\sqrt{|S|}\kappa_{\varepsilon}^{4}r^{2}\log^{4}n$.
Consider the setting where $r\asymp1$, $N\asymp T$, and $\bm{\Sigma}_{\varepsilon}$
is well-conditioned (i.e., $\kappa_{\varepsilon}\asymp1$). Then,
the SNR condition becomes $\theta\gg\sqrt{|S|}\log^{4}n$, illustrating
the effect of the subset size $|S|$ on the growth rate required for
SNR $\theta$. The $\sqrt{|S|}$ factor appears in the SNR because we use the Chi-square distribution as the null distribution to accommodate the general case when $|S|$ is small. When $|S|\asymp1$, regardless of the log factors,
the SNR condition is equivalent to $\theta\gg1$, which match the
minimal condition required for the consistent estimations of the left
and right singular subspaces if there is no correlation in the noise
matrix \citep[e.g.,][]{yan2021inference}.
\citet{BaiNg2006factor} studied the factor specification test under
the whole period case, i.e., $S=\{1,2,\ldots,T\}$. They derived the
approximate cumulative distribution function (CDF) for their test
statistic, but they did not provide a finite sample error bound for
it and their results were under the strong factor model. In comparison,
our results give a precise characterization of the error and adapt
to the weak factor model. In our case, both the low SNR in the weak
factor model and the subset $S$ in the hypothesis $H_{0}$ make it
challenging to design the test statistics. The subspace perturbation
bounds in Theorem \ref{Thm UV 1st approx row-wise error} enable us
to conquer these difficulties and establish non-asymptotic analysis
of our test statistic for any arbitrary subset $S$.
\subsubsection{Estimation of the noise covariance matrix $\bm{\Sigma}_{\varepsilon}$\label{sec:noise cov estimation}}
We accommodate the case where the cross-sectional dimension $N$ is
much larger than the temporal dimension $T$. In such case, estimating
the high-dimensional covariance matrix $\bm{\Sigma}_{\varepsilon}$
becomes challenging yet is vital for numerous statistical tests, including
the factor specification test and the two-sample test for betas discussed
in the next section. Following the approaches in, e.g., \citet{bickel2008covariance,fan2011covFsee,POET2013},
we assume that the error covariance matrix $\bm{\Sigma}_{\varepsilon}$
is sparse in a suitable sense and use the adaptive thresholding method
to estimate $\bm{\Sigma}_{\varepsilon}$. It's important to note that
the assumed sparsity of $\bm{\Sigma}_{\varepsilon}$ is solely for
our statistical tests and is not a required condition to establish
our main theories on the subspace perturbation bounds in Theorem \ref{Thm UV 1st approx row-wise error}.
\begin{assumption} \label{Assump_noise_cov_sparse}For some $q\in[0,1)$,
there exists a constant $s(\bm{\Sigma}_{\varepsilon})>0$ such that
\[
\max_{1\leq i\leq N}\sum_{j=1}^{N}((\bm{\Sigma}_{\varepsilon})_{i,i}(\bm{\Sigma}_{\varepsilon})_{j,j})^{(1-q)/2}|(\bm{\Sigma}_{\varepsilon})_{i,j}|^{q}\leq s(\bm{\Sigma}_{\varepsilon}).
\]
\end{assumption}
The assumption is slightly weaker than those in, e.g., \citet{bickel2008covariance,POET2013},
where they assumed that $\lambda_{\min}(\bm{\Sigma}_{\varepsilon})\asymp1$,
$\lambda_{\max}(\bm{\Sigma}_{\varepsilon})\asymp1$, and $\max_{1\leq i\leq N}\sum_{j=1}^{N}|(\bm{\Sigma}_{\varepsilon})_{i,j}|^{q}\leq s_{0}$
for some sparsity parameter $s_{0}$. In particular, for $q=0$, it
constrains on the maximum number of nonzero elements of $\bm{\Sigma}_{\varepsilon}$.
This sparsity assumption for the noise covariance $\bm{\Sigma}_{\varepsilon}$
is also natural in economics and finance. As the latent common factors
largely explain the co-movements in the panel data, the correlation
among individual asset's idiosyncratic noises should be close to zero.
A specific example of the sparse structure of $\bm{\Sigma}_{\varepsilon}$
arises from the remaining sector effects \citep{gaglia2016panelData}.
We estimate the idiosyncratic noise matrix using $\widehat{\bm{E}}=\bm{X}-\widehat{\bm{B}}\widehat{\bm{F}}^{\top}$,
where $\widehat{\bm{B}}$ and $\widehat{\bm{F}}$ are the PCA estimators
of factor loadings and factors defined in Section \ref{sec:F B inference estimate}.
Then, the pilot estimator of the covariance matrix $\bm{\Sigma}_{\varepsilon}$
is given by $\widehat{\bm{\bm{\Sigma}}}_{\varepsilon}=T^{-1}\widehat{\bm{E}}\widehat{\bm{E}}^{\top}$.
We will show in Lemma \ref{Lemma noise cov matrix estimate} later
that $\max_{1\leq i,j\leq N}\bigl|(\widehat{\bm{\bm{\Sigma}}}_{\varepsilon}-\bm{\Sigma}_{\varepsilon})_{i,j}\bigl|\lesssim\epsilon_{N,T}\sqrt{(\bm{\Sigma}_{\varepsilon})_{i,i}(\bm{\Sigma}_{\varepsilon})_{j,j}}$,
where $\epsilon_{N,T}>0$ is negligible as $N,T\rightarrow\infty$.
Then we apply an adaptive thresholding method \citep{bickel2008covariance}
to obtain the sparse covariance matrix estimator $\widehat{\bm{\bm{\Sigma}}}_{\varepsilon}^{\tau}$
as follows,
\[
(\widehat{\bm{\bm{\Sigma}}}_{\varepsilon}^{\tau})_{i,j}:=\Biggl\{\begin{array}{cc}
(\widehat{\bm{\bm{\Sigma}}}_{\varepsilon})_{i,j}, & \qquad\text{if}\text{ }i=j,\\
h((\widehat{\bm{\bm{\Sigma}}}_{\varepsilon})_{i,j},\tau_{i,j}), & \qquad\text{if}\text{ }i\neq j,
\end{array}
\]
where $h(z,\tau)$ is a thresholding function with the threshold value
$\tau$. Here, the threshold value $\tau_{i,j}$ is set adaptively
to $\tau_{i,j}=C\epsilon_{N,T}\sqrt{(\widehat{\bm{\bm{\Sigma}}}_{\varepsilon})_{i,i}(\widehat{\bm{\bm{\Sigma}}}_{\varepsilon})_{j,j}}$
with some large constant $C>0$. The idea behind the adaptive thresholding
is examining the sample correlation matrix and retaining entries exceeding
$C\epsilon_{N,T}$ in magnitude. The specific examples of the function
$h$ include the hard-thresholding function $h(z,\tau)=z1_{\{|z|<\tau\}}$,
among other common thresholding functions like the soft-thresholding
function and SCAD \citep{FanLi2001SCAD}. In general, we require that
the thresholding function $h(z,\tau)$ to satisfy: (i) $h(z,\tau)=0\text{ if }|z|<\tau$;
(ii) $|h(z,\tau)-z|<\tau$. The estimation error is given as follows.
\begin{lemma}\label{Lemma noise cov matrix estimate}Suppose that
all the assumptions in Corollary \ref{corollary:error bound B F}
hold, and that there exists a sufficiently small constant $\epsilon_{N,T}>0$
such that
\begin{equation}
\max_{1\leq i,j\leq N}\bigl(\frac{1}{\sqrt{T}}(1+(1+\gamma_{\varepsilon,j})\gamma_{i}+(1+\gamma_{\varepsilon,i})\gamma_{j})+(\beta_{i}\gamma_{j}+\beta_{j}\gamma_{i})(\frac{n}{\theta^{2}T}+\frac{1}{\theta\sqrt{T}})\bigr)\cdot r\log^{2}n\leq\epsilon_{N,T},\label{noise cov error condition}
\end{equation}
where
\[
\beta_{i}:=\frac{\bigl\Vert\bm{b}_{i}\bigl\Vert_{2}}{\sqrt{(\bm{\Sigma}_{\varepsilon})_{i,i}}},\qquad\gamma_{i}:=\frac{\sigma_{r}\bigl\Vert\bar{\bm{U}}_{i,\cdot}\bigl\Vert_{2}}{\sqrt{(\bm{\Sigma}_{\varepsilon})_{i,i}}}\qquad\text{and}\qquad\gamma_{\varepsilon,i}:=\frac{\bigl\Vert(\bm{\Sigma}_{\varepsilon}^{1/2})_{i,\cdot}\bigl\Vert_{1}}{\bigl\Vert(\bm{\Sigma}_{\varepsilon}^{1/2})_{i,\cdot}\bigl\Vert_{2}}.
\]
Then,
we have that, with probability at least $1-O(n^{-10})$, it holds
$\max_{1\leq i,j\leq N}((\bm{\Sigma}_{\varepsilon})_{i,i}(\bm{\Sigma}_{\varepsilon})_{j,j})^{-1/2}\bigl|(\widehat{\bm{\bm{\Sigma}}}_{\varepsilon}-\bm{\Sigma}_{\varepsilon})_{i,j}\bigl|\leq C_{0}\epsilon_{N,T}$,
for a universal constant $C_{0}>0$. Further, under Assumption \ref{Assump_noise_cov_sparse},
the generalized thresholding estimator $\widehat{\bm{\bm{\Sigma}}}_{\varepsilon}^{\tau}$
satisfies that, there exists some universal constant $C_{0}>0$ such
that, with probability at least $1-O(n^{-10})$,
\[
\big\Vert\widehat{\bm{\bm{\Sigma}}}_{\varepsilon}^{\tau}-\bm{\Sigma}_{\varepsilon}\big\Vert_{2}\leq C_{0}(\epsilon_{N,T})^{1-q}s(\bm{\Sigma}_{\varepsilon}).
\]
\end{lemma}
The above error bounds for the estimators $\widehat{\bm{\bm{\Sigma}}}_{\varepsilon}$
and $\widehat{\bm{\bm{\Sigma}}}_{\varepsilon}^{\tau}$ differ from
the existing literature owing to our row-wise subspace perturbation
bounds in Theorem \ref{Thm UV 1st approx row-wise error}. Let us
interpret the condition (\ref{noise cov error condition}) under some
specific settings. Consider the setting where $r\asymp1$, $N\asymp T$,
$\bar{\bm{U}}$ is $\mu$-incoherent,
$\Vert\bm{B}\Vert_{2,\infty}\lesssim1$
and $\bm{\Sigma}_{\varepsilon}$ is well-behaved such that $\gamma_{\varepsilon,i}\asymp1$,
$(\bm{\Sigma}_{\varepsilon})_{i,i}\asymp1$, and $\Vert\bm{\Sigma}_{\varepsilon}\Vert_{2}\asymp1$.
Then, the condition (\ref{noise cov error condition}) is equivalent
to $T\geq(\sigma_{r}\log^{2}n)/\epsilon_{N,T}$, which requires the
sample size $T$ is sufficiently large. According to Lemma \ref{Lemma noise cov matrix estimate},
fulfilling our assumption on the estimation error $\Vert\widehat{\bm{\Sigma}}_{\varepsilon}^{\tau}-\bm{\Sigma}_{\varepsilon}\Vert_{2}$
required for Theorem \ref{Thm factor test plug-in Chi-sq} merely
requires that the estimation error $\epsilon_{N,T}$ satisfies
\[
\epsilon_{N,T}\ll\bigl((\left\Vert \bm{\Sigma}_{\varepsilon}\right\Vert _{2}/s(\bm{\Sigma}_{\varepsilon}))|S|^{-1/2}\kappa_{\varepsilon}^{-4}r^{-1}\log^{-2}n\bigr)^{1/(1-q)}.
\]
In subsequent applications, we will directly use $\epsilon_{N,T}$,
instead of $\widehat{\bm{\bm{\Sigma}}}_{\varepsilon}^{\tau}$, to
express assumptions on the estimation error of the noise covariance
matrix.
\subsection{Test for structural breaks in betas}
The discussion so far has studied the factor models with constant
betas. However, in many cases, there is a possibility of time-variant
betas, and it is important to test for structural changes in betas
\citep[e.g.,][]{StockWatson2009beta_break,breitung2011beta_break,chen2014beta_break,han2015beta_break}.
We test for two time periods $\Gamma_{1}=\{t_{1}+1,t_{1}+2,\ldots,t_{1}+T_{1}\}$
and $\Gamma_{2}=\{t_{2}+1,t_{2}+2,\ldots,t_{2}+T_{2}\}$ that are
possibly not consecutive, i.e., $t_{1}+T_{1}\leq t_{2}+1$. The factor
models in these two periods are subject to a structural break in beta: for each given individual unit $i$,
\begin{align*}
x_{i,t} & =\bm{f}_{t}^{\top}\bm{b}_{i}^{1}+\varepsilon_{i,t}\qquad\text{for}\qquad t\in\Gamma_{1},\\
x_{i,t} & =\bm{f}_{t}^{\top}\bm{b}_{i}^{2}+\varepsilon_{i,t}\qquad\text{for}\qquad t\in\Gamma_{2},
\end{align*}
where $i$ is the index of the cross-sectional unit of interest. Here,
we assume that the factors are the same, but the betas can be different.
For example, we might ask if a external shock such as the 2008 financial crises has caused the changes in factor loadings before and after the shock with a transition period $(T_1, T_2)$.
The data consist of two panels, $\bm{X}^{1}$ and $\bm{X}^{2}$, during $\Gamma_{1}$ and $\Gamma_{2}$, respectively, where $\bm{X}^{1}\in\mathbb{R}^{N\times T_{1}}$
and $\bm{X}^{2}\in\mathbb{R}^{N\times T_{2}}$.
We test the hypothesis as follows:
\[
H_{0}:\ \bm{b}_{i}^{1}=\bm{b}_{i}^{2}\ \leftrightarrow\ H_{1}:\ \bm{b}_{i}^{1}\neq\bm{b}_{i}^{2}.
\]
Different from our approach, \citet{chen2014beta_break,han2015beta_break}
tested for the entire factor loading, with their null hypothesis being
$H_{0}:\ \bm{b}_{i}^{1}=\bm{b}_{i}^{2}\text{ for }\forall i$. \citet{StockWatson2009beta_break,breitung2011beta_break}
also studied the test for a given $i$ like ours, while \citet{StockWatson2009beta_break}
focused on empirical studies and \citet{breitung2011beta_break} developed
and established theories for the test statistics. Our work differs
in that, as we will show below, the effectiveness of our test statistic
does not require the pervasiveness assumption needed in \citet{breitung2011beta_break}.
In addition, our formulation allows the gap between the two periods
$\Gamma_{1}$ and $\Gamma_{2}$, which can model the special transition periods
we purposely, e.g., financial crisis, enabling us to test for
structural breaks before and after these special periods.
To construct the test statistic, we first merge the two panels $\bm{X}^{1}$
and $\bm{X}^{2}$ into a $N\times(T_{1}+T_{2})$-dimensional data
matrix $\bm{X}=(\bm{X}^{1},\bm{X}^{2})$. Next, we apply SVD as described
in Algorithm \ref{alg:PCA-SVD} to obtain an estimator of the factors:
$\widehat{\bm{F}}=\sqrt{T}\widehat{\bm{V}}$, where $T=T_{1}+T_{2}$
and $\widehat{\bm{U}}\widehat{\bm{\Sigma}}\widehat{\bm{V}}^{\top}$
is the truncated rank-$r$ SVD of $T^{-1/2}\bm{X}$ with $\widehat{\bm{U}}\in\mathbb{R}^{N\times r}$
and $\widehat{\bm{V}}\in\mathbb{R}^{T\times r}$. Then, we split the
estimated factors into $\widehat{\bm{F}}=((\widehat{\bm{F}}^{1})^{\top},(\widehat{\bm{F}}^{2})^{\top})^{\top}$
according to two time periods $\Gamma_{1}$ and $\Gamma_{2}$, where
$\widehat{\bm{F}}^{j}\in\mathbb{R}^{T_{j}\times r}$. Subsequently,
we obtain the estimator for $\bm{b}_{i}^{j}$ by regressing $\bm{X}_{i,\cdot}^{j}$
on $\widehat{\bm{F}}^{j}$:
\[
\widehat{\bm{b}}_{i}^{j}=((\widehat{\bm{F}}^{j})^{\top}\widehat{\bm{F}}^{j})^{-1}(\widehat{\bm{F}}^{j})^{\top}(\bm{X}_{i,\cdot}^{j})^{\top}\qquad\text{for}\qquad j=1,2.
\]
Finally, under null hypothesis $H_{0}$, we show that $\widehat{\bm{b}}_{i}^{1}-\widehat{\bm{b}}_{i}^{2}$
is approximately Gaussian via the first-order approximation results
in Theorem \ref{Thm UV 1st approx row-wise error}, and construct
the test statistic using the plug-in estimator of the covariance matrix.
\begin{theorem}\label{Thm beta structure test}Suppose that the assumptions
in Theorem \ref{Thm UV 1st approx row-wise error} hold, and the covariance
estimator $\widehat{\bm{\bm{\Sigma}}}_{\varepsilon}^{\tau}$ satisfies
the conditions as in Lemma \ref{Lemma noise cov matrix estimate}.
For any given target error level $\delta>0$, assume there exists
some universal constants $c_{0},C_{0},C_{U}>0$ such that $\min(T_{1},T_{2})\geq C_{0}(r+\log n)$,
$\Vert\bm{U}_{i,\cdot}\Vert_{2}\leq C_{U}\sqrt{T/(T_{1}T_{2})}$,
$T\geq C_{0}^{2}C_{U}^{2}\delta^{-2}\kappa_{\varepsilon}r\log^{3}n$,
$\vartheta_{i}\geq C_{0}C_{U}^{2}\delta^{-1}T^{-1/2}\kappa_{\varepsilon}^{2}r\log n$,
\[
\theta\geq C_{0}\{C_{U}^{2}\frac{1}{\delta}(\kappa_{\varepsilon}\frac{n}{T}+\kappa_{\varepsilon}^{2}\sqrt{\frac{n}{T}}\Vert\bm{U}_{i,\cdot}\Vert_{2})+\sqrt{T_{1}T_{2}/T}\frac{1}{\delta}\kappa_{\varepsilon}\}r\log^{3/2}n,
\]
and $C_{0}r^{3/2}\Vert\bm{w}\Vert_{\infty}/\Vert\bm{w}\Vert_{2}\leq\delta$,
where $\bm{w}$ is a vector defined as $\bm{w}:=[\bm{\Sigma}_{\varepsilon}^{1/2}(\bm{I}_{N}+\bm{U}\bm{U}^{\top})]_{\cdot,i}$;
for the parameter $\epsilon_{N,T}$ (cf. Lemma \ref{Lemma noise cov matrix estimate})
that captures the estimation error of the noise covariance matrix
$\bm{\Sigma}_{\varepsilon}$, assume that $\epsilon_{N,T}\leq c_{0}\delta/(\kappa_{\varepsilon}^{2}r\log n)$
and $\epsilon_{N,T}\leq c_{0}[\delta\Vert\bm{\Sigma}_{\varepsilon}\Vert_{2}/(s(\bm{\Sigma}_{\varepsilon})\Vert\bm{U}_{i,\cdot}\Vert_{2}\kappa_{\varepsilon}^{2}r\log n)]^{1/(1-q)}$.
Then, under the null hypothesis $H_{0}:\ \bm{b}_{i}^{1}=\bm{b}_{i}^{2}$,
we have that
\[
\left|\mathbb{P}\left(\mathscr{\widehat{B}}_{i}\leq\chi_{1-\alpha}^{2}(r)\right)-(1-\alpha)\right|\leq\delta,
\]
where the test statistic $\mathscr{\widehat{B}}_{i}$ is given by
\[
\mathscr{\widehat{B}}_{i}:=\frac{1}{T\widehat{\varphi}_{i}}(\widehat{\bm{b}}_{i}^{1}-\widehat{\bm{b}}_{i}^{2})^{\top}[\prod_{j=1}^{2}(\widehat{\bm{F}}^{j})^{\top}\widehat{\bm{F}}^{j}](\widehat{\bm{b}}_{i}^{1}-\widehat{\bm{b}}_{i}^{2}),
\]
with $\widehat{\varphi}_{i}:=(\widehat{\bm{\Sigma}}_{\varepsilon}^{\tau}+\widehat{\bm{U}}\widehat{\bm{U}}^{\top}\widehat{\bm{\Sigma}}_{\varepsilon}^{\tau}+\widehat{\bm{\Sigma}}_{\varepsilon}^{\tau}\widehat{\bm{U}}\widehat{\bm{U}}^{\top}+\widehat{\bm{U}}\widehat{\bm{U}}^{\top}\widehat{\bm{\Sigma}}_{\varepsilon}^{\tau}\widehat{\bm{U}}\widehat{\bm{U}}^{\top})_{i,i}$.
\end{theorem}\begin{remark}If all the entries of the noise matrix
$\bm{Z}$ in (\ref{noise matrix formula}) are Gaussian, then we do
not need the assumption that $C_{0}r^{3/2}\Vert\bm{w}\Vert_{\infty}/\Vert\bm{w}\Vert_{2}\leq\delta$,
where $\bm{w}=[\bm{\Sigma}_{\varepsilon}^{1/2}(\bm{I}_{N}+\bm{U}\bm{U}^{\top})]_{\cdot,i}$.\end{remark}
To illustrate our assumptions, we consider the following setting similar
to that discussed after Theorem \ref{Thm UV 1st approx row-wise error}:~$r\asymp1$,
$T_{1}\asymp T\asymp N$, $T_{1}\gg T_{2}$, $\bm{\Sigma}_{\varepsilon}$
is well-conditioned and its sparsity parameter satisfies $s(\bm{\Sigma}_{\varepsilon})\lesssim\Vert\bm{\Sigma}_{\varepsilon}\Vert_{2}$,
and the column subspace of $\bm{B}$ is incoherent. Then, fulfilling
our assumptions merely requires that, $\theta\gg\log^{3/2}n$ and
$\vartheta_{i}\gg T^{-1/2}\log n$ for the SNR, $T\gg\log^{3}n$ for
the sample size, and $\epsilon_{N,T}\ll\min(\log^{-1}n,(N^{1/2}/\log n)^{1/(1-q)})$
for the estimation error of $\bm{\Sigma}_{\varepsilon}$. In particular,
our SNR conditions are less restrictive than prior work \citep[e.g.,][]{breitung2011beta_break}
that needed the pervasiveness assumption, which required the SNR to
grow as $\theta\gg N^{1/4}$. Our results adapt to the weak factor
model with the cross-sectional correlations, and to the best of our
knowledge, no prior work has developed the test statistics for the
structural break test under this case.
\subsection{The two-sample test for betas}
In finance and economics, besides the latent common factors that drive
the co-movements of asset returns, the factor loadings, also known
as betas, are important as well, which measure the sensitivity of
asset return to the movements of the factors. Consider the example
that the panel data $\bm{X}$ is the stock return, where the $i$-th
row of $\bm{X}$ represents the time series data of the $i$-th stock.
Then, the factor loadings $\bm{B}$ assess the risk exposure of these
stocks to the latent common factors $\bm{f}_{t}$, with the $i$-th
row $\bm{b}_{i}$ being the $i$-th stock's beta.
For any distinct $i$ and $j$, we aim to test the hypothesis that
$\bm{b}_{i}=\bm{b}_{j}$, i.e., if the $i$-th and the $j$-th stocks
have the same risk exposure on the common factors. It is a statistical
approach to evaluate the similarity in risk structure between two
stocks. Our test statistic is similar to the idea of two-sample test:
we show that $\widehat{\bm{B}}_{i,\cdot}-\widehat{\bm{B}}_{j,\cdot}$
is approximately Gaussian and then derive a Chi-square test statistic.
In Theorem \ref{Thm factor test plug-in Chi-sq} below, we construct
the test statistic and show its validity.
\begin{theorem}\label{Thm two-sample test B}Suppose that the assumptions
in Theorem \ref{Thm UV 1st approx row-wise error} hold. Assume that
$s(\bm{\Sigma}_{\varepsilon})/\lambda_{\min}(\bm{\Sigma}_{\varepsilon})\leq C_{\varepsilon}$
for some universal constant $C_{\varepsilon}$, and the covariance
estimator $\widehat{\bm{\bm{\Sigma}}}_{\varepsilon}^{\tau}$ satisfies
the conditions as in Lemma \ref{Lemma noise cov matrix estimate}.
For any given target error level $\delta>0$, assume there exists
some universal constants $c_{0},C_{0}>0$ such that $T\geq C_{0}r^{2}\delta^{-2}\log n$,
\begin{subequations}
\begin{equation}
\max(\Vert\bar{\bm{U}}_{i,\cdot}\Vert_{2},\Vert\bar{\bm{U}}_{j,\cdot}\Vert_{2})\leq c_{0}\delta(\frac{n}{\theta\sqrt{T}}+1)^{-1}\kappa_{\varepsilon}^{-1/2}r^{1/2}\log^{-1}n,\label{B two-sample test SNR 1}
\end{equation}
and
\begin{equation}
\min(\vartheta_{i},\vartheta_{j})\geq C_{0}\delta^{-1}\sqrt{\kappa_{\varepsilon}}\frac{\theta n}{\vartheta\sqrt{T}}(\big\Vert\bar{\bm{U}}\big\Vert_{2,\infty}+\frac{1}{\sqrt{n}})\sqrt{r}\log^{3/2}n;\label{B two-sample test SNR 2}
\end{equation}
\end{subequations}
for the parameter $\epsilon_{N,T}$ (cf. Lemma \ref{Lemma noise cov matrix estimate})
that captures the estimation error of the noise covariance matrix
$\bm{\Sigma}_{\varepsilon}$, assume that $\epsilon_{N,T}\leq c_{0}(\delta r^{3}\log^{2}n)^{1/(1-q)}$.
Then, for any $i,j\in\{1,2,\ldots,N\}$ satisfying $i\neq j$, under
the null hypothesis $H_{0}:\ \bm{b}_{i}=\bm{b}_{j}$, we have that
\[
\left|\mathbb{P}\left(\mathfrak{\widehat{\mathcal{B}}}_{ij}\leq\chi_{1-\alpha}^{2}(r)\right)-(1-\alpha)\right|\leq\delta,
\]
where the test statistic $\mathfrak{\widehat{\mathcal{B}}}_{ij}$
is given by
\[
\mathfrak{\widehat{\mathcal{B}}}_{ij}:=T((\widehat{\bm{\bm{\Sigma}}}_{\varepsilon}^{\tau})_{i,i}+(\widehat{\bm{\bm{\Sigma}}}_{\varepsilon}^{\tau})_{j,j}-2(\widehat{\bm{\bm{\Sigma}}}_{\varepsilon}^{\tau})_{i,j})^{-1}\big\Vert\widehat{\bm{B}}_{i,\cdot}-\widehat{\bm{B}}_{j,\cdot}\big\Vert_{2}^{2}.
\]
\end{theorem}
\begin{remark}If all the entries of the noise matrix $\bm{Z}$ in
(\ref{noise matrix formula}) are Gaussian, then we do not need the
assumption that $T\geq C_{0}r^{2}\delta^{-2}\log n$ on the sample
size $T$.\end{remark}
To illustrate the SNR conditions (\ref{B two-sample test SNR 1})--(\ref{B two-sample test SNR 2}),
we consider the setting discussed after Theorem \ref{Thm UV 1st approx row-wise error},
i.e.,~$r\asymp1$, $N\asymp T$, $\bm{\Sigma}_{\varepsilon}$ is
well-conditioned, and the column subspace of $\bm{B}$ is incoherent.
Then, fulfilling our assumptions (\ref{B two-sample test SNR 1})--(\ref{B two-sample test SNR 2})
merely requires that $\theta\gg\log n$ and $\theta^{-1}\vartheta\min(\vartheta_{i},\vartheta_{j})\gg\log^{3/2}n$.
These SNR conditions are less restrictive than the prior work \citep[e.g.,][]{BaiNg2023PCA,Jiang2023PCA,ChoiMing2024PCA}
that required the SNR to grow with a polynomial rate of $n$.
To the best of our knowledge, no prior work has developed the test
statistics for this two-sample test of betas. Our results are applicable
to the weak factor model with the cross-sectional correlations. The
subspace perturbation bounds in Theorem \ref{Thm UV 1st approx row-wise error}
pave the way for us to statistically assess the similarity for arbitrary
two rows of factor loadings $\bm{B}$.
\subsection{Statistical inference for the systematic risks}
In the factor model $x_{i,t}=\bm{b}_{i}^{\top}\bm{f}_{t}+\varepsilon_{i,t}$,
the risk associated with stock return $x_{i,t}$ is decomposed into
two parts \citep[e.g.,][]{Bai2003ECTA} --- systematic risk from
the common component $\bm{b}_{i}^{\top}\bm{f}_{t}$ and idiosyncratic
risk from the noise $\varepsilon_{i,t}$. Systematic risk, often referred
to as market risk, is integral in financial economics as it represents
the inherent risk affecting the entire market or market segment. This
risk, driven by broader economic forces such as inflation, political
events, and changes in interest rates, comes from the risk factors
$\bm{f}_{t}$ that explain the systematic co-movements and impacts
all the stocks.
A standard metric of the systematic risk is the variance of $\bm{b}_{i}^{\top}\bm{f}_{t}$,
which is given by $\operatorname{{\rm Var}}(\bm{b}_{i}^{\top}\bm{f}_{t})=
\Vert\bm{b}_{i}\Vert_{2}^{2}$. Our focus is on constructing the confidence interval (CI) for systematic
risk $\Vert\bm{b}_{i}\Vert_{2}^{2}$. Similar to the idea we conduct
the inference for beta $\bm{b}_{i}$ in Corollary \ref{Thm B inference}
where the estimator of $\bm{b}_{i}$ is $\widehat{\bm{B}}_{i,\cdot}$,
we show that $\Vert\widehat{\bm{B}}_{i,\cdot}\Vert_{2}^{2}-\Vert\bm{b}_{i}\Vert_{2}^{2}$
is approximately Gaussian and then construct the CI for $\Vert\bm{b}_{i}\Vert_{2}^{2}$.
We present the CI and its validity in Theorem \ref{Thm B row norm}.
\begin{theorem}\label{Thm B row norm}Suppose that the assumptions
in Theorem \ref{Thm UV 1st approx row-wise error} hold. Assume that
covariance estimator $\widehat{\bm{\bm{\Sigma}}}_{\varepsilon}^{\tau}$
satisfies the conditions as in Lemma \ref{Lemma noise cov matrix estimate}.
For any given target error level $\delta>0$, assume there exist some
universal constants $c_{0},C_{0},C_{U}>0$ such that $T\geq C_{0}\delta^{-2}\log n$,
$\Vert\bar{\bm{U}}\Vert_{2,\infty}\leq C_{U}\theta/\vartheta$, $\Vert\bar{\bm{U}}\Vert_{2,\infty}\leq C_{U}\sqrt{T}\vartheta/n$,
$\Vert\bar{\bm{U}}\Vert_{2,\infty}\leq C_{U}\sqrt{T}\theta^{2}/(\vartheta n)$,
$\Vert\bm{b}_{i}\Vert_{2}^{2}/(\bm{\Sigma}_{\varepsilon})_{i,i}\leq c_{0}\delta^{2}r^{-1}\log^{-1}n$,
\begin{subequations}
\begin{equation}
\vartheta_{i}\geq C_{0}\delta^{-1}\frac{\sqrt{\kappa_{\varepsilon}}\theta}{\vartheta_{i}\sqrt{T}}\frac{1}{\Vert\bar{\bm{U}}_{i,\cdot}\Vert_{2}}r\log^{2}n,\qquad\vartheta\geq C_{0}\delta^{-1}\frac{\sqrt{\kappa_{\varepsilon}}\theta}{\vartheta_{i}\sqrt{T}}(\sqrt{n}+n\Vert\bar{\bm{U}}\Vert_{2,\infty})r\log^{2}n\label{B rownorm CI SNR 1}
\end{equation}
\begin{equation}
\vartheta_{i}\geq C_{0}\delta^{-1}\sqrt{\kappa_{\varepsilon}}(\frac{1}{\sqrt{T}}+\frac{n}{\theta T})(1+\frac{1}{\Vert\bar{\bm{U}}_{i,\cdot}\Vert_{2}})r\log n,\qquad\text{and}\qquad\Vert\bar{\bm{U}}_{i,\cdot}\Vert_{2}\leq c_{0}\delta\kappa_{\varepsilon}^{-1/2}(1+\frac{n}{\theta\sqrt{T}})^{-1};\label{B rownorm CI SNR 2}
\end{equation}
\end{subequations}
for the parameter $\epsilon_{N,T}$ (cf. Lemma \ref{Lemma noise cov matrix estimate})
that captures the estimation error of the noise covariance matrix
$\bm{\Sigma}_{\varepsilon}$, assume that $\epsilon_{N,T}\leq c_{0}\delta r^{-1}\log^{-1/2}n$.
Then we have that, for any $1\leq i\leq N$,
\[
\left\vert \mathbb{P}(\Vert\bm{b}_{i}\Vert_{2}^{2}\in\text{CI}_{i}^{\bm{B},(1-\alpha)})-(1-\alpha)\right\vert \leq\delta,
\]
where the confidence interval is constructed as
\[
\text{CI}_{i}^{\bm{B},(1-\alpha)}:=[\big\Vert\widehat{\bm{B}}_{i,\cdot}\big\Vert_{2}^{2}-\widehat{\sigma}_{B,i}z_{1-\frac{1}{2}\alpha},\big\Vert\widehat{\bm{B}}_{i,\cdot}\big\Vert_{2}^{2}+\widehat{\sigma}_{B,i}z_{1-\frac{1}{2}\alpha}],\qquad\text{with}\qquad\widehat{\sigma}_{B,i}:=\frac{2}{\sqrt{T}}\sqrt{(\widehat{\bm{\bm{\Sigma}}}_{\varepsilon}^{\tau})_{i,i}}\big\Vert\widehat{\bm{B}}_{i,\cdot}\big\Vert_{2},
\]
and $z_{p}$ is the $p$-quantile of the standard Gaussian $\mathcal{N}(0,1)$.
\end{theorem}
\begin{remark}If all the entries of the noise matrix $\bm{Z}$ in
(\ref{noise matrix formula}) are Gaussian, then we do not need the
assumption that $T\geq C_{0}\delta^{-2}\log n$ on the sample size
$T$.\end{remark}
We now explain the assumptions under the setting discussed after Theorem
\ref{Thm two-sample test B}. In this case, to make (\ref{B rownorm CI SNR 1})
hold, it suffices to assume that $\theta^{-1}\vartheta_{i}\vartheta\gg\log^{2}n$
and $\theta^{-1}\vartheta_{i}^{2}\gg\log^{2}n$; to make (\ref{B rownorm CI SNR 2})
hold, it suffices to assume that $\theta\gg1$, $\vartheta_{i}\theta\gg\log n$,
and $\vartheta_{i}\gg T^{-1/2}\log n$. All these conditions require
only polynomial growth rates of $\log n$ for the SNRs $\vartheta_{i}$,
$\vartheta$, and $\theta$. The assumption on $\Vert\bar{\bm{U}}\Vert_{2,\infty}$
is the same with that we required in Corollary \ref{corollary:error bound B F}
to prove the estimation error for factor loading.
On one hand, as shown in Theorem \ref{Thm B row norm}, the CI width
$\widehat{\sigma}_{B,i}$ is proportional to the square root of the
product of noise level estimator $(\widehat{\bm{\bm{\Sigma}}}_{\varepsilon}^{\tau})_{i,i}$
and systematic risk estimator $\Vert\widehat{\bm{B}}_{i,\cdot}\Vert_{2}^{2}$.
On the other hand, in the proof of Theorem \ref{Thm B row norm},
we show that the bias of systematic risk estimator $\Vert\widehat{\bm{B}}_{i,\cdot}\Vert_{2}^{2}$
is proportional to $\Vert\bm{b}_{i}\Vert_{2}^{2}$. Thus, the assumption
that $\Vert\bm{b}_{i}\Vert_{2}^{2}/(\bm{\Sigma}_{\varepsilon})_{i,i}\leq c_{0}\delta^{2}r^{-1}\log^{-1}n$
is essential because the validity of inference hinges on the dominance
of the CI width over the bias. Consequently, we assume that the ratio
between systematic risk $\Vert\bm{b}_{i}\Vert_{2}^{2}$ and noise
level $(\bm{\Sigma}_{\varepsilon})_{i,i}$ cannot be too large. The
row-wise subspace perturbation bounds in Theorem \ref{Thm UV 1st approx row-wise error}
facilitate us to conduct statistical inference for the systematic
risk of any given stock in the panel data.
\section{Numerical experiments\label{sec:Numerical-experiments}}
In this section, we conduct Monte Carlo simulations to demonstrate
our inferential theories for the PC estimators in the weak factor
models. Additionally, our empirical studies reveal that the testing
results based on our test statistics surprisingly align with the economic
cycles and financial crisis periods.
\subsection{Monte Carlo simulations}
To make our simulations similar to the real applications, we use the
standard Fama-French three-factor model:
\[
x_{i,t}=\bm{b}_{i}^{\top}\bm{f}_{t}+\varepsilon_{i,t},\qquad1\leq i\leq N,1\leq t\leq T,
\]
where the dimension of the latent factor $\bm{f}_{t}$
is set to $r=3$. The idiosyncratic noise $\varepsilon_{i,t}$ exhibits
cross-sectional correlations, and the noise covariance matrix $\bm{\Sigma}_{\varepsilon}$
is sparse.
We generate the factor loadings $\{\bm{b}_{i}\}_{i=1}^{N}$, the factors
$\{\bm{f}_{t}\}_{t=1}^{T}$, and the noise terms $\{\bm{\varepsilon}_{t}\}_{t=1}^{T}$
independently from $\mathcal{N}(0,\bm{\Sigma}_{b})$, $\mathcal{N}(0,s_{f}\bm{\Sigma}_{f})$,
and $\mathcal{N}(0,\bm{\Sigma}_{\varepsilon})$, respectively, where
$\bm{\varepsilon}_{t}=(\varepsilon_{1,t},\ldots,\varepsilon_{N,t})^{\top}$.
Both $\bm{\Sigma}_{b}$ and $\bm{\Sigma}_{f}$ are set to $r\times r$
identity matrices. We generate $\bm{\Sigma}_{\varepsilon}$ as a block-diagonal
matrix $\bm{\Sigma}_{\varepsilon}=\text{\ensuremath{\mathsf{diag}}}(\bm{A}_{1},\bm{A}_{2},\ldots,\bm{A}_{J})$,
where the number of blocks is set to $J=20$. We set $N=300$ and
$T=200$. For each $\bm{A}_{i}$, we construct it as an equi-correlation
matrix $\bm{A}_{i}=(1-\rho_{i})\bm{I}_{m}+\rho_{i}\bm{1}_{m}\bm{1}_{m}^{\top}$,
where the block size $m$ is set to $m=N/J=15$, $\bm{1}_{m}$ is
the $m$-dimensional vector of ones, and $\rho_{i}$ is drawn from
a uniform distribution on $[0,0.5]$. To validate our test statistic
for two-sample test of betas, we set $\bm{b}_{2}$ equal to $\bm{b}_{1}$
after generating the loading matrix $\bm{B}$. This slight modification
allows us to examine our test statistics under the null hypothesis
$\bm{b}_{1}=\bm{b}_{2}$. In our simulation results, we report the
values of $\theta=\sigma_{r}/\Vert\bm{\Sigma}_{\varepsilon}^{1/2}\Vert_{2}$
to reflect different levels of SNR.
First, we demonstrate the practical validity of the confidence regions
constructed using Corollaries \ref{Thm F inference}--\ref{Thm B inference}
and Theorem \ref{Thm B row norm}, for the factors, betas (i.e., factor
loadings), and systematic risks, respectively.
\begin{itemize}
\item For factors (resp. betas), we construct 95\% confidence regions by
substituting the asymptotic covariance matrix $\bm{\Sigma}_{V,t}$
(resp. $\bm{\Sigma}_{B,i}$) with their consistent estimators as commented
after Corollary \ref{Thm B inference}. We define $\widehat{\mathsf{Cov}}_{F}(t)$
(resp. $\widehat{\mathsf{Cov}}_{B}(i)$) as the empirical probability
that the constructed confidence region covers $\bm{f}_{t}^{\top}\bm{R}_{F}$
(resp. $\bm{b}_{i}^{\top}\bm{R}_{B}$) over 200 Monte Carlo trials,
where $\bm{R}_{F}$ (resp. $\bm{R}_{B}$) is the rotation matrix defined
in Corollary \ref{Thm F inference} (resp. Corollary \ref{Thm B inference}).
\item For systematic risks, similarly, we define $\widehat{\mathsf{Cov}}_{\ell_{2}}(i)$
as the empirical probability that the 95\% confidence interval $\text{CI}_{i}^{\bm{b},0.95}$
constructed via Theorem \ref{Thm B row norm} covers $\Vert\bm{b}_{i}\Vert_{2}^{2}$
over 200 Monte Carlo trials.
\end{itemize}
Finally, we compute the mean and standard deviation for $\{\widehat{\mathsf{Cov}}_{F}(t)\}_{t=1}^{T}$,
$\{\widehat{\mathsf{Cov}}_{B}(i)\}_{i=1}^{N}$, and $\{\widehat{\mathsf{Cov}}_{\ell_{2}}(i)\}_{i=1}^{N}$,
then present these as $\mathsf{Mean}(\widehat{\mathsf{Cov}})$ and
$\mathsf{Std}(\widehat{\mathsf{Cov}})$, with the results reported
in Table \ref{table:CI cover}. As indicated in Table \ref{table:CI cover},
the coverage probabilities are close to 0.95 according to the mean
value, and exhibit stability across different rows of $\bm{F}$ and
$\bm{B}$ as evidenced by low standard deviation. As SNR $\theta$
goes down, the slight slippage of the coverage probabilities of the
systematic risks reconciles with our comments after Theorem \ref{Thm B row norm}
that, $\Vert\bm{b}_{i}\Vert_{2}^{2}/(\bm{\Sigma}_{\varepsilon})_{i,i}$,
which is close to SNR, reflects the ratio between the bias and the
CI width and cannot be too large. These favorable numerical results
persist even under low SNR, supporting our inferential theory under
the weak factor model.
\begin{table}[H]
\caption{Empirical coverage rates of 95\%-CI\label{table:CI cover}}
\centering
\begin{tabular}{c|c|c|c|c|c|c}
\hline
$(N,T)=(300,200)$ & \multicolumn{2}{c|}{Factor} & \multicolumn{2}{c|}{Beta} & \multicolumn{2}{c}{Systematic Risk}\tabularnewline
\hline
SNR $\theta$ & $\mathsf{Mean}(\widehat{\mathsf{Cov}})$ & $\mathsf{Std}(\widehat{\mathsf{Cov}})$ & $\mathsf{Mean}(\widehat{\mathsf{Cov}})$ & $\mathsf{Std}(\widehat{\mathsf{Cov}})$ & $\mathsf{Mean}(\widehat{\mathsf{Cov}})$ & $\mathsf{Std}(\widehat{\mathsf{Cov}})$\tabularnewline
\hline
$4.5$ & $0.9383$ & $0.0172$ & $0.9325$ & $0.0171$ & $0.9071$ & $0.0400$\tabularnewline
\hline
$3.5$ & $0.9298$ & $0.0190$ & $0.9264$ & $0.0184$ & $0.9192$ & $0.0323$\tabularnewline
\hline
$2.5$ & $0.9045$ & $0.0210$ & $0.9103$ & $0.0215$ & $0.9244$ & $0.0292$\tabularnewline
\hline
\end{tabular}
\end{table}
Next, we demonstrate the effectiveness of our test statistics in Theorems
\ref{Thm factor test plug-in Chi-sq} and \ref{Thm two-sample test B}
by showing their satisfactory size and power.
\begin{itemize}
\item For the null hypothesis $H_{0}:\text{ }\bm{v}\in col(\bm{F}_{S,\cdot})$
in the factor specification test of Theorem \ref{Thm factor test plug-in Chi-sq},
the time index subset $S$ is chosen as $S=\{\frac{1}{2}T+1,\frac{1}{2}T+2,\ldots,\frac{1}{2}T+L\}$
with $|S|=L=12$, and the observed factors $\bm{v}$ is set as
\[
\bm{v}=\bm{F}_{S,\cdot}\bm{w}+\delta\bm{g}.
\]
The vector $\bm{w}$ is set as $\bm{w}=(w_{1},w_{2},w_{3})^{\top}=(1,1,0.5)^{\top}$,
and the parameter $\delta$ controls the deviation of the alternatives
from the null distribution. The vector $\bm{g}$ is constructed as
follows: a vector $\bm{u}$ is drawn from the standard Gaussian $\mathcal{N}(0,\bm{I}_{|S|})$
and the projection residual $\bm{u}^{\perp}=(\bm{I}_{|S|}-\bm{P}_{\bm{F}_{S,\cdot}})\bm{u}$
is computed, where $\bm{P}_{\bm{F}_{S,\cdot}}$ is the projection
matrix on the column space of $\bm{F}_{S,\cdot}$. Then, we set $\bm{g}=2\bm{u}^{\perp}/\Vert\bm{u}^{\perp}\Vert_{2}\cdot\Vert\bm{F}_{S,\cdot}\Vert_{\mathrm{F}}\Vert\bm{w}\Vert_{2}$.
The formulation of $\bm{g}$ ensures that the signal strength from
the column space of $\bm{F}_{S,\cdot}$, which is captured by $\Vert\bm{F}_{S,\cdot}\Vert_{\mathrm{F}}\Vert\bm{w}\Vert_{2}$,
is balanced with that from the space orthogonal to $\bm{F}_{S,\cdot}$.
The null hypothesis $H_{0}$ should not be rejected when $\delta=0$,
and should be rejected when $\delta>0$.
\item For the null hypothesis $H_{0}:\ \bm{b}_{i}^{1}=\bm{b}_{i}^{2}$ in
the structural break test of betas in Theorem \ref{Thm beta structure test},
the cross-sectional unit is set as $i=1$, and the time subsets are
set as $\Gamma_{1}=\{1,2,\ldots,T_{1}\}$ and $\Gamma_{2}=\{T_{1}+1,T_{1}+2,\ldots,T_{1}+T_{2}\}$
with $T_{1}=T_{2}=T/2=100$. The beta $\bm{b}_{i}^{1}$ on period
$\Gamma_{1}$ is generated by the aforementioned procedure, while
the beta $\bm{b}_{i}^{2}$ on period $\Gamma_{2}$ is set as
\[
\bm{b}_{i}^{2}=\bm{b}_{i}^{1}+\Delta\Vert\bm{b}_{i}^{1}\Vert_{2}\mathbf{1},
\]
where $\mathbf{1}=(1,1,1)^{\top}$. The null hypothesis $H_{0}$ should
not be rejected when $\Delta=0$, and should be rejected when $\Delta>0$.
\item For the null hypothesis $H_{0}:\text{ }\bm{b}_{i}=\bm{b}_{j}$ in
the two-sample test of betas in Theorem \ref{Thm two-sample test B},
we set $i=1$ and study two cases for $j=2$ and $j=3$ respectively.
According to our simulation setup, the null hypothesis $H_{0}$ should
not be rejected when $j=2$, and should be rejected when $j=3$.
\end{itemize}
Tables \ref{table:factor tests}--\ref{table:beta tests} report
the empirical rejections rates at 5\% significance level over 200
Monte Carlo trials. Table \ref{table:factor tests} (resp. \ref{table:beta tests})
shows that for the test statistics in Theorem \ref{Thm factor test plug-in Chi-sq}
(resp. Theorems \ref{Thm beta structure test} and \ref{Thm two-sample test B}),
the results are favorable, exhibiting appropriate size and power,
even under a weak signal setup where the SNR $\theta$ is small.
\begin{table}[H]
\caption{Empirical rejection rates at level 5\% for the factor tests\label{table:factor tests}}
\centering
\begin{tabular}{c|c|c|c|c|c}
\hline
$(N,T)=(300,200)$ & \multicolumn{5}{c}{Factor specification test}\tabularnewline
\hline
SNR $\theta$ & $\delta=0$ & $\delta=0.25$ & $\delta=0.5$ & $\delta=0.75$ & $\delta=1$ \tabularnewline
\hline
$5.5$ & $0.000$ & $0.000$ & $0.745$ & $0.990$ & $1.000$ \tabularnewline
\hline
$5.0$ & $0.000$ & $0.000$ & $0.505$ & $0.945$ & $0.995$ \tabularnewline
\hline
$4.5$ & $0.000$ & $0.000$ & $0.250$ & $0.895$ & $0.980$ \tabularnewline
\hline
\end{tabular}
\end{table}
\begin{table}[H]
\caption{Empirical rejection rates at level 5\% for the beta tests\label{table:beta tests}}
\centering
\begin{tabular}{c|c|c|c|c|c|c|c}
\hline
$(N,T)=(300,200)$ & \multicolumn{5}{c|}{Structural break test} & \multicolumn{2}{c}{Two-sample test}\tabularnewline
\hline
SNR $\theta$ & $\Delta=0$ & $\Delta=0.25$ & $\Delta=0.5$ & $\Delta=0.75$ & $\Delta=1$ & $j=2$ & $j=3$\tabularnewline
\hline
$5.5$ & $0.030$ & $0.510$ & $0.990$ & $1.000$ & $1.000$ & $0.045$ & $1.000$\tabularnewline
\hline
$5.0$ & $0.030$ & $0.445$ & $0.970$ & $1.000$ & $1.000$ & $0.050$ & $1.000$\tabularnewline
\hline
$4.5$ & $0.035$ & $0.355$ & $0.945$ & $1.000$ & $1.000$ & $0.050$ & $1.000$\tabularnewline
\hline
\end{tabular}
\end{table}
\subsection{Empirical studies}
We analyze the monthly returns data of the S\&P 500 constituents from
the CRSP database for the period from January 1995 to March 2024.
We apply the factor specification test in Theorem \ref{Thm factor test plug-in Chi-sq}
and the structural break test for betas in Theorem \ref{Thm beta structure test}
to the stock returns.
First, we consider the factor specification test. The observed factors
we study are the Fama-French three factors: market (MKT), size (SMB),
and value (HML), denoted as $\bm{v}^{(1)}$, $\bm{v}^{(2)}$, and
$\bm{v}^{(3)}$, respectively. We obtain the time series data for
these factors from Kenneth French's website. We conduct our tests
via a rolling window approach, moving a 60-month window $[t+1,t+T]$
($T=60$) through the dataset. To mitigate the survival bias, we keep
the time series that have no more than 50\% missing values in each
window, i.e., the number of missing values is less than $T/2$, and
fill the missing values by the median of each time series. For each
window, the formed data matrix $\bm{X}^{t}$ is an $N\times T$ matrix,
where the cross-sectional dimension $N$ varies with $t$. We assume
that $\bm{X}^{t}$ satisfies the factor model as in (\ref{factor model matrix form}),
i.e.,
\[
\bm{X}^{t}=\bm{B}^{t}(\bm{F}^{t})^{\top}+\bm{E}^{t}.
\]
Here, the superscript $t$ is added to each matrix to emphasize that
we apply the PCA method across different time windows $[t+1,t+T]$.
The number of factors is fixed as $r=3$. We conduct the factor specification
test in Theorem \ref{Thm factor test plug-in Chi-sq} to test the
null hypothesis
\[
\bm{v}_{S^{t},\cdot}^{(i)}=\bm{F}_{S^{t},\cdot}^{t}\bm{w}^{t,(i)},
\]
for the three factors, corresponding to $i=1,2,3$, respectively.
The time index subset $S^{t}$ is set as $S^{t}=[t+T-L+1,t+T]$ with
$|S^{t}|=L=12$.
The above procedure implies that, for each 12-month period $S^{t}=[t+T-L+1,t+T]$,
to test if the observed factors are in the column space of the latent
factors, we look back and utilize a broader historical data window
$[t+1,t+T]$ ($T=60$) to estimate the latent common factors $\bm{F}^{t}$.
Then we test the null hypothesis for this specific 12-month period
$S^{t}$, a subset at the end of the whole window $[t+1,t+T]$. Finally,
we plot in Figure \ref{fig:factor test} the test statistics against
the time index $t$ for each factor, underscoring the 95\% critical
value. As highlighted by \citet{FanLiaoYao2015power}, this rolling
window manner not only utilizes the up-to-date information in the
equity universe, but also alleviates the impacts of time-varying betas
and sampling biases.
\begin{figure}[H]
\centering
\begin{tabular}{c}
\includegraphics[scale=0.8]{figures/FF_3f_factor_test}\tabularnewline
\end{tabular}
\caption{The evolution of the test statistics\label{fig:factor test}}
\end{figure}
In Figure \ref{fig:factor test}, our findings indicate that during
the financial crisis of 2007-2009, the null hypothesis that the size
factor SMB lies in the latent factors' column space, is rejected at
95\% confidence level. This suggests a diminished importance and reduced
explanatory power of the size factor SMB for stock return data in
this period. Note that the sizes of stocks can still be important
or even more important than that in the normal period, but not necessarily
captured by the variable SMB. Similar interpretations hold for the
value factor HML around 2009. During the COVID period around 2019,
both the size factor SMB and the value factor HML exhibit a loss in
explanatory power. The spike for the size factor SMB during 2022 is
probably due to the war between Russia and Ukraine that started in
February 2022, which is an unexpected economic shock to the stock
market. Notably, the market portfolio maintains its explanatory strength
throughout, indicating its stability and resilience as an explanatory
variable, even during distinct economic cycles.
Next, we consider the structural break test for betas for individual stocks. We test whether
the betas have changed before and after the three economic recessions
covered by the time horizon of our data -- the Early 2000s Recession
(Mar. 2001--Nov. 2001) due to the dot com bubble, the 2008 Great Recession due to housing bubble and
financial crisis (Dec. 2007--Jun. 2009), and the COVID-19 Recession (Feb. 2020--Apr. 2020). Here, the
start and the end of each recession are according to the NBER's Business
Cycle Dating Committee.
For each recession period $[t_{\text{start}},t_{\text{end}}]$, we
first take the data $\bm{X}^{1}$ and $\bm{X}^{2}$ lying in the time
window $[t_{\text{start}}-T_{1},t_{\text{start}}-1]$ and $[t_{\text{end}}+1,t_{\text{end}}+T_{2}]$,
respectively, where $T_{1}=T_{2}=60$. Next, we merge the two panels
to get $\bm{X}=(\bm{X}^{1},\bm{X}^{2})$, and fill the missing values
in the same manner as in the factor specification test. We assume
that $\bm{X}$ satisfies the factor model as follows
\begin{align*}
x_{i,t} & =\bm{f}_{t}^{\top}\bm{b}_{i}^{1}+\varepsilon_{i,t}\qquad\text{for}\qquad t\in\{t_{\text{start}}-T_{1},t_{\text{start}}-T_{1}+1,\ldots,t_{\text{start}}-1\},\\
x_{i,t} & =\bm{f}_{t}^{\top}\bm{b}_{i}^{2}+\varepsilon_{i,t}\qquad\text{for}\qquad t\in\{t_{\text{end}}+1,t_{\text{end}}+2,\ldots,t_{\text{end}}+T_{2}\}.
\end{align*}
The number of factors is chosen by a scree plot of the eigenvalues
of $\bm{X}$, which is $4$, $3$, and $3$ for the three recessions,
respectively. We test the hypothesis $H_{0}:\ \bm{b}_{i}^{1}=\bm{b}_{i}^{2}\ \leftrightarrow\ H_{1}:\ \bm{b}_{i}^{1}\neq\bm{b}_{i}^{2}$
for each cross-sectional unit $i$ using the test statistic in Theorem
\ref{Thm beta structure test}. To report the test results, we group
the stocks into 11 sectors by Global Industrial Classification Standard
(GICS), and then count the numbers of stocks in each sector that reject
the null at the 95\% confidence level. The test results for the three
recessions are illustrated in Figures \ref{fig:beta test 1}, \ref{fig:beta test 2},
and \ref{fig:beta test 3}, respectively.
\begin{figure}[H]
\begin{tabular}{l}
\qquad{}\includegraphics[scale=0.65]{figures/beta_test_2001-03}\tabularnewline
\end{tabular}
\caption{Test for breaks in betas for Early 2000s Recession (total number of
stocks is in bracket)\label{fig:beta test 1}}
\end{figure}
In Figure \ref{fig:beta test 1}, we observe that the sectors with
the highest number of stocks experiencing structural breaks in betas
are Financials, Information Technology, Industrials, and Utilities.
The impact on the Information Technology sector is likely due to the
dot-com bubble burst, which was one of the triggers of the Early 2000s
Recession. Another significant cause of this recession was the 9/11
attacks. Such severe economic shocks are possible reasons why typically
stable sectors like Industrials and Utilities were also affected.
\begin{figure}[H]
\begin{tabular}{l}
\qquad{}\includegraphics[scale=0.65]{figures/beta_test_2007-12}\tabularnewline
\end{tabular}
\caption{Test for breaks in betas for 2008 Great Recession (total number of
stocks is in bracket)\label{fig:beta test 2}}
\end{figure}
In Figure \ref{fig:beta test 2}, the 2007-2009 financial crisis,
marked by the subprime mortgage crisis and the collapse of the United
States housing bubble, affected many sectors. The financial sector
experienced a strong impact due to direct exposure to mortgage-backed
securities and other related financial instruments. The crisis led
to the failure or collapse of many of the United States' largest financial
institutions. Even though our analysis is inevitably influenced by
survival bias, as we can only analyze stocks that existed before and
after the crisis, we still observe that the financial sector had the
most affected stocks. The impact in Real Estate reflects the direct
consequences of the housing bubble burst. For sectors related to consumer
spending, such as Consumer Discretionary, Consumer Staples, and Health
Care, the shocks can be attributed to reduced consumer spending due
to increased unemployment and economic uncertainty.
\begin{figure}[H]
\begin{tabular}{l}
\qquad{}\includegraphics[scale=0.65]{figures/beta_test_2020-02}\tabularnewline
\end{tabular}
\caption{Test for breaks in betas for COVID-19 Recession (total number of stocks
is in bracket)\label{fig:beta test 3}}
\end{figure}
In Figure \ref{fig:beta test 3}, the economic effects of the pandemic
can be seen in many affected sectors, such as Real Estate, Industrials,
Health Care, and Financials. The significant changes in Real Estate
likely reflect the uncertainties brought about by lockdowns and health
crises, resulting in severe fluctuations in property values and rent
payments. The impact on the Health Care sector may be tied to the
heightened demand for medical services and supplies, alongside volatility
in biotechnology investments. Notably, Utilities and Energy showed
smaller changes, suggesting relative stability in these sectors despite
overall market volatility. During the 2008 Great Recession, these
sectors also demonstrated similar stability. This resilience could
be attributed to the essential nature of services provided by these
sectors, making them less susceptible to economic disruptions.
\section{Other related works\label{sec:related work}}
The factor model is an important topic in finance and economics. The
early econometric studies on factor model can be dated back to \citet{forni2000factor,StockWatson2002PCA}.
Most previous works on factor model assumes that all factors are strong
or pervasive, that is, the SNR grows at rate $\sqrt{N}$ \citep[e.g.,][]{BaiNg2002ECTA,Bai2003ECTA,POET2013}.
When the SNR grows at a rate slower than $\sqrt{N}$, the model is
often called the weak factor model, which has been a popular research
topic in recent years. The method for determining the number of factors
under the weak factor model has been studied by a few papers \citep[e.g.,][]{Onatski2009number,Onatski2010number,frey2022weakFactorNumber}.
Several recent works have pursued the estimation and inference in
the weak factor model. To name a few examples, \citet{BaiNg2023PCA,Jiang2023PCA,ChoiMing2024PCA}
studied the consistency and asymptotic normality of PC estimators;
\citet{YoTa2022sWF_estimate,YoTa2022sWF_infer,wei2023sWF} studied
the sparsity-induced weak factor models where the low SNR is due to
the sparsity of the factor loading matrix; \citet{BKP2021weakFactor}
proposed an estimator of factor strength and established its theoretical
guarantee; \citet{Onatski2012PCAweak} showed that the PC estimators
are inconsistent in the extreme case (also known as the super-weak
factor model) where the SNR is $O(1)$.
The PCA is one of the most popular methods for the factor model, and
has been studied in many papers mentioned previously \citep[e.g.,][]{StockWatson2002PCA,BaiNg2002ECTA,Bai2003ECTA,Onatski2012PCAweak,POET2013,BaiNg2023PCA,Jiang2023PCA,ChoiMing2024PCA}.
Among the enormous literature, \citet{BaiNg2023PCA,Jiang2023PCA,ChoiMing2024PCA}
are the most recent and closest to our paper, since we all focus on
the PC estimators, especially on the inference side, under the weak
factor model. Besides the theoretical analysis, the variants of PCA
under the weak factor model have also been applied to empirical asset
pricing, macroeconomic forecasting, and many other important problems
in finance \citep[e.g.,][]{gigXiu2021testWeak,GiglioXiu2023weakFactorPredict}.
The estimation of factor model is closely related to the low-rank
matrix denoising in the statistical machine learning community. In
recent years, studying the factor model from the view of low-rank
matrix denoising has provided lots of exciting findings and understandings
(see \citet{FanLiLiao2021ARFEfactor,yan2021inference} for comprehensive
reviews), and our theory is partly inspired by this low-rank structure
of factor model.
From a technical perspective, our estimation procedure is a spectral
method, and is related to previous studies of spectral methods on
PCA and subspace estimation \citep{abbe2020ell_p,cai2019subspace,chen2020bridging,yan2021inference,zhou2023deflated}.
Our analysis relies on the leave-one-out techniques that have found
wide applications in analyzing spectral estimators and nonconvex optimization
algorithms \citep{el2015impact,abbe2017entrywise,ma2017implicit,chen2019noisy,chen2020convex};
see the recent monograph \citet{chen2021Monograph} for more details.
In addition, the problem and analysis in this paper is also related
to past works on inference for other low-rank models \citep{chen2019inference,xia2021statistical,chernozhukov2021inference,choi2023inference,choi2024inference,yan2024entrywise}.
\section{Conclusions and discussions\label{sec:Discussion}}
In this paper, we establish a novel theory for PCA under the weak
factor model, offering significant advancements on the inference of
PCA over the existing literature. The weak factor model removes the
pervasiveness assumption that requires of SNR growing at a $\sqrt{N}$
rate, where $N$ is the cross-sectional dimension. Our theory covers
both estimation and inference of factors and factor loadings. Notably,
in the regime $N\asymp T$, where $T$ is the temporal dimension,
we show that the asymptotic normality of PC estimators holds as long
as the SNR grows faster than a polynomial rate of $\log N$, while
the previous work required a polynomial rate of $N$. The optimality
of our theory is in the sense that, the required growth rate of the
SNR for consistency aligns with that for asymptotic normality, differing
only by some logarithmic factors. Our theory paves the way to design
easy-to-implement test statistics for practical applications, e.g.,
factor specification test, structural break test for betas, and build
confidence regions for crucial model parameters, e.g., betas and systematic
risks. We validate our statistical methods through extensive Monte
Carlo simulations, and conduct empirical studies to find noteworthy
correlations between our test results and specific economic cycles.
While the current scope of our theory is considerably wide-ranging,
it is possible to further widen it and there are lots of topics that
are worth pursuing. For instance, how to extend the theory to the
setting where each observed variables have only fouth moment via robustification
of the covariance input \citep{fan2021robust}? How to generalize
our inferential theory to the case where the time-serial correlation
also appears in the idiosyncratic noise? If the panel data is missing
at random, are the PC estimators still valid under the weak factor
model, and how to conduct statistical inference in this case? What
benefit could our new theory bring to forecasting methods based on
the factor-augmented regression? Among many directions, these topics
can be investigated in future research.
\section*{Acknowledgements}
J.~Fan is supported in part by the NSF grants DMS-2053832 and DMS-2210833, and ONR grant N00014-22-1-2340. Y.~Yan is supported in part by the Norbert Wiener Postdoctoral Fellowship
from MIT.