EconBase
← Back to paper

A Unified Framework for Estimation of High-dimensional Conditional Factor Models

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

120,656 characters

A Unified Framework for Estimation of High-dimensional Conditional Factor Models



\begin{bibunit}
\def\spacingset#1{\renewcommand{\baselinestretch}
{#1}\small\normalsize} \spacingset{1.8}

\pdfbookmark[1]{Title}{title}

\if11
{
  \title{\bf A Unified Framework for Estimation of High-dimensional Conditional Factor Models}
  \author{Qihui Chen\\ School of Management and Economics\\ The Chinese University of Hong Kong, Shenzhen\\ [email removed]}
  \date{}
  \maketitle
} \fi

\if01
{
  \bigskip
  \bigskip
  \bigskip
\begin{center}
    {\LARGE\bf A Unified Framework for Estimation of High-dimensional Conditional Factor Models}
\end{center}
  \medskip
} \fi


\begin{abstract}
This paper presents a general framework for estimating high-dimensional conditional latent factor models via constrained nuclear norm regularization. We establish large sample properties of the estimators and provide efficient algorithms for their computation. To improve practical applicability, we propose a cross-validation procedure for selecting the regularization parameter. Our framework unifies the estimation of various conditional factor models, enabling the derivation of new asymptotic results while addressing limitations of existing methods, which are often model-specific or restrictive. Empirical analyses of the cross section of individual US stock returns suggest that imposing homogeneity improves the model's out-of-sample predictability, with our new method outperforming existing alternatives.
\end{abstract}


\noindent
{\it Keywords:} Constrained nuclear norm regularization, asset pricing, characteristics, macro state variables, factor zoo
\vfill

\newpage
\section{Introduction}\label{Sec1}
\setlength{\abovedisplayskip}{4pt}
\setlength{\belowdisplayskip}{4pt}
\setlength{\abovedisplayshortskip}{4pt}
\setlength{\belowdisplayshortskip}{4pt}

In empirical asset pricing, a central question revolves around understanding why different assets yield varying average returns. Conditional factor models offer a comprehensive framework for integrating conditional information to address this inquiry \citep{Gagliardinietal_Timevarying_2016, Gagliardinietal_EstimationConditionalFactor_2019}. This paper delves into the investigation of a high-dimensional conditional factor model defined as follows:
\begin{align}\label{Eqn: Model}
\hspace{-0.2cm}y_{it} = \alpha_{it} + \beta_{it}^{\prime}f_{t} + \varepsilon_{it} \hspace{-0.05cm}\text{ with }  \hspace{-0.05cm}\alpha_{it} = a_i^{\prime}x_{it} \hspace{-0.05cm}\text{ and } \hspace{-0.05cm}\beta_{it} = B_i^{\prime} x_{it}, i\hspace{-0.05cm}=\hspace{-0.05cm}1,\ldots,N,t\hspace{-0.05cm}=\hspace{-0.05cm}1,\ldots,T.
\end{align}
Here, $y_{it}$ denotes the excess return of asset $i$ in time period $t$, $f_{t}$ represents a $K\times 1$ vector of \textit{unobserved latent} factors, $\alpha_{it}$ characterizes a pricing error, $\beta_{it}$ denotes a $K\times 1$ vector of risk exposures, $\varepsilon_{it}$ stands for an error term, $x_{it}$ is a $p\times 1$ vector of pre-specified explanatory variables known at the beginning of time period $t$ (such as constants, sieve transformations of asset characteristics, sieve transformations of macro state variables, and their interactions), and $a_i$ and $B_i$ are $p\times1$ vector and $p\times K$ matrix of unknown coefficients, respectively. This model captures time-variation in the risk exposures (i.e., $B_i^{\prime} x_{it}$) and the pricing error (i.e., $a_i^{\prime} x_{it}$) through their associations with $x_{it}$, while also allowing for the distinction between ``risk'' and ``mispricing'' explanations regarding the role of $x_{it}$ in predicting returns, thereby contributing to resolving the ongoing ``characteristics versus covariance'' debate \citep{DanielTitman_Characteristics_1997}. Moreover, given that $K$ can be significantly smaller than $p$, the model facilitates the condensation of information from a large dimension of $x_{it}$ into a smaller number of factors, thereby mitigating the so-called ``factor zoo'' that proliferate in the literature \citep{Cochrane_Presidential_2011}. However, the estimation of the model encounters at least two challenges: i) $\{f_t\}_{t\leq T}$ are unknown and unobservable; ii) the dimension of the unknown parameters $\{a_i\}_{i\leq N}$, $\{B_i\}_{i\leq N}$, and $\{f_t\}_{t\leq T}$ is high.

The model nests various factor models in the literature. Unlike homogeneous versions of conditional factor models \citep{Parketal_FactorDynamics_2009,Kellyetal_Characteristics_2019,Chenetal_SeimiparametricFactor_2021}, our model allows for heterogeneity of $a_i$ and $B_i$ across assets. Consequently, our model nests classical factor models \citep{Ross_APT_1976,ChamberlainRothschild_FactorStuctures_1982} where $x_{it} =1$ and $a_i=0$; semiparametric factor models \citep{Connoretal_EfficientFFFactor_2012,Fanetal_ProjectedPCA_2016,Kimetal_Arbitrage_2019} where $x_{it}$ comprises a constant and sieve transformations of asset's time-invariant characteristics, with homogeneity of $a_i$ and $B_i$ across assets for non-constant explanatory variables; and state-varying factor models \citep{PelgerXiong_State-varying_2019} where $x_{it}$ encompasses a constant and sieve transformations of macro state variables, with $a_i=0$. Unlike \citet{Gagliardinietal_Timevarying_2016}, our model does not necessitate observable $f_t$ and accommodates the presence of arbitrage and large $p$, referred to as the unconstrained conditional factor model.

We provide a general framework for the estimation of high-dimensional conditional factor models. Specifically, we develop a nuclear norm regularized estimation of the model in \eqref{Eqn: Model} with \textit{constraints} on $\{a_{i},B_{i}\}_{i\leq N}$. The estimation procedure comprises two steps: first, estimating an $Np\times T$ reduced rank matrix composed of block matrices $\{a_i + B_{i}f_t\}_{i\leq N,t\leq T}$ using nuclear norm regularization under the constraints; then, extracting estimators of $K$, $\{a_i\}_{i\leq N}$, $\{B_i\}_{i\leq N}$, and $\{f_t\}_{t\leq T}$ from the estimated matrix using eigenvalue decomposition. We establish asymptotic properties of the estimators under a restricted strong convexity condition. Our framework allows for both $p\to\infty$ and $K\to\infty$ and may accommodate the presence of missing values, which are prevalent in stock return datasets.

The general framework enables the estimation of the aforementioned nested models in a unified manner, overcoming limitations of existing methods that are often model-specific or restrictive.\footnote{For example, \citet{ConnorKorajczyk_Performance_1986}, \citet{StockWatson_PCA_2002}, and \cite{BaiNg_NumberofFactors_2002} estimate the classical factor model by principal component analysis (PCA), while \citet{Fanetal_largecovariance_2013} use a principal orthogonal complement thresholding method. \citet{Fanetal_ProjectedPCA_2016} propose a projected-PCA for the semiparametric factor model based on linear sieves, while \citet{Fanetal_StructuralDeep_2022} employ neural networks. \citet{PelgerXiong_State-varying_2019} estimate the state-varying factor by a local version of PCA based on kernel smoothing. \citet{Chenetal_SeimiparametricFactor_2021} develop a regressed-PCA for the homogenous conditional factor model; \citet{Parketal_FactorDynamics_2009} propose a Newton-Raphson algorithm; \citet{Kellyetal_Characteristics_2019} propose
an alternating least squares procedure; \citet{Guetal_Autoencoder_2021} propose an autoencoder method. \citet{Gagliardinietal_Timevarying_2016} require observable factors for estimating a conditional factor model with no arbitrage.\label{foot1}} By tailoring the general theory for each model and providing simple primitive conditions, we make several contributions. First, we offer a novel estimation approach for the homogeneous conditional factor model, allowing $p$ to grow as fast as $N$. Second, we accommodate time-varying characteristics, nonzero pricing errors, and non-noisy intercepts in both pricing errors and risk exposures for the semiparametric factor model. Third, our estimator is capable of consistently estimating the factor space in the state-varying factor model. Fourth, to the best of our knowledge, our paper is the first to provide an estimation method for the unconstrained conditional factor model.


To enhance the practical applicability of our estimation procedure, we offer two contributions. Firstly, we present an efficient computing algorithm for finding the constrained nuclear norm regularized estimator of the reduced rank matrix in each model. This contribution is particularly valuable as constrained nuclear norm regularization involves high-dimensional constrained nonsmooth convex minimization, where computational efficiency is crucial for practical implementation. Secondly, we propose a cross-validation (CV) procedure to determine the optimal regularization parameter and validate its effectiveness through a series of Monte Carlo simulations. This contribution is essential because the choice of regularization parameter significantly impacts the estimates, and a systematic method for its selection is necessary to ensure robustness and reliability of the results. Our simulation studies demonstrate that the finite sample performance of our estimators, using the CV-chosen regularization parameter, is satisfactory and encouraging. Our simulations also demonstrate the superiority of our estimators compared to existing ones. We apply our unified framework to analyze the cross section of individual stock returns in the US market. Our analysis reveals that imposing homogeneity of $a_i$ and $B_i$ across assets enhances the model's out-of-sample predictability, with our method outperforming existing approaches.


Nuclear norm regularization has been extensively employed for estimating reduced-rank matrices in the statistical literature, primarily focusing on estimating the reduced-rank matrix itself. For instance, \citet{NegahbanWainwright_Scaling_2011} investigate \textit{unconstrained} nuclear norm regularized estimation of trace linear regression models under a restricted strong convexity condition; \citet{RohdeTsybakov_Lowrank_2011} examine the same problem under a restricted isometry condition; \citet{Fanetal_Generalized_2019} study generalized trace regression models. Our work diverges from these studies in several key aspects. Firstly, we require constrained nuclear norm regularization, which entails extending  existing methodologies to accommodate constraints. Secondly, our parameters of interest are $K$, $\{a_i\}_{i\leq N}$, $\{B_i\}_{i\leq N}$, and $\{f_t\}_{t\leq T}$, rather than the reduced-rank matrix. This distinction introduces an additional step in the estimation procedure to estimate these parameters from the reduced-rank matrix.

There have been several recent studies in the econometric literature that utilize unconstrained nuclear norm regularization. \citet{BaiNg_RankRegularized_2019} use it to enhance estimation of the classical factor model. \citet{MoonWeidner_Nuclear_2018} leverage it to improve estimation of panel data models with interactive fixed effects. \citet{Chernozhukovetal_LowRank_2018} employ it to estimate panel data models with heterogeneous coefficients. \citet{Atheyetal_CompletionCausal_2021} adopt this approach in treatment effect estimation. For more examples, see \citet{MoonWeidner_Nuclear_2018}. To the best of our knowledge, the use of constrained nuclear norm regularization in estimating conditional factor models has not been studied previously.

The literature on the cross section of asset returns is extensive; here we focus on conditional factor models. While our paper emphasizes models with latent factors, a substantial portion of empirical asset pricing research relies on pre-specified observable factors. These factors are often constructed using portfolio-sorting approaches, such as those outlined in \citet{FamaFrench_Commonrisk_1993}, based on asset characteristics.\footnote{Notable works in this area include studies by \citet{Shanken_Intertemporal_1990}, \citet{FersonHarvey_Variation_1991,FersonHarvey_Conditioning_1999}, \citet{LettauLudvigson_Consumption_2001}, and \citet{Gagliardinietal_Timevarying_2016}, among others. For a comprehensive review, see \citet{Gagliardinietal_EstimationConditionalFactor_2019}.} This approach encounters challenges related to the ``characteristics versus covariances'' debate and the ``factor zoo'' problem. We contribute to the literature by presenting a unified method for estimating conditional factor models without the need for pre-specified factors, which are well-suited for addressing the debate and problem \citep{Kellyetal_Characteristics_2019,Chenetal_SeimiparametricFactor_2021}.

The structure of the paper is outlined as follows. Section \ref{Sec2} presents several nested models. Section \ref{Sec3} outlines the general estimation framework. Section \ref{Sec4} establishes the asymptotic properties of the estimators. Section \ref{Sec5} tailors the general theory for each model. Section \ref{Sec6} presents simulation studies. Section \ref{Sec7} analyzes the cross section of individual US stock returns. Finally, Section \ref{Sec8} provides a brief conclusion. The Supplementary Appendix presents proofs of main results, computing algorithms, additional discussions, and additional simulations.

\section{Nested Models}\label{Sec2}
Our model in \eqref{Eqn: Model} nests many factor models in the literature.

\begin{ex}[Classical Factor Models]\label{Ex: Classical}
The arbitrage pricing theory by \citet{Ross_APT_1976} and \citet{ChamberlainRothschild_FactorStuctures_1982} gives rise to the following model:
\begin{align}\label{Eqn: Classical}
y_{it} = \lambda_i^{\prime}f_t + e_{it},
\end{align}
where $\lambda_i$ represents an unknown vector of risk exposures and $e_{it}$ is the idiosyncratic component. Our model encompasses \eqref{Eqn: Classical} where $x_{it} =1$, $a_{i} = 0$, $B_{i} = \lambda_{i}^{\prime}$, and $\varepsilon_{it} = e_{it}$.
\end{ex}

\begin{ex}[Semiparametric Factor Models]\label{Ex: SemiTin}
The model examined by \citet{Connoretal_EfficientFFFactor_2012}, \citet{Fanetal_ProjectedPCA_2016}, and \citet{Kimetal_Arbitrage_2019} is as follows:\footnote{\citet{Fanetal_ProjectedPCA_2016} assume that $\phi(\cdot)=0$ and $\mu_i=0$, \citet{Connoretal_EfficientFFFactor_2012} additionally assume that $\Phi(\cdot)$ are univariate functions and $\lambda_{i}=0$, and \citet{Kimetal_Arbitrage_2019}  assume that $\mu_i=0$ and $\lambda_i=0$. }
\begin{align}\label{Eqn: SemiTin}
y_{it} = \phi(z_i)+\mu_i + (\Phi(z_i) + \lambda_{i})^{\prime}f_t + e_{it},
\end{align}
where $z_i$ represents a vector of asset's time-invariant characteristics, $\phi(\cdot)$ and $\Phi(\cdot)$ are unknown functions, $\mu_i$ and $\lambda_{i}$ are unknown scalar and vector (intercepts in the pricing errors and risk exposures, which are usually interpreted as the components that are not explained by the characteristics), and $e_{it}$ is the idiosyncratic component. Using sieve methods, $\phi(z_i) = \phi^{\prime}h(z_i) +\delta(z_i)$ and $\Phi(z_i) = \Phi^{\prime}h(z_i) +\Delta(z_i)$, where $h(z_i)$ denotes a vector of basis functions of $z_i$ (excluding constants), $\phi$ and $\Phi$ represent unknown vector and matrix of coefficients, and $\delta(z_i)$ and $\Delta(z_i)$ are negligible sieve approximation errors. Our model nests \eqref{Eqn: SemiTin} where $x_{it} = (1,h(z_i)^{\prime})^{\prime}$, $a_i = (\mu_i,\phi^{\prime})^{\prime}$, $B_i = (\lambda_{i},\Phi^{\prime})^{\prime}$, and $\varepsilon_{it} = e_{it} + \delta(z_i) + \Delta(z_{i})^{\prime}f_t$. Thus, the rows of $a_i$ and $B_i$ corresponding to $h(z_i)$ are homogenous across $i$, meaning that the coefficients for non-constant explanatory variables are homogenous across assets.
\end{ex}

\begin{ex}[State-varying Factor Models]\label{Ex: Statevarying}
\citet{PelgerXiong_State-varying_2019} examine the following model:
\begin{align}\label{Eqn: Statevarying}
y_{it} = \Phi_i(z_{t})^{\prime}f_{t} + e_{it},
\end{align}
where $z_{t}$ represents a vector of constant and macro state variables known at the beginning of time period $t$, $\Phi_i(\cdot)$ is a vector of unknown functions, and $e_{it}$ is the idiosyncratic component. Employing sieve methods, $\Phi_i(z_t) = \Phi_i^{\prime}h(z_t) +\Delta_{i}(z_t)$, where $h(z_t)$ denotes a vector of basis functions of $z_t$ (which may include a constant), $\Phi_i$ is an unknown matrix of coefficients, and $\Delta_i(z_t)$ is a vector of negligible sieve approximation errors. Our model encompasses \eqref{Eqn: Statevarying} where $x_{it} = h(z_t)$, $a_i = 0$, $B_i = \Phi_i$, and $\varepsilon_{it} = e_{it} + \Delta_i(z_{t})^{\prime}f_t$.
\end{ex}

\begin{ex}[Homogeneous Conditional Factor Models]\label{Ex: HomoConditional}
\citet{Parketal_FactorDynamics_2009}, \citet{Kellyetal_Characteristics_2019}, and \citet{Chenetal_SeimiparametricFactor_2021} propose the following model:\footnote{\citet{Kellyetal_Characteristics_2019} assume that $\phi(\cdot)$ and $\Phi(\cdot)$ are linear functions.}
\begin{align}\label{Eqn: HomoConditional}
y_{it} = \phi_0(z_{it}) + \Phi_0(z_{it})^{\prime}f_{t} + e_{it},
\end{align}
where $z_{it}$ represents a vector of constant and asset characteristics known at the beginning of time period $t$, $\phi_0(\cdot)$ and $\Phi_0(\cdot)$ are unknown functions, and $e_{it}$ is the idiosyncratic component. Employing sieve methods, $\phi_0(z_{it}) = \phi_0^{\prime}h(z_{it}) +\delta(z_{it})$ and $\Phi_0(z_{it}) = \Phi_0^{\prime}h(z_{it}) +\Delta(z_{it})$, where $h(z_{it})$ denotes a vector of basis functions of $z_{it}$ (which may include a constant), $\phi_0$ and $\Phi_0$ are unknown vector and matrix of coefficients, and $\delta(z_{it})$ and $\Delta(z_{it})$ are negligible sieve approximation errors. Our model nests \eqref{Eqn: HomoConditional} where $x_{it} = h(z_{it})$, $a_i = \phi_0$, $B_i = \Phi_0$, and $\varepsilon_{it} = e_{it} + \delta(z_{it}) + \Delta(z_{it})^{\prime}f_t$. Thus, $a_i$ and $B_i$ are homogenous across $i$, meaning that the coefficients of explanatory variables are homogenous across assets.
\end{ex}

\begin{ex}[Unconstrained Conditional Factor Models]\label{Ex: ConitionalObs}
In the absence of arbitrage opportunities, \citet{Gagliardinietal_Timevarying_2016} propose the following model:
\begin{align}\label{Eqn: ConitionalObs}
y_{it} =z_{t}^{\prime}\Psi_i z_{t} + z_{it}^{\prime} \Upsilon_i z_{t} + z_{t}^{\prime}\Lambda_i f_t + z_{it}^{\prime}\Xi_i f_t + e_{it},
\end{align}
where $z_{t}$ represents a vector of constant and macro state variables known at the beginning of time period $t$, $z_{it}$ is a vector of asset characteristics known at the beginning of time period $t$, $\Psi_i$, $\Upsilon_i$, $\Lambda_i$, and $\Xi_i$ are unknown matrices of coefficients satisfying certain no arbitrage constraints, and $e_{it}$ is the idiosyncratic component. Our model encompasses \eqref{Eqn: ConitionalObs} without the no arbitrage constraints where $x_{it}$ consists of quadratic transformations of $z_t$ and $z_{it}$, $a_i$ and $B_i$ are functions of $\Phi_i$, $\Psi_i$, $\Upsilon_i$, and $\Lambda_i$, and $\varepsilon_{it} = e_{it}$. In contrast to their estimation method, which relies on observable $f_t$, our estimation procedure treats $f_t$ as latent factors and allows for the presence of arbitrage and large $p$.
\end{ex}

\section{Estimation Strategy}\label{Sec3}
For the convenience of the reader, we gather standard pieces of notation here, which will be utilized throughout the paper. We denote a $k\times k$ identity matrix as $I_k$. The Euclidean norm of a column vector $x$ is represented by $\|x\|$. For a symmetric matrix $A$, we denote its trace as $\mathrm{tr}(A)$, its smallest and largest eigenvalues as $\lambda_{\min}(A)$ and $\lambda_{\max}(A)$. The operator norm of a matrix $A$ is denoted by $\|A\|_2$, its Frobenius norm by  $\|A\|_{F}$, and its vectorization by $\mathrm{vec}(A)$. The Kronecker product of matrices $C$ and $D$ is denoted as $C \otimes D$. Unless specified, asymptotic statements in the paper shall be understood to hold as $N\to\infty$ with fixed $T$ or as $(N,T)\to\infty$, whenever appropriate.

We begin by reformulating the model in \eqref{Eqn: Model} using vectors/matrices. Let $a\equiv (a_1^{\prime},a_2^{\prime},\ldots, $ $a_{N}^{\prime})^{\prime}$ which is an $Np\times 1$ vector of unknown coefficients, $B \equiv (B_1^{\prime},B_2^{\prime},\ldots, B_{N}^{\prime})^{\prime}$ which is an $Np\times K$ matrix of unknown coefficients, and $F \equiv(f_1,f_2,\ldots, f_{T})^{\prime}$ which is a $T\times K$ matrix of latent factors. Let $\Pi$ be an $Np\times T$ unknown parameter matrix that collects the product of $(a_i, B_i)$ and $(1,f_t^{\prime})^{\prime}$, defined as
\begin{align}\label{Eqn: BigMatrixParameter}
\Pi\equiv \left(
            \begin{array}{c}
              (a_1, B_1) \\
              (a_2, B_2) \\
              \vdots \\
              (a_N, B_N) \\
            \end{array}
          \right) \left(
                    \begin{array}{cccc}
                      \left(
                        \begin{array}{c}
                          1 \\
                          f_1 \\
                        \end{array}
                      \right),
                       & \left(
                        \begin{array}{c}
                          1 \\
                          f_2 \\
                        \end{array}
                      \right), & \cdots, & \left(
                        \begin{array}{c}
                          1 \\
                          f_T \\
                        \end{array}
                      \right) \\
                    \end{array}
                  \right) \equiv a 1^{\prime}_{T} + BF^{\prime},
\end{align}
where $1_{T}$ is a $T\times 1$ vector of ones. Let $X_{it}\equiv (e_{N,i}\otimes x_{it}) e_{T,t}^{\prime}$ be an $Np\times T$ observed data matrix of $x_{it}$, where $e_{N,i}$ is the $i$th column of $I_N$ and $e_{T,t}$ is the $t$th column of $I_T$. Then $x_{it}^{\prime}a_i + x_{it}^{\prime}B_{i}f_{t} = \mathrm{tr}(X_{it}^{\prime}\Pi)$, so \eqref{Eqn: Model} can be succinctly expressed as
\begin{align}\label{Eqn: ModelRef}
y_{it} = \mathrm{tr}(X_{it}^{\prime}\Pi) + \varepsilon_{it}.
\end{align}
Since $\Pi$ has at most rank $K+1$, \eqref{Eqn: ModelRef} can be viewed as a trace linear regression model with reduced rank coefficient matrix $\Pi$ \citep{NegahbanWainwright_Scaling_2011, RohdeTsybakov_Lowrank_2011}. Thus, we first estimate $\Pi$ by using the nuclear norm regularization \citep{Fazel_MatrixRank_2002}, which employs the nuclear norm penalty as a surrogate function to enforce the reduced rank constraint. Our estimator of $\Pi$ is given by
\begin{align}\label{Eqn: NuclearNEst}
\hat{\Pi} = \operatornamewithlimits{arg\,min}_{\Gamma\in\mathcal{S}\subset\mathbf{R}^{Np\times T}}\frac{1}{2}\sum_{i=1}^{N}\sum_{t=1}^{T}(y_{it}-\mathrm{tr}(X_{it}^{\prime}\Gamma))^{2} +\lambda_{NT} \|\Gamma\|_{\ast},
\end{align}
where $\mathcal{S}\subset\mathbf{R}^{Np\times T}$ is convex, $\|\Gamma\|_{\ast}$ is the nuclear norm of $\Gamma$, and $\lambda_{NT}>0$ is a regularization parameter.\footnote{The nuclear norm of $\Gamma$ is $\|\Gamma\|_{\ast}=\sum_{j=1}^{\min\{Np,T\}}\sigma_{j}(\Gamma)$, corresponding to the sum of its singular values, where $\sigma_{j}(\Gamma)$'s are the singular values of $\Gamma$. The nuclear norm of $\Gamma$ is the convex envelope of the rank of $\Gamma$ over the set of matrices with spectral norm no greater than one; see, for example, \citet{Rechtetal_Guaranteed_2010}. } \textit{In particular, by introducing $\mathcal{S}$, which can be strictly smaller than $\mathbf{R}^{Np\times T}$, we can enforce the constraints of $\Pi$ induced by those of $a$ and $B$ \textemdash a critical aspect that has not been explored in the existing literature\textemdash enabling the estimation of various models within a unified framework.}  We set $\mathcal{S}=\mathbf{R}^{Np\times T}$ in Examples \ref{Ex: Classical}, \ref{Ex: Statevarying}, and \ref{Ex: ConitionalObs}, $\mathcal{S}=\mathcal{D}_{M}$ for $0<M<\infty$ (where $\mathcal{D}_{M}$ is given in \eqref{Eqn: Ex2: 1}) in Example \ref{Ex: SemiTin},  and $\mathcal{S} = \{1_{N}\otimes \Gamma: \Gamma\in \mathbf{R}^{p\times T}\}$ in Example \ref{Ex: HomoConditional}; see Section \ref{Sec5} for details. In the latter two cases,  $\mathcal{S}$ is strictly smaller than $\mathbf{R}^{Np\times T}$. Since \eqref{Eqn: NuclearNEst} involves constrained nonsmooth convex minimization, $\hat{\Pi}$ generally does not have an analytical closed form. Although several algorithms are available for solving convex minimization problems with a nuclear norm \citep{VandenbergheBoyd_Semidefinite_1996, Bertsekas_NonlinearProgramming_1999, LiuVandenberghe_Interior_2010, Maetal_Fixedpoint_2011}, they are not suitable for the high-dimensional settings with constraints in our context. In Appendix \ref{App: E}, we provide an efficient computing algorithm for each setting in Examples \ref{Ex: Classical}-\ref{Ex: ConitionalObs}.

We next proceed to derive estimators for $K$, $a$, $B$, and  $F$ from the nuclear norm regularized estimator $\hat{\Pi}$. Let $\hat{K}$, $\hat{a}$, $\hat{B}$, and $\hat{F}$ denote these estimators. Define $M_{T}\equiv I_{T} - 1_{T}1_{T}^{\prime}/T$. Since $\Pi M_{T} = BF^{\prime}M_{T}$, we can obtain $\hat{K}$ and $\hat{B}$ from the eigenvalues and eigenvectors of $\hat{\Pi} M_{T}\hat{\Pi}^{\prime}$. Specifically, $\hat{K}$ is given by
\begin{align}\label{Eqn: KEst}
\hat{K} = \sum_{j=1}^{Np}1\{\lambda_{j}(\hat{\Pi} M_{T}\hat{\Pi}^{\prime})\geq \delta_{NT}\},
\end{align}
where $\lambda_{j}(A)$ denotes the $j$th largest eigenvalue of $A$ and $\delta_{NT}>0$ is a threshold value. If $\hat{K} = 0$, $\hat{a} = {\hat{\Pi}1_{T}}/{T}$, $\hat{B} =0 $, and $\hat{F}=0$; otherwise we proceed as follows. To estimate $B$, we use the following normalization: $B^{\prime}B/N=I_K$ and $F^{\prime} M_{T} F/T$ being diagonal with diagonal entries in descending order. Then the columns of $\hat{B}/\sqrt{N}$ are given by the eigenvectors of $\hat{\Pi} M_{T}\hat{\Pi}^{\prime}$ corresponding to its largest $\hat{K}$ eigenvalues. To estimate $a$ and $F$, we impose the following condition: $a^{\prime} B = 0$. Since $a = (I_{Np}-B(B^{\prime}B)^{-1}B^{\prime})\Pi1_{T}/T$ and $F = \Pi^{\prime}B(B^{\prime}B)^{-1}$, we thus obtain
\begin{align}\label{Eqn: aFEst}
\hat{a} = \left(I_{Np}-\frac{\hat{B}\hat{B}^{\prime}}{N}\right)\frac{\hat{\Pi}1_{T}}{T} \text{ and } \hat{F} =  \frac{\hat{\Pi}^{\prime}\hat{B}}{N}.
\end{align}
It it noteworthy that in Examples \ref{Ex: SemiTin} and \ref{Ex: HomoConditional} there is no need to enforce the homogeneity restriction of $a$ and $B$ in extracting $\hat{a}$ and $\hat{B}$ from $\hat{\Pi}$ again to ensure the same homogeneity structure of $\hat{a}$ and $\hat{B}$; see Sections \ref{Sec52} and \ref{Sec53} for details.

Our estimation procedure is adaptable to accommodate the presence of missing values. In this case, the double summations in \eqref{Eqn: NuclearNEst} must be replaced with summations over non-missing data. This amounts to redefining the observations as $y_{it}m_{it}$ and $x_{it}m_{it}$, and the error term as $\varepsilon_{it}m_{it}$, where $m_{it} = 0$ when $y_{it}$ or $x_{it}$ are missing, and $1$ otherwise. Since $\hat{\Pi}$ accommodates missing values, we can employ a CV approach to choose the regularization parameter $\lambda_{NT}$ in \eqref{Eqn: NuclearNEst}. Specifically, we first randomly divide the observations into $L$ folds with observations indexed by $\{\mathcal{I}_{\ell}\}_{\ell\leq L}$, where $\mathcal{I}_{\ell}$ comprises observation indices in the $\ell$th fold, $\{\mathcal{I}_{\ell}\}_{\ell\leq L}$ are mutually exclusive, and $\cup_{\ell\leq L}\mathcal{I}_{\ell}=\mathcal{I}\equiv \{1,2,\cdots,N\}\times \{1,2,\cdots, T\}$. Rolling $\ell$ from $1$ to $L$, we then leave observations $\{(y_{it},x_{it}): (i,t)\in\mathcal{I}_{\ell}\}$ out, use observations $\{(y_{it},x_{it}): (i,t)\in\mathcal{I}/\mathcal{I}_{\ell}\}$ for training, and calculate the out-of-sample mean square error $\mathrm{MSE}_\ell$ for observations $\{(y_{it},x_{it}): (i,t)\in\mathcal{I}_{\ell}\}$. Finally, we choose $\lambda_{NT}$ by minimizing the average out-of-sample mean square error $\sum_{\ell=1}^{L}\mathrm{MSE}_\ell/L$.

\begin{rem}\label{Rem: ALS}
Enforcing the rank constraint directly is perhaps the most intuitive approach to incorporate the reduced-rank structure. This leads to the following problem:
\begin{align}\label{Eqn: LeastS}
\min_{{c_i}\in\mathbf{R}^{p},{D_i}\in\mathbf{R}^{p\times K},{g}_t\in\mathbf{R}^{K}}\frac{1}{2}\sum_{i=1}^{N}\sum_{t=1}^{T}(y_{it} - x_{it}^{\prime}c_i - x_{it}^{\prime}D_ig_t)^{2}.
\end{align}
However, solving \eqref{Eqn: LeastS} poses at least two challenges.\footnote{Enforcing the constraints of $a$ and $B$ in Examples \ref{Ex: SemiTin} and \ref{Ex: HomoConditional} does not resolve the challenges.} Firstly, it requires knowledge of $K$, which must be estimated prior to solving the problem. Secondly, \eqref{Eqn: LeastS} is nonconvex and its solution lacks an analytical closed form. These challenges not only complicate the design of computational algorithms to find the solution but also hinder derivation of its asymptotic properties. One potential approach to address the second challenge is alternating least squares; however, it may suffer from non-convergence issues due to the nonconvexity of \eqref{Eqn: LeastS} \citep{GolubVanLoan_MatrixComputation_2013, Chietal_Nonconvex_2019}. In contrast, the problem in \eqref{Eqn: NuclearNEst} is convex, and our estimators can be numerically solved efficiently without requiring prior knowledge of $K$, complemented by the asymptotic properties derived in Sections \ref{Sec4} and \ref{Sec5}.
\end{rem}

\section{Asymptotic Analysis}\label{Sec4}
In this section, we conduct an asymptotic analysis for our estimators in a general setup. Specifically, we establish consistency of $\hat{K}$ and a rate of convergence of $\hat{\Pi}$, $\hat{a}, \hat{B}$, and $\hat{F}$.

We begin by introducing the so-called ``restricted strong convexity'' condition \citep{Negahbanetal_UnifiedFramework_2012}. This condition ensures that the quadratic loss function in \eqref{Eqn: NuclearNEst} is strictly convex over a restricted set of ``low-rank'' matrices. To describe the set, we define some notation. Let $\Pi = U \Sigma V^{\prime}$ be a singular value decomposition of $\Pi$, where $U$ and $V$ are $Np\times Np$ and $T\times T$ orthonormal matrices, and $\Sigma$ is a diagonal matrix with singular values of $\Pi$ in the diagonal in descending order. Write $U = (U_1,U_2)$ and $V = (V_1,V_2)$, where the columns of $U_2$ and $V_2$ are singular vectors corresponding to the zero singular values of $\Pi$. For any $Np \times T$ matrix $\Delta$, let $\mathcal{P}(\Delta)\equiv U_2U_2^{\prime}\Delta  V_2V_2^{\prime}$ and $\mathcal{M}(\Delta)\equiv \Delta - \mathcal{P}(\Delta)$. Heuristically, $\mathcal{M}(\Delta)$ can be thought of as the projection of $\Delta$ onto the ``low-rank'' space of $\Pi$, and  $\mathcal{P}(\Delta)$ is the projection of $\Delta$ onto its orthogonal space. The restricted set of ``low-rank'' matrices is
\begin{align}\label{Eqn: RestrictedSet}
\mathcal{C}\equiv \{\Delta\in\mathcal{S}\ominus \mathcal{S}: \|\mathcal{P}(\Delta)\|_{\ast}\leq 3\|\mathcal{M}(\Delta)\|_{\ast}\},
\end{align}
where $\mathcal{S}\ominus \mathcal{S}$ is the Minkowski difference between $\mathcal{S}$ and $\mathcal{S}$, that is, $\mathcal{S}\ominus \mathcal{S} = \{\Gamma_1-\Gamma_2: \Gamma_1,\Gamma_2\in\mathcal{S}\}$. We impose the restricted strong convexity condition as follows.

\begin{ass}\label{Ass: NuclearN}
(i) Assume that $\Pi\in\mathcal{S}\subset \mathbf{R}^{Np\times T}$. For any $\Delta\in\mathcal{S}\ominus \mathcal{S}$, the following decomposition holds:
\[\sum_{i=1}^{N}\sum_{t=1}^{T}|\mathrm{tr}(X_{it}^{\prime}\Delta)|^2 = \mathcal{Q}_{NT}(\Delta) + \mathcal{L}_{NT}(\Delta)\]
such that for some constant $0<\kappa<\infty$,
\[\mathcal{Q}_{NT}(\Delta) \geq \kappa \|\Delta\|_{F}^{2} \text{ for all } \Delta\in\mathcal{C},\]
and for some $r_{NT}>0$,
\[|\mathcal{L}_{NT}(\Delta)|\leq r_{NT} \|\Delta\|_{\ast}  \text{ for all } \Delta\in\mathcal{S}\ominus \mathcal{S}.\]
(ii) The following condition holds:
\[\left|\sum_{i=1}^{N}\sum_{t=1}^{T}\mathrm{tr}(\varepsilon_{it} X_{it}^{\prime}\Delta)\right|\leq \frac{1}{2}r_{NT} \|\Delta\|_{\ast} \text{ for all } \Delta\in\mathcal{S}\ominus \mathcal{S}.\]
\end{ass}

Assumption \ref{Ass: NuclearN} is weaker than the conditions of Corollary 1 in \citet{NegahbanWainwright_Scaling_2011}, which require $\mathcal{S}=\mathbf{R}^{Np \times T}$ and $\mathcal{L}_{NT}(\cdot)=0$, and are too restrictive in Examples \ref{Ex: SemiTin} and \ref{Ex: HomoConditional}. We refer to the condition: ``$\mathcal{Q}_{NT}(\Delta) \geq \kappa \|\Delta\|_{F}^{2}$ for all $\Delta\in\mathcal{C}$'' as the restricted strong convexity condition. Allowing $\mathcal{L}_{NT}(\cdot)\neq 0$ facilitates providing easy-to-verify primitive conditions for the restricted strong convexity condition in Example \ref{Ex: SemiTin}. The rate $r_{NT}$ plays an important role in determining the convergence rate of $\hat{\Pi}$, thus determining how fast $p$ and $K$ can grow.

\begin{ass}\label{Ass: NuclearKaBF}
There exist some constants $0<d_{\min}\leq d_{\max}<\infty$ such that: (i) $d_{\min}<\lambda_{\min}(B'B/N)\leq \lambda_{\max}(B'B/N)<d_{\max}$ for large $N$; (ii) $\max_{t\leq T}\|f_{t}\|< d_{\max}$; (iii) $\lambda_{\min}(F'M_{T}F/T)> d_{\min}$; (iv) $a^{\prime}a/N<d_{\max}$; (v) $a^{\prime}B = 0$.
\end{ass}

For the sake of clarity in presentation, we assume that $\{a_{i},B_{i}\}_{i\leq N}$ and $\{f_t\}_{t\leq T}$ are non-random. In other words, all stochastic statements are implicitly conditional on their realization. Assumption \ref{Ass: NuclearKaBF}(i) resembles the pervasive condition in \citet{StockWatson_PCA_2002} and \citet{BaiNg_NumberofFactors_2002}, which necessitates that $\{f_t\}_{t\leq T}$ are strong factors. Assumptions \ref{Ass: NuclearKaBF}(iv) and (v) are identification conditions for $a$, see Appendix \ref{App: F: 2} for discussion; similar assumptions are also used in \citet{Chenetal_SeimiparametricFactor_2021}. While Assumptions \ref{Ass: NuclearN} and \ref{Ass: NuclearKaBF} consist of high-level conditions for the general setup, in Section \ref{Sec5} we provide primitive conditions for each setting in Examples \ref{Ex: Classical}-\ref{Ex: ConitionalObs}.

\begin{thm}\label{Thm: NuclearNRate}
Suppose Assumption \ref{Ass: NuclearN} holds. Let $\hat{\Pi}$, $\hat{K}$, $\hat{a}, \hat{B}$, and $\hat{F}$ be given in \eqref{Eqn: NuclearNEst}-\eqref{Eqn: aFEst}. Assume that $0<K<\min\{Np,T\}-1$ and $\lambda_{NT}\geq 2r_{NT}$. (i) Then
\begin{align*}
\|\hat{\Pi}  - \Pi\|_{F}\leq \frac{3\sqrt{2(K+1)}\lambda_{NT}}{\kappa}.
\end{align*}
(ii) Suppose Assumption \ref{Ass: NuclearKaBF} also holds. Assume that $\delta_{NT}/(NT)\to 0$ and $\delta_{NT}/(K\lambda^{2}_{NT})\to \infty $. Let $H \equiv (F^{\prime}M_T\hat{F})(\hat{F}^{\prime}M_T\hat{F})^{-1}$. Then
\begin{align*}
P(\hat{K}= K) &\to 1,\\
\|\hat{a} - a\|&=O_{p}\left(\frac{\sqrt{K}\lambda_{NT}}{\sqrt{T}}\right),\notag\\
\|\hat{B} - B H\|_{F}&=O_{p}\left(\frac{\sqrt{K}\lambda_{NT}}{\sqrt{T}}\right),\notag\\
\|\hat{F}-F(H^{\prime})^{-1}\|_{F}&=O_{p}\left(\frac{\sqrt{K}\lambda_{NT}}{\sqrt{N}}\right).
\end{align*}
\end{thm}

Theorem \ref{Thm: NuclearNRate}(i) gives a deterministic statement about the estimation error of $\hat{\Pi}$, extending Corollary 1 of \citet{NegahbanWainwright_Scaling_2011} by allowing $\mathcal{L}_{NT}(\cdot)\neq 0$ and constraints on $\Pi$ (i.e., $\mathcal{S}\neq \mathbf{R}^{Np \times T}$) in addition to the reduced-rank constraint. While Assumption \ref{Ass: NuclearN} and $\lambda_{NT}\geq 2r_{NT}$ may not hold deterministically, they often hold with probability approaching one, as verified in Section \ref{Sec5}. In such cases, the result of Theorem \ref{Thm: NuclearNRate}(i) holds with probability approaching one, and the results of Theorem \ref{Thm: NuclearNRate}(ii) persist. Due to identification issues, $B$ and $F$ can only be consistently estimated up to a rotational transformation, as commonly encountered in high-dimensional factor analyses. The asymptotic results hold as $N\to\infty$ with fixed $T$ or as $(N,T)\to\infty$, as appropriate. Theorem \ref{Thm: NuclearNRate} is a theory for the general setup under high-level assumptions, which is applicable for each setting in Examples \ref{Ex: Classical}-\ref{Ex: ConitionalObs}. In Section \ref{Sec5}, we tailor Theorem \ref{Thm: NuclearNRate} for each model by providing low-level sufficient assumptions that are easier to verify. In all cases, $p$ and $K$ are permitted to grow with $N$ or $(N,T)$ for the consistency of the estimators, and the presence of missing values is allowed.

\section{Revisiting Nested Models}\label{Sec5}
For simplicity of notation, we continue to use $x_{it}$ representing the vector of explanatory variables in all models, rather than each model's specific notation in Section \ref{Sec2}.

\subsection{Examples \ref{Ex: Classical}, \ref{Ex: Statevarying}, and \ref{Ex: ConitionalObs}}\label{Sec51}

Our objective is to estimate $a$, $B$, $F$, and $K$.\footnote{For simplicity of presentation, we continue to use $a$ and $B$ representing the coefficients of interest in all three examples, rather than each example's specific notation in Section \ref{Sec2}, and ignore the sieve approximation error in Example \ref{Ex: Statevarying} (so $\Delta_{i}(\cdot) = 0$). This allows us to unify results in one theorem. One may account for the sieve approximation error as similar to Corollaries \ref{Cor: Ex2: 1} and \ref{Cor: Ex4: 1}.} No constraints are imposed on $a$ and $B$ and we set $\mathcal{S} = \mathbf{R}^{Np\times T}$ in \eqref{Eqn: NuclearNEst}. In the scenario  when $x_{it}=1$, we can obtain an analytical closed form for $\hat{\Pi}$. Let $Y$ be an $N\times T$ matrix with the $it$th entry $y_{it}$. Consider the singular value decomposition $Y = U\Sigma V^{\prime}$, where $U$ and $V$ are $N\times N$ and $T\times T$ orthonormal matrices and $\Sigma$ is an $N\times T$ diagonal matrix with singular values $\sigma_{j}(Y)$'s in the diagonal in descending order. For $x>0$, define $\Sigma_{x}$ be an $N\times T$ diagonal matrix with $\max\{0,\sigma_j(Y)-x\}$ in descending order. Consequently, $\hat{\Pi} = U\Sigma_{\lambda_{NT}/2} V^{\prime}$, as described in \citet{Caietal_SVD_2010} and \citet{Maetal_Fixedpoint_2011}. However, an analytical closed form is not available for general cases. An efficient algorithm for finding $\hat{\Pi}$ is provided in Appendix \ref{App: E}.

To provide primitive conditions, we impose the following assumptions.

\begin{ass}\label{Ass: Ex135: 1}
(i) There exists some constant $0<\kappa<\infty$ such that
\[\sum_{i=1}^{N}\sum_{t=1}^{T}|\mathrm{tr}(X_{it}^{\prime}\Delta)|^2\geq \kappa \|\Delta\|_{F}^{2} \text{ for all } \Delta\in\mathcal{D},\]
where $\mathcal{D}\equiv\{\Delta\in\mathbf{R}^{Np\times T}: \|\mathcal{P}(\Delta)\|_{\ast}\leq 3\|\mathcal{M}(\Delta)\|_{\ast}\}$. (ii) $\{(x_{1t}^{\prime}e_{1t},x_{2t}^{\prime}e_{2t},\ldots,x_{Nt}^{\prime}e_{Nt})'\}_{t\leq T}$ is a sequence of independent sub-Gaussian vectors.\footnote{Independence is not necessary here and also in Assumptions \ref{Ass: Eqn: Ex2: 1}(iv), (v) and \ref{Ass: Eqn: Ex4: 1}(iii). We may allow for weak dependence over $t$; see Lemma \ref{Lem: UseD1}. }
\end{ass}

A condition similar to Assumption \ref{Ass: Ex135: 1} has been imposed in \citet{MoonWeidner_Nuclear_2018} and \citet{Chernozhukovetal_LowRank_2018}. We apply Theorem \ref{Thm: NuclearNRate} to obtain the following corollary.

\begin{cor}\label{Cor: Ex135: 1}
Suppose Assumption \ref{Ass: Ex135: 1}(ii) holds. Let $\hat{\Pi}$, $\hat{K}$, $\hat{a}, \hat{B}$, and $\hat{F}$ be given in \eqref{Eqn: NuclearNEst}-\eqref{Eqn: aFEst} with $\mathcal{S} = \mathbf{R}^{Np\times T}$ and $\lambda_{NT}=\sqrt{(Np+T)\log N}$. Assume that $0<K<\min\{Np,T\}-1$. (i) If $x_{it} =1$ or Assumption \ref{Ass: Ex135: 1}(i) holds, then as $(N,T)\to\infty$,
\begin{align*}
\frac{1}{\sqrt{NT}}\|\hat{\Pi}  - \Pi\|_{F}= O_{p}\left(\sqrt{\frac{K(Np+T)\log N}{NT}}\right).
\end{align*}
(ii) Suppose Assumptions \ref{Ass: NuclearKaBF}(i)-(iii) additionally hold. Assume that as $(N,T)\to\infty$, \hspace{-0.05cm}$\delta_{NT}/(NT)\hspace{-0.05cm}\to \hspace{-0.05cm}0$ \hspace{-0.05cm}and\hspace{-0.05cm} $\delta_{NT}/[K(Np+T)\log N]\to \hspace{-0.05cm}\infty\hspace{-0.05cm} $. \hspace{-0.1cm}Let $H \hspace{-0.1cm}\equiv \hspace{-0.1cm}(F^{\prime}M_T\hat{F})(\hat{F}^{\prime}M_T\hat{F})^{-1}$. If $a=0$ or Assumptions \ref{Ass: NuclearKaBF}(iv)-(v) hold, then as $(N,T)\to\infty$,
\begin{align*}
P(\hat{K}= K) &\to 1,\\
\frac{1}{\sqrt{N}}\|\hat{a} - a\|&=O_{p}\left(\sqrt{\frac{K(Np+T)\log N}{NT}}\right),\notag\\
\frac{1}{\sqrt{N}}\|\hat{B} - B H\|_{F}&=O_{p}\left(\sqrt{\frac{K(Np+T)\log N}{NT}}\right),\notag\\
\frac{1}{\sqrt{T}}\|\hat{F}-F(H^{\prime})^{-1}\|_{F}&=O_{p}\left(\sqrt{\frac{K(Np+T)\log N}{NT}}\right).
\end{align*}
\end{cor}

Corollary \ref{Cor: Ex135: 1} requires large $N$ and large $T$. In particular, $K(Np+T)\log N=o(NT)$ is required for the consistency of the estimators. This implies that $p$ is allowed to grow as $(N,T)\to\infty$.
While the result for Example \ref{Ex: Classical} is well-documented in the literature, the results for Examples \ref{Ex: Statevarying} and \ref{Ex: ConitionalObs} are novel. Distinct from \citet{PelgerXiong_State-varying_2019}, we offer an estimator capable of consistently estimate $F$ up to a common rotational transformation, which is not state-specific. In other words, we can consistently estimate the factor space. Moreover, our method allows for large $p$. In contrast to \citet{Gagliardinietal_Timevarying_2016}, we provide an estimation approach that does not necessitate observable $f_t$ and permits the presence of arbitrage and large $p$. Notably, there is no available method for estimating the unconstrained conditional latent factor model in the literature.




\subsection{Example \ref{Ex: SemiTin}}\label{Sec52}
Our objective is to estimate $\mu\equiv (\mu_1,\mu_2,\ldots,\mu_N)^{\prime}$, $\Lambda\equiv (\lambda_{1}, \lambda_{2},\ldots,\lambda_{N})^{\prime}$, $\phi$, $\Phi$, $F$, and $K$. Since $a = (\mu_1, \phi^{\prime}, \mu_2, \phi^{\prime},\ldots, \mu_N, \phi^{\prime})^{\prime}$ and $B = (\lambda_1, \Phi^{\prime}, \lambda_2, \Phi^{\prime}, \ldots, \lambda_N, \Phi^{\prime})^\prime$, we have $\Pi =a 1^{\prime}_{T} + BF^{\prime}= ((\pi_1,\Pi^{\ast\prime}), (\pi_2,\Pi^{\ast\prime}),\ldots, (\pi_N,\Pi^{\ast\prime}))^{\prime}$, where $\pi_i\equiv\mu_i1_{T}+ F\lambda_i$ and $\Pi^{\ast} \equiv \phi1_{T}^{\prime} + \Phi F^{\prime}$, which are $ T\times 1$ vector and $(p-1)\times T$ matrix, respectively. Then we set
\begin{align}\label{Eqn: Ex2: 1}
\hspace{-0.25cm}\mathcal{S} = \mathcal{D}_M \equiv \left\{\left(
                    \begin{array}{c}
                      \gamma_1^{\prime} \\
                      \Gamma^{\ast} \\
                      \gamma_2^{\prime} \\
                      \Gamma^{\ast} \\
                      \vdots \\
                      \gamma_N^{\prime} \\
                      \Gamma^{\ast} \\
                    \end{array}
                  \right)
: \left(
    \begin{array}{c}
    \gamma_1^{\prime} \\
    \gamma_2^{\prime} \\
      \vdots \\
    \gamma_N^{\prime} \\
    \end{array}
  \right)
\in\mathbf{R}^{N\times T}, \Gamma^{\ast}\in \mathbf{R}^{(p-1)\times T} \text{ and } \|\Gamma^{\ast}\|_{\max}\leq M\right\}
\end{align}
for $0<M<\infty$ in \eqref{Eqn: NuclearNEst}, where $\|\Gamma^{\ast}\|_{\max}$ denotes the largest absolute value of the entries of $\Gamma^{\ast}$.\footnote{\label{foot7}Imposing $\|\Gamma^{\ast}\|_{\infty}\leq M$ facilitates providing easy-to-verify primitive conditions for Assumption \ref{Ass: NuclearN}(i).} Since $\hat{\Pi}\in\mathcal{D}_{M}$, we can write $\hat{\Pi} = ((\hat{\pi}_1,\hat{\Pi}^{\ast\prime}), (\hat{\pi}_2,\hat{\Pi}^{\ast\prime}),\ldots, (\hat{\pi}_N,\hat{\Pi}^{\ast\prime}))^{\prime}$, where $\hat{\pi}_i$ is an estimator of $\pi_i$ and $\hat{\Pi}^{\ast}$ is an estimator of $\Pi^{\ast}$. Let ${\Pi}^{\diamond}\equiv({\pi}_{1},{\pi}_{2}, \ldots, {\pi}_{N})^{\prime}$ and $\hat{\Pi}^{\diamond}\equiv(\hat{\pi}_{1},\hat{\pi}_{2}, \ldots, \hat{\pi}_{N})^{\prime}$. An efficient algorithm for finding $\hat{\Pi}^{\diamond}$ and $\hat{\Pi}^{\ast}$ is provided in Appendix \ref{App: E}.

By Lemma \ref{Lem: TechD2} (iv) and simple algebra, we can write
\begin{align}\label{Eqn: Ex2: Estimators}
\hat{a} = ((\hat{\mu}_{1},\hat{\phi}^{\prime}), (\hat{\mu}_{2},\hat{\phi}^{\prime}),\ldots, (\hat{\mu}_{N},\hat{\phi}^{\prime}))^{\prime} \text{ and } \hat{B} = ((\hat{\lambda}_{1},\hat{\Phi}^{\prime}),(\hat{\lambda}_{2},\hat{\Phi}^{\prime}),\ldots,(\hat{\lambda}_{N},\hat{\Phi}^{\prime}))^{\prime},
\end{align}
where $\hat{\mu}_i$ is a scalar, $\hat{\phi}$ is a $(p-1)\times 1$ vector, $\hat{\lambda}_i$ is a $\hat{K} \times 1$ vector, and $\hat{\Phi}$ is a $(p-1)\times \hat{K}$ matrix. Thus, $\hat{a}$ and $\hat{B}$ share the same homogeneity structure with $a$ and $B$, respectively. It is not necessary to enforce the homogeneity restriction of $a$ and $B$ in extracting $\hat{a}$ and $\hat{B}$ from $\hat{\Pi}$ to ensure the same homogeneity structure, as the homogeneity structure of $\hat{\Pi}$ inherited from $a$ and $B$ automatically passes to $\hat{a}$ and $\hat{B}$. We define the estimators of  ${\mu}$, $\Lambda$, $\phi$, and $\Phi$ as  $\hat{\mu}\equiv (\hat{\mu}_1, \hat{\mu}_2, \ldots, \hat{\mu}_{N})^{\prime}$, $\hat{\Lambda} \equiv (\hat{\lambda}_{1},\hat{\lambda}_{2},\ldots, \hat{\lambda}_{N})^{\prime}$, $\hat{\phi}$, and $\hat{\Phi}$, respectively.

A convergence rate for $\hat{\Pi}^{\diamond}$, $\hat{\Pi}^{\ast}$, $\hat{\mu}$, $\hat{\Lambda}$, $\hat{\phi}$, and $\hat{\Phi}$ follows immediately from Theorem \ref{Thm: NuclearNRate}, as we have $\|\hat{\Pi}  - \Pi\|^{2}_{F} = \|\hat{\Pi}^{\diamond}-{\Pi}^{\diamond}\|^{2}_{F} + N\|\hat{\Pi}^{\ast}  - \Pi^{\ast}\|^{2}_{F}$, $\|\hat{a} - a\|^{2} = \|\hat{\mu} - \mu\|^2 + N\|\hat{\phi} - \phi\|^2$, and $\|\hat{B} - BH\|_{F}^{2} = \|\hat{\Lambda} - \Lambda H\|_{F}^{2} + N\|\hat{\Phi} - \Phi H\|^2_{F}$. To provide primitive conditions, we impose the following assumptions.

\begin{ass}\label{Ass: Eqn: Ex2: 1}
(i) Write $x_{it} = (1,x_{it}^{\ast\prime})^{\prime}$.\footnote{We allow for time-varying characteristics, so we write $x_{it}$ rather than $x_i$.} There are positive constants $c_{\min}$ and $c_{\max}$ such that: with probability approaching one as $(N,T)\to\infty$,
\[c_{\min} \leq \min_{t\leq T}\lambda_{\min}\left(\frac{1}{N}\sum_{i=1}^{N}x^{\ast}_{it}x_{it}^{\ast\prime}\right)\leq \max_{t\leq T}\lambda_{\max}\left(\frac{1}{N}\sum_{i=1}^{N}x^{\ast}_{it}x_{it}^{\ast\prime}\right)\leq c_{\max} .\]
(ii) $\max_{t\leq T}\hspace{-0.1cm}\|\phi + \Phi f_{t}\|_{\infty}$ is bounded. \hspace{-0.1cm}(iii) $\max_{t\leq T}\hspace{-0.1cm}E[\|\sum_{i=1}^{N}x^{\ast}_{it}e_{it}/\hspace{-0.1cm}\sqrt{Np}\|^{2}]$ is bounded. (iv) $\{(x_{1t}^{\ast\prime},x_{2t}^{\ast\prime},\ldots,x_{Nt}^{\ast\prime})'\}_{t\leq T}$ is a sequence of independent sub-Gaussian vectors. (v) $\{(e_{1t},e_{2t},$ $\ldots,e_{Nt})'\}_{t\leq T}$ is a sequence of independent sub-Gaussian vectors. (vi) $\sup_{z}|\delta(z)| = O(p^{-s})$ and $\sup_{z}\|\Delta(z)\| = O(p^{-s})$ for some constant $s>0$.
\end{ass}

\begin{ass}\label{Ass: Eqn: Ex2: 2}
There are constants $0<d_{\min}\leq d_{\max}<\infty$ such that: (i) $\lambda_{\min}(\Phi'\Phi+\Lambda^{\prime}\Lambda/N)>d_{\min}$; (ii) $\lambda_{\max}(\Phi'\Phi)\hspace{-0.05cm}<\hspace{-0.05cm}d_{\max}/2$; (iii) $\lambda_{\max}(\Lambda^{\prime}\Lambda/N)\hspace{-0.05cm}<\hspace{-0.05cm}d_{\max}/2$; (iv) $\max_{t\leq T}\|f_{t}\|$ $< d_{\max}$; (v) $\lambda_{\min}(F'M_{T}F/T)> d_{\min}$; (vi) $\|\phi\|^{2}<d_{\max}/2$; (vii) $\|\mu\|^{2}/N<d_{\max}/2$; (viii) $\phi^{\prime}\Phi = 0$; (ix) $\mu^{\prime}\Lambda = 0$.
\end{ass}

Assumption \ref{Ass: Eqn: Ex2: 1} involves no multicolinearity, finite moments, weak dependence, and small sieve approximation errors, all of which are standard in the literature. Conditions similar to Assumption \ref{Ass: Eqn: Ex2: 1}(i), (iii), and (vi) have been imposed in \citet{Fanetal_ProjectedPCA_2016}. We apply Theorem \ref{Thm: NuclearNRate} to obtain the following corollary.

\begin{cor}\label{Cor: Ex2: 1}
Suppose Assumption \ref{Ass: Eqn: Ex2: 1} holds. Let $\hat{\Pi}$ be given in \eqref{Eqn: NuclearNEst} with $\mathcal{S} = \mathcal{D}_M$ and $\lambda_{NT}= [M\sqrt{(Np^2+ Tp)} + \sqrt{NT} p^{-s}]\sqrt{\log N} $. Let $\hat{\Pi}^{\diamond}$ and $\hat{\Pi}^{\ast}$ be given below \eqref{Eqn: Ex2: 1}.  Let  $\hat{K}$, $\hat{F}$, $\hat{\mu}$, $\hat{\Lambda}$, $\hat{\phi}$, and $\hat{\Phi}$ be given in \eqref{Eqn: KEst}, \eqref{Eqn: aFEst}, and \eqref{Eqn: Ex2: Estimators}. Assume that $0<K<\min\{N+p-1,T\}-1$. (i) Then as $(N,T)\to\infty$,
\begin{align*}
\frac{1}{\sqrt{NT}}\|\hat{\Pi}^{\diamond} - \Pi^{\diamond}\|_{F}&= O_{p}\left(M\sqrt{\frac{K(Np^{2}+Tp)\log N}{NT}} + \frac{\sqrt{K \log N}}{p^{s}}\right),\notag\\
\frac{1}{\sqrt{T}}\|\hat{\Pi}^{\ast}  - \Pi^{\ast}\|_{F}&= O_{p}\left(M\sqrt{\frac{K(Np^{2}+Tp)\log N}{NT}}+ \frac{\sqrt{K \log N}}{p^{s}}\right).
\end{align*}
(ii) Suppose Assumption \ref{Ass: Eqn: Ex2: 2} additionally holds. \hspace{-0.1cm}Assume that as $(N,T)\hspace{-0.05cm}\to\infty$, $\delta_{NT}/(NT)$ $\to 0$ and $\delta_{NT}/\{K[M^{2}(Np^{2}+Tp)+NT p^{-2s}]\log N \}\to \infty $. Let $H \equiv (F^{\prime}M_T\hat{F})(\hat{F}^{\prime}M_T\hat{F})^{-1}$. Then as $(N,T)\to\infty$,
\begin{align*}
P(\hat{K}= K) &\to 1,\\
\frac{1}{\sqrt{N}}\|\hat{\mu} - \mu\|&=O_{p}\left(M\sqrt{\frac{K(Np^{2}+Tp)\log N}{NT}}+ \frac{\sqrt{K \log N}}{p^{s}}\right),\notag\\
\frac{1}{\sqrt{N}}\|\hat{\Lambda} - \Lambda H\|_{F}&=O_{p}\left(M\sqrt{\frac{K(Np^{2}+Tp)\log N}{NT}}+ \frac{\sqrt{K \log N}}{p^{s}}\right),\notag\\
\|\hat{\phi} - \phi\|&=O_{p}\left(M\sqrt{\frac{K(Np^{2}+Tp)\log N}{NT}}+ \frac{\sqrt{K \log N}}{p^{s}}\right),\notag\\
\|\hat{\Phi} - \Phi H\|_{F}&=O_{p}\left(M\sqrt{\frac{K(Np^{2}+Tp)\log N}{NT}}+ \frac{\sqrt{K \log N}}{p^{s}}\right),\notag\\
\frac{1}{\sqrt{T}}\|\hat{F}-F(H^{\prime})^{-1}\|_{F}&=O_{p}\left(M\sqrt{\frac{K(Np^{2}+Tp)\log N}{NT}}+ \frac{\sqrt{K \log N}}{p^{s}}\right).
\end{align*}
\end{cor}

The slower rate in Corollary \ref{Cor: Ex2: 1} compared to Corollary \ref{Cor: Ex135: 1} is attributed to the reliance on a set of easier-to-verify conditions, namely Assumption \ref{Ass: Eqn: Ex2: 1}, rather than a version of Assumption \ref{Ass: Ex135: 1}. However, it is noteworthy that the rate can be improved to $O_{p}(\sqrt{K(N+p+T)\log N/(NT)} + \sqrt{K \log N}/p^{s})$ under Assumption \ref{Ass: Ex135: 1}.  The second term $\sqrt{K \log N}/p^{s}$ arises from sieve approximation errors. Our results differ from \citet{Fanetal_ProjectedPCA_2016} in several aspects. First, we allow for $\mu_i\neq 0 $ and $\phi\neq 0$, which are crucial to capture pricing errors in asset pricing. Second, we permit $x_{it}$ to vary over $t$, a critical feature in asset pricing as many stock characteristics (e.g., book to market ratio and momentum) change from month to month. Our simulations in Appendix \ref{App: G: sub4} show that \citet{Fanetal_ProjectedPCA_2016}'s projected-PCA fails in the presence of time-varying $x_{it}$. Third, we do not require that $\lambda_i$ has zero mean and weak cross-sectional dependence (in such cases $\lambda_i$ can be interpreted as a vector of noises), which is barely justified in practice. We allow for non-noisy intercepts $\mu_i$ and $\lambda_i$ in pricing errors and risk exposures. Fourth, we allow $K\to\infty$. In addition, our results extend  \citet{Chenetal_SeimiparametricFactor_2021} by allowing for the heterogeneity of $\mu_i$ and $\lambda_i$ across $i$.



\subsection{Example \ref{Ex: HomoConditional}}\label{Sec53}
Our objective is to estimate $\phi_0$, $\Phi_0$, $F$, and $K$. Since $a = 1_{N}\otimes \phi_0$ and $B = 1_{N}\otimes \Phi_0$, we have $\Pi =a 1^{\prime}_{T} + BF^{\prime}= 1_{N} \otimes \Pi_0$, where $\Pi_0\equiv \phi_0 1^{\prime}_{T} + \Phi_0 F^{\prime}$, which is a $p\times T$ matrix. Then we set $\mathcal{S} = \{1_{N}\otimes \Gamma: \Gamma\in \mathbf{R}^{p\times T}\}$ in \eqref{Eqn: NuclearNEst}. Since $\hat{\Pi}\in\mathcal{S}$, we can write $\hat\Pi = 1_{N} \otimes \hat{\Pi}_0$, where $\hat{\Pi}_0$ is an estimator of ${\Pi}_0$. An efficient algorithm for finding $\hat{\Pi}_0$ is provided in Appendix \ref{App: E}.

By Lemma \ref{Lem: TechD1}(iv) and simple algebra, we can write
\begin{align}\label{Eqn: Ex4: Estimators}
\hat{a} = 1_{N} \otimes \hat{\phi}_0 \text{ and } \hat{B} = 1_{N} \otimes\hat{\Phi}_0,
\end{align}
where $\hat{\phi}_0$ is a $p\times 1$ vector and $\hat{\Phi}_0$ is a $p\times \hat{K}$ matrix. For the same reason as in Example \ref{Ex: SemiTin}, there is no need to enforce the homogeneity restriction of $a$ and $B$ in extracting $\hat{a}$ and $\hat{B}$ from $\hat{\Pi}$. We define the estimators of $\phi_0$ and $\Phi_0$ as $\hat{\phi}_0$ and $\hat{\Phi}_0$, respectively.

A convergence rate for $\hat{\Pi}_0$, $\hat{\phi}_0$, and $\hat{\Phi}_0$ follows immediately from Theorem \ref{Thm: NuclearNRate}, as we have $\|\hat{\Pi}  - \Pi\|_{F} = \sqrt{N}\|\hat{\Pi}_0  - \Pi_0\|_{F}$, $\|\hat{a} - a\| = \sqrt{N}\|\hat{\phi}_0 - \phi_0\| $, and $\|\hat{B} - BH\|_{F} = \sqrt{N}\|\hat{\Phi}_0 - \Phi_0H\|_{F}$. To provide primitive conditions, we impose the following assumptions.

\begin{ass}\label{Ass: Eqn: Ex4: 1}
(i) There are positive constants $c_{\min}$ and $c_{\max}$ such that: with probability approaching one as $N\to\infty$ with fixed $T$ or as $(N,T)\to\infty$,
\[c_{\min} \leq \min_{t\leq T}\lambda_{\min}\left(\frac{1}{N}\sum_{i=1}^{N}x_{it}x_{it}^{\prime}\right)\leq \max_{t\leq T}\lambda_{\max}\left(\frac{1}{N}\sum_{i=1}^{N}x_{it}x_{it}^{\prime}\right)\leq c_{\max} .\]
(ii) $E[\|\sum_{i=1}^{N}x_{it}e_{it}/\sqrt{Np}\|^2]$ is bounded for each $t\leq T$. (iii) $\{\sum_{i=1}^{N}x_{it}e_{it}/\sqrt{N}\}_{t\leq T}$ is a sequence of independent sub-Gaussian vectors. (iv) $\sup_{z}|\delta(z)| = O(p^{-s})$ and $\sup_{z}\|\Delta(z)\| = O(p^{-s})$ for some constant $s>0$.
\end{ass}

\begin{ass}\label{Ass: Eqn: Ex4: 2}
There are constants $0<d_{\min}\leq d_{\max}<\infty$ such that: (i) $d_{\min}<\lambda_{\min}(\Phi_0^{\prime}\Phi_0)\leq \lambda_{\max}(\Phi_0^{\prime}\Phi_0)<d_{\max}$; (ii) $\max_{t\leq T}\|f_{t}\|< d_{\max}$; (iii) $\lambda_{\min}(F'M_{T}F/T)> d_{\min}$; (iv) $\|\phi_0\|^2<d_{\max}$; (v) $\phi_0^{\prime}\Phi_0 = 0$.
\end{ass}

Assumption \ref{Ass: Eqn: Ex4: 1} involves no multicolinearity, finite moments, weak dependence, and small sieve approximation errors, all of which are standard in the literature. Assumptions \ref{Ass: Eqn: Ex4: 1}(i), (ii), (iv), and \ref{Ass: Eqn: Ex4: 2} have been imposed in \citet{Chenetal_SeimiparametricFactor_2021}. We apply Theorem \ref{Thm: NuclearNRate} to obtain the following corollary.

\begin{cor}\label{Cor: Ex4: 1}
Suppose Assumptions \ref{Ass: Eqn: Ex4: 1}(i), (ii), and (iv) hold. Let $\hat\Pi$ be given in \eqref{Eqn: NuclearNEst} with $\mathcal{S} = \{1_{N}\otimes \Gamma: \Gamma\in \mathbf{R}^{p\times T}\}$ and $\lambda_{NT}=(\sqrt{p+T} +\sqrt{NT}p^{-s})\sqrt{\log N}$. Let $\hat{\Pi}_0 $ be given above \eqref{Eqn: Ex4: Estimators}. Let $\hat{K}$, $\hat{F}$, $\hat{\phi}_0$, and $\hat{\Phi}_0$ be given in \eqref{Eqn: KEst}, \eqref{Eqn: aFEst}, and \eqref{Eqn: Ex4: Estimators}. Assume $0<K<\min\{p,T\}-1$. (i) Then as $N\to\infty$ with fixed $T$,
\begin{align*}
\frac{1}{\sqrt{T}}\|\hat{\Pi}_0  - \Pi_0\|_{F}= O_{p}\left(\sqrt{\frac{K(p+T)\log N}{NT}} + \frac{\sqrt{K \log N}}{p^{s}}\right).
\end{align*}
(ii) Suppose Assumption \ref{Ass: Eqn: Ex4: 2} additionally holds. Assume that as $N\to\infty$ with fixed $T$, $\delta_{NT}/(NT)\hspace{-0.05cm}\to \hspace{-0.05cm}0$ and $\delta_{NT}/[K(p+T+ NTp^{-2s})\log N]\hspace{-0.05cm}\to \hspace{-0.05cm}\infty$. \hspace{-0.1cm}Let $H \hspace{-0.1cm}\equiv \hspace{-0.1cm}(F^{\prime}M_T\hat{F})(\hat{F}^{\prime}M_T\hat{F})^{-1}$. Then as $N\to\infty$ with fixed $T$,
\begin{align*}
P(\hat{K}= K) &\to 1,\\
\|\hat{\phi}_0 - \phi_0\|&=O_{p}\left(\sqrt{\frac{K(p+T)\log N}{NT}} + \frac{\sqrt{K \log N}}{p^{s}}\right),\notag\\
\|\hat{\Phi}_0 - \Phi_0 H\|_{F}&=O_{p}\left(\sqrt{\frac{K(p+T)\log N}{NT}} + \frac{\sqrt{K \log N}}{p^{s}}\right),\notag\\
\frac{1}{\sqrt{T}}\|\hat{F}-F(H^{\prime})^{-1}\|_{F}&=O_{p}\left(\sqrt{\frac{K(p+T)\log N}{NT}} + \frac{\sqrt{K \log N}}{p^{s}}\right).
\end{align*}
(iii) If Assumption \ref{Ass: Eqn: Ex4: 1}(ii) is replaced with Assumption \ref{Ass: Eqn: Ex4: 1}(iii), then (i) and (ii) continue to hold by replacing ``as $N\to\infty$ with fixed $T$'' with ``as $(N,T)\to\infty$'' in all places.
\end{cor}

Corollary \ref{Cor: Ex4: 1} establishes a convergence rate of $\hat{\Pi}_0$, $\hat{K}$, $\hat{\phi}_0, \hat{\Phi}_0$, and $\hat{F}$ either under large $N$ with fixed $T$ or scenarios with both large $N$ and large $T$. In particular, $K(p+T)\log N = o(NT)$ is required for the consistency. This implies that $p$ can be as large as $N$ up to $\log N$. Such a result represents a significant improvement from similar results in \citet{Chenetal_SeimiparametricFactor_2021}, which require that $p$ grows at a rate slower than $N^{1/3}$. Our simulations in Appendix \ref{App: G: sub4} show that \citet{Chenetal_SeimiparametricFactor_2021}'s regressed-PCA exhibits poor performance when $p$ is close to $N$. The rate $\sqrt{K \log N}/p^{s}$ arises from sieve approximation errors. In addition, our framework accommodates the scenario where $K$ tends to infinity and allows for weak cross-sectional dependence of $x_{it}$.


\section{Simulation Studies}\label{Sec6}
In this section, we conduct Monte Carlo simulations to investigate the finite sample performance of our estimators. We consider settings with $p=37$, $N = 500, 1000, 2000$, and $T = 250, 500$, which are comparable with those in the empirical analysis in Section \ref{Sec7}.

We consider three different data generating processes (DGPs), which correspond to the settings in Examples \ref{Ex: SemiTin}, \ref{Ex: HomoConditional},  and \ref{Ex: ConitionalObs}. In all three DGPs, we let
\begin{align}\label{Eqn: Covariates}
x_{it,1}=1, x_{it,2}=\sigma_{t}u_{it,1}, x_{it,3} = 0.3x_{i(t-1),3} + u_{it,2}, x_{it,4}=u_{it,3}, \ldots, x_{it,37}=u_{it,36},
\end{align}
where $u_{it} = (u_{it,1},u_{it,2},\ldots, u_{it,36})^{\prime}$ are i.i.d. $N(0,I_{36})$ across both $i$ and $t$, $\sigma_{t}$'s are i.i.d. $U(1,2)$ over $t$, and $x_{i0,3}$'s are i.i.d. $N(0,1)$ across $i$. Let $x_{it} = (x_{it,1},x_{it,2},\ldots,x_{it,37})^{\prime}$, hence $p=37$. Let $f_{t} = 0.3 f_{t-1} + \eta_t$, where $\eta_t$'s are i.i.d. $N(1_2,I_2)$ over $t$ and $f_0\sim N(1_2/0.7,I_2/0.91)$, resulting in $K=2$. The errors $\varepsilon_{it}$'s be i.i.d. $N(0,4)$ across both $i$ and $t$. In the first DGP (DGP1), we assume
\begin{align}\label{Eqn: alphabeta1}
a_i &= \left(
         \begin{array}{ccccccccccc}
             0 & \theta_{1i} & 0 & 0 & \theta_{2i} & 0 & 0 &\cdots & \theta_{12i} & 0 & 0 \\
         \end{array}
       \right)^{\prime} \text{ and }
\notag\\
B_i &=\left(
       \begin{array}{ccccccccccc}
         0 & 0 & \varrho_{1i} & 0 & 0 & \varrho_{2i} & 0 & \cdots & 0 & \varrho_{12i} & 0\\
         \varrho_{13i} & 0 & 0 &  \varrho_{14i} & 0 & 0 & \varrho_{15i} & \cdots & 0 & 0 & \varrho_{25i}\\
       \end{array}
     \right)^{\prime},
\end{align}
where $\theta_{ji}$'s are i.i.d. $N(0,1/4)$ across both $i$ and $j=1,2,\ldots, 12$ and $\varrho_{ji}$'s are i.i.d. $U(1/3,1)$ across both $i$ and $j=1,2,\ldots, 25$. In DGP1, $a_i$ and $B_i$ are heterogenous across $i$, which is the setting in Example \ref{Ex: ConitionalObs}. We are interested in $a$, $B$, $F$, and $K$. In the second DGP (DGP2), we assume
\begin{align}\label{Eqn: alphabeta2}
a_i &= \left(
         \begin{array}{cc}
            \mu_i & \phi^{\prime} \\
         \end{array}
       \right)^{\prime} = \left(
         \begin{array}{ccccccccccc}
            0 & 1/2 & 0 & 0 & 1/2 & 0 & 0 & \cdots & 1/2 & 0 & 0 \\
         \end{array}
       \right)^{\prime} \text{ and }
\notag\\
B_i &=\left(
         \begin{array}{cc}
            \lambda_i & \Phi^{\prime} \\
         \end{array}
       \right)^{\prime} = \left(
       \begin{array}{ccccccccccc}
         0 & 0 & 2/3 & 0 & 0 & 2/3 & 0 & \cdots & 0 & 2/3 & 0\\
         \vartheta_{i} &  0 & 0 & 2/3 & 0 & 0 & 2/3 & \cdots &  0 & 0 & 2/3 \\
       \end{array}
     \right)^{\prime},
\end{align}
where $\vartheta_i$'s are i.i.d. $U(1,3)$ across $i$. In DGP2, the rows of $a_i$ and $B_i$ corresponding to the nonconstant part of $x_{it}$ are homogenous across $i$, which is the setting in Example \ref{Ex: SemiTin}.  We are interested in $\mu$, $\phi$, $\Lambda$, $\Phi$, $F$, and $K$.
In the third DGP (DGP3), we assume
\begin{align}\label{Eqn: alphabeta3}
a_i &=\phi_0= \left(
         \begin{array}{ccccccccccc}
          0 & 1/2 & 0 & 0 & 1/2 & 0 & 0 & \cdots & 1/2 & 0& 0 \\
         \end{array}
       \right)^{\prime} \text{ and }
\notag\\
B_i &=\Phi_0 = \left(
       \begin{array}{ccccccccccc}
         0 & 0 & 2/3 & 0 & 0 & 2/3 & 0 & \cdots & 0 & 2/3 & 0\\
         2/3 & 0 & 0 &  2/3 & 0 & 0 & 2/3 & \cdots & 0 & 0 & 2/3\\
       \end{array}
     \right)^{\prime}.
\end{align}
In DGP3, $a_i$ and $B_i$ are homogenous across $i$, which is the setting in Example \ref{Ex: HomoConditional}. We are interested in $\phi_0$, $\Phi_0$, $F$, and $K$. Here, $u_{it}$'s, $\sigma_{t}$'s, $x_{i0,3}$'s, $\eta_t$'s, $f_0$, $\theta_{ji}$'s, $\varrho_{ji}$'s, $\vartheta_{i}$'s,  and $\varepsilon_{it}$'s are mutually independent. We generate $y_{it}$ according to the model in \eqref{Eqn: Model}.

For DGP1, we implement the estimators in \eqref{Eqn: NuclearNEst}-\eqref{Eqn: aFEst} with $\mathcal{S} = \mathbf{R}^{Np\times T}$. We assess the performance of $\hat{\Pi}$, $\hat{a}, \hat{B}$, $\hat{F}$, and $\hat{K}$. By Corollary \ref{Cor: Ex135: 1}, we set $\lambda_{NT} = c\sqrt{(Np+T)\log N}$ and  $\delta_{NT} = 2(Np+T)\log N$ for some $c>0$. For DGP2, we implement the estimators in \eqref{Eqn: NuclearNEst}-\eqref{Eqn: aFEst} and \eqref{Eqn: Ex2: Estimators} as well as below \eqref{Eqn: Ex2: 1} with $\mathcal{S} = \mathcal{D}_\infty$. We evaluate the performance of $\hat{\Pi}^{\diamond}$, $\hat{\Pi}^{\ast}$, $\hat{\mu}$, $\hat\Lambda$, $\hat{\phi}, \hat{\Phi}$, $\hat{F}$, and $\hat{K}$. By the discussion after Corollary \ref{Cor: Ex2: 1}, we set $\lambda_{NT} = c\sqrt{(N+p+T)\log N}$ and  $\delta_{NT} = 2(N+p+T)\log N$ for some $c>0$. For DGP3, we implement the estimators in \eqref{Eqn: NuclearNEst}-\eqref{Eqn: aFEst} and \eqref{Eqn: Ex4: Estimators} with $\mathcal{S} = \{1_{N}\otimes \Gamma: \Gamma\in \mathbf{R}^{p\times T}\}$. We evaluate the performance of $\hat{\Pi}_0$, $\hat{\phi}_0, \hat{\Phi}_0$, $\hat{F}$, and $\hat{K}$. By Corollary \ref{Cor: Ex4: 1}, we set $\lambda_{NT} = c\sqrt{(p+T)\log N}$ and $\delta_{NT} = 2\sqrt{N}(p+T)\log(N)$ for some $c>0$.

To determine the optimal value of $c$, we employ the $5$-fold CV approach, as outlined in Section \ref{Sec3} with $L=5$. The mean square errors of the regularized estimators ($\hat{\Pi}$, $(\hat{\Pi}^{\diamond\prime}, \sqrt{N}\hat{\Pi}^{\ast\prime})$, and $\hat{\Pi}_0$) are assessed both with fixed values of $c$ and using the CV method, with $c$ confined to $[0,2]$.\footnote{Specifically, we consider the grid set $\{0, 0.05, 0.1, 0.2, \ldots, 0.9,1,1.5,2\}$.} All simulation results are based on $200$ simulation replications. The main findings are as follows.
\begin{itemize}
  \item Nuclear norm regularization significantly enhances the performance of the estimators. In DGP1 and DPG2 (see Figures \ref{Fig: DGP1} and \ref{Fig: DGP2}), the mean square error of the unregularized estimator (i.e., $c=0$) remains relatively constant as both $N$ and $T$ increase (the value stays constantly around 40 in DGP1 and above 10 in DGP2), indicating potential inconsistency. Conversely, applying appropriate nuclear norm regularization (e.g., $c=1$) not only reduces the mean square error for each combination of $(N,T)$ but also drive the error towards zero as both $N$ and $T$ increase (e.g., the value for $c=1$ is getting closer to zero as $N$ and $T$ increase). This suggests that regularized estimators with a well-chosen $c$ value are consistent, aligning with Corollaries \ref{Cor: Ex135: 1} and \ref{Cor: Ex2: 1}. In DGP3 (see Figure \ref{Fig: DGP3}), although the mean square error of the unregularized estimator decreases with increasing $N$ (note that the scale of the vertical axis changes across columns of graphs), a properly chosen $c$ value (e.g., $c=0.6$ or $0.7$) leads to smaller errors. Thus, the simulations underscore the crucial role of nuclear norm regularization.
  \item The regularized estimators exhibit high sensitivity to the choice of $c$. For instance, selecting $c>2$ in DGP3 can result in a larger mean square error than that of the unregularized estimator across all $(N,T)$ combinations (as seen in Figure \ref{Fig: DGP3}). Therefore, careful consideration is essential when choosing $c$ in practice.
  \item The CV approach proves effective in selecting $c$ to minimize mean square error. Across all the three DGPs, the mean square error of the regularized estimator using the CV-selected $c$ value closely approximates the smallest error obtained with fixed $c$ values (as evidenced by the blue line closely tracking the lowest point of the dash-dotted line in Figures \ref{Fig: DGP1}-\ref{Fig: DGP3}), irrespective of $(N,T)$ combinations.
\end{itemize}
Overall, these findings highlight the importance of nuclear norm regularization, the sensitivity of estimators to $c$, and the efficacy of the CV approach in selecting an optimal $c$ value for minimizing mean square error.


We proceed to assess the performance of estimators other than $\hat{\Pi}$, $\hat{\Pi}^{\diamond}$, $\hat{\Pi}^{\ast}$, and $\hat{\Pi}_0$ by utilizing the CV-selected value of $c$. Tables \ref{Tab: MSEDGP1}-\ref{Tab: MSEDGP3} present their mean square errors or correct rates. The main findings are summarized as follows.
First, the number factor estimators (i.e., $\hat{K}$) consistently perform well across all cases, only one correct rate falling below $100\%$. This indicates their reliability in estimating the number of factors. Second, in DPG1 and DPG2, all mean square errors decrease as both $N$ and $T$ increase, consistent with Corollaries \ref{Cor: Ex135: 1} and \ref{Cor: Ex2: 1}. Similarly, in DGP3, mean square errors decrease as $N$ increases, indicating consistency as $N\to\infty$, aligning with Corollary \ref{Cor: Ex4: 1}. Third, increasing $N$ consistently reduces the mean square errors of the factor estimators (i.e., $\hat{F}$) across all cases, while increasing $T$ may not have a similar effect. In addition, increasing either $N$ or $T$ tends to reduce the mean square errors of $\hat{\Pi}^{\ast}$, $\hat{\phi}$, $\hat{\Phi}$, $\hat{\phi}_0$, and $\hat{\Phi}_0$. While this phenomenon is not explicitly explained by Corollaries \ref{Cor: Ex2: 1} and \ref{Cor: Ex4: 1}, it may be attributed to the homogeneity of the estimands (${\Pi}^{\ast}$, ${\phi}$, ${\Phi}$, ${\phi}_0$, and ${\Phi}_0$). In conclusion, our estimators exhibit promising performance in finite sample settings. The same findings are observed in settings with sparse $a$ and $B$, as well as in scenarios with small $p$, $N$, and $T$; see Appendix \ref{App: G} for additional simulation results. Furthermore, Appendix \ref{App: G} demonstrates the superiority of our estimators compared to existing ones.

\section{Empirical Analysis}\label{Sec7}
In this section, we analyze the cross section of individual stock returns in the US market using the same dataset as in \citet{Chenetal_SeimiparametricFactor_2021}, originally derived from \citet{Freybergeretal_Dissecting_2017}. The dataset comprises monthly returns and 36 characteristics of 12,813 individual US stocks spanning from September 1968 to May 2014. Due to a significant proportion of missing values in many stocks, we opt to exclude stocks with a sample length less than 200 to ensure that the proportion of missing values remains manageable. This results in an unbalanced panel with $N = 2,121$ and $T = 549$. Each time period includes at least 580 stocks with observations on both returns and the 36 characteristics, while each stock has observations in at least 200 time periods. Additionally, we transform the values of each characteristic to relative ranking values within the range $[-0.5, 0.5]$ in each time period.

In our analysis, we consider six different model specifications. The first three specifications, denoted S1, S2, and S3, include $x_{it}$ comprising a constant and the 36 characteristics. The remaining three specifications, denoted S4, S5, and S6, involve $x_{it}$ consisting of a constant and linear B-splines of 18 characteristics with one internal knot, as studied in \citet{Chenetal_SeimiparametricFactor_2021}. Refer to their paper for the 18 characteristics. In S1 and S4, we explore an unconstrained conditional factor model (corresponding to the setup in Example \ref{Ex: ConitionalObs}), where $a_{i}$ and $B_i$ can vary heterogeneously across $i$. For S2 and S5, we investigate a semiparametric conditional factor model (corresponding to the setup in Example \ref{Ex: SemiTin}), where the rows of $a_i$ and $B_i$ corresponding to the nonconstant explanatory variables in $x_{it}$ are constrained to be homogeneous. Lastly, S3 and S6 examine a homogeneous conditional factor model (corresponding to the setup in Example \ref{Ex: HomoConditional}), where $a_i$ and $B_i$ are constrained to be homogeneous. We estimate the models for $K=1,2,\ldots, 10$ by using our new method and select the regularization parameter using the 5-fold CV approach as outlined in Section \ref{Sec3}. Specifically, we set $\lambda_{NT} = c\sqrt{(Np+T)\log N}$ for S1 and S4, $\lambda_{NT} = c\sqrt{(N+p+T)\log N}$ for S2 and S5, $\lambda_{NT} = c\sqrt{(p+T)\log N}$ for S3 and S6, and choose $c$ from the set $\{0, 0.01, 0.02, 0.05, 0.1, 0.2, 0.5,1,2,5\}/100$. For a comparison, we also evaluate \citet{Chenetal_SeimiparametricFactor_2021}'s regressed-PCA method, denoted R1 and R2, alongside the homogeneous conditional factor models (S3 and S6).

To assess the performance of the models, we adopt various goodness-of-fit measures. First, we consider different types of in-sample $R^2$ measures:
\begin{align}
R^2 & = 1-\frac{\sum_{i, t}(y_{it}- x_{it}^{\prime}\hat{a}_{i} - x_{it}^{\prime}\hat{B}_{i}\hat{f}_t)^2}{\sum_{i, t} y_{it}^{2}}, \label{Eqn: R21}\\
R^2_{T,N} & = 1 - \frac{1}{N} \sum_{i} \frac{\sum_{t}(y_{it}- x_{it}^{\prime}\hat{a}_{i} - x_{it}^{\prime}\hat{B}_{i}\hat{f}_t)^2}{\sum_{t}y_{it}^{2}},\label{Eqn: R22}\\
R^2_{N,T} & = 1 - \frac{1}{T} \sum_{t} \frac{\sum_{i}(y_{it}- x_{it}^{\prime}\hat{a}_{i} - x_{it}^{\prime}\hat{B}_{i}\hat{f}_t)^2}{\sum_{i}y_{it}^{2}},\label{Eqn: R23}
\end{align}
where $\hat{a}\equiv (\hat{a}_1^{\prime},\hat{a}_2^{\prime},\ldots, \hat{a}_{N}^{\prime})^{\prime}$, $\hat{B} \equiv (\hat{B}_1^{\prime},\hat{B}_2^{\prime},\ldots, \hat{B}_{N}^{\prime})^{\prime}$, and $\hat{F} \equiv(\hat{f}_1,\hat{f}_2,\ldots, \hat{f}_{T})^{\prime}$. Here, the first one is total $R^2$, measuring the overall explanatory power of the models. The second one measures the cross-sectional average of time series $R^2$ across all stocks, reflecting the ability of the models to capture common variation in asset returns. The third one measures the time series average of cross-sectional $R^2$, which is the one of interest for evaluating the models' ability to explain the cross-section of average returns.
Second, we assess out-of-sample prediction. For $t\geq 300$, we utilize the data up to $t-1$ for estimation and obtain estimators, say $\hat{a}_{it}$, $\hat{B}_{it}$, and $\hat{F}_{t}\equiv (\hat{f}^{(t)}_{1},\hat{f}^{(t)}_{2},\ldots, \hat{f}^{(t)}_{t-1})^{\prime}$. The out-of-sample prediction of $y_{it}$ is then computed as $x_{it}^{\prime}\hat{a}_{it} - x_{it}^{\prime}\hat{B}_{it}\hat{\lambda}_t$, where $\hat{\lambda}_{t} = \sum_{s\leq t-1}\hat{f}^{(t)}_s/(t-1)$. Analogously, we define  three types of out-of-sample predictive $R^2$'s by replacing $\hat{a}_i$, $\hat{B}_i$ and $\hat{f}_t$ with $\hat{a}_{it}$, $\hat{B}_{it}$ and $\hat{\lambda}_t$:
\begin{align}
R^{2}_{O} & = 1-\frac{\sum_{i, t\geq 300}(y_{it}- x_{it}^{\prime}\hat{a}_{it} - x_{it}^{\prime}\hat{B}_{it}\hat{\lambda}_t)^2}{\sum_{i, t\geq 300} y_{it}^{2}}, \label{Eqn: R21predictive}\\
R^2_{T,N,O} & = 1 - \frac{1}{N} \sum_{i} \frac{\sum_{t\geq 300}(y_{it}- x_{it}^{\prime}\hat{a}_{it} - x_{it}^{\prime}\hat{B}_{it}\hat{\lambda}_t)^2}{\sum_{t\geq 300}y_{it}^{2}},\label{Eqn: R22predictive}\\
R^2_{N,T,O} & = 1 - \frac{1}{T-299} \sum_{t\geq 300} \frac{\sum_{i}(y_{it}- x_{it}^{\prime}\hat{a}_{it} - x_{it}^{\prime}\hat{B}_{it}\hat{\lambda}_t)^2}{\sum_{i}y_{it}^{2}}.\label{Eqn: R23predictive}
\end{align}

The results depicted in Figure \ref{Fig: Emprical1} yield several key observations. Firstly, the in-sample $R^2$ values of our methods (S1, S2, S3, S4, S5, and S6) exhibit an increasing trend as the number of factors $K$ rises, while the out-of-sample $R^2$ metrics remain unaffected by changes in $K$. This constancy arises from the fact that $\hat{\lambda} = \sum_{t\leq T}\hat{f}_t/T = \hat{F}^{\prime}1_{T}/T = \hat{B}^{\prime}\hat{\Pi}1_{T}/(NT)$, $\hat{a} + \hat{B}\hat{\lambda} = \hat{\Pi}1_{T}/T$, rendering the out-of-sample predictions of $y_{it}$ independent of $K$. Secondly, among the linear models (S1, S2, and S3), S1 consistently outperforms others in terms of in-sample $R^2$ values across all tested values of $K$. Conversely, S3 emerges as the top performer in out-of-sample $R^2$ metrics for all configurations. This suggests that enforcing  homogeneity of $a_i$ and $B_i$ across $i$ may improve the model's out-of-sample predictability, despite potentially compromising the in-sample fit. Similarly, for the spline models (S4, S5, and S6), enforcing homgogeneity of $a_i$ and $B_i$ across $i$ yields improvements in out-of-sample predictability. Thirdly, S5 and S6 demonstrate superior out-of-sample performance compared to S2 and S3, respectively. This underscores the potential benefits of incorporating spline transformations of characteristics, emphasizing the significance of capturing nonlinear relationships. Lastly, the importance of nonlinearity is also observed for the regressed-PCA method; R2 has larger out-of-sample $R^2$ values than R1. However, S3 and S6 exhibit better both in-sample and out-of-sample performance than R1 and R2, respectively. This implies that our method outperforms the regressed-PCA method. In conclusion, while S1 exhibits the most favorable in-sample performance, S6 stands out for its superior out-of-sample predictive capabilities.

\section{Concluding Remarks}\label{Sec8}
In this paper, we introduced a nuclear norm regularized estimation approach for high-dimensional conditional factor models and established large sample properties of the estimators. Our method provides a unified framework for estimating various conditional factor models, facilitating the derivation of new asymptotic results while addressing the limitations of existing methods, which are often model-specific or restrictive. We applied this method to analyze the cross section of individual US stock returns, uncovering potential improvements in out-of-sample performance by enforcing homogeneity of $a_i$ and $B_i$ across $i$. Our results also show that the proposed method outperforms existing alternatives.

In asset pricing, addressing key inference problems\textemdash such as testing for zero pricing errors and conducting specification tests for risk exposure functions\textemdash is crucial for evaluating and comparing factor models. Previous studies, including \citet{XiaYuan_InfMatrix_2019}, \citet{Chenetal_Uncertainty_2019}, and  \citet{Chernozhukovetal_InferenceLowRank_2021}, have investigated debiasing techniques in trace linear regression models with $p=1$ and $a_i=0$, with applications to matrix completion, PCA with missing data, and heterogeneous treatment effects. However, these methods are not applicable to our framework, which accommodates large $p$ and $a_i\neq 0$, as is often the case in asset pricing. Developing a general inferential method within this framework presents an intriguing avenue for future research.



\spacingset{1.75}

\addcontentsline{toc}{section}{References}
\putbib
\end{bibunit}

\pgfplotstableread{
c	d1	d2	d3	d4	d5	d6	d10	d20	d30	d40	d50	d60
0	43.14426637	43.33627923	43.50645384	41.59508512	41.76457064	41.97837082	4.169800373	3.996109432	3.850109822	1.821119546	1.686205982	1.594316174
0.05	32.08304342	32.03894122	32.06602237	29.30623291	28.95795319	29.00921376	4.169800373	3.996109432	3.850109822	1.821119546	1.686205982	1.594316174
0.1	23.54430716	23.36835109	23.36198416	20.10658389	19.39373486	19.32940007	4.169800373	3.996109432	3.850109822	1.821119546	1.686205982	1.594316174
0.2	12.23444796	11.98405134	11.9600371	8.540870773	7.7832915	7.672400098	4.169800373	3.996109432	3.850109822	1.821119546	1.686205982	1.594316174
0.3	6.595638549	6.415623745	6.394547331	3.49027754	3.068663151	2.992567961	4.169800373	3.996109432	3.850109822	1.821119546	1.686205982	1.594316174
0.4	4.395456642	4.232468312	4.179969769	1.928527141	1.73600281	1.659238317	4.169800373	3.996109432	3.850109822	1.821119546	1.686205982	1.594316174
0.5	4.059342044	3.845888286	3.73065423	1.80548605	1.680349044	1.564735503	4.169800373	3.996109432	3.850109822	1.821119546	1.686205982	1.594316174
0.6	4.639974264	4.49352774	4.390555511	2.427700319	2.35902577	2.207595137	4.169800373	3.996109432	3.850109822	1.821119546	1.686205982	1.594316174
0.7	5.390295758	5.250412363	5.115788629	2.804322671	2.608810821	2.519225758	4.169800373	3.996109432	3.850109822	1.821119546	1.686205982	1.594316174
0.8	5.778735496	5.546078172	5.365294694	2.653498884	2.361544595	2.238620608	4.169800373	3.996109432	3.850109822	1.821119546	1.686205982	1.594316174
0.9	5.729392572	5.442801024	5.289179373	2.366988798	2.176505816	2.187699483	4.169800373	3.996109432	3.850109822	1.821119546	1.686205982	1.594316174
1	5.565479792	5.356919464	5.407878306	2.306518885	2.395930596	2.601045541	4.169800373	3.996109432	3.850109822	1.821119546	1.686205982	1.594316174
1.5	7.71740683	8.151418736	8.582979736	3.929517516	4.16301821	4.405527265	4.169800373	3.996109432	3.850109822	1.821119546	1.686205982	1.594316174
2	10.73069694	11.37161622	12.00254787	5.752482952	6.111023504	6.498363794	4.169800373	3.996109432	3.850109822	1.821119546	1.686205982	1.594316174
}\dgpa

\begin{figure}[htbp]
\centering
\begin{subfigure}[b]{0.32\textwidth}
\centering
\resizebox{\linewidth}{3cm}{
\begin{tikzpicture}
\begin{axis}[
    legend style={draw=none},
    grid = minor,
    xmax=2,xmin=0,
    ymax=40,ymin=0,
    xtick={0,1,2},
    ytick={20,40},
    title={$N=500,T=250$},
    tick label style={/pgf/number format/fixed},
legend style={at={(0.2,0.9)},anchor=north,
    row sep = 3pt}]
\addplot[smooth,tension=0.5,color=black, line width=0.75pt,dashdotted] table[x = c,y=d1] from \dgpa;
\addplot[smooth,tension=0.5,no markers, color=blue, line width=0.75pt] table[x = c,y=d10] from \dgpa;
\legend{\footnotesize Fixed $c$, \footnotesize CV}
\end{axis}
\end{tikzpicture}}
\end{subfigure}
\begin{subfigure}[b]{0.32\textwidth}
\centering
\resizebox{\linewidth}{3cm}{
\begin{tikzpicture}
\begin{axis}[
    legend style={draw=none},
    grid = minor,
    xmax=2,xmin=0,
    ymax=40,ymin=0,
    xtick={0,1,2},
    ytick={20,40},
    title={$N=1000,T=250$},
    tick label style={/pgf/number format/fixed},
legend style={at={(0.2,0.9)},anchor=north,
    row sep = 3pt}]
\addplot[smooth,tension=0.5,color=black, line width=0.75pt,dashdotted] table[x = c,y=d2] from \dgpa;
\addplot[smooth,tension=0.5,no markers, color=blue, line width=0.75pt] table[x = c,y=d20] from \dgpa;
\legend{\footnotesize Fixed $c$, \footnotesize CV}
\end{axis}
\end{tikzpicture}}
\end{subfigure}
\begin{subfigure}[b]{0.32\textwidth}
\centering
\resizebox{\linewidth}{3cm}{
\begin{tikzpicture}
\begin{axis}[
    legend style={draw=none},
    grid = minor,
    xmax=2,xmin=0,
    ymax=40,ymin=0,
    xtick={0,1,2},
    ytick={20,40},
    title={$N=2000,T=250$},
    tick label style={/pgf/number format/fixed},
legend style={at={(0.2,0.9)},anchor=north,
    row sep = 3pt}]
\addplot[smooth,tension=0.5,color=black, line width=0.75pt,dashdotted] table[x = c,y=d3] from \dgpa;
\addplot[smooth,tension=0.5,no markers, color=blue, line width=0.75pt] table[x = c,y=d30] from \dgpa;
\legend{\footnotesize Fixed $c$, \footnotesize CV}
\end{axis}
\end{tikzpicture}}
\end{subfigure}

\begin{subfigure}[b]{0.32\textwidth}
\centering
\resizebox{\linewidth}{3cm}{
\begin{tikzpicture}
\begin{axis}[
    legend style={draw=none},
    grid = minor,
    xmax=2,xmin=0,
    ymax=40,ymin=0,
    xtick={0,1,2},
    ytick={20,40},
    title={$N=500,T=500$},
    tick label style={/pgf/number format/fixed},
legend style={at={(0.2,0.9)},anchor=north,
    row sep = 3pt}]
\addplot[smooth,tension=0.5,color=black, line width=0.75pt,dashdotted] table[x = c,y=d4] from \dgpa;
\addplot[smooth,tension=0.5,no markers, color=blue, line width=0.75pt] table[x = c,y=d40] from \dgpa;
\legend{\footnotesize Fixed $c$, \footnotesize CV}
\end{axis}
\end{tikzpicture}}
\end{subfigure}
\begin{subfigure}[b]{0.32\textwidth}
\centering
\resizebox{\linewidth}{3cm}{
\begin{tikzpicture}
\begin{axis}[
    legend style={draw=none},
    grid = minor,
    xmax=2,xmin=0,
    ymax=40,ymin=0,
    xtick={0,1,2},
    ytick={20,40},
    title={$N=1000,T=500$},
    tick label style={/pgf/number format/fixed},
legend style={at={(0.2,0.9)},anchor=north,
    row sep = 3pt}]
\addplot[smooth,tension=0.5,color=black, line width=0.75pt,dashdotted] table[x = c,y=d5] from \dgpa;
\addplot[smooth,tension=0.5,no markers, color=blue, line width=0.75pt] table[x = c,y=d50] from \dgpa;
\legend{\footnotesize Fixed $c$, \footnotesize CV}
\end{axis}
\end{tikzpicture}}
\end{subfigure}
\begin{subfigure}[b]{0.32\textwidth}
\centering
\resizebox{\linewidth}{3cm}{
\begin{tikzpicture}
\begin{axis}[
    legend style={draw=none},
    grid = minor,
    xmax=2,xmin=0,
    ymax=40,ymin=0,
    xtick={0,1,2},
    ytick={20,40},
    title={$N=2000,T=500$},
    tick label style={/pgf/number format/fixed},
legend style={at={(0.2,0.9)},anchor=north,
    row sep = 3pt}]
\addplot[smooth,tension=0.5,color=black, line width=0.75pt,dashdotted] table[x = c,y=d6] from \dgpa;
\addplot[smooth,tension=0.5,no markers, color=blue, line width=0.75pt] table[x = c,y=d60] from \dgpa;
\legend{\footnotesize Fixed $c$, \footnotesize CV}
\end{axis}
\end{tikzpicture}}
\end{subfigure}
\caption{Mean square errors of $\hat{\Pi}$ when using fixed $c$ and CV: DGP1}\label{Fig: DGP1}
\end{figure}

\pgfplotstableread{
c	d1	d2	d3	d4	d5	d6  d10	d20	d30	d40	d50	d60
0	24.55839437	24.16194186	23.96283721	23.83049028	23.47120882	23.28934132	0.204140592	0.150309638	0.150082636	0.133793707	0.109463196	0.07745866
0.05	3.571337437	3.578098296	3.576812316	3.502824797	3.478312025	3.482357528	0.204140592	0.150309638	0.150082636	0.133793707	0.109463196	0.07745866
0.1	2.970929389	3.000093571	3.010431768	2.942819685	2.922942773	2.935466205	0.204140592	0.150309638	0.150082636	0.133793707	0.109463196	0.07745866
0.2	2.142551935	2.128320948	2.11247134	2.097298787	2.037780035	2.022078369	0.204140592	0.150309638	0.150082636	0.133793707	0.109463196	0.07745866
0.3	1.489770499	1.434977475	1.394806768	1.446744207	1.361641069	1.310895211	0.204140592	0.150309638	0.150082636	0.133793707	0.109463196	0.07745866
0.4	1.001448585	0.915966195	0.850249751	0.962615026	0.866175217	0.789574242	0.204140592	0.150309638	0.150082636	0.133793707	0.109463196	0.07745866
0.5	0.653469242	0.556268875	0.475153265	0.61705589	0.52233315	0.439336609	0.204140592	0.150309638	0.150082636	0.133793707	0.109463196	0.07745866
0.6	0.420501601	0.328901216	0.254188009	0.38371882	0.300425231	0.227704386	0.204140592	0.150309638	0.150082636	0.133793707	0.109463196	0.07745866
0.7	0.27854637	0.20363607	0.150144267	0.23857852	0.171805881	0.118650583	0.204140592	0.150309638	0.150082636	0.133793707	0.109463196	0.07745866
0.8	0.204489441	0.15025438	0.120881442	0.159256028	0.109669903	0.077293761	0.204140592	0.150309638	0.150082636	0.133793707	0.109463196	0.07745866
0.9	0.177889305	0.141664271	0.128880179	0.125945557	0.090375619	0.073115778	0.204140592	0.150309638	0.150082636	0.133793707	0.109463196	0.07745866
1	0.181336614	0.155644406	0.147303773	0.1217567	0.094184922	0.08247404	0.204140592	0.150309638	0.150082636	0.133793707	0.109463196	0.07745866
1.5	0.315932678	0.281357944	0.270269968	0.209020799	0.169102144	0.150959432	0.204140592	0.150309638	0.150082636	0.133793707	0.109463196	0.07745866
2	0.509283795	0.457452371	0.442318719	0.337662003	0.274968678	0.247084728	0.204140592	0.150309638	0.150082636	0.133793707	0.109463196	0.07745866
}\dgpb


\begin{figure}[htbp]
\centering
\begin{subfigure}[b]{0.32\textwidth}
\centering
\resizebox{\linewidth}{3cm}{
\begin{tikzpicture}
\begin{axis}[
    grid = minor,
    xmax=2,xmin=0,
    ymax=10,ymin=0,
    xtick={0,1,2},
    ytick={5,10},
    title={$N=500,T=250$},
    tick label style={/pgf/number format/fixed},
    legend style={draw=none, at={(0.2,0.9)},anchor=north,
    row sep = 3pt}]
\addplot[smooth,tension=0.1,color=black, line width=0.75pt,dashdotted] table[x = c,y=d1] from \dgpb;
\addplot[smooth,tension=0.1,no markers, color=blue, line width=0.75pt] table[x = c,y=d10] from \dgpb;
\legend{\footnotesize Fixed $c$, \footnotesize CV}
\end{axis}
\end{tikzpicture}}
\end{subfigure}
\begin{subfigure}[b]{0.32\textwidth}
\centering
\resizebox{\linewidth}{3cm}{
\begin{tikzpicture}
\begin{axis}[
    grid = minor,
    xmax=2,xmin=0,
    ymax=10,ymin=0,
    xtick={0,1,2},
    ytick={5,10},
    title={$N=1000,T=250$},
    tick label style={/pgf/number format/fixed},
    legend style={draw=none, at={(0.2,0.9)},anchor=north,
    row sep = 3pt}]
\addplot[smooth,tension=0.1,color=black, line width=0.75pt,dashdotted] table[x = c,y=d2] from \dgpb;
\addplot[smooth,tension=0.1,no markers, color=blue, line width=0.75pt] table[x = c,y=d20] from \dgpb;
\legend{\footnotesize Fixed $c$, \footnotesize CV}
\end{axis}
\end{tikzpicture}}
\end{subfigure}
\begin{subfigure}[b]{0.32\textwidth}
\centering
\resizebox{\linewidth}{3cm}{
\begin{tikzpicture}
\begin{axis}[
    grid = minor,
    xmax=2,xmin=0,
    ymax=10,ymin=0,
    xtick={0,1,2},
    ytick={5,10},
    title={$N=2000,T=250$},
    tick label style={/pgf/number format/fixed},
    legend style={draw=none, at={(0.2,0.9)},anchor=north,
    row sep = 3pt}]
\addplot[smooth,tension=0.1,color=black, line width=0.75pt,dashdotted] table[x = c,y=d3] from \dgpb;
\addplot[smooth,tension=0.1,no markers, color=blue, line width=0.75pt] table[x = c,y=d30] from \dgpb;
\legend{\footnotesize Fixed $c$, \footnotesize CV}
\end{axis}
\end{tikzpicture}}
\end{subfigure}

\begin{subfigure}[b]{0.32\textwidth}
\centering
\resizebox{\linewidth}{3cm}{
\begin{tikzpicture}
\begin{axis}[
    grid = minor,
    xmax=2,xmin=0,
    ymax=10,ymin=0,
    xtick={0,1,2},
    ytick={5,10},
    title={$N=500,T=500$},
    tick label style={/pgf/number format/fixed},
    legend style={draw=none, at={(0.2,0.9)},anchor=north,
    row sep = 3pt}]
\addplot[smooth,tension=0.1,color=black, line width=0.75pt,dashdotted] table[x = c,y=d4] from \dgpb;
\addplot[smooth,tension=0.1,no markers, color=blue, line width=0.75pt] table[x = c,y=d40] from \dgpb;
\legend{\footnotesize Fixed $c$, \footnotesize CV}
\end{axis}
\end{tikzpicture}}
\end{subfigure}
\begin{subfigure}[b]{0.32\textwidth}
\centering
\resizebox{\linewidth}{3cm}{
\begin{tikzpicture}
\begin{axis}[
    grid = minor,
    xmax=2,xmin=0,
    ymax=10,ymin=0,
    xtick={0,1,2},
    ytick={5,10},
    title={$N=1000,T=500$},
    tick label style={/pgf/number format/fixed},
    legend style={draw=none, at={(0.2,0.9)},anchor=north,
    row sep = 3pt}]
\addplot[smooth,tension=0.1,color=black, line width=0.75pt,dashdotted] table[x = c,y=d5] from \dgpb;
\addplot[smooth,tension=0.1,no markers, color=blue, line width=0.75pt] table[x = c,y=d50] from \dgpb;
\legend{\footnotesize Fixed $c$, \footnotesize CV}
\end{axis}
\end{tikzpicture}}
\end{subfigure}
\begin{subfigure}[b]{0.32\textwidth}
\centering
\resizebox{\linewidth}{3cm}{
\begin{tikzpicture}
\begin{axis}[
    grid = minor,
    xmax=2,xmin=0,
    ymax=10,ymin=0,
    xtick={0,1,2},
    ytick={5,10},
    title={$N=2000,T=500$},
    tick label style={/pgf/number format/fixed},
    legend style={draw=none, at={(0.2,0.9)},anchor=north,
    row sep = 3pt}]
\addplot[smooth,tension=0.1,color=black, line width=0.75pt,dashdotted] table[x = c,y=d6] from \dgpb;
\addplot[smooth,tension=0.1,no markers, color=blue, line width=0.75pt] table[x = c,y=d60] from \dgpb;
\legend{\footnotesize Fixed $c$, \footnotesize CV}
\end{axis}
\end{tikzpicture}}
\end{subfigure}
\caption{Mean square errors of $(\hat{\Pi}^{\diamond\prime}, \sqrt{N}\hat{\Pi}^{\ast\prime})$ when using fixed $c$ and CV: DGP2}\label{Fig: DGP2}
\end{figure}

\pgfplotstableread{
c	d1	d2	d3	d4	d5	d6	d10	d20	d30	d40	d50	d60
0	0.315186412	0.151644977	0.074242933	0.315283062	0.151342761	0.074203126	0.053402803	0.02746442	0.01344317	0.048648191	0.025929072	0.012648708
0.05	0.274626226	0.132247697	0.06456282	0.275495043	0.132342158	0.064718108	0.053402803	0.02746442	0.01344317	0.048648191	0.025929072	0.012648708
0.1	0.238040197	0.114630754	0.055761851	0.239448271	0.115010093	0.056056517	0.053402803	0.02746442	0.01344317	0.048648191	0.025929072	0.012648708
0.2	0.176131736	0.084579088	0.040755953	0.17797218	0.085208616	0.0411667	0.053402803	0.02746442	0.01344317	0.048648191	0.025929072	0.012648708
0.3	0.128212659	0.061191003	0.029141246	0.129704323	0.061663814	0.029457328	0.053402803	0.02746442	0.01344317	0.048648191	0.025929072	0.012648708
0.4	0.093138731	0.044170528	0.020841481	0.093585177	0.044103773	0.020858001	0.053402803	0.02746442	0.01344317	0.048648191	0.025929072	0.012648708
0.5	0.069865494	0.033221388	0.0157499	0.068656402	0.032279573	0.015298531	0.053402803	0.02746442	0.01344317	0.048648191	0.025929072	0.012648708
0.6	0.056914134	0.027646015	0.013456611	0.054046877	0.02592295	0.01265732	0.053402803	0.02746442	0.01344317	0.048648191	0.025929072	0.012648708
0.7	0.05207068	0.026234379	0.013270325	0.048375435	0.024216932	0.012382856	0.053402803	0.02746442	0.01344317	0.048648191	0.025929072	0.012648708
0.8	0.053097026	0.027734333	0.014472316	0.049233867	0.025705647	0.013595737	0.053402803	0.02746442	0.01344317	0.048648191	0.025929072	0.012648708
0.9	0.058026694	0.031014026	0.01644803	0.054154238	0.028934535	0.01550308	0.053402803	0.02746442	0.01344317	0.048648191	0.025929072	0.012648708
1	0.065243985	0.035206175	0.018792668	0.061072756	0.032884238	0.017706274	0.053402803	0.02746442	0.01344317	0.048648191	0.025929072	0.012648708
1.5	0.116005859	0.063515101	0.034433462	0.108498852	0.05938283	0.032383674	0.053402803	0.02746442	0.01344317	0.048648191	0.025929072	0.012648708
2	0.187507243	0.103322792	0.056400679	0.175321907	0.096647806	0.052983099	0.053402803	0.02746442	0.01344317	0.048648191	0.025929072	0.012648708
}\dgpc


\begin{figure}[htbp]
\centering
\begin{subfigure}[b]{0.32\textwidth}
\centering
\resizebox{\linewidth}{3cm}{
\begin{tikzpicture}
\begin{axis}[
    legend style={draw=none},
    grid = minor,
    xmax=2,xmin=0,
    ymax=0.4,ymin=0,
    xtick={0,1,2},
    ytick={0.2,0.4},
    title={$N=500,T=250$},
    tick label style={/pgf/number format/fixed},
legend style={at={(0.2,0.9)},anchor=north,
    row sep = 3pt}]
\addplot[smooth,tension=0.5,color=black, line width=0.75pt,dashdotted] table[x = c,y=d1] from \dgpc;
\addplot[smooth,tension=0.5,no markers, color=blue, line width=0.75pt] table[x = c,y=d10] from \dgpc;
\legend{\footnotesize Fixed $c$, \footnotesize CV}
\end{axis}
\end{tikzpicture}}
\end{subfigure}
\begin{subfigure}[b]{0.32\textwidth}
\centering
\resizebox{\linewidth}{3cm}{
\begin{tikzpicture}
\begin{axis}[
    legend style={draw=none},
    grid = minor,
    xmax=2,xmin=0,
    ymax=0.2,ymin=0,
    xtick={0,1,2},
    ytick={0.1,0.2},
    title={$N=1000,T=250$},
    tick label style={/pgf/number format/fixed},
legend style={at={(0.2,0.9)},anchor=north,
    row sep = 3pt}]
\addplot[smooth,tension=0.5,color=black, line width=0.75pt,dashdotted] table[x = c,y=d2] from \dgpc;
\addplot[smooth,tension=0.5,no markers, color=blue, line width=0.75pt] table[x = c,y=d20] from \dgpc;
\legend{\footnotesize Fixed $c$, \footnotesize CV}
\end{axis}
\end{tikzpicture}}
\end{subfigure}
\begin{subfigure}[b]{0.32\textwidth}
\centering
\resizebox{\linewidth}{3cm}{
\begin{tikzpicture}
\begin{axis}[
    legend style={draw=none},
    grid = minor,
    xmax=2,xmin=0,
    ymax=0.1,ymin=0,
    xtick={0,1,2},
    ytick={0.05,0.1},
    title={$N=2000,T=250$},
    tick label style={/pgf/number format/fixed},
legend style={at={(0.2,0.9)},anchor=north,
    row sep = 3pt}]
\addplot[smooth,tension=0.5,color=black, line width=0.75pt,dashdotted] table[x = c,y=d3] from \dgpc;
\addplot[smooth,tension=0.5,no markers, color=blue, line width=0.75pt] table[x = c,y=d30] from \dgpc;
\legend{\footnotesize Fixed $c$, \footnotesize CV}
\end{axis}
\end{tikzpicture}}
\end{subfigure}

\begin{subfigure}[b]{0.32\textwidth}
\centering
\resizebox{\linewidth}{3cm}{
\begin{tikzpicture}
\begin{axis}[
    legend style={draw=none},
    grid = minor,
    xmax=2,xmin=0,
    ymax=0.4,ymin=0,
    xtick={0,1,2},
    ytick={0.2,0.4},
    title={$N=500,T=500$},
    tick label style={/pgf/number format/fixed},
legend style={at={(0.2,0.9)},anchor=north,
    row sep = 3pt}]
\addplot[smooth,tension=0.5,color=black, line width=0.75pt,dashdotted] table[x = c,y=d4] from \dgpc;
\addplot[smooth,tension=0.5,no markers, color=blue, line width=0.75pt] table[x = c,y=d40] from \dgpc;
\legend{\footnotesize Fixed $c$, \footnotesize CV}
\end{axis}
\end{tikzpicture}}
\end{subfigure}
\begin{subfigure}[b]{0.32\textwidth}
\centering
\resizebox{\linewidth}{3cm}{
\begin{tikzpicture}
\begin{axis}[
    legend style={draw=none},
    grid = minor,
    xmax=2,xmin=0,
    ymax=0.2,ymin=0,
    xtick={0,1,2},
    ytick={0.1,0.2},
    title={$N=1000,T=500$},
    tick label style={/pgf/number format/fixed},
legend style={at={(0.2,0.9)},anchor=north,
    row sep = 3pt}]
\addplot[smooth,tension=0.5,color=black, line width=0.75pt,dashdotted] table[x = c,y=d5] from \dgpc;
\addplot[smooth,tension=0.5,no markers, color=blue, line width=0.75pt] table[x = c,y=d50] from \dgpc;
\legend{\footnotesize Fixed $c$, \footnotesize CV}
\end{axis}
\end{tikzpicture}}
\end{subfigure}
\begin{subfigure}[b]{0.32\textwidth}
\centering
\resizebox{\linewidth}{3cm}{
\begin{tikzpicture}
\begin{axis}[
    legend style={draw=none},
    grid = minor,
    xmax=2,xmin=0,
    ymax=0.1,ymin=0,
    xtick={0,1,2},
    ytick={0.05,0.1},
    title={$N=2000,T=500$},
    tick label style={/pgf/number format/fixed},
legend style={at={(0.2,0.9)},anchor=north,
    row sep = 3pt}]
\addplot[smooth,tension=0.5,color=black, line width=0.75pt,dashdotted] table[x = c,y=d6] from \dgpc;
\addplot[smooth,tension=0.5,no markers, color=blue, line width=0.75pt] table[x = c,y=d60] from \dgpc;
\legend{\footnotesize Fixed $c$, \footnotesize CV}
\end{axis}
\end{tikzpicture}}
\end{subfigure}
\caption{Mean square errors of $\hat{\Pi}_0$ when using fixed $c$ and CV: DGP3}\label{Fig: DGP3}
\end{figure}


\setlength{\tabcolsep}{18pt}
\begin{table}[htbp]
\centering
\resizebox{0.99\textwidth}{!}{
\begin{threeparttable}
\caption{Mean square errors of $\hat{\Pi}$, $\hat{a}$, $\hat{B}$, and $\hat{F}$, and correct rates of $\hat{K}$: DGP1\tnote{\dag}}\label{Tab: MSEDGP1}
\begin{tabular}{cccccccccc}
\hline\hline
&(N,T)&&$\hat{\Pi}$&$\hat{a}$&$\hat{B}$&$\hat{F}$&&$\hat{K}$&\\
\cline{2-9}
&$(500,250)$    &&4.170&2.295&0.853&0.183&&0.950&\\
&$(1000,250)$   &&3.996&2.233&0.800&0.171&&1.000&\\
&$(2000,250)$   &&3.850&2.188&0.759&0.154&&1.000&\\
\cline{2-9}
&$(500,500)$   &&1.821&1.641&0.243&0.088&&1.000&\\
&$(1000,500)$  &&1.686&1.595&0.222&0.066&&1.000&\\
&$(2000,500)$  &&1.584&1.543&0.210&0.053&&1.000&\\
\hline\hline
\end{tabular}
\begin{tablenotes}
      \small
      \item[\dag] The mean square errors of $\hat{\Pi}$, $\hat{a}$ , $\hat{B}$, and $\hat{F}$ are given by $\sum_{\ell=1}^{200}\|\hat{\Pi}^{(\ell)}-\Pi\|_{F}^2/200NT$, $\sum_{\ell=1}^{200}\|\hat{a}^{(\ell)}-a\|^2/200N$, $\sum_{\ell=1}^{200}\|\hat{B}^{(\ell)}-BH^{(\ell)}\|_{F}^2/200N$ and $\sum_{\ell=1}^{200}\|\hat{F}^{(\ell)}- F(H^{{(\ell)}\prime})^{-1}\|_{F}^2/{200T}$, where $\hat{\Pi}^{(\ell)}$, $\hat{a}^{(\ell)}$, $\hat{B}^{(\ell)}$, and $\hat{F}^{(\ell)}$ are estimates in the $\ell$th simulation replication, and $H^{(\ell)}\equiv (F^{\prime}M_T\hat{F}^{(\ell)})(\hat{F}^{{(\ell)}\prime}M_T\hat{F}^{(\ell)})^{-1}$ is a rotational transformation matrix. The value of $c$ is chosen from $\{0, 0.05, 0.1, 0.2, \ldots, 0.9,1,1.5,2\}$ by using the 5-fold CV method as outlined in Section \ref{Sec3}.
    \end{tablenotes}
\end{threeparttable}
}
\end{table}


\setlength{\tabcolsep}{12pt}
\begin{table}[htbp]
\centering
\resizebox{0.99\textwidth}{!}{
\begin{threeparttable}
\caption{Mean square errors of $\hat{\Pi}^{\diamond}$, $\hat{\Pi}^{\ast}$, $\hat{\mu}$, $\hat\Lambda$, $\hat{\phi}, \hat{\Phi}$, and $\hat{F}$, and correct rates of $\hat{K}$: DGP2\tnote{\dag}}\label{Tab: MSEDGP2}
\begin{tabular}{ccccccccccccc}
\hline\hline
&(N,T)&&$\hat{\Pi}^{\diamond}$&$\hat{\Pi}^{\ast}$&$\hat{\mu}$&$\hat\Lambda$&$\hat{\phi}$&$\hat{\Phi}$&$\hat{F}$&&$\hat{K}$&\\
\cline{2-12}
&$(500,250)$     &&0.108&0.096&0.061&0.005&0.157&0.009&0.038&&1.000&\\
&$(1000,250)$    &&0.077&0.073&0.062&0.005&0.133&0.008&0.028&&1.000&\\
&$(2000,250)$    &&0.095&0.055&0.065&0.005&0.104&0.006&0.020&&1.000&\\
\cline{2-12}
&$(500,500)$    &&0.060&0.074&0.031&0.003&0.109&0.006&0.032&&1.000&\\
&$(1000,500)$   &&0.061&0.048&0.033&0.002&0.076&0.004&0.020&&1.000&\\
&$(2000,500)$   &&0.040&0.038&0.033&0.002&0.065&0.004&0.014&&1.000&\\
\hline\hline
\end{tabular}
\begin{tablenotes}
      \small
      \item[\dag] The mean square errors of $\hat{\Pi}^{\diamond}$, $\hat{\Pi}^{\ast}$, $\hat{\mu}$, $\hat\Lambda$, $\hat{\phi}, \hat{\Phi}$, and $\hat{F}$ are given by $\sum_{\ell=1}^{200}\|\hat{\Pi}^{\diamond(\ell)}-\Pi^{\diamond}\|_{F}^2/200NT$, $\sum_{\ell=1}^{200}\|\hat{\Pi}^{\ast(\ell)}-\Pi^{\ast}\|_{F}^2/200T$,$\sum_{\ell=1}^{200}\|\hat{\mu}^{(\ell)}-\mu\|^2/200N$, $\sum_{\ell=1}^{200}\|\hat{\Lambda}^{(\ell)}-\Lambda H^{(\ell)}\|_{F}^2/200N$, $\sum_{\ell=1}^{200}\|\hat{\phi}^{(\ell)}-\phi\|^2/200$, $\sum_{\ell=1}^{200}\|\hat{\Phi}^{(\ell)}-\Phi H^{(\ell)}\|^2/200$ and $\sum_{\ell=1}^{200}\|\hat{F}^{(\ell)}- F(H^{{(\ell)}\prime})^{-1}\|_{F}^2/{200T}$, where $\hat{\Pi}^{\diamond(\ell)}$, $\hat{\Pi}^{\ast(\ell)}$, $\hat{\mu}^{(\ell)}$, $\hat{\Lambda}^{(\ell)}$, $\hat{\phi}^{(\ell)}$, $\hat{\Phi}^{(\ell)}$, and $\hat{F}^{(\ell)}$ are estimates in the $\ell$th simulation replication, and $H^{(\ell)}\equiv (F^{\prime}M_T\hat{F}^{(\ell)})(\hat{F}^{{(\ell)}\prime}M_T\hat{F}^{(\ell)})^{-1}$ is a rotational transformation matrix. The value of $c$ is chosen from $\{0, 0.05, 0.1, 0.2, \ldots, 0.9,1,1.5,2\}$ by using the 5-fold CV method as outlined in Section \ref{Sec3}.
    \end{tablenotes}
\end{threeparttable}
}
\end{table}

\setlength{\tabcolsep}{18pt}
\begin{table}[htbp]
\centering
\resizebox{0.99\textwidth}{!}{
\begin{threeparttable}
\caption{Mean square errors of $\hat{\Pi}_0$, $\hat{\phi}_0$, $\hat{\Phi}_0$, and $\hat{F}$ ($\times 10^{-2}$), and correct rates of $\hat{K}$: DGP3\tnote{\dag}}\label{Tab: MSEDGP3}
\begin{tabular}{cccccccccc}
\hline\hline
&(N,T)&&$\hat{\Pi}_0$&$\hat{\phi}_0$&$\hat{\Phi}_0$&$\hat{F}$&&$\hat{K}$&\\
\cline{2-9}
&$(500,250)$    &&5.340&4.007&0.271&2.224&&1.000&\\
&$(1000,250)$   &&2.746&1.785&0.121&1.124&&1.000&\\
&$(2000,250)$   &&1.344&0.974&0.065&0.580&&1.000&\\
\cline{2-9}
&$(500,500)$   &&4.865&3.482&0.234&2.187&&1.000&\\
&$(1000,500)$  &&2.594&1.477&0.099&1.064&&1.000&\\
&$(2000,500)$  &&1.265&0.810&0.054&0.559&&1.000&\\
\hline\hline
\end{tabular}
\begin{tablenotes}
      \small
      \item[\dag] The mean square errors of $\hat{\Pi}_0$, $\hat{\phi}_0$ , $\hat{\Phi}_0$, and $\hat{F}$ are given by $\sum_{\ell=1}^{200}\|\hat{\Pi}_0^{(\ell)}-\Pi_0\|_{F}^2/200T$, $\sum_{\ell=1}^{200}\|\hat{\phi}_0^{(\ell)}-\phi\|^2/200$, $\sum_{\ell=1}^{200}\|\hat{\Phi}_0^{(\ell)}-\Phi H^{(\ell)}\|_{F}^2/200$ and $\sum_{\ell=1}^{200}\|\hat{F}^{(\ell)}- F(H^{{(\ell)}\prime})^{-1}\|_{F}^2/{200T}$, where $\hat{\Pi}_0^{(\ell)}$, $\hat{\phi}_0^{(\ell)}$, $\hat{\Phi}_0^{(\ell)}$, and $\hat{F}^{(\ell)}$ are estimates in the $\ell$th simulation replication, and $H^{(\ell)}\equiv (F^{\prime}M_T\hat{F}^{(\ell)})(\hat{F}^{{(\ell)}\prime}M_T\hat{F}^{(\ell)})^{-1}$ is a rotational transformation matrix. The value of $c$ is chosen from $\{0, 0.05, 0.1, 0.2, \ldots, 0.9,1,1.5,2\}$ by using the 5-fold CV method as outlined in Section \ref{Sec3}.
    \end{tablenotes}
\end{threeparttable}
}
\end{table}

\pgfplotstableread{
k	x1	x2	x3	x4	x5	x6	x7	x8	x9	x10	x11	x12
1	19.78397619	16.0496283	22.03292447	19.73344125	15.9938032	21.98069301	0.425273488	0.311054379	0.573209396	0.050534939	0.055825105	0.05223146
2	23.0014627	17.73511348	24.76457163	22.9527652	17.67916671	24.71132158	0.425273488	0.311054379	0.573209396	0.0486975	0.055946769	0.053250046
3	25.0268463	19.28833767	26.56515449	24.97852174	19.23405106	26.51237451	0.425273488	0.311054379	0.573209396	0.048324558	0.054286606	0.052779978
4	26.60842464	20.53462093	27.95754052	26.56039318	20.48145688	27.90540099	0.425273488	0.311054379	0.573209396	0.048031457	0.053164058	0.052139525
5	27.82976075	20.96279622	28.53417785	27.78323551	20.91129621	28.48331217	0.425273488	0.311054379	0.573209396	0.046525246	0.051500006	0.050865675
6	28.88765005	21.47669566	29.37062689	28.84371213	21.42903585	29.32555334	0.425273488	0.311054379	0.573209396	0.043937918	0.047659804	0.045073547
7	29.7444564	22.12044479	30.12003611	29.71221227	22.0872535	30.08730266	0.425273488	0.311054379	0.573209396	0.032244139	0.033191291	0.032733452
8	30.59685879	22.65088575	30.85109124	30.56593078	22.62033581	30.8204353	0.425273488	0.311054379	0.573209396	0.030928004	0.030549947	0.030655941
9	31.39774395	23.13030608	31.48915112	31.36997237	23.10203218	31.46277561	0.425273488	0.311054379	0.573209396	0.027771583	0.028273903	0.026375514
10	32.21667299	23.51575945	31.98281204	32.18937347	23.48885673	31.95690381	0.425273488	0.311054379	0.573209396	0.02729952	0.026902723	0.025908226
}\empa

\pgfplotstableread{
k	x1	x2	x3	x4	x5	x6	x7	x8	x9
1	18.76569193	15.4511116	21.01375672	18.68086381	15.36174128	20.91939556	0.431590505	0.291440272	0.60915766
2	21.52130469	16.83284757	23.43484188	21.43867791	16.74233284	23.33699895	0.431590505	0.291440272	0.60915766
3	23.33659308	18.03321005	24.83143466	23.25387779	17.94206342	24.73359332	0.431590505	0.291440272	0.60915766
4	24.79577503	19.14549108	26.14152398	24.71289728	19.05490673	26.04440973	0.431590505	0.291440272	0.60915766
5	25.96589867	19.72514425	26.85217679	25.88448709	19.63622678	26.75547804	0.431590505	0.291440272	0.60915766
6	26.96752425	20.17916008	27.58643068	26.88682413	20.0923324	27.49223402	0.431590505	0.291440272	0.60915766
7	27.82991203	20.68437029	28.36175757	27.74897961	20.59804742	28.2672491	0.431590505	0.291440272	0.60915766
8	28.66900065	21.04614467	28.84771355	28.58994915	20.96091933	28.75562084	0.431590505	0.291440272	0.60915766
9	29.40747168	21.5512486	29.58116539	29.33651189	21.47932124	29.50298396	0.431590505	0.291440272	0.60915766
10  30.15945774	21.88639719	30.05057626	30.08881386	21.81445376	29.97273641	0.431590505	0.291440272	0.60915766
}\empb


\pgfplotstableread{
k	x1	x2	x3	x4	x5	x6	x7	x8	x9	x10	x11	x12
1	18.20072892	15.26663258	19.16387575	17.82355183	14.79504704	18.7366448	0.834675501	0.556749775	1.07	0.377177084	0.471585541	0.427230945
2	19.47372843	16.14815566	20.48679072	19.08796113	15.67154105	20.04761936	0.834675501	0.556749775	1.07	0.385767298	0.476614612	0.439171365
3	21.12188542	17.6823269	22.84629057	20.7800035	17.19922972	22.41672007	0.834675501	0.556749775	1.07	0.341881917	0.48309718	0.429570499
4	22.54645031	18.7879371	24.6562571	22.21722146	18.3402633	24.25647883	0.834675501	0.556749775	1.07	0.329228846	0.447673799	0.399778265
5	22.88209261	19.15667531	25.0520791	22.67599493	18.89002953	24.78176698	0.834675501	0.556749775	1.07	0.206097677	0.266645782	0.270312118
6	23.34220596	19.59630166	25.6214229	23.22021748	19.45253597	25.50152092	0.834675501	0.556749775	1.07	0.12198848	0.14376569	0.119901987
7	23.66748845	19.94946119	26.04318542	23.54643426	19.80776749	25.92288627	0.834675501	0.556749775	1.07	0.121054191	0.141693703	0.120299146
8	23.91039198	20.20174572	26.27570259	23.88275982	20.16815315	26.23980135	0.834675501	0.556749775	1.07	0.027632156	0.033592571	0.035901235
9	24.06037721	20.3581503	26.44275535	24.03348048	20.32607118	26.40716528	0.834675501	0.556749775	1.07	0.026896735	0.032079122	0.035590068
10	24.22196944	20.53850597	26.61357558	24.19339303	20.50467011	26.5743102	0.834675501	0.556749775	1.07	0.028576409	0.033835869	0.039265382
}\empc

\pgfplotstableread{
k	x1	x2	x3  x7	x8	x9
1	2.4151775	0.270744452	2.11800967	0.794638557	0.462437015	0.978870287
2	3.61323705	0.858619263	3.441763274	0.794638557	0.462437015	0.978870287
3	4.559943135	1.324080283	4.437781153	0.794638557	0.462437015	0.978870287
4	4.765972671	1.656090297	4.623180378	0.794638557	0.462437015	0.978870287
5	8.747233654	4.616366892	8.737732062	0.794638557	0.462437015	0.978870287
6	11.37130751	7.207234276	11.08379845	0.794638557	0.462437015	0.978870287
7	13.40334322	9.618156033	13.06671806	0.794638557	0.462437015	0.978870287
8	16.95898943	13.07758168	16.97741614	0.794638557	0.462437015	0.978870287
9	19.23457836	15.66545377	19.6535871	0.794638557	0.462437015	0.978870287
10  20.03926363	16.49919903	20.35534774	0.794638557	0.462437015	0.978870287
}\empregressedpca

\pgfplotstableread{
k	x1	x2	x3	x4	x5	x6	x7	x8	x9	x10	x11	x12
1	18.76673286	15.30625399	20.95104112	18.73792622	15.2719071	20.91618163	0.404136198	0.304530481	0.577825976	0.028806637	0.034346886	0.034859486
2	21.6242784	16.72190856	23.50696865	21.59610295	16.68812602	23.47142938	0.404136198	0.304530481	0.577825976	0.028175446	0.033782536	0.035539267
3	23.34308699	17.90151103	24.95455185	23.31523509	17.8699816	24.91962685	0.404136198	0.304530481	0.577825976	0.0278519	0.031529431	0.034924996
4	24.59322027	18.79766041	26.07427425	24.56672625	18.76828182	26.04197985	0.404136198	0.304530481	0.577825976	0.02649402	0.02937859	0.032294399
5	25.63147032	19.16174473	26.5829391	25.60611835	19.13341529	26.55128872	0.404136198	0.304530481	0.577825976	0.025351971	0.028329446	0.031650379
6	26.47450801	19.50841809	27.24773126	26.45065155	19.48133058	27.2192703	0.404136198	0.304530481	0.577825976	0.023856461	0.027087509	0.028460957
7	27.21814799	19.96035984	27.88220827	27.19451124	19.93430927	27.85354374	0.404136198	0.304530481	0.577825976	0.023636755	0.026050576	0.02866453
8	27.88940635	20.31018598	28.44764624	27.87457921	20.29377989	28.43109794	0.404136198	0.304530481	0.577825976	0.014827139	0.016406092	0.0165483
9	28.48601845	20.60910367	28.89216548	28.47179514	20.59521179	28.87680388	0.404136198	0.304530481	0.577825976	0.014223309	0.013891884	0.015361596
10	29.0886451	20.88821919	29.35572407	29.07768941	20.87695932	29.34472081	0.404136198	0.304530481	0.577825976	0.010955691	0.011259878	0.011003255
}\empaspline

\pgfplotstableread{
k	x1	x2	x3	x4	x5	x6	x7	x8	x9
1	18.21376464	15.08225972	19.26981639	18.14018907	14.97738242	19.17183983	0.466663562	0.30740805	0.682441868
2	21.39172538	16.8779238	23.21987308	21.32001502	16.79694204	23.12958975	0.466663562	0.30740805	0.682441868
3	23.21036393	18.09507539	24.6618819	23.13834286	18.01638289	24.57154117	0.466663562	0.30740805	0.682441868
4	24.60064279	19.09498298	25.87211299	24.52912277	19.01716806	25.78375924	0.466663562	0.30740805	0.682441868
5	25.77610958	19.68258797	26.58531642	25.7068148	19.60723891	26.49836989	0.466663562	0.30740805	0.682441868
6	26.8378879	20.15756436	27.44367613	26.77045494	20.0850586	27.36156266	0.466663562	0.30740805	0.682441868
7	27.75956697	20.68667443	28.41417755	27.70003269	20.62507016	28.34869038	0.466663562	0.30740805	0.682441868
8	28.60887668	21.04634712	28.93592615	28.55213651	20.98544788	28.87342997	0.466663562	0.30740805	0.682441868
9	29.37462613	21.40000969	29.43111884	29.31848225	21.33905417	29.36910247	0.466663562	0.30740805	0.682441868
10  30.11243206	21.9021766	30.1198585	30.05793753	21.84589042	30.06123271	0.466663562	0.30740805	0.682441868
}\empbspline


\pgfplotstableread{
k	x1	x2	x3	x4	x5	x6	x7	x8	x9	x10	x11	x12
1	13.10776141	9.670879077	12.33359277	12.69000145	9.123067886	11.87889585	0.864456072	0.604413358	1.080307192	0.417759955	0.547811191	0.454696914
2	15.01725349	11.70886333	13.86303032	14.59759032	11.17794811	13.39901714	0.864456072	0.604413358	1.080307192	0.419663172	0.530915224	0.464013174
3	17.81232534	14.6544419	17.69184302	17.42751736	14.1291686	17.21712697	0.864456072	0.604413358	1.080307192	0.384807981	0.525273295	0.474716052
4	18.60071732	15.08747128	18.68605936	18.21071773	14.70728754	18.25149473	0.864456072	0.604413358	1.080307192	0.389999586	0.380183732	0.434564631
5	20.44039553	16.53566919	21.32885478	20.16753269	16.20164243	21.02411215	0.864456072	0.604413358	1.080307192	0.272862839	0.334026752	0.30474263
6	20.87068851	16.99828725	21.72060871	20.74406607	16.8352111	21.55560507	0.864456072	0.604413358	1.080307192	0.126622439	0.163076146	0.165003644
7	21.90544273	18.04922017	23.28705198	21.69665669	17.79859404	22.98446656	0.864456072	0.604413358	1.080307192	0.208786038	0.250626122	0.302585418
8	22.72418108	18.75859548	24.22033442	22.5712626	18.57500396	24.00162082	0.864456072	0.604413358	1.080307192	0.152918472	0.18359152	0.218713602
9	23.04429504	19.11544664	24.58025596	22.8526714	18.89543952	24.27495593	0.864456072	0.604413358	1.080307192	0.191623645	0.220007119	0.305300034
10	23.44119958	19.64001011	25.17929301	23.28802816	19.43136666	24.93887761	0.864456072	0.604413358	1.080307192	0.153171413	0.208643453	0.240415398
}\empcspline

\pgfplotstableread{
k	x1	x2	x3  x7	x8	x9
1	3.730955782	1.457065701	3.445342152	0.841677528	0.514120961	1.016947396
2	9.803144174	6.509265921	9.005506164	0.841677528	0.514120961	1.016947396
3	11.3797279	7.958427723	10.11507774	0.841677528	0.514120961	1.016947396
4	15.31450372	11.87442339	14.35138896	0.841677528	0.514120961	1.016947396
5	16.14869052	12.70308811	15.42907235	0.841677528	0.514120961	1.016947396
6	16.54821647	13.03980857	15.86031005	0.841677528	0.514120961	1.016947396
7	17.96624755	14.09265656	17.85728909	0.841677528	0.514120961	1.016947396
8	18.28952	14.42483044	18.17414921	0.841677528	0.514120961	1.016947396
9	19.18600499	15.32249453	19.28049809	0.841677528	0.514120961	1.016947396
10	19.50205769	15.62131374	19.54298715	0.841677528	0.514120961	1.016947396
}\empregressedpcaspline

\begin{figure}[H]
\centering
\begin{subfigure}[b]{0.32\textwidth}
\centering
\resizebox{\linewidth}{!}{
\pgfplotsset{title style={at={(0.5,0.91)}}}
\begin{tikzpicture}
\begin{axis}[
        ybar,
    height=13cm,
    width=13cm,
    enlarge y limits=false,
    axis lines*=left,
    xmax=10,xmin=1,
    ymax=40,ymin=0,
    xtick={1,2,3,4,5,6,7,8,9,10},
    ytick={0,20,40},
        title={$R^{2}$},
     legend style={at={(0.51,0.9)},
        anchor=north,legend columns=-1,
        /tikz/every even column/.append style={column sep=0.5cm}
        },
    ]
\addplot[smooth,tension=0.3, mark=star,color=blue, line width=1.2pt] table[x = k,y=x1] from \empa;
\addplot[smooth,tension=0.3, mark=square,color=red, line width=1.2pt] table[x = k,y=x1] from \empb;
\addplot[smooth,tension=0.3, mark=triangle,color=black, line width=1.2pt] table[x = k,y=x1] from \empc;
\addplot[smooth,tension=0.3, mark=star,color=blue, line width=1.2pt,dashdotted] table[x = k,y=x1] from \empaspline;
\addplot[smooth,tension=0.3, mark=square,color=red, line width=1.2pt,dashdotted] table[x = k,y=x1] from \empbspline;
\addplot[smooth,tension=0.3, mark=triangle,color=black, line width=1.2pt,dashdotted] table[x = k,y=x1] from \empcspline;
\addplot[smooth,tension=0.3, mark=o,color=green, line width=1.2pt] table[x = k,y=x1] from \empregressedpca;
\addplot[smooth,tension=0.3, mark=o,color=green, line width=1.2pt,dashdotted] table[x = k,y=x1] from \empregressedpcaspline;
\legend{S1,S2,S3,S4,S5,S6,R1,R2}
  \end{axis}
\end{tikzpicture}}
\end{subfigure}
\begin{subfigure}[b]{0.32\textwidth}
\centering
\resizebox{\linewidth}{!}{
\pgfplotsset{title style={at={(0.5,0.91)}}}
\begin{tikzpicture}
\begin{axis}[
        ybar,
    height=13cm,
    width=13cm,
    enlarge y limits=false,
    axis lines*=left,
    xmax=10,xmin=1,
    ymax=40,ymin=0,
    xtick={1,2,3,4,5,6,7,8,9,10},
    ytick={0,20,40},
        title={$R^{2}_{T,N}$},
     legend style={at={(0.51,0.9)},
        anchor=north,legend columns=-1,
        /tikz/every even column/.append style={column sep=0.5cm}
        },
    ]
\addplot[smooth,tension=0.3, mark=star,color=blue, line width=1.2pt] table[x = k,y=x3] from \empa;
\addplot[smooth,tension=0.3, mark=square,color=red, line width=1.2pt] table[x = k,y=x3] from \empb;
\addplot[smooth,tension=0.3, mark=triangle,color=black, line width=1.2pt] table[x = k,y=x3] from \empc;
\addplot[smooth,tension=0.3, mark=star,color=blue, line width=1.2pt,dashdotted] table[x = k,y=x3] from \empaspline;
\addplot[smooth,tension=0.3, mark=square,color=red, line width=1.2pt,dashdotted] table[x = k,y=x3] from \empbspline;
\addplot[smooth,tension=0.3, mark=triangle,color=black, line width=1.2pt,dashdotted] table[x = k,y=x3] from \empcspline;
\addplot[smooth,tension=0.3, mark=o,color=green, line width=1.2pt] table[x = k,y=x3] from \empregressedpca;
\addplot[smooth,tension=0.3, mark=o,color=green, line width=1.2pt,dashdotted] table[x = k,y=x3] from \empregressedpcaspline;
\legend{S1,S2,S3,S4,S5,S6,R1,R2}
  \end{axis}
\end{tikzpicture}}
\end{subfigure}
\begin{subfigure}[b]{0.32\textwidth}
\centering
\resizebox{\linewidth}{!}{
\pgfplotsset{title style={at={(0.5,0.91)}}}
\begin{tikzpicture}
\begin{axis}[
        ybar,
    height=13cm,
    width=13cm,
    enlarge y limits=false,
    axis lines*=left,
    xmax=10,xmin=1,
    ymax=40,ymin=0,
    xtick={1,2,3,4,5,6,7,8,9,10},
    ytick={0,20,40},
        title={$R^{2}_{N,T}$},
     legend style={at={(0.51,0.9)},
        anchor=north,legend columns=-1,
        /tikz/every even column/.append style={column sep=0.5cm}
        },
    ]
\addplot[smooth,tension=0.3, mark=star,color=blue, line width=1.2pt] table[x = k,y=x2] from \empa;
\addplot[smooth,tension=0.3, mark=square,color=red, line width=1.2pt] table[x = k,y=x2] from \empb;
\addplot[smooth,tension=0.3, mark=triangle,color=black, line width=1.2pt] table[x = k,y=x2] from \empc;
\addplot[smooth,tension=0.3, mark=star,color=blue, line width=1.2pt,dashdotted] table[x = k,y=x2] from \empaspline;
\addplot[smooth,tension=0.3, mark=square,color=red, line width=1.2pt,dashdotted] table[x = k,y=x2] from \empbspline;
\addplot[smooth,tension=0.3, mark=triangle,color=black, line width=1.2pt,dashdotted] table[x = k,y=x2] from \empcspline;
\addplot[smooth,tension=0.3, mark=o,color=green, line width=1.2pt] table[x = k,y=x2] from \empregressedpca;
\addplot[smooth,tension=0.3, mark=o,color=green, line width=1.2pt,dashdotted] table[x = k,y=x2] from \empregressedpcaspline;
\legend{S1,S2,S3,S4,S5,S6,R1,R2}
  \end{axis}
\end{tikzpicture}}
\end{subfigure}


\begin{subfigure}[b]{0.32\textwidth}
\centering
\resizebox{\linewidth}{!}{
\pgfplotsset{title style={at={(0.5,0.91)}}}
\begin{tikzpicture}
\begin{axis}[
        ybar,
    height=13cm,
    width=13cm,
    enlarge y limits=false,
    axis lines*=left,
    xmax=10,xmin=1,
    ymax=1,ymin=0.2,
    xtick={1,2,3,4,5,6,7,8,9,10},
    ytick={0.2,1},
        title={$R^{2,O}$},
     legend style={at={(0.51,0.12)},
        anchor=north,legend columns=-1,
        /tikz/every even column/.append style={column sep=0.5cm}
        },
    ]
\addplot[smooth,tension=0.3, mark=star,color=blue, line width=1.2pt] table[x = k,y=x7] from \empa;
\addplot[smooth,tension=0.3, mark=square,color=red, line width=1.2pt] table[x = k,y=x7] from \empb;
\addplot[smooth,tension=0.3, mark=triangle,color=black, line width=1.2pt] table[x = k,y=x7] from \empc;
\addplot[smooth,tension=0.3, mark=star,color=blue, line width=1.2pt,dashdotted] table[x = k,y=x7] from \empaspline;
\addplot[smooth,tension=0.3, mark=square,color=red, line width=1.2pt,dashdotted] table[x = k,y=x7] from \empbspline;
\addplot[smooth,tension=0.3, mark=triangle,color=black, line width=1.2pt,dashdotted] table[x = k,y=x7] from \empcspline;
\addplot[smooth,tension=0.3, mark=o,color=green, line width=1.2pt] table[x = k,y=x7] from \empregressedpca;
\addplot[smooth,tension=0.3, mark=o,color=green, line width=1.2pt,dashdotted] table[x = k,y=x7] from \empregressedpcaspline;
\legend{S1,S2,S3,S4,S5,S6,R1,R2}
  \end{axis}
\end{tikzpicture}}
\end{subfigure}
\begin{subfigure}[b]{0.32\textwidth}
\centering
\resizebox{\linewidth}{!}{
\pgfplotsset{title style={at={(0.5,0.91)}}}
\begin{tikzpicture}
\begin{axis}[
        ybar,
    height=13cm,
    width=13cm,
    enlarge y limits=false,
    axis lines*=left,
    xmax=10,xmin=1,
    ymax=1.2,ymin=0.2,
    xtick={1,2,3,4,5,6,7,8,9,10},
    ytick={0.2,1.2},
        title={$R^{2}_{T,N,O}$},
     legend style={at={(0.51,0.12)},
        anchor=north,legend columns=-1,
        /tikz/every even column/.append style={column sep=0.5cm}
        },
    ]
\addplot[smooth,tension=0.3, mark=star,color=blue, line width=1.2pt] table[x = k,y=x9] from \empa;
\addplot[smooth,tension=0.3, mark=square,color=red, line width=1.2pt] table[x = k,y=x9] from \empb;
\addplot[smooth,tension=0.3, mark=triangle,color=black, line width=1.2pt] table[x = k,y=x9] from \empc;
\addplot[smooth,tension=0.3, mark=star,color=blue, line width=1.2pt,dashdotted] table[x = k,y=x9] from \empaspline;
\addplot[smooth,tension=0.3, mark=square,color=red, line width=1.2pt,dashdotted] table[x = k,y=x9] from \empbspline;
\addplot[smooth,tension=0.3, mark=triangle,color=black, line width=1.2pt,dashdotted] table[x = k,y=x9] from \empcspline;
\addplot[smooth,tension=0.3, mark=o,color=green, line width=1.2pt] table[x = k,y=x9] from \empregressedpca;
\addplot[smooth,tension=0.3, mark=o,color=green, line width=1.2pt,dashdotted] table[x = k,y=x9] from \empregressedpcaspline;
\legend{S1,S2,S3,S4,S5,S6,R1,R2}
  \end{axis}
\end{tikzpicture}}
\end{subfigure}
\begin{subfigure}[b]{0.32\textwidth}
\centering
\resizebox{\linewidth}{!}{
\pgfplotsset{title style={at={(0.5,0.91)}}}
\begin{tikzpicture}
\begin{axis}[
        ybar,
    height=13cm,
    width=13cm,
    enlarge y limits=false,
    axis lines*=left,
    xmax=10,xmin=1,
    ymax=0.7,ymin=0.2,
    xtick={1,2,3,4,5,6,7,8,9,10},
    ytick={0.2,0.7},
        title={$R^{2}_{N,T,O}$},
     legend style={at={(0.51,0.12)},
        anchor=north,legend columns=-1,
        /tikz/every even column/.append style={column sep=0.5cm}
        },
    ]
\addplot[smooth,tension=0.3, mark=star,color=blue, line width=1.2pt] table[x = k,y=x8] from \empa;
\addplot[smooth,tension=0.3, mark=square,color=red, line width=1.2pt] table[x = k,y=x8] from \empb;
\addplot[smooth,tension=0.3, mark=triangle,color=black, line width=1.2pt] table[x = k,y=x8] from \empc;
\addplot[smooth,tension=0.3, mark=star,color=blue, line width=1.2pt,dashdotted] table[x = k,y=x8] from \empaspline;
\addplot[smooth,tension=0.3, mark=square,color=red, line width=1.2pt,dashdotted] table[x = k,y=x8] from \empbspline;
\addplot[smooth,tension=0.3, mark=triangle,color=black, line width=1.2pt,dashdotted] table[x = k,y=x8] from \empcspline;
\addplot[smooth,tension=0.3, mark=o,color=green, line width=1.2pt] table[x = k,y=x8] from \empregressedpca;
\addplot[smooth,tension=0.3, mark=o,color=green, line width=1.2pt,dashdotted] table[x = k,y=x8] from \empregressedpcaspline;
\legend{S1,S2,S3,S4,S5,S6,R1,R2}
  \end{axis}
\end{tikzpicture}}
\end{subfigure}
\caption{In-sample and out-of-sample $R^2$'s} \label{Fig: Emprical1}
\end{figure}

\clearpage
\newpage


\begin{bibunit}
\def\spacingset#1{\renewcommand{\baselinestretch}
{#1}\small\normalsize} \spacingset{1.8}
\setlength{\abovedisplayskip}{4pt}
\setlength{\belowdisplayskip}{4pt}
\setlength{\abovedisplayshortskip}{4pt}
\setlength{\belowdisplayshortskip}{4pt}