EconBase
← Back to paper

A Unified Framework for Estimation of High-dimensional Conditional Factor Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

120,662 characters · 0 sections · 80 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

A Unified Framework for Estimation of High-dimensional Conditional Factor Models

bibunit\def\spacingset#1{ {#1}} \spacingset{1.8} \pdfbookmark[1]{Title}{title} \if11 \fi \if01 { \begin{center} {\bf A Unified Framework for Estimation of High-dimensional Conditional Factor Models} \end{center} } \fi \begin{abstract} This paper presents a general framework for estimating high-dimensional conditional latent factor models via constrained nuclear norm regularization. We establish large sample properties of the estimators and provide efficient algorithms for their computation. To improve practical applicability, we propose a cross-validation procedure for selecting the regularization parameter. Our framework unifies the estimation of various conditional factor models, enabling the derivation of new asymptotic results while addressing limitations of existing methods, which are often model-specific or restrictive. Empirical analyses of the cross section of individual US stock returns suggest that imposing homogeneity improves the model's out-of-sample predictability, with our new method outperforming existing alternatives. \end{abstract} {\it Keywords:} Constrained nuclear norm regularization, asset pricing, characteristics, macro state variables, factor zoo \section{Introduction} {4pt} {4pt} {4pt} {4pt} In empirical asset pricing, a central question revolves around understanding why different assets yield varying average returns. Conditional factor models offer a comprehensive framework for integrating conditional information to address this inquiry Gagliardinietal_Timevarying_2016, Gagliardinietal_EstimationConditionalFactor_2019. This paper delves into the investigation of a high-dimensional conditional factor model defined as follows: \begin{align} y_{it} = \alpha_{it} + \beta_{it}^{\prime}f_{t} + \varepsilon_{it} with \alpha_{it} = a_i^{\prime}x_{it} and \beta_{it} = B_i^{\prime} x_{it}, i=1,\ldots,N,t=1,\ldots,T. \end{align} Here, $y_{it}$ denotes the excess return of asset $i$ in time period $t$, $f_{t}$ represents a $K\times 1$ vector of unobserved latent factors, $\alpha_{it}$ characterizes a pricing error, $\beta_{it}$ denotes a $K\times 1$ vector of risk exposures, $\varepsilon_{it}$ stands for an error term, $x_{it}$ is a $p\times 1$ vector of pre-specified explanatory variables known at the beginning of time period $t$ (such as constants, sieve transformations of asset characteristics, sieve transformations of macro state variables, and their interactions), and $a_i$ and $B_i$ are $p\times1$ vector and $p\times K$ matrix of unknown coefficients, respectively. This model captures time-variation in the risk exposures (i.e., $B_i^{\prime} x_{it}$) and the pricing error (i.e., $a_i^{\prime} x_{it}$) through their associations with $x_{it}$, while also allowing for the distinction between “risk” and “mispricing” explanations regarding the role of $x_{it}$ in predicting returns, thereby contributing to resolving the ongoing “characteristics versus covariance” debate DanielTitman_Characteristics_1997. Moreover, given that $K$ can be significantly smaller than $p$, the model facilitates the condensation of information from a large dimension of $x_{it}$ into a smaller number of factors, thereby mitigating the so-called “factor zoo” that proliferate in the literature Cochrane_Presidential_2011. However, the estimation of the model encounters at least two challenges: i) $\{f_t\}_{t\leq T}$ are unknown and unobservable; ii) the dimension of the unknown parameters $\{a_i\}_{i\leq N}$, $\{B_i\}_{i\leq N}$, and $\{f_t\}_{t\leq T}$ is high. The model nests various factor models in the literature. Unlike homogeneous versions of conditional factor models Parketal_FactorDynamics_2009,Kellyetal_Characteristics_2019,Chenetal_SeimiparametricFactor_2021, our model allows for heterogeneity of $a_i$ and $B_i$ across assets. Consequently, our model nests classical factor models Ross_APT_1976,ChamberlainRothschild_FactorStuctures_1982 where $x_{it} =1$ and $a_i=0$; semiparametric factor models Connoretal_EfficientFFFactor_2012,Fanetal_ProjectedPCA_2016,Kimetal_Arbitrage_2019 where $x_{it}$ comprises a constant and sieve transformations of asset's time-invariant characteristics, with homogeneity of $a_i$ and $B_i$ across assets for non-constant explanatory variables; and state-varying factor models PelgerXiong_State-varying_2019 where $x_{it}$ encompasses a constant and sieve transformations of macro state variables, with $a_i=0$. Unlike Gagliardinietal_Timevarying_2016, our model does not necessitate observable $f_t$ and accommodates the presence of arbitrage and large $p$, referred to as the unconstrained conditional factor model. We provide a general framework for the estimation of high-dimensional conditional factor models. Specifically, we develop a nuclear norm regularized estimation of the model in (ref) with constraints on $\{a_{i},B_{i}\}_{i\leq N}$. The estimation procedure comprises two steps: first, estimating an $Np\times T$ reduced rank matrix composed of block matrices $\{a_i + B_{i}f_t\}_{i\leq N,t\leq T}$ using nuclear norm regularization under the constraints; then, extracting estimators of $K$, $\{a_i\}_{i\leq N}$, $\{B_i\}_{i\leq N}$, and $\{f_t\}_{t\leq T}$ from the estimated matrix using eigenvalue decomposition. We establish asymptotic properties of the estimators under a restricted strong convexity condition. Our framework allows for both $p\to\infty$ and $K\to\infty$ and may accommodate the presence of missing values, which are prevalent in stock return datasets. The general framework enables the estimation of the aforementioned nested models in a unified manner, overcoming limitations of existing methods that are often model-specific or restrictive.\footnote{For example, ConnorKorajczyk_Performance_1986, StockWatson_PCA_2002, and BaiNg_NumberofFactors_2002 estimate the classical factor model by principal component analysis (PCA), while Fanetal_largecovariance_2013 use a principal orthogonal complement thresholding method. Fanetal_ProjectedPCA_2016 propose a projected-PCA for the semiparametric factor model based on linear sieves, while Fanetal_StructuralDeep_2022 employ neural networks. PelgerXiong_State-varying_2019 estimate the state-varying factor by a local version of PCA based on kernel smoothing. Chenetal_SeimiparametricFactor_2021 develop a regressed-PCA for the homogenous conditional factor model; Parketal_FactorDynamics_2009 propose a Newton-Raphson algorithm; Kellyetal_Characteristics_2019 propose an alternating least squares procedure; Guetal_Autoencoder_2021 propose an autoencoder method. Gagliardinietal_Timevarying_2016 require observable factors for estimating a conditional factor model with no arbitrage.} By tailoring the general theory for each model and providing simple primitive conditions, we make several contributions. First, we offer a novel estimation approach for the homogeneous conditional factor model, allowing $p$ to grow as fast as $N$. Second, we accommodate time-varying characteristics, nonzero pricing errors, and non-noisy intercepts in both pricing errors and risk exposures for the semiparametric factor model. Third, our estimator is capable of consistently estimating the factor space in the state-varying factor model. Fourth, to the best of our knowledge, our paper is the first to provide an estimation method for the unconstrained conditional factor model. To enhance the practical applicability of our estimation procedure, we offer two contributions. Firstly, we present an efficient computing algorithm for finding the constrained nuclear norm regularized estimator of the reduced rank matrix in each model. This contribution is particularly valuable as constrained nuclear norm regularization involves high-dimensional constrained nonsmooth convex minimization, where computational efficiency is crucial for practical implementation. Secondly, we propose a cross-validation (CV) procedure to determine the optimal regularization parameter and validate its effectiveness through a series of Monte Carlo simulations. This contribution is essential because the choice of regularization parameter significantly impacts the estimates, and a systematic method for its selection is necessary to ensure robustness and reliability of the results. Our simulation studies demonstrate that the finite sample performance of our estimators, using the CV-chosen regularization parameter, is satisfactory and encouraging. Our simulations also demonstrate the superiority of our estimators compared to existing ones. We apply our unified framework to analyze the cross section of individual stock returns in the US market. Our analysis reveals that imposing homogeneity of $a_i$ and $B_i$ across assets enhances the model's out-of-sample predictability, with our method outperforming existing approaches. Nuclear norm regularization has been extensively employed for estimating reduced-rank matrices in the statistical literature, primarily focusing on estimating the reduced-rank matrix itself. For instance, NegahbanWainwright_Scaling_2011 investigate unconstrained nuclear norm regularized estimation of trace linear regression models under a restricted strong convexity condition; RohdeTsybakov_Lowrank_2011 examine the same problem under a restricted isometry condition; Fanetal_Generalized_2019 study generalized trace regression models. Our work diverges from these studies in several key aspects. Firstly, we require constrained nuclear norm regularization, which entails extending existing methodologies to accommodate constraints. Secondly, our parameters of interest are $K$, $\{a_i\}_{i\leq N}$, $\{B_i\}_{i\leq N}$, and $\{f_t\}_{t\leq T}$, rather than the reduced-rank matrix. This distinction introduces an additional step in the estimation procedure to estimate these parameters from the reduced-rank matrix. There have been several recent studies in the econometric literature that utilize unconstrained nuclear norm regularization. BaiNg_RankRegularized_2019 use it to enhance estimation of the classical factor model. MoonWeidner_Nuclear_2018 leverage it to improve estimation of panel data models with interactive fixed effects. Chernozhukovetal_LowRank_2018 employ it to estimate panel data models with heterogeneous coefficients. Atheyetal_CompletionCausal_2021 adopt this approach in treatment effect estimation. For more examples, see MoonWeidner_Nuclear_2018. To the best of our knowledge, the use of constrained nuclear norm regularization in estimating conditional factor models has not been studied previously. The literature on the cross section of asset returns is extensive; here we focus on conditional factor models. While our paper emphasizes models with latent factors, a substantial portion of empirical asset pricing research relies on pre-specified observable factors. These factors are often constructed using portfolio-sorting approaches, such as those outlined in FamaFrench_Commonrisk_1993, based on asset characteristics.\footnote{Notable works in this area include studies by Shanken_Intertemporal_1990, FersonHarvey_Variation_1991,FersonHarvey_Conditioning_1999, LettauLudvigson_Consumption_2001, and Gagliardinietal_Timevarying_2016, among others. For a comprehensive review, see Gagliardinietal_EstimationConditionalFactor_2019.} This approach encounters challenges related to the “characteristics versus covariances” debate and the “factor zoo” problem. We contribute to the literature by presenting a unified method for estimating conditional factor models without the need for pre-specified factors, which are well-suited for addressing the debate and problem Kellyetal_Characteristics_2019,Chenetal_SeimiparametricFactor_2021. The structure of the paper is outlined as follows. Section (ref) presents several nested models. Section (ref) outlines the general estimation framework. Section (ref) establishes the asymptotic properties of the estimators. Section (ref) tailors the general theory for each model. Section (ref) presents simulation studies. Section (ref) analyzes the cross section of individual US stock returns. Finally, Section (ref) provides a brief conclusion. The Supplementary Appendix presents proofs of main results, computing algorithms, additional discussions, and additional simulations. \section{Nested Models} Our model in (ref) nests many factor models in the literature. \begin{ex}[Classical Factor Models] The arbitrage pricing theory by Ross_APT_1976 and ChamberlainRothschild_FactorStuctures_1982 gives rise to the following model: \begin{align} y_{it} = \lambda_i^{\prime}f_t + e_{it}, \end{align} where $\lambda_i$ represents an unknown vector of risk exposures and $e_{it}$ is the idiosyncratic component. Our model encompasses (ref) where $x_{it} =1$, $a_{i} = 0$, $B_{i} = \lambda_{i}^{\prime}$, and $\varepsilon_{it} = e_{it}$. \end{ex} \begin{ex}[Semiparametric Factor Models] The model examined by Connoretal_EfficientFFFactor_2012, Fanetal_ProjectedPCA_2016, and Kimetal_Arbitrage_2019 is as follows:\footnote{Fanetal_ProjectedPCA_2016 assume that $\phi(\cdot)=0$ and $\mu_i=0$, Connoretal_EfficientFFFactor_2012 additionally assume that $\Phi(\cdot)$ are univariate functions and $\lambda_{i}=0$, and Kimetal_Arbitrage_2019 assume that $\mu_i=0$ and $\lambda_i=0$. } \begin{align} y_{it} = \phi(z_i)+\mu_i + (\Phi(z_i) + \lambda_{i})^{\prime}f_t + e_{it}, \end{align} where $z_i$ represents a vector of asset's time-invariant characteristics, $\phi(\cdot)$ and $\Phi(\cdot)$ are unknown functions, $\mu_i$ and $\lambda_{i}$ are unknown scalar and vector (intercepts in the pricing errors and risk exposures, which are usually interpreted as the components that are not explained by the characteristics), and $e_{it}$ is the idiosyncratic component. Using sieve methods, $\phi(z_i) = \phi^{\prime}h(z_i) +\delta(z_i)$ and $\Phi(z_i) = \Phi^{\prime}h(z_i) +\Delta(z_i)$, where $h(z_i)$ denotes a vector of basis functions of $z_i$ (excluding constants), $\phi$ and $\Phi$ represent unknown vector and matrix of coefficients, and $\delta(z_i)$ and $\Delta(z_i)$ are negligible sieve approximation errors. Our model nests (ref) where $x_{it} = (1,h(z_i)^{\prime})^{\prime}$, $a_i = (\mu_i,\phi^{\prime})^{\prime}$, $B_i = (\lambda_{i},\Phi^{\prime})^{\prime}$, and $\varepsilon_{it} = e_{it} + \delta(z_i) + \Delta(z_{i})^{\prime}f_t$. Thus, the rows of $a_i$ and $B_i$ corresponding to $h(z_i)$ are homogenous across $i$, meaning that the coefficients for non-constant explanatory variables are homogenous across assets. \end{ex} \begin{ex}[State-varying Factor Models] PelgerXiong_State-varying_2019 examine the following model: \begin{align} y_{it} = \Phi_i(z_{t})^{\prime}f_{t} + e_{it}, \end{align} where $z_{t}$ represents a vector of constant and macro state variables known at the beginning of time period $t$, $\Phi_i(\cdot)$ is a vector of unknown functions, and $e_{it}$ is the idiosyncratic component. Employing sieve methods, $\Phi_i(z_t) = \Phi_i^{\prime}h(z_t) +\Delta_{i}(z_t)$, where $h(z_t)$ denotes a vector of basis functions of $z_t$ (which may include a constant), $\Phi_i$ is an unknown matrix of coefficients, and $\Delta_i(z_t)$ is a vector of negligible sieve approximation errors. Our model encompasses (ref) where $x_{it} = h(z_t)$, $a_i = 0$, $B_i = \Phi_i$, and $\varepsilon_{it} = e_{it} + \Delta_i(z_{t})^{\prime}f_t$. \end{ex} \begin{ex}[Homogeneous Conditional Factor Models] Parketal_FactorDynamics_2009, Kellyetal_Characteristics_2019, and Chenetal_SeimiparametricFactor_2021 propose the following model:\footnote{Kellyetal_Characteristics_2019 assume that $\phi(\cdot)$ and $\Phi(\cdot)$ are linear functions.} \begin{align} y_{it} = \phi_0(z_{it}) + \Phi_0(z_{it})^{\prime}f_{t} + e_{it}, \end{align} where $z_{it}$ represents a vector of constant and asset characteristics known at the beginning of time period $t$, $\phi_0(\cdot)$ and $\Phi_0(\cdot)$ are unknown functions, and $e_{it}$ is the idiosyncratic component. Employing sieve methods, $\phi_0(z_{it}) = \phi_0^{\prime}h(z_{it}) +\delta(z_{it})$ and $\Phi_0(z_{it}) = \Phi_0^{\prime}h(z_{it}) +\Delta(z_{it})$, where $h(z_{it})$ denotes a vector of basis functions of $z_{it}$ (which may include a constant), $\phi_0$ and $\Phi_0$ are unknown vector and matrix of coefficients, and $\delta(z_{it})$ and $\Delta(z_{it})$ are negligible sieve approximation errors. Our model nests (ref) where $x_{it} = h(z_{it})$, $a_i = \phi_0$, $B_i = \Phi_0$, and $\varepsilon_{it} = e_{it} + \delta(z_{it}) + \Delta(z_{it})^{\prime}f_t$. Thus, $a_i$ and $B_i$ are homogenous across $i$, meaning that the coefficients of explanatory variables are homogenous across assets. \end{ex} \begin{ex}[Unconstrained Conditional Factor Models] In the absence of arbitrage opportunities, Gagliardinietal_Timevarying_2016 propose the following model: \begin{align} y_{it} =z_{t}^{\prime}\Psi_i z_{t} + z_{it}^{\prime} \Upsilon_i z_{t} + z_{t}^{\prime}\Lambda_i f_t + z_{it}^{\prime}\Xi_i f_t + e_{it}, \end{align} where $z_{t}$ represents a vector of constant and macro state variables known at the beginning of time period $t$, $z_{it}$ is a vector of asset characteristics known at the beginning of time period $t$, $\Psi_i$, $\Upsilon_i$, $\Lambda_i$, and $\Xi_i$ are unknown matrices of coefficients satisfying certain no arbitrage constraints, and $e_{it}$ is the idiosyncratic component. Our model encompasses (ref) without the no arbitrage constraints where $x_{it}$ consists of quadratic transformations of $z_t$ and $z_{it}$, $a_i$ and $B_i$ are functions of $\Phi_i$, $\Psi_i$, $\Upsilon_i$, and $\Lambda_i$, and $\varepsilon_{it} = e_{it}$. In contrast to their estimation method, which relies on observable $f_t$, our estimation procedure treats $f_t$ as latent factors and allows for the presence of arbitrage and large $p$. \end{ex} \section{Estimation Strategy} For the convenience of the reader, we gather standard pieces of notation here, which will be utilized throughout the paper. We denote a $k\times k$ identity matrix as $I_k$. The Euclidean norm of a column vector $x$ is represented by $\|x\|$. For a symmetric matrix $A$, we denote its trace as $\mathrm{tr}(A)$, its smallest and largest eigenvalues as $\lambda_{\min}(A)$ and $\lambda_{\max}(A)$. The operator norm of a matrix $A$ is denoted by $\|A\|_2$, its Frobenius norm by $\|A\|_{F}$, and its vectorization by $\mathrm{vec}(A)$. The Kronecker product of matrices $C$ and $D$ is denoted as $C \otimes D$. Unless specified, asymptotic statements in the paper shall be understood to hold as $N\to\infty$ with fixed $T$ or as $(N,T)\to\infty$, whenever appropriate. We begin by reformulating the model in (ref) using vectors/matrices. Let $a\equiv (a_1^{\prime},a_2^{\prime},\ldots, $ $a_{N}^{\prime})^{\prime}$ which is an $Np\times 1$ vector of unknown coefficients, $B \equiv (B_1^{\prime},B_2^{\prime},\ldots, B_{N}^{\prime})^{\prime}$ which is an $Np\times K$ matrix of unknown coefficients, and $F \equiv(f_1,f_2,\ldots, f_{T})^{\prime}$ which is a $T\times K$ matrix of latent factors. Let $\Pi$ be an $Np\times T$ unknown parameter matrix that collects the product of $(a_i, B_i)$ and $(1,f_t^{\prime})^{\prime}$, defined as \begin{align} \Pi\equiv \left( \begin{array}{c} (a_1, B_1) \\ (a_2, B_2) \\ \vdots \\ (a_N, B_N) \\ \end{array} \right) \left( \begin{array}{cccc} \left( \begin{array}{c} 1 \\ f_1 \\ \end{array} \right), & \left( \begin{array}{c} 1 \\ f_2 \\ \end{array} \right), & \cdots, & \left( \begin{array}{c} 1 \\ f_T \\ \end{array} \right) \\ \end{array} \right) \equiv a 1^{\prime}_{T} + BF^{\prime}, \end{align} where $1_{T}$ is a $T\times 1$ vector of ones. Let $X_{it}\equiv (e_{N,i}\otimes x_{it}) e_{T,t}^{\prime}$ be an $Np\times T$ observed data matrix of $x_{it}$, where $e_{N,i}$ is the $i$th column of $I_N$ and $e_{T,t}$ is the $t$th column of $I_T$. Then $x_{it}^{\prime}a_i + x_{it}^{\prime}B_{i}f_{t} = \mathrm{tr}(X_{it}^{\prime}\Pi)$, so (ref) can be succinctly expressed as \begin{align} y_{it} = \mathrm{tr}(X_{it}^{\prime}\Pi) + \varepsilon_{it}. \end{align} Since $\Pi$ has at most rank $K+1$, (ref) can be viewed as a trace linear regression model with reduced rank coefficient matrix $\Pi$ NegahbanWainwright_Scaling_2011, RohdeTsybakov_Lowrank_2011. Thus, we first estimate $\Pi$ by using the nuclear norm regularization Fazel_MatrixRank_2002, which employs the nuclear norm penalty as a surrogate function to enforce the reduced rank constraint. Our estimator of $\Pi$ is given by \begin{align} \hat{\Pi} = \operatornamewithlimits{arg\,min}_{\Gamma\in\mathcal{S}\subset\mathbf{R}^{Np\times T}}\frac{1}{2}\sum_{i=1}^{N}\sum_{t=1}^{T}(y_{it}-\mathrm{tr}(X_{it}^{\prime}\Gamma))^{2} +\lambda_{NT} \|\Gamma\|_{\ast}, \end{align} where $\mathcal{S}\subset\mathbf{R}^{Np\times T}$ is convex, $\|\Gamma\|_{\ast}$ is the nuclear norm of $\Gamma$, and $\lambda_{NT}>0$ is a regularization parameter.\footnote{The nuclear norm of $\Gamma$ is $\|\Gamma\|_{\ast}=\sum_{j=1}^{\min\{Np,T\}}\sigma_{j}(\Gamma)$, corresponding to the sum of its singular values, where $\sigma_{j}(\Gamma)$'s are the singular values of $\Gamma$. The nuclear norm of $\Gamma$ is the convex envelope of the rank of $\Gamma$ over the set of matrices with spectral norm no greater than one; see, for example, Rechtetal_Guaranteed_2010. } In particular, by introducing $\mathcal{S}$, which can be strictly smaller than $\mathbf{R}^{Np\times T}$, we can enforce the constraints of $\Pi$ induced by those of $a$ and $B$ \textemdash a critical aspect that has not been explored in the existing literature\textemdash enabling the estimation of various models within a unified framework. We set $\mathcal{S}=\mathbf{R}^{Np\times T}$ in Examples (ref), (ref), and (ref), $\mathcal{S}=\mathcal{D}_{M}$ for $0<M<\infty$ (where $\mathcal{D}_{M}$ is given in (ref)) in Example (ref), and $\mathcal{S} = \{1_{N}\otimes \Gamma: \Gamma\in \mathbf{R}^{p\times T}\}$ in Example (ref); see Section (ref) for details. In the latter two cases, $\mathcal{S}$ is strictly smaller than $\mathbf{R}^{Np\times T}$. Since (ref) involves constrained nonsmooth convex minimization, $\hat{\Pi}$ generally does not have an analytical closed form. Although several algorithms are available for solving convex minimization problems with a nuclear norm VandenbergheBoyd_Semidefinite_1996, Bertsekas_NonlinearProgramming_1999, LiuVandenberghe_Interior_2010, Maetal_Fixedpoint_2011, they are not suitable for the high-dimensional settings with constraints in our context. In Appendix (ref), we provide an efficient computing algorithm for each setting in Examples (ref)-(ref). We next proceed to derive estimators for $K$, $a$, $B$, and $F$ from the nuclear norm regularized estimator $\hat{\Pi}$. Let $\hat{K}$, $\hat{a}$, $\hat{B}$, and $\hat{F}$ denote these estimators. Define $M_{T}\equiv I_{T} - 1_{T}1_{T}^{\prime}/T$. Since $\Pi M_{T} = BF^{\prime}M_{T}$, we can obtain $\hat{K}$ and $\hat{B}$ from the eigenvalues and eigenvectors of $\hat{\Pi} M_{T}\hat{\Pi}^{\prime}$. Specifically, $\hat{K}$ is given by \begin{align} \hat{K} = \sum_{j=1}^{Np}1\{\lambda_{j}(\hat{\Pi} M_{T}\hat{\Pi}^{\prime})\geq \delta_{NT}\}, \end{align} where $\lambda_{j}(A)$ denotes the $j$th largest eigenvalue of $A$ and $\delta_{NT}>0$ is a threshold value. If $\hat{K} = 0$, $\hat{a} = {\hat{\Pi}1_{T}}/{T}$, $\hat{B} =0 $, and $\hat{F}=0$; otherwise we proceed as follows. To estimate $B$, we use the following normalization: $B^{\prime}B/N=I_K$ and $F^{\prime} M_{T} F/T$ being diagonal with diagonal entries in descending order. Then the columns of $\hat{B}/\sqrt{N}$ are given by the eigenvectors of $\hat{\Pi} M_{T}\hat{\Pi}^{\prime}$ corresponding to its largest $\hat{K}$ eigenvalues. To estimate $a$ and $F$, we impose the following condition: $a^{\prime} B = 0$. Since $a = (I_{Np}-B(B^{\prime}B)^{-1}B^{\prime})\Pi1_{T}/T$ and $F = \Pi^{\prime}B(B^{\prime}B)^{-1}$, we thus obtain \begin{align} \hat{a} = \left(I_{Np}-\frac{\hat{B}\hat{B}^{\prime}}{N}\right)\frac{\hat{\Pi}1_{T}}{T} \text{ and } \hat{F} = \frac{\hat{\Pi}^{\prime}\hat{B}}{N}. \end{align} It it noteworthy that in Examples (ref) and (ref) there is no need to enforce the homogeneity restriction of $a$ and $B$ in extracting $\hat{a}$ and $\hat{B}$ from $\hat{\Pi}$ again to ensure the same homogeneity structure of $\hat{a}$ and $\hat{B}$; see Sections (ref) and (ref) for details. Our estimation procedure is adaptable to accommodate the presence of missing values. In this case, the double summations in (ref) must be replaced with summations over non-missing data. This amounts to redefining the observations as $y_{it}m_{it}$ and $x_{it}m_{it}$, and the error term as $\varepsilon_{it}m_{it}$, where $m_{it} = 0$ when $y_{it}$ or $x_{it}$ are missing, and $1$ otherwise. Since $\hat{\Pi}$ accommodates missing values, we can employ a CV approach to choose the regularization parameter $\lambda_{NT}$ in (ref). Specifically, we first randomly divide the observations into $L$ folds with observations indexed by $\{\mathcal{I}_{\ell}\}_{\ell\leq L}$, where $\mathcal{I}_{\ell}$ comprises observation indices in the $\ell$th fold, $\{\mathcal{I}_{\ell}\}_{\ell\leq L}$ are mutually exclusive, and $\cup_{\ell\leq L}\mathcal{I}_{\ell}=\mathcal{I}\equiv \{1,2,\cdots,N\}\times \{1,2,\cdots, T\}$. Rolling $\ell$ from $1$ to $L$, we then leave observations $\{(y_{it},x_{it}): (i,t)\in\mathcal{I}_{\ell}\}$ out, use observations $\{(y_{it},x_{it}): (i,t)\in\mathcal{I}/\mathcal{I}_{\ell}\}$ for training, and calculate the out-of-sample mean square error $\mathrm{MSE}_\ell$ for observations $\{(y_{it},x_{it}): (i,t)\in\mathcal{I}_{\ell}\}$. Finally, we choose $\lambda_{NT}$ by minimizing the average out-of-sample mean square error $\sum_{\ell=1}^{L}\mathrm{MSE}_\ell/L$. \begin{rem} Enforcing the rank constraint directly is perhaps the most intuitive approach to incorporate the reduced-rank structure. This leads to the following problem: \begin{align} \min_{{c_i}\in\mathbf{R}^{p},{D_i}\in\mathbf{R}^{p\times K},{g}_t\in\mathbf{R}^{K}}\frac{1}{2}\sum_{i=1}^{N}\sum_{t=1}^{T}(y_{it} - x_{it}^{\prime}c_i - x_{it}^{\prime}D_ig_t)^{2}. \end{align} However, solving (ref) poses at least two challenges.\footnote{Enforcing the constraints of $a$ and $B$ in Examples (ref) and (ref) does not resolve the challenges.} Firstly, it requires knowledge of $K$, which must be estimated prior to solving the problem. Secondly, (ref) is nonconvex and its solution lacks an analytical closed form. These challenges not only complicate the design of computational algorithms to find the solution but also hinder derivation of its asymptotic properties. One potential approach to address the second challenge is alternating least squares; however, it may suffer from non-convergence issues due to the nonconvexity of (ref) GolubVanLoan_MatrixComputation_2013, Chietal_Nonconvex_2019. In contrast, the problem in (ref) is convex, and our estimators can be numerically solved efficiently without requiring prior knowledge of $K$, complemented by the asymptotic properties derived in Sections (ref) and (ref). \end{rem} \section{Asymptotic Analysis} In this section, we conduct an asymptotic analysis for our estimators in a general setup. Specifically, we establish consistency of $\hat{K}$ and a rate of convergence of $\hat{\Pi}$, $\hat{a}, \hat{B}$, and $\hat{F}$. We begin by introducing the so-called “restricted strong convexity” condition Negahbanetal_UnifiedFramework_2012. This condition ensures that the quadratic loss function in (ref) is strictly convex over a restricted set of “low-rank” matrices. To describe the set, we define some notation. Let $\Pi = U \Sigma V^{\prime}$ be a singular value decomposition of $\Pi$, where $U$ and $V$ are $Np\times Np$ and $T\times T$ orthonormal matrices, and $\Sigma$ is a diagonal matrix with singular values of $\Pi$ in the diagonal in descending order. Write $U = (U_1,U_2)$ and $V = (V_1,V_2)$, where the columns of $U_2$ and $V_2$ are singular vectors corresponding to the zero singular values of $\Pi$. For any $Np \times T$ matrix $\Delta$, let $\mathcal{P}(\Delta)\equiv U_2U_2^{\prime}\Delta V_2V_2^{\prime}$ and $\mathcal{M}(\Delta)\equiv \Delta - \mathcal{P}(\Delta)$. Heuristically, $\mathcal{M}(\Delta)$ can be thought of as the projection of $\Delta$ onto the “low-rank” space of $\Pi$, and $\mathcal{P}(\Delta)$ is the projection of $\Delta$ onto its orthogonal space. The restricted set of “low-rank” matrices is \begin{align} \mathcal{C}\equiv \{\Delta\in\mathcal{S}\ominus \mathcal{S}: \|\mathcal{P}(\Delta)\|_{\ast}\leq 3\|\mathcal{M}(\Delta)\|_{\ast}\}, \end{align} where $\mathcal{S}\ominus \mathcal{S}$ is the Minkowski difference between $\mathcal{S}$ and $\mathcal{S}$, that is, $\mathcal{S}\ominus \mathcal{S} = \{\Gamma_1-\Gamma_2: \Gamma_1,\Gamma_2\in\mathcal{S}\}$. We impose the restricted strong convexity condition as follows. \begin{ass} (i) Assume that $\Pi\in\mathcal{S}\subset \mathbf{R}^{Np\times T}$. For any $\Delta\in\mathcal{S}\ominus \mathcal{S}$, the following decomposition holds: \[\sum_{i=1}^{N}\sum_{t=1}^{T}|\mathrm{tr}(X_{it}^{\prime}\Delta)|^2 = \mathcal{Q}_{NT}(\Delta) + \mathcal{L}_{NT}(\Delta)\] such that for some constant $0<\kappa<\infty$, \[\mathcal{Q}_{NT}(\Delta) \geq \kappa \|\Delta\|_{F}^{2} \text{ for all } \Delta\in\mathcal{C},\] and for some $r_{NT}>0$, \[|\mathcal{L}_{NT}(\Delta)|\leq r_{NT} \|\Delta\|_{\ast} \text{ for all } \Delta\in\mathcal{S}\ominus \mathcal{S}.\] (ii) The following condition holds: \[\left|\sum_{i=1}^{N}\sum_{t=1}^{T}\mathrm{tr}(\varepsilon_{it} X_{it}^{\prime}\Delta)\right|\leq \frac{1}{2}r_{NT} \|\Delta\|_{\ast} \text{ for all } \Delta\in\mathcal{S}\ominus \mathcal{S}.\] \end{ass} Assumption (ref) is weaker than the conditions of Corollary 1 in NegahbanWainwright_Scaling_2011, which require $\mathcal{S}=\mathbf{R}^{Np \times T}$ and $\mathcal{L}_{NT}(\cdot)=0$, and are too restrictive in Examples (ref) and (ref). We refer to the condition: “$\mathcal{Q}_{NT}(\Delta) \geq \kappa \|\Delta\|_{F}^{2}$ for all $\Delta\in\mathcal{C}$” as the restricted strong convexity condition. Allowing $\mathcal{L}_{NT}(\cdot)\neq 0$ facilitates providing easy-to-verify primitive conditions for the restricted strong convexity condition in Example (ref). The rate $r_{NT}$ plays an important role in determining the convergence rate of $\hat{\Pi}$, thus determining how fast $p$ and $K$ can grow. \begin{ass} There exist some constants $0<d_{\min}\leq d_{\max}<\infty$ such that: (i) $d_{\min}<\lambda_{\min}(B'B/N)\leq \lambda_{\max}(B'B/N)<d_{\max}$ for large $N$; (ii) $\max_{t\leq T}\|f_{t}\|< d_{\max}$; (iii) $\lambda_{\min}(F'M_{T}F/T)> d_{\min}$; (iv) $a^{\prime}a/N<d_{\max}$; (v) $a^{\prime}B = 0$. \end{ass} For the sake of clarity in presentation, we assume that $\{a_{i},B_{i}\}_{i\leq N}$ and $\{f_t\}_{t\leq T}$ are non-random. In other words, all stochastic statements are implicitly conditional on their realization. Assumption (ref)(i) resembles the pervasive condition in StockWatson_PCA_2002 and BaiNg_NumberofFactors_2002, which necessitates that $\{f_t\}_{t\leq T}$ are strong factors. Assumptions (ref)(iv) and (v) are identification conditions for $a$, see Appendix (ref) for discussion; similar assumptions are also used in Chenetal_SeimiparametricFactor_2021. While Assumptions (ref) and (ref) consist of high-level conditions for the general setup, in Section (ref) we provide primitive conditions for each setting in Examples (ref)-(ref). \begin{thm} Suppose Assumption (ref) holds. Let $\hat{\Pi}$, $\hat{K}$, $\hat{a}, \hat{B}$, and $\hat{F}$ be given in (ref)-(ref). Assume that $0<K<\min\{Np,T\}-1$ and $\lambda_{NT}\geq 2r_{NT}$. (i) Then \begin{align*} \|\hat{\Pi} - \Pi\|_{F}\leq \frac{3\sqrt{2(K+1)}\lambda_{NT}}{\kappa}. \end{align*} (ii) Suppose Assumption (ref) also holds. Assume that $\delta_{NT}/(NT)\to 0$ and $\delta_{NT}/(K\lambda^{2}_{NT})\to \infty $. Let $H \equiv (F^{\prime}M_T\hat{F})(\hat{F}^{\prime}M_T\hat{F})^{-1}$. Then \begin{align*} P(\hat{K}= K) &\to 1,\\ \|\hat{a} - a\|&=O_{p}\left(\frac{\sqrt{K}\lambda_{NT}}{\sqrt{T}}\right),\notag\\ \|\hat{B} - B H\|_{F}&=O_{p}\left(\frac{\sqrt{K}\lambda_{NT}}{\sqrt{T}}\right),\notag\\ \|\hat{F}-F(H^{\prime})^{-1}\|_{F}&=O_{p}\left(\frac{\sqrt{K}\lambda_{NT}}{\sqrt{N}}\right). \end{align*} \end{thm} Theorem (ref)(i) gives a deterministic statement about the estimation error of $\hat{\Pi}$, extending Corollary 1 of NegahbanWainwright_Scaling_2011 by allowing $\mathcal{L}_{NT}(\cdot)\neq 0$ and constraints on $\Pi$ (i.e., $\mathcal{S}\neq \mathbf{R}^{Np \times T}$) in addition to the reduced-rank constraint. While Assumption (ref) and $\lambda_{NT}\geq 2r_{NT}$ may not hold deterministically, they often hold with probability approaching one, as verified in Section (ref). In such cases, the result of Theorem (ref)(i) holds with probability approaching one, and the results of Theorem (ref)(ii) persist. Due to identification issues, $B$ and $F$ can only be consistently estimated up to a rotational transformation, as commonly encountered in high-dimensional factor analyses. The asymptotic results hold as $N\to\infty$ with fixed $T$ or as $(N,T)\to\infty$, as appropriate. Theorem (ref) is a theory for the general setup under high-level assumptions, which is applicable for each setting in Examples (ref)-(ref). In Section (ref), we tailor Theorem (ref) for each model by providing low-level sufficient assumptions that are easier to verify. In all cases, $p$ and $K$ are permitted to grow with $N$ or $(N,T)$ for the consistency of the estimators, and the presence of missing values is allowed. \section{Revisiting Nested Models} For simplicity of notation, we continue to use $x_{it}$ representing the vector of explanatory variables in all models, rather than each model's specific notation in Section (ref). \subsection{Examples (ref), (ref), and (ref)} Our objective is to estimate $a$, $B$, $F$, and $K$.\footnote{For simplicity of presentation, we continue to use $a$ and $B$ representing the coefficients of interest in all three examples, rather than each example's specific notation in Section (ref), and ignore the sieve approximation error in Example (ref) (so $\Delta_{i}(\cdot) = 0$). This allows us to unify results in one theorem. One may account for the sieve approximation error as similar to Corollaries (ref) and (ref).} No constraints are imposed on $a$ and $B$ and we set $\mathcal{S} = \mathbf{R}^{Np\times T}$ in (ref). In the scenario when $x_{it}=1$, we can obtain an analytical closed form for $\hat{\Pi}$. Let $Y$ be an $N\times T$ matrix with the $it$th entry $y_{it}$. Consider the singular value decomposition $Y = U\Sigma V^{\prime}$, where $U$ and $V$ are $N\times N$ and $T\times T$ orthonormal matrices and $\Sigma$ is an $N\times T$ diagonal matrix with singular values $\sigma_{j}(Y)$'s in the diagonal in descending order. For $x>0$, define $\Sigma_{x}$ be an $N\times T$ diagonal matrix with $\max\{0,\sigma_j(Y)-x\}$ in descending order. Consequently, $\hat{\Pi} = U\Sigma_{\lambda_{NT}/2} V^{\prime}$, as described in Caietal_SVD_2010 and Maetal_Fixedpoint_2011. However, an analytical closed form is not available for general cases. An efficient algorithm for finding $\hat{\Pi}$ is provided in Appendix (ref). To provide primitive conditions, we impose the following assumptions. \begin{ass} (i) There exists some constant $0<\kappa<\infty$ such that \[\sum_{i=1}^{N}\sum_{t=1}^{T}|\mathrm{tr}(X_{it}^{\prime}\Delta)|^2\geq \kappa \|\Delta\|_{F}^{2} \text{ for all } \Delta\in\mathcal{D},\] where $\mathcal{D}\equiv\{\Delta\in\mathbf{R}^{Np\times T}: \|\mathcal{P}(\Delta)\|_{\ast}\leq 3\|\mathcal{M}(\Delta)\|_{\ast}\}$. (ii) $\{(x_{1t}^{\prime}e_{1t},x_{2t}^{\prime}e_{2t},\ldots,x_{Nt}^{\prime}e_{Nt})'\}_{t\leq T}$ is a sequence of independent sub-Gaussian vectors.\footnote{Independence is not necessary here and also in Assumptions (ref)(iv), (v) and (ref)(iii). We may allow for weak dependence over $t$; see Lemma (ref). } \end{ass} A condition similar to Assumption (ref) has been imposed in MoonWeidner_Nuclear_2018 and Chernozhukovetal_LowRank_2018. We apply Theorem (ref) to obtain the following corollary. \begin{cor} Suppose Assumption (ref)(ii) holds. Let $\hat{\Pi}$, $\hat{K}$, $\hat{a}, \hat{B}$, and $\hat{F}$ be given in (ref)-(ref) with $\mathcal{S} = \mathbf{R}^{Np\times T}$ and $\lambda_{NT}=\sqrt{(Np+T)\log N}$. Assume that $0<K<\min\{Np,T\}-1$. (i) If $x_{it} =1$ or Assumption (ref)(i) holds, then as $(N,T)\to\infty$, \begin{align*} \frac{1}{\sqrt{NT}}\|\hat{\Pi} - \Pi\|_{F}= O_{p}\left(\sqrt{\frac{K(Np+T)\log N}{NT}}\right). \end{align*} (ii) Suppose Assumptions (ref)(i)-(iii) additionally hold. Assume that as $(N,T)\to\infty$, $\delta_{NT}/(NT)\hspace{-0.05cm}\to \hspace{-0.05cm}0$ and $\delta_{NT}/[K(Np+T)\log N]\to \hspace{-0.05cm}\infty\hspace{-0.05cm} $. Let $H \hspace{-0.1cm}\equiv \hspace{-0.1cm}(F^{\prime}M_T\hat{F})(\hat{F}^{\prime}M_T\hat{F})^{-1}$. If $a=0$ or Assumptions (ref)(iv)-(v) hold, then as $(N,T)\to\infty$, \begin{align*} P(\hat{K}= K) &\to 1,\\ \frac{1}{\sqrt{N}}\|\hat{a} - a\|&=O_{p}\left(\sqrt{\frac{K(Np+T)\log N}{NT}}\right),\notag\\ \frac{1}{\sqrt{N}}\|\hat{B} - B H\|_{F}&=O_{p}\left(\sqrt{\frac{K(Np+T)\log N}{NT}}\right),\notag\\ \frac{1}{\sqrt{T}}\|\hat{F}-F(H^{\prime})^{-1}\|_{F}&=O_{p}\left(\sqrt{\frac{K(Np+T)\log N}{NT}}\right). \end{align*} \end{cor} Corollary (ref) requires large $N$ and large $T$. In particular, $K(Np+T)\log N=o(NT)$ is required for the consistency of the estimators. This implies that $p$ is allowed to grow as $(N,T)\to\infty$. While the result for Example (ref) is well-documented in the literature, the results for Examples (ref) and (ref) are novel. Distinct from PelgerXiong_State-varying_2019, we offer an estimator capable of consistently estimate $F$ up to a common rotational transformation, which is not state-specific. In other words, we can consistently estimate the factor space. Moreover, our method allows for large $p$. In contrast to Gagliardinietal_Timevarying_2016, we provide an estimation approach that does not necessitate observable $f_t$ and permits the presence of arbitrage and large $p$. Notably, there is no available method for estimating the unconstrained conditional latent factor model in the literature. \subsection{Example (ref)} Our objective is to estimate $\mu\equiv (\mu_1,\mu_2,\ldots,\mu_N)^{\prime}$, $\Lambda\equiv (\lambda_{1}, \lambda_{2},\ldots,\lambda_{N})^{\prime}$, $\phi$, $\Phi$, $F$, and $K$. Since $a = (\mu_1, \phi^{\prime}, \mu_2, \phi^{\prime},\ldots, \mu_N, \phi^{\prime})^{\prime}$ and $B = (\lambda_1, \Phi^{\prime}, \lambda_2, \Phi^{\prime}, \ldots, \lambda_N, \Phi^{\prime})^\prime$, we have $\Pi =a 1^{\prime}_{T} + BF^{\prime}= ((\pi_1,\Pi^{\ast\prime}), (\pi_2,\Pi^{\ast\prime}),\ldots, (\pi_N,\Pi^{\ast\prime}))^{\prime}$, where $\pi_i\equiv\mu_i1_{T}+ F\lambda_i$ and $\Pi^{\ast} \equiv \phi1_{T}^{\prime} + \Phi F^{\prime}$, which are $ T\times 1$ vector and $(p-1)\times T$ matrix, respectively. Then we set \begin{align} \mathcal{S} = \mathcal{D}_M \equiv \left\{\left( \begin{array}{c} \gamma_1^{\prime} \\ \Gamma^{\ast} \\ \gamma_2^{\prime} \\ \Gamma^{\ast} \\ \vdots \\ \gamma_N^{\prime} \\ \Gamma^{\ast} \\ \end{array} \right) : \left( \begin{array}{c} \gamma_1^{\prime} \\ \gamma_2^{\prime} \\ \vdots \\ \gamma_N^{\prime} \\ \end{array} \right) \in\mathbf{R}^{N\times T}, \Gamma^{\ast}\in \mathbf{R}^{(p-1)\times T} \text{ and } \|\Gamma^{\ast}\|_{\max}\leq M\right\} \end{align} for $0<M<\infty$ in (ref), where $\|\Gamma^{\ast}\|_{\max}$ denotes the largest absolute value of the entries of $\Gamma^{\ast}$.\footnote{Imposing $\|\Gamma^{\ast}\|_{\infty}\leq M$ facilitates providing easy-to-verify primitive conditions for Assumption (ref)(i).} Since $\hat{\Pi}\in\mathcal{D}_{M}$, we can write $\hat{\Pi} = ((\hat{\pi}_1,\hat{\Pi}^{\ast\prime}), (\hat{\pi}_2,\hat{\Pi}^{\ast\prime}),\ldots, (\hat{\pi}_N,\hat{\Pi}^{\ast\prime}))^{\prime}$, where $\hat{\pi}_i$ is an estimator of $\pi_i$ and $\hat{\Pi}^{\ast}$ is an estimator of $\Pi^{\ast}$. Let ${\Pi}^{\diamond}\equiv({\pi}_{1},{\pi}_{2}, \ldots, {\pi}_{N})^{\prime}$ and $\hat{\Pi}^{\diamond}\equiv(\hat{\pi}_{1},\hat{\pi}_{2}, \ldots, \hat{\pi}_{N})^{\prime}$. An efficient algorithm for finding $\hat{\Pi}^{\diamond}$ and $\hat{\Pi}^{\ast}$ is provided in Appendix (ref). By Lemma (ref) (iv) and simple algebra, we can write \begin{align} \hat{a} = ((\hat{\mu}_{1},\hat{\phi}^{\prime}), (\hat{\mu}_{2},\hat{\phi}^{\prime}),\ldots, (\hat{\mu}_{N},\hat{\phi}^{\prime}))^{\prime} \text{ and } \hat{B} = ((\hat{\lambda}_{1},\hat{\Phi}^{\prime}),(\hat{\lambda}_{2},\hat{\Phi}^{\prime}),\ldots,(\hat{\lambda}_{N},\hat{\Phi}^{\prime}))^{\prime}, \end{align} where $\hat{\mu}_i$ is a scalar, $\hat{\phi}$ is a $(p-1)\times 1$ vector, $\hat{\lambda}_i$ is a $\hat{K} \times 1$ vector, and $\hat{\Phi}$ is a $(p-1)\times \hat{K}$ matrix. Thus, $\hat{a}$ and $\hat{B}$ share the same homogeneity structure with $a$ and $B$, respectively. It is not necessary to enforce the homogeneity restriction of $a$ and $B$ in extracting $\hat{a}$ and $\hat{B}$ from $\hat{\Pi}$ to ensure the same homogeneity structure, as the homogeneity structure of $\hat{\Pi}$ inherited from $a$ and $B$ automatically passes to $\hat{a}$ and $\hat{B}$. We define the estimators of ${\mu}$, $\Lambda$, $\phi$, and $\Phi$ as $\hat{\mu}\equiv (\hat{\mu}_1, \hat{\mu}_2, \ldots, \hat{\mu}_{N})^{\prime}$, $\hat{\Lambda} \equiv (\hat{\lambda}_{1},\hat{\lambda}_{2},\ldots, \hat{\lambda}_{N})^{\prime}$, $\hat{\phi}$, and $\hat{\Phi}$, respectively. A convergence rate for $\hat{\Pi}^{\diamond}$, $\hat{\Pi}^{\ast}$, $\hat{\mu}$, $\hat{\Lambda}$, $\hat{\phi}$, and $\hat{\Phi}$ follows immediately from Theorem (ref), as we have $\|\hat{\Pi} - \Pi\|^{2}_{F} = \|\hat{\Pi}^{\diamond}-{\Pi}^{\diamond}\|^{2}_{F} + N\|\hat{\Pi}^{\ast} - \Pi^{\ast}\|^{2}_{F}$, $\|\hat{a} - a\|^{2} = \|\hat{\mu} - \mu\|^2 + N\|\hat{\phi} - \phi\|^2$, and $\|\hat{B} - BH\|_{F}^{2} = \|\hat{\Lambda} - \Lambda H\|_{F}^{2} + N\|\hat{\Phi} - \Phi H\|^2_{F}$. To provide primitive conditions, we impose the following assumptions. \begin{ass} (i) Write $x_{it} = (1,x_{it}^{\ast\prime})^{\prime}$.\footnote{We allow for time-varying characteristics, so we write $x_{it}$ rather than $x_i$.} There are positive constants $c_{\min}$ and $c_{\max}$ such that: with probability approaching one as $(N,T)\to\infty$, \[c_{\min} \leq \min_{t\leq T}\lambda_{\min}\left(\frac{1}{N}\sum_{i=1}^{N}x^{\ast}_{it}x_{it}^{\ast\prime}\right)\leq \max_{t\leq T}\lambda_{\max}\left(\frac{1}{N}\sum_{i=1}^{N}x^{\ast}_{it}x_{it}^{\ast\prime}\right)\leq c_{\max} .\] (ii) $\max_{t\leq T}\hspace{-0.1cm}\|\phi + \Phi f_{t}\|_{\infty}$ is bounded. (iii) $\max_{t\leq T}\hspace{-0.1cm}E[\|\sum_{i=1}^{N}x^{\ast}_{it}e_{it}/\hspace{-0.1cm}\sqrt{Np}\|^{2}]$ is bounded. (iv) $\{(x_{1t}^{\ast\prime},x_{2t}^{\ast\prime},\ldots,x_{Nt}^{\ast\prime})'\}_{t\leq T}$ is a sequence of independent sub-Gaussian vectors. (v) $\{(e_{1t},e_{2t},$ $\ldots,e_{Nt})'\}_{t\leq T}$ is a sequence of independent sub-Gaussian vectors. (vi) $\sup_{z}|\delta(z)| = O(p^{-s})$ and $\sup_{z}\|\Delta(z)\| = O(p^{-s})$ for some constant $s>0$. \end{ass} \begin{ass} There are constants $0<d_{\min}\leq d_{\max}<\infty$ such that: (i) $\lambda_{\min}(\Phi'\Phi+\Lambda^{\prime}\Lambda/N)>d_{\min}$; (ii) $\lambda_{\max}(\Phi'\Phi)\hspace{-0.05cm}<\hspace{-0.05cm}d_{\max}/2$; (iii) $\lambda_{\max}(\Lambda^{\prime}\Lambda/N)\hspace{-0.05cm}<\hspace{-0.05cm}d_{\max}/2$; (iv) $\max_{t\leq T}\|f_{t}\|$ $< d_{\max}$; (v) $\lambda_{\min}(F'M_{T}F/T)> d_{\min}$; (vi) $\|\phi\|^{2}<d_{\max}/2$; (vii) $\|\mu\|^{2}/N<d_{\max}/2$; (viii) $\phi^{\prime}\Phi = 0$; (ix) $\mu^{\prime}\Lambda = 0$. \end{ass} Assumption (ref) involves no multicolinearity, finite moments, weak dependence, and small sieve approximation errors, all of which are standard in the literature. Conditions similar to Assumption (ref)(i), (iii), and (vi) have been imposed in Fanetal_ProjectedPCA_2016. We apply Theorem (ref) to obtain the following corollary. \begin{cor} Suppose Assumption (ref) holds. Let $\hat{\Pi}$ be given in (ref) with $\mathcal{S} = \mathcal{D}_M$ and $\lambda_{NT}= [M\sqrt{(Np^2+ Tp)} + \sqrt{NT} p^{-s}]\sqrt{\log N} $. Let $\hat{\Pi}^{\diamond}$ and $\hat{\Pi}^{\ast}$ be given below (ref). Let $\hat{K}$, $\hat{F}$, $\hat{\mu}$, $\hat{\Lambda}$, $\hat{\phi}$, and $\hat{\Phi}$ be given in (ref), (ref), and (ref). Assume that $0<K<\min\{N+p-1,T\}-1$. (i) Then as $(N,T)\to\infty$, \begin{align*} \frac{1}{\sqrt{NT}}\|\hat{\Pi}^{\diamond} - \Pi^{\diamond}\|_{F}&= O_{p}\left(M\sqrt{\frac{K(Np^{2}+Tp)\log N}{NT}} + \frac{\sqrt{K \log N}}{p^{s}}\right),\notag\\ \frac{1}{\sqrt{T}}\|\hat{\Pi}^{\ast} - \Pi^{\ast}\|_{F}&= O_{p}\left(M\sqrt{\frac{K(Np^{2}+Tp)\log N}{NT}}+ \frac{\sqrt{K \log N}}{p^{s}}\right). \end{align*} (ii) Suppose Assumption (ref) additionally holds. Assume that as $(N,T)\hspace{-0.05cm}\to\infty$, $\delta_{NT}/(NT)$ $\to 0$ and $\delta_{NT}/\{K[M^{2}(Np^{2}+Tp)+NT p^{-2s}]\log N \}\to \infty $. Let $H \equiv (F^{\prime}M_T\hat{F})(\hat{F}^{\prime}M_T\hat{F})^{-1}$. Then as $(N,T)\to\infty$, \begin{align*} P(\hat{K}= K) &\to 1,\\ \frac{1}{\sqrt{N}}\|\hat{\mu} - \mu\|&=O_{p}\left(M\sqrt{\frac{K(Np^{2}+Tp)\log N}{NT}}+ \frac{\sqrt{K \log N}}{p^{s}}\right),\notag\\ \frac{1}{\sqrt{N}}\|\hat{\Lambda} - \Lambda H\|_{F}&=O_{p}\left(M\sqrt{\frac{K(Np^{2}+Tp)\log N}{NT}}+ \frac{\sqrt{K \log N}}{p^{s}}\right),\notag\\ \|\hat{\phi} - \phi\|&=O_{p}\left(M\sqrt{\frac{K(Np^{2}+Tp)\log N}{NT}}+ \frac{\sqrt{K \log N}}{p^{s}}\right),\notag\\ \|\hat{\Phi} - \Phi H\|_{F}&=O_{p}\left(M\sqrt{\frac{K(Np^{2}+Tp)\log N}{NT}}+ \frac{\sqrt{K \log N}}{p^{s}}\right),\notag\\ \frac{1}{\sqrt{T}}\|\hat{F}-F(H^{\prime})^{-1}\|_{F}&=O_{p}\left(M\sqrt{\frac{K(Np^{2}+Tp)\log N}{NT}}+ \frac{\sqrt{K \log N}}{p^{s}}\right). \end{align*} \end{cor} The slower rate in Corollary (ref) compared to Corollary (ref) is attributed to the reliance on a set of easier-to-verify conditions, namely Assumption (ref), rather than a version of Assumption (ref). However, it is noteworthy that the rate can be improved to $O_{p}(\sqrt{K(N+p+T)\log N/(NT)} + \sqrt{K \log N}/p^{s})$ under Assumption (ref). The second term $\sqrt{K \log N}/p^{s}$ arises from sieve approximation errors. Our results differ from Fanetal_ProjectedPCA_2016 in several aspects. First, we allow for $\mu_i\neq 0 $ and $\phi\neq 0$, which are crucial to capture pricing errors in asset pricing. Second, we permit $x_{it}$ to vary over $t$, a critical feature in asset pricing as many stock characteristics (e.g., book to market ratio and momentum) change from month to month. Our simulations in Appendix (ref) show that Fanetal_ProjectedPCA_2016's projected-PCA fails in the presence of time-varying $x_{it}$. Third, we do not require that $\lambda_i$ has zero mean and weak cross-sectional dependence (in such cases $\lambda_i$ can be interpreted as a vector of noises), which is barely justified in practice. We allow for non-noisy intercepts $\mu_i$ and $\lambda_i$ in pricing errors and risk exposures. Fourth, we allow $K\to\infty$. In addition, our results extend Chenetal_SeimiparametricFactor_2021 by allowing for the heterogeneity of $\mu_i$ and $\lambda_i$ across $i$. \subsection{Example (ref)} Our objective is to estimate $\phi_0$, $\Phi_0$, $F$, and $K$. Since $a = 1_{N}\otimes \phi_0$ and $B = 1_{N}\otimes \Phi_0$, we have $\Pi =a 1^{\prime}_{T} + BF^{\prime}= 1_{N} \otimes \Pi_0$, where $\Pi_0\equiv \phi_0 1^{\prime}_{T} + \Phi_0 F^{\prime}$, which is a $p\times T$ matrix. Then we set $\mathcal{S} = \{1_{N}\otimes \Gamma: \Gamma\in \mathbf{R}^{p\times T}\}$ in (ref). Since $\hat{\Pi}\in\mathcal{S}$, we can write $\hat\Pi = 1_{N} \otimes \hat{\Pi}_0$, where $\hat{\Pi}_0$ is an estimator of ${\Pi}_0$. An efficient algorithm for finding $\hat{\Pi}_0$ is provided in Appendix (ref). By Lemma (ref)(iv) and simple algebra, we can write \begin{align} \hat{a} = 1_{N} \otimes \hat{\phi}_0 \text{ and } \hat{B} = 1_{N} \otimes\hat{\Phi}_0, \end{align} where $\hat{\phi}_0$ is a $p\times 1$ vector and $\hat{\Phi}_0$ is a $p\times \hat{K}$ matrix. For the same reason as in Example (ref), there is no need to enforce the homogeneity restriction of $a$ and $B$ in extracting $\hat{a}$ and $\hat{B}$ from $\hat{\Pi}$. We define the estimators of $\phi_0$ and $\Phi_0$ as $\hat{\phi}_0$ and $\hat{\Phi}_0$, respectively. A convergence rate for $\hat{\Pi}_0$, $\hat{\phi}_0$, and $\hat{\Phi}_0$ follows immediately from Theorem (ref), as we have $\|\hat{\Pi} - \Pi\|_{F} = \sqrt{N}\|\hat{\Pi}_0 - \Pi_0\|_{F}$, $\|\hat{a} - a\| = \sqrt{N}\|\hat{\phi}_0 - \phi_0\| $, and $\|\hat{B} - BH\|_{F} = \sqrt{N}\|\hat{\Phi}_0 - \Phi_0H\|_{F}$. To provide primitive conditions, we impose the following assumptions. \begin{ass} (i) There are positive constants $c_{\min}$ and $c_{\max}$ such that: with probability approaching one as $N\to\infty$ with fixed $T$ or as $(N,T)\to\infty$, \[c_{\min} \leq \min_{t\leq T}\lambda_{\min}\left(\frac{1}{N}\sum_{i=1}^{N}x_{it}x_{it}^{\prime}\right)\leq \max_{t\leq T}\lambda_{\max}\left(\frac{1}{N}\sum_{i=1}^{N}x_{it}x_{it}^{\prime}\right)\leq c_{\max} .\] (ii) $E[\|\sum_{i=1}^{N}x_{it}e_{it}/\sqrt{Np}\|^2]$ is bounded for each $t\leq T$. (iii) $\{\sum_{i=1}^{N}x_{it}e_{it}/\sqrt{N}\}_{t\leq T}$ is a sequence of independent sub-Gaussian vectors. (iv) $\sup_{z}|\delta(z)| = O(p^{-s})$ and $\sup_{z}\|\Delta(z)\| = O(p^{-s})$ for some constant $s>0$. \end{ass} \begin{ass} There are constants $0<d_{\min}\leq d_{\max}<\infty$ such that: (i) $d_{\min}<\lambda_{\min}(\Phi_0^{\prime}\Phi_0)\leq \lambda_{\max}(\Phi_0^{\prime}\Phi_0)<d_{\max}$; (ii) $\max_{t\leq T}\|f_{t}\|< d_{\max}$; (iii) $\lambda_{\min}(F'M_{T}F/T)> d_{\min}$; (iv) $\|\phi_0\|^2<d_{\max}$; (v) $\phi_0^{\prime}\Phi_0 = 0$. \end{ass} Assumption (ref) involves no multicolinearity, finite moments, weak dependence, and small sieve approximation errors, all of which are standard in the literature. Assumptions (ref)(i), (ii), (iv), and (ref) have been imposed in Chenetal_SeimiparametricFactor_2021. We apply Theorem (ref) to obtain the following corollary. \begin{cor} Suppose Assumptions (ref)(i), (ii), and (iv) hold. Let $\hat\Pi$ be given in (ref) with $\mathcal{S} = \{1_{N}\otimes \Gamma: \Gamma\in \mathbf{R}^{p\times T}\}$ and $\lambda_{NT}=(\sqrt{p+T} +\sqrt{NT}p^{-s})\sqrt{\log N}$. Let $\hat{\Pi}_0 $ be given above (ref). Let $\hat{K}$, $\hat{F}$, $\hat{\phi}_0$, and $\hat{\Phi}_0$ be given in (ref), (ref), and (ref). Assume $0<K<\min\{p,T\}-1$. (i) Then as $N\to\infty$ with fixed $T$, \begin{align*} \frac{1}{\sqrt{T}}\|\hat{\Pi}_0 - \Pi_0\|_{F}= O_{p}\left(\sqrt{\frac{K(p+T)\log N}{NT}} + \frac{\sqrt{K \log N}}{p^{s}}\right). \end{align*} (ii) Suppose Assumption (ref) additionally holds. Assume that as $N\to\infty$ with fixed $T$, $\delta_{NT}/(NT)\hspace{-0.05cm}\to \hspace{-0.05cm}0$ and $\delta_{NT}/[K(p+T+ NTp^{-2s})\log N]\hspace{-0.05cm}\to \hspace{-0.05cm}\infty$. Let $H \hspace{-0.1cm}\equiv \hspace{-0.1cm}(F^{\prime}M_T\hat{F})(\hat{F}^{\prime}M_T\hat{F})^{-1}$. Then as $N\to\infty$ with fixed $T$, \begin{align*} P(\hat{K}= K) &\to 1,\\ \|\hat{\phi}_0 - \phi_0\|&=O_{p}\left(\sqrt{\frac{K(p+T)\log N}{NT}} + \frac{\sqrt{K \log N}}{p^{s}}\right),\notag\\ \|\hat{\Phi}_0 - \Phi_0 H\|_{F}&=O_{p}\left(\sqrt{\frac{K(p+T)\log N}{NT}} + \frac{\sqrt{K \log N}}{p^{s}}\right),\notag\\ \frac{1}{\sqrt{T}}\|\hat{F}-F(H^{\prime})^{-1}\|_{F}&=O_{p}\left(\sqrt{\frac{K(p+T)\log N}{NT}} + \frac{\sqrt{K \log N}}{p^{s}}\right). \end{align*} (iii) If Assumption (ref)(ii) is replaced with Assumption (ref)(iii), then (i) and (ii) continue to hold by replacing “as $N\to\infty$ with fixed $T$” with “as $(N,T)\to\infty$” in all places. \end{cor} Corollary (ref) establishes a convergence rate of $\hat{\Pi}_0$, $\hat{K}$, $\hat{\phi}_0, \hat{\Phi}_0$, and $\hat{F}$ either under large $N$ with fixed $T$ or scenarios with both large $N$ and large $T$. In particular, $K(p+T)\log N = o(NT)$ is required for the consistency. This implies that $p$ can be as large as $N$ up to $\log N$. Such a result represents a significant improvement from similar results in Chenetal_SeimiparametricFactor_2021, which require that $p$ grows at a rate slower than $N^{1/3}$. Our simulations in Appendix (ref) show that Chenetal_SeimiparametricFactor_2021's regressed-PCA exhibits poor performance when $p$ is close to $N$. The rate $\sqrt{K \log N}/p^{s}$ arises from sieve approximation errors. In addition, our framework accommodates the scenario where $K$ tends to infinity and allows for weak cross-sectional dependence of $x_{it}$. \section{Simulation Studies} In this section, we conduct Monte Carlo simulations to investigate the finite sample performance of our estimators. We consider settings with $p=37$, $N = 500, 1000, 2000$, and $T = 250, 500$, which are comparable with those in the empirical analysis in Section (ref). We consider three different data generating processes (DGPs), which correspond to the settings in Examples (ref), (ref), and (ref). In all three DGPs, we let \begin{align} x_{it,1}=1, x_{it,2}=\sigma_{t}u_{it,1}, x_{it,3} = 0.3x_{i(t-1),3} + u_{it,2}, x_{it,4}=u_{it,3}, \ldots, x_{it,37}=u_{it,36}, \end{align} where $u_{it} = (u_{it,1},u_{it,2},\ldots, u_{it,36})^{\prime}$ are i.i.d. $N(0,I_{36})$ across both $i$ and $t$, $\sigma_{t}$'s are i.i.d. $U(1,2)$ over $t$, and $x_{i0,3}$'s are i.i.d. $N(0,1)$ across $i$. Let $x_{it} = (x_{it,1},x_{it,2},\ldots,x_{it,37})^{\prime}$, hence $p=37$. Let $f_{t} = 0.3 f_{t-1} + \eta_t$, where $\eta_t$'s are i.i.d. $N(1_2,I_2)$ over $t$ and $f_0\sim N(1_2/0.7,I_2/0.91)$, resulting in $K=2$. The errors $\varepsilon_{it}$'s be i.i.d. $N(0,4)$ across both $i$ and $t$. In the first DGP (DGP1), we assume \begin{align} a_i &= \left( \begin{array}{ccccccccccc} 0 & \theta_{1i} & 0 & 0 & \theta_{2i} & 0 & 0 &\cdots & \theta_{12i} & 0 & 0 \\ \end{array} \right)^{\prime} \text{ and } \notag\\ B_i &=\left( \begin{array}{ccccccccccc} 0 & 0 & \varrho_{1i} & 0 & 0 & \varrho_{2i} & 0 & \cdots & 0 & \varrho_{12i} & 0\\ \varrho_{13i} & 0 & 0 & \varrho_{14i} & 0 & 0 & \varrho_{15i} & \cdots & 0 & 0 & \varrho_{25i}\\ \end{array} \right)^{\prime}, \end{align} where $\theta_{ji}$'s are i.i.d. $N(0,1/4)$ across both $i$ and $j=1,2,\ldots, 12$ and $\varrho_{ji}$'s are i.i.d. $U(1/3,1)$ across both $i$ and $j=1,2,\ldots, 25$. In DGP1, $a_i$ and $B_i$ are heterogenous across $i$, which is the setting in Example (ref). We are interested in $a$, $B$, $F$, and $K$. In the second DGP (DGP2), we assume \begin{align} a_i &= \left( \begin{array}{cc} \mu_i & \phi^{\prime} \\ \end{array} \right)^{\prime} = \left( \begin{array}{ccccccccccc} 0 & 1/2 & 0 & 0 & 1/2 & 0 & 0 & \cdots & 1/2 & 0 & 0 \\ \end{array} \right)^{\prime} \text{ and } \notag\\ B_i &=\left( \begin{array}{cc} \lambda_i & \Phi^{\prime} \\ \end{array} \right)^{\prime} = \left( \begin{array}{ccccccccccc} 0 & 0 & 2/3 & 0 & 0 & 2/3 & 0 & \cdots & 0 & 2/3 & 0\\ \vartheta_{i} & 0 & 0 & 2/3 & 0 & 0 & 2/3 & \cdots & 0 & 0 & 2/3 \\ \end{array} \right)^{\prime}, \end{align} where $\vartheta_i$'s are i.i.d. $U(1,3)$ across $i$. In DGP2, the rows of $a_i$ and $B_i$ corresponding to the nonconstant part of $x_{it}$ are homogenous across $i$, which is the setting in Example (ref). We are interested in $\mu$, $\phi$, $\Lambda$, $\Phi$, $F$, and $K$. In the third DGP (DGP3), we assume \begin{align} a_i &=\phi_0= \left( \begin{array}{ccccccccccc} 0 & 1/2 & 0 & 0 & 1/2 & 0 & 0 & \cdots & 1/2 & 0& 0 \\ \end{array} \right)^{\prime} \text{ and } \notag\\ B_i &=\Phi_0 = \left( \begin{array}{ccccccccccc} 0 & 0 & 2/3 & 0 & 0 & 2/3 & 0 & \cdots & 0 & 2/3 & 0\\ 2/3 & 0 & 0 & 2/3 & 0 & 0 & 2/3 & \cdots & 0 & 0 & 2/3\\ \end{array} \right)^{\prime}. \end{align} In DGP3, $a_i$ and $B_i$ are homogenous across $i$, which is the setting in Example (ref). We are interested in $\phi_0$, $\Phi_0$, $F$, and $K$. Here, $u_{it}$'s, $\sigma_{t}$'s, $x_{i0,3}$'s, $\eta_t$'s, $f_0$, $\theta_{ji}$'s, $\varrho_{ji}$'s, $\vartheta_{i}$'s, and $\varepsilon_{it}$'s are mutually independent. We generate $y_{it}$ according to the model in (ref). For DGP1, we implement the estimators in (ref)-(ref) with $\mathcal{S} = \mathbf{R}^{Np\times T}$. We assess the performance of $\hat{\Pi}$, $\hat{a}, \hat{B}$, $\hat{F}$, and $\hat{K}$. By Corollary (ref), we set $\lambda_{NT} = c\sqrt{(Np+T)\log N}$ and $\delta_{NT} = 2(Np+T)\log N$ for some $c>0$. For DGP2, we implement the estimators in (ref)-(ref) and (ref) as well as below (ref) with $\mathcal{S} = \mathcal{D}_\infty$. We evaluate the performance of $\hat{\Pi}^{\diamond}$, $\hat{\Pi}^{\ast}$, $\hat{\mu}$, $\hat\Lambda$, $\hat{\phi}, \hat{\Phi}$, $\hat{F}$, and $\hat{K}$. By the discussion after Corollary (ref), we set $\lambda_{NT} = c\sqrt{(N+p+T)\log N}$ and $\delta_{NT} = 2(N+p+T)\log N$ for some $c>0$. For DGP3, we implement the estimators in (ref)-(ref) and (ref) with $\mathcal{S} = \{1_{N}\otimes \Gamma: \Gamma\in \mathbf{R}^{p\times T}\}$. We evaluate the performance of $\hat{\Pi}_0$, $\hat{\phi}_0, \hat{\Phi}_0$, $\hat{F}$, and $\hat{K}$. By Corollary (ref), we set $\lambda_{NT} = c\sqrt{(p+T)\log N}$ and $\delta_{NT} = 2\sqrt{N}(p+T)\log(N)$ for some $c>0$. To determine the optimal value of $c$, we employ the $5$-fold CV approach, as outlined in Section (ref) with $L=5$. The mean square errors of the regularized estimators ($\hat{\Pi}$, $(\hat{\Pi}^{\diamond\prime}, \sqrt{N}\hat{\Pi}^{\ast\prime})$, and $\hat{\Pi}_0$) are assessed both with fixed values of $c$ and using the CV method, with $c$ confined to $[0,2]$.\footnote{Specifically, we consider the grid set $\{0, 0.05, 0.1, 0.2, \ldots, 0.9,1,1.5,2\}$.} All simulation results are based on $200$ simulation replications. The main findings are as follows. \begin{itemize} • Nuclear norm regularization significantly enhances the performance of the estimators. In DGP1 and DPG2 (see Figures (ref) and (ref)), the mean square error of the unregularized estimator (i.e., $c=0$) remains relatively constant as both $N$ and $T$ increase (the value stays constantly around 40 in DGP1 and above 10 in DGP2), indicating potential inconsistency. Conversely, applying appropriate nuclear norm regularization (e.g., $c=1$) not only reduces the mean square error for each combination of $(N,T)$ but also drive the error towards zero as both $N$ and $T$ increase (e.g., the value for $c=1$ is getting closer to zero as $N$ and $T$ increase). This suggests that regularized estimators with a well-chosen $c$ value are consistent, aligning with Corollaries (ref) and (ref). In DGP3 (see Figure (ref)), although the mean square error of the unregularized estimator decreases with increasing $N$ (note that the scale of the vertical axis changes across columns of graphs), a properly chosen $c$ value (e.g., $c=0.6$ or $0.7$) leads to smaller errors. Thus, the simulations underscore the crucial role of nuclear norm regularization. • The regularized estimators exhibit high sensitivity to the choice of $c$. For instance, selecting $c>2$ in DGP3 can result in a larger mean square error than that of the unregularized estimator across all $(N,T)$ combinations (as seen in Figure (ref)). Therefore, careful consideration is essential when choosing $c$ in practice. • The CV approach proves effective in selecting $c$ to minimize mean square error. Across all the three DGPs, the mean square error of the regularized estimator using the CV-selected $c$ value closely approximates the smallest error obtained with fixed $c$ values (as evidenced by the blue line closely tracking the lowest point of the dash-dotted line in Figures (ref)-(ref)), irrespective of $(N,T)$ combinations. \end{itemize} Overall, these findings highlight the importance of nuclear norm regularization, the sensitivity of estimators to $c$, and the efficacy of the CV approach in selecting an optimal $c$ value for minimizing mean square error. We proceed to assess the performance of estimators other than $\hat{\Pi}$, $\hat{\Pi}^{\diamond}$, $\hat{\Pi}^{\ast}$, and $\hat{\Pi}_0$ by utilizing the CV-selected value of $c$. Tables (ref)-(ref) present their mean square errors or correct rates. The main findings are summarized as follows. First, the number factor estimators (i.e., $\hat{K}$) consistently perform well across all cases, only one correct rate falling below $100\%$. This indicates their reliability in estimating the number of factors. Second, in DPG1 and DPG2, all mean square errors decrease as both $N$ and $T$ increase, consistent with Corollaries (ref) and (ref). Similarly, in DGP3, mean square errors decrease as $N$ increases, indicating consistency as $N\to\infty$, aligning with Corollary (ref). Third, increasing $N$ consistently reduces the mean square errors of the factor estimators (i.e., $\hat{F}$) across all cases, while increasing $T$ may not have a similar effect. In addition, increasing either $N$ or $T$ tends to reduce the mean square errors of $\hat{\Pi}^{\ast}$, $\hat{\phi}$, $\hat{\Phi}$, $\hat{\phi}_0$, and $\hat{\Phi}_0$. While this phenomenon is not explicitly explained by Corollaries (ref) and (ref), it may be attributed to the homogeneity of the estimands (${\Pi}^{\ast}$, ${\phi}$, ${\Phi}$, ${\phi}_0$, and ${\Phi}_0$). In conclusion, our estimators exhibit promising performance in finite sample settings. The same findings are observed in settings with sparse $a$ and $B$, as well as in scenarios with small $p$, $N$, and $T$; see Appendix (ref) for additional simulation results. Furthermore, Appendix (ref) demonstrates the superiority of our estimators compared to existing ones. \section{Empirical Analysis} In this section, we analyze the cross section of individual stock returns in the US market using the same dataset as in Chenetal_SeimiparametricFactor_2021, originally derived from Freybergeretal_Dissecting_2017. The dataset comprises monthly returns and 36 characteristics of 12,813 individual US stocks spanning from September 1968 to May 2014. Due to a significant proportion of missing values in many stocks, we opt to exclude stocks with a sample length less than 200 to ensure that the proportion of missing values remains manageable. This results in an unbalanced panel with $N = 2,121$ and $T = 549$. Each time period includes at least 580 stocks with observations on both returns and the 36 characteristics, while each stock has observations in at least 200 time periods. Additionally, we transform the values of each characteristic to relative ranking values within the range $[-0.5, 0.5]$ in each time period. In our analysis, we consider six different model specifications. The first three specifications, denoted S1, S2, and S3, include $x_{it}$ comprising a constant and the 36 characteristics. The remaining three specifications, denoted S4, S5, and S6, involve $x_{it}$ consisting of a constant and linear B-splines of 18 characteristics with one internal knot, as studied in Chenetal_SeimiparametricFactor_2021. Refer to their paper for the 18 characteristics. In S1 and S4, we explore an unconstrained conditional factor model (corresponding to the setup in Example (ref)), where $a_{i}$ and $B_i$ can vary heterogeneously across $i$. For S2 and S5, we investigate a semiparametric conditional factor model (corresponding to the setup in Example (ref)), where the rows of $a_i$ and $B_i$ corresponding to the nonconstant explanatory variables in $x_{it}$ are constrained to be homogeneous. Lastly, S3 and S6 examine a homogeneous conditional factor model (corresponding to the setup in Example (ref)), where $a_i$ and $B_i$ are constrained to be homogeneous. We estimate the models for $K=1,2,\ldots, 10$ by using our new method and select the regularization parameter using the 5-fold CV approach as outlined in Section (ref). Specifically, we set $\lambda_{NT} = c\sqrt{(Np+T)\log N}$ for S1 and S4, $\lambda_{NT} = c\sqrt{(N+p+T)\log N}$ for S2 and S5, $\lambda_{NT} = c\sqrt{(p+T)\log N}$ for S3 and S6, and choose $c$ from the set $\{0, 0.01, 0.02, 0.05, 0.1, 0.2, 0.5,1,2,5\}/100$. For a comparison, we also evaluate Chenetal_SeimiparametricFactor_2021's regressed-PCA method, denoted R1 and R2, alongside the homogeneous conditional factor models (S3 and S6). To assess the performance of the models, we adopt various goodness-of-fit measures. First, we consider different types of in-sample $R^2$ measures: \begin{align} R^2 & = 1-\frac{\sum_{i, t}(y_{it}- x_{it}^{\prime}\hat{a}_{i} - x_{it}^{\prime}\hat{B}_{i}\hat{f}_t)^2}{\sum_{i, t} y_{it}^{2}}, \\ R^2_{T,N} & = 1 - \frac{1}{N} \sum_{i} \frac{\sum_{t}(y_{it}- x_{it}^{\prime}\hat{a}_{i} - x_{it}^{\prime}\hat{B}_{i}\hat{f}_t)^2}{\sum_{t}y_{it}^{2}},\\ R^2_{N,T} & = 1 - \frac{1}{T} \sum_{t} \frac{\sum_{i}(y_{it}- x_{it}^{\prime}\hat{a}_{i} - x_{it}^{\prime}\hat{B}_{i}\hat{f}_t)^2}{\sum_{i}y_{it}^{2}}, \end{align} where $\hat{a}\equiv (\hat{a}_1^{\prime},\hat{a}_2^{\prime},\ldots, \hat{a}_{N}^{\prime})^{\prime}$, $\hat{B} \equiv (\hat{B}_1^{\prime},\hat{B}_2^{\prime},\ldots, \hat{B}_{N}^{\prime})^{\prime}$, and $\hat{F} \equiv(\hat{f}_1,\hat{f}_2,\ldots, \hat{f}_{T})^{\prime}$. Here, the first one is total $R^2$, measuring the overall explanatory power of the models. The second one measures the cross-sectional average of time series $R^2$ across all stocks, reflecting the ability of the models to capture common variation in asset returns. The third one measures the time series average of cross-sectional $R^2$, which is the one of interest for evaluating the models' ability to explain the cross-section of average returns. Second, we assess out-of-sample prediction. For $t\geq 300$, we utilize the data up to $t-1$ for estimation and obtain estimators, say $\hat{a}_{it}$, $\hat{B}_{it}$, and $\hat{F}_{t}\equiv (\hat{f}^{(t)}_{1},\hat{f}^{(t)}_{2},\ldots, \hat{f}^{(t)}_{t-1})^{\prime}$. The out-of-sample prediction of $y_{it}$ is then computed as $x_{it}^{\prime}\hat{a}_{it} - x_{it}^{\prime}\hat{B}_{it}\hat{\lambda}_t$, where $\hat{\lambda}_{t} = \sum_{s\leq t-1}\hat{f}^{(t)}_s/(t-1)$. Analogously, we define three types of out-of-sample predictive $R^2$'s by replacing $\hat{a}_i$, $\hat{B}_i$ and $\hat{f}_t$ with $\hat{a}_{it}$, $\hat{B}_{it}$ and $\hat{\lambda}_t$: \begin{align} R^{2}_{O} & = 1-\frac{\sum_{i, t\geq 300}(y_{it}- x_{it}^{\prime}\hat{a}_{it} - x_{it}^{\prime}\hat{B}_{it}\hat{\lambda}_t)^2}{\sum_{i, t\geq 300} y_{it}^{2}}, \\ R^2_{T,N,O} & = 1 - \frac{1}{N} \sum_{i} \frac{\sum_{t\geq 300}(y_{it}- x_{it}^{\prime}\hat{a}_{it} - x_{it}^{\prime}\hat{B}_{it}\hat{\lambda}_t)^2}{\sum_{t\geq 300}y_{it}^{2}},\\ R^2_{N,T,O} & = 1 - \frac{1}{T-299} \sum_{t\geq 300} \frac{\sum_{i}(y_{it}- x_{it}^{\prime}\hat{a}_{it} - x_{it}^{\prime}\hat{B}_{it}\hat{\lambda}_t)^2}{\sum_{i}y_{it}^{2}}. \end{align} The results depicted in Figure (ref) yield several key observations. Firstly, the in-sample $R^2$ values of our methods (S1, S2, S3, S4, S5, and S6) exhibit an increasing trend as the number of factors $K$ rises, while the out-of-sample $R^2$ metrics remain unaffected by changes in $K$. This constancy arises from the fact that $\hat{\lambda} = \sum_{t\leq T}\hat{f}_t/T = \hat{F}^{\prime}1_{T}/T = \hat{B}^{\prime}\hat{\Pi}1_{T}/(NT)$, $\hat{a} + \hat{B}\hat{\lambda} = \hat{\Pi}1_{T}/T$, rendering the out-of-sample predictions of $y_{it}$ independent of $K$. Secondly, among the linear models (S1, S2, and S3), S1 consistently outperforms others in terms of in-sample $R^2$ values across all tested values of $K$. Conversely, S3 emerges as the top performer in out-of-sample $R^2$ metrics for all configurations. This suggests that enforcing homogeneity of $a_i$ and $B_i$ across $i$ may improve the model's out-of-sample predictability, despite potentially compromising the in-sample fit. Similarly, for the spline models (S4, S5, and S6), enforcing homgogeneity of $a_i$ and $B_i$ across $i$ yields improvements in out-of-sample predictability. Thirdly, S5 and S6 demonstrate superior out-of-sample performance compared to S2 and S3, respectively. This underscores the potential benefits of incorporating spline transformations of characteristics, emphasizing the significance of capturing nonlinear relationships. Lastly, the importance of nonlinearity is also observed for the regressed-PCA method; R2 has larger out-of-sample $R^2$ values than R1. However, S3 and S6 exhibit better both in-sample and out-of-sample performance than R1 and R2, respectively. This implies that our method outperforms the regressed-PCA method. In conclusion, while S1 exhibits the most favorable in-sample performance, S6 stands out for its superior out-of-sample predictive capabilities. \section{Concluding Remarks} In this paper, we introduced a nuclear norm regularized estimation approach for high-dimensional conditional factor models and established large sample properties of the estimators. Our method provides a unified framework for estimating various conditional factor models, facilitating the derivation of new asymptotic results while addressing the limitations of existing methods, which are often model-specific or restrictive. We applied this method to analyze the cross section of individual US stock returns, uncovering potential improvements in out-of-sample performance by enforcing homogeneity of $a_i$ and $B_i$ across $i$. Our results also show that the proposed method outperforms existing alternatives. In asset pricing, addressing key inference problems\textemdash such as testing for zero pricing errors and conducting specification tests for risk exposure functions\textemdash is crucial for evaluating and comparing factor models. Previous studies, including XiaYuan_InfMatrix_2019, Chenetal_Uncertainty_2019, and Chernozhukovetal_InferenceLowRank_2021, have investigated debiasing techniques in trace linear regression models with $p=1$ and $a_i=0$, with applications to matrix completion, PCA with missing data, and heterogeneous treatment effects. However, these methods are not applicable to our framework, which accommodates large $p$ and $a_i\neq 0$, as is often the case in asset pricing. Developing a general inferential method within this framework presents an intriguing avenue for future research. \spacingset{1.75} \addcontentsline{toc}{section}{References} \putbib

\pgfplotstableread{ c d1 d2 d3 d4 d5 d6 d10 d20 d30 d40 d50 d60 0 43.14426637 43.33627923 43.50645384 41.59508512 41.76457064 41.97837082 4.169800373 3.996109432 3.850109822 1.821119546 1.686205982 1.594316174 0.05 32.08304342 32.03894122 32.06602237 29.30623291 28.95795319 29.00921376 4.169800373 3.996109432 3.850109822 1.821119546 1.686205982 1.594316174 0.1 23.54430716 23.36835109 23.36198416 20.10658389 19.39373486 19.32940007 4.169800373 3.996109432 3.850109822 1.821119546 1.686205982 1.594316174 0.2 12.23444796 11.98405134 11.9600371 8.540870773 7.7832915 7.672400098 4.169800373 3.996109432 3.850109822 1.821119546 1.686205982 1.594316174 0.3 6.595638549 6.415623745 6.394547331 3.49027754 3.068663151 2.992567961 4.169800373 3.996109432 3.850109822 1.821119546 1.686205982 1.594316174 0.4 4.395456642 4.232468312 4.179969769 1.928527141 1.73600281 1.659238317 4.169800373 3.996109432 3.850109822 1.821119546 1.686205982 1.594316174 0.5 4.059342044 3.845888286 3.73065423 1.80548605 1.680349044 1.564735503 4.169800373 3.996109432 3.850109822 1.821119546 1.686205982 1.594316174 0.6 4.639974264 4.49352774 4.390555511 2.427700319 2.35902577 2.207595137 4.169800373 3.996109432 3.850109822 1.821119546 1.686205982 1.594316174 0.7 5.390295758 5.250412363 5.115788629 2.804322671 2.608810821 2.519225758 4.169800373 3.996109432 3.850109822 1.821119546 1.686205982 1.594316174 0.8 5.778735496 5.546078172 5.365294694 2.653498884 2.361544595 2.238620608 4.169800373 3.996109432 3.850109822 1.821119546 1.686205982 1.594316174 0.9 5.729392572 5.442801024 5.289179373 2.366988798 2.176505816 2.187699483 4.169800373 3.996109432 3.850109822 1.821119546 1.686205982 1.594316174 1 5.565479792 5.356919464 5.407878306 2.306518885 2.395930596 2.601045541 4.169800373 3.996109432 3.850109822 1.821119546 1.686205982 1.594316174 1.5 7.71740683 8.151418736 8.582979736 3.929517516 4.16301821 4.405527265 4.169800373 3.996109432 3.850109822 1.821119546 1.686205982 1.594316174 2 10.73069694 11.37161622 12.00254787 5.752482952 6.111023504 6.498363794 4.169800373 3.996109432 3.850109822 1.821119546 1.686205982 1.594316174 }\dgpa

figure[figure omitted — 4,180 chars of source]

\pgfplotstableread{ c d1 d2 d3 d4 d5 d6 d10 d20 d30 d40 d50 d60 0 24.55839437 24.16194186 23.96283721 23.83049028 23.47120882 23.28934132 0.204140592 0.150309638 0.150082636 0.133793707 0.109463196 0.07745866 0.05 3.571337437 3.578098296 3.576812316 3.502824797 3.478312025 3.482357528 0.204140592 0.150309638 0.150082636 0.133793707 0.109463196 0.07745866 0.1 2.970929389 3.000093571 3.010431768 2.942819685 2.922942773 2.935466205 0.204140592 0.150309638 0.150082636 0.133793707 0.109463196 0.07745866 0.2 2.142551935 2.128320948 2.11247134 2.097298787 2.037780035 2.022078369 0.204140592 0.150309638 0.150082636 0.133793707 0.109463196 0.07745866 0.3 1.489770499 1.434977475 1.394806768 1.446744207 1.361641069 1.310895211 0.204140592 0.150309638 0.150082636 0.133793707 0.109463196 0.07745866 0.4 1.001448585 0.915966195 0.850249751 0.962615026 0.866175217 0.789574242 0.204140592 0.150309638 0.150082636 0.133793707 0.109463196 0.07745866 0.5 0.653469242 0.556268875 0.475153265 0.61705589 0.52233315 0.439336609 0.204140592 0.150309638 0.150082636 0.133793707 0.109463196 0.07745866 0.6 0.420501601 0.328901216 0.254188009 0.38371882 0.300425231 0.227704386 0.204140592 0.150309638 0.150082636 0.133793707 0.109463196 0.07745866 0.7 0.27854637 0.20363607 0.150144267 0.23857852 0.171805881 0.118650583 0.204140592 0.150309638 0.150082636 0.133793707 0.109463196 0.07745866 0.8 0.204489441 0.15025438 0.120881442 0.159256028 0.109669903 0.077293761 0.204140592 0.150309638 0.150082636 0.133793707 0.109463196 0.07745866 0.9 0.177889305 0.141664271 0.128880179 0.125945557 0.090375619 0.073115778 0.204140592 0.150309638 0.150082636 0.133793707 0.109463196 0.07745866 1 0.181336614 0.155644406 0.147303773 0.1217567 0.094184922 0.08247404 0.204140592 0.150309638 0.150082636 0.133793707 0.109463196 0.07745866 1.5 0.315932678 0.281357944 0.270269968 0.209020799 0.169102144 0.150959432 0.204140592 0.150309638 0.150082636 0.133793707 0.109463196 0.07745866 2 0.509283795 0.457452371 0.442318719 0.337662003 0.274968678 0.247084728 0.204140592 0.150309638 0.150082636 0.133793707 0.109463196 0.07745866 }\dgpb

figure[figure omitted — 4,135 chars of source]

\pgfplotstableread{ c d1 d2 d3 d4 d5 d6 d10 d20 d30 d40 d50 d60 0 0.315186412 0.151644977 0.074242933 0.315283062 0.151342761 0.074203126 0.053402803 0.02746442 0.01344317 0.048648191 0.025929072 0.012648708 0.05 0.274626226 0.132247697 0.06456282 0.275495043 0.132342158 0.064718108 0.053402803 0.02746442 0.01344317 0.048648191 0.025929072 0.012648708 0.1 0.238040197 0.114630754 0.055761851 0.239448271 0.115010093 0.056056517 0.053402803 0.02746442 0.01344317 0.048648191 0.025929072 0.012648708 0.2 0.176131736 0.084579088 0.040755953 0.17797218 0.085208616 0.0411667 0.053402803 0.02746442 0.01344317 0.048648191 0.025929072 0.012648708 0.3 0.128212659 0.061191003 0.029141246 0.129704323 0.061663814 0.029457328 0.053402803 0.02746442 0.01344317 0.048648191 0.025929072 0.012648708 0.4 0.093138731 0.044170528 0.020841481 0.093585177 0.044103773 0.020858001 0.053402803 0.02746442 0.01344317 0.048648191 0.025929072 0.012648708 0.5 0.069865494 0.033221388 0.0157499 0.068656402 0.032279573 0.015298531 0.053402803 0.02746442 0.01344317 0.048648191 0.025929072 0.012648708 0.6 0.056914134 0.027646015 0.013456611 0.054046877 0.02592295 0.01265732 0.053402803 0.02746442 0.01344317 0.048648191 0.025929072 0.012648708 0.7 0.05207068 0.026234379 0.013270325 0.048375435 0.024216932 0.012382856 0.053402803 0.02746442 0.01344317 0.048648191 0.025929072 0.012648708 0.8 0.053097026 0.027734333 0.014472316 0.049233867 0.025705647 0.013595737 0.053402803 0.02746442 0.01344317 0.048648191 0.025929072 0.012648708 0.9 0.058026694 0.031014026 0.01644803 0.054154238 0.028934535 0.01550308 0.053402803 0.02746442 0.01344317 0.048648191 0.025929072 0.012648708 1 0.065243985 0.035206175 0.018792668 0.061072756 0.032884238 0.017706274 0.053402803 0.02746442 0.01344317 0.048648191 0.025929072 0.012648708 1.5 0.116005859 0.063515101 0.034433462 0.108498852 0.05938283 0.032383674 0.053402803 0.02746442 0.01344317 0.048648191 0.025929072 0.012648708 2 0.187507243 0.103322792 0.056400679 0.175321907 0.096647806 0.052983099 0.053402803 0.02746442 0.01344317 0.048648191 0.025929072 0.012648708 }\dgpc

figure[figure omitted — 4,202 chars of source]

{18pt}

table[table omitted — 1,579 chars of source]

{12pt}

table[table omitted — 2,204 chars of source]

{18pt}

table[table omitted — 1,664 chars of source]

\pgfplotstableread{ k x1 x2 x3 x4 x5 x6 x7 x8 x9 x10 x11 x12 1 19.78397619 16.0496283 22.03292447 19.73344125 15.9938032 21.98069301 0.425273488 0.311054379 0.573209396 0.050534939 0.055825105 0.05223146 2 23.0014627 17.73511348 24.76457163 22.9527652 17.67916671 24.71132158 0.425273488 0.311054379 0.573209396 0.0486975 0.055946769 0.053250046 3 25.0268463 19.28833767 26.56515449 24.97852174 19.23405106 26.51237451 0.425273488 0.311054379 0.573209396 0.048324558 0.054286606 0.052779978 4 26.60842464 20.53462093 27.95754052 26.56039318 20.48145688 27.90540099 0.425273488 0.311054379 0.573209396 0.048031457 0.053164058 0.052139525 5 27.82976075 20.96279622 28.53417785 27.78323551 20.91129621 28.48331217 0.425273488 0.311054379 0.573209396 0.046525246 0.051500006 0.050865675 6 28.88765005 21.47669566 29.37062689 28.84371213 21.42903585 29.32555334 0.425273488 0.311054379 0.573209396 0.043937918 0.047659804 0.045073547 7 29.7444564 22.12044479 30.12003611 29.71221227 22.0872535 30.08730266 0.425273488 0.311054379 0.573209396 0.032244139 0.033191291 0.032733452 8 30.59685879 22.65088575 30.85109124 30.56593078 22.62033581 30.8204353 0.425273488 0.311054379 0.573209396 0.030928004 0.030549947 0.030655941 9 31.39774395 23.13030608 31.48915112 31.36997237 23.10203218 31.46277561 0.425273488 0.311054379 0.573209396 0.027771583 0.028273903 0.026375514 10 32.21667299 23.51575945 31.98281204 32.18937347 23.48885673 31.95690381 0.425273488 0.311054379 0.573209396 0.02729952 0.026902723 0.025908226 }\empa

\pgfplotstableread{ k x1 x2 x3 x4 x5 x6 x7 x8 x9 1 18.76569193 15.4511116 21.01375672 18.68086381 15.36174128 20.91939556 0.431590505 0.291440272 0.60915766 2 21.52130469 16.83284757 23.43484188 21.43867791 16.74233284 23.33699895 0.431590505 0.291440272 0.60915766 3 23.33659308 18.03321005 24.83143466 23.25387779 17.94206342 24.73359332 0.431590505 0.291440272 0.60915766 4 24.79577503 19.14549108 26.14152398 24.71289728 19.05490673 26.04440973 0.431590505 0.291440272 0.60915766 5 25.96589867 19.72514425 26.85217679 25.88448709 19.63622678 26.75547804 0.431590505 0.291440272 0.60915766 6 26.96752425 20.17916008 27.58643068 26.88682413 20.0923324 27.49223402 0.431590505 0.291440272 0.60915766 7 27.82991203 20.68437029 28.36175757 27.74897961 20.59804742 28.2672491 0.431590505 0.291440272 0.60915766 8 28.66900065 21.04614467 28.84771355 28.58994915 20.96091933 28.75562084 0.431590505 0.291440272 0.60915766 9 29.40747168 21.5512486 29.58116539 29.33651189 21.47932124 29.50298396 0.431590505 0.291440272 0.60915766 10 30.15945774 21.88639719 30.05057626 30.08881386 21.81445376 29.97273641 0.431590505 0.291440272 0.60915766 }\empb

\pgfplotstableread{ k x1 x2 x3 x4 x5 x6 x7 x8 x9 x10 x11 x12 1 18.20072892 15.26663258 19.16387575 17.82355183 14.79504704 18.7366448 0.834675501 0.556749775 1.07 0.377177084 0.471585541 0.427230945 2 19.47372843 16.14815566 20.48679072 19.08796113 15.67154105 20.04761936 0.834675501 0.556749775 1.07 0.385767298 0.476614612 0.439171365 3 21.12188542 17.6823269 22.84629057 20.7800035 17.19922972 22.41672007 0.834675501 0.556749775 1.07 0.341881917 0.48309718 0.429570499 4 22.54645031 18.7879371 24.6562571 22.21722146 18.3402633 24.25647883 0.834675501 0.556749775 1.07 0.329228846 0.447673799 0.399778265 5 22.88209261 19.15667531 25.0520791 22.67599493 18.89002953 24.78176698 0.834675501 0.556749775 1.07 0.206097677 0.266645782 0.270312118 6 23.34220596 19.59630166 25.6214229 23.22021748 19.45253597 25.50152092 0.834675501 0.556749775 1.07 0.12198848 0.14376569 0.119901987 7 23.66748845 19.94946119 26.04318542 23.54643426 19.80776749 25.92288627 0.834675501 0.556749775 1.07 0.121054191 0.141693703 0.120299146 8 23.91039198 20.20174572 26.27570259 23.88275982 20.16815315 26.23980135 0.834675501 0.556749775 1.07 0.027632156 0.033592571 0.035901235 9 24.06037721 20.3581503 26.44275535 24.03348048 20.32607118 26.40716528 0.834675501 0.556749775 1.07 0.026896735 0.032079122 0.035590068 10 24.22196944 20.53850597 26.61357558 24.19339303 20.50467011 26.5743102 0.834675501 0.556749775 1.07 0.028576409 0.033835869 0.039265382 }\empc

\pgfplotstableread{ k x1 x2 x3 x7 x8 x9 1 2.4151775 0.270744452 2.11800967 0.794638557 0.462437015 0.978870287 2 3.61323705 0.858619263 3.441763274 0.794638557 0.462437015 0.978870287 3 4.559943135 1.324080283 4.437781153 0.794638557 0.462437015 0.978870287 4 4.765972671 1.656090297 4.623180378 0.794638557 0.462437015 0.978870287 5 8.747233654 4.616366892 8.737732062 0.794638557 0.462437015 0.978870287 6 11.37130751 7.207234276 11.08379845 0.794638557 0.462437015 0.978870287 7 13.40334322 9.618156033 13.06671806 0.794638557 0.462437015 0.978870287 8 16.95898943 13.07758168 16.97741614 0.794638557 0.462437015 0.978870287 9 19.23457836 15.66545377 19.6535871 0.794638557 0.462437015 0.978870287 10 20.03926363 16.49919903 20.35534774 0.794638557 0.462437015 0.978870287 }\empregressedpca

\pgfplotstableread{ k x1 x2 x3 x4 x5 x6 x7 x8 x9 x10 x11 x12 1 18.76673286 15.30625399 20.95104112 18.73792622 15.2719071 20.91618163 0.404136198 0.304530481 0.577825976 0.028806637 0.034346886 0.034859486 2 21.6242784 16.72190856 23.50696865 21.59610295 16.68812602 23.47142938 0.404136198 0.304530481 0.577825976 0.028175446 0.033782536 0.035539267 3 23.34308699 17.90151103 24.95455185 23.31523509 17.8699816 24.91962685 0.404136198 0.304530481 0.577825976 0.0278519 0.031529431 0.034924996 4 24.59322027 18.79766041 26.07427425 24.56672625 18.76828182 26.04197985 0.404136198 0.304530481 0.577825976 0.02649402 0.02937859 0.032294399 5 25.63147032 19.16174473 26.5829391 25.60611835 19.13341529 26.55128872 0.404136198 0.304530481 0.577825976 0.025351971 0.028329446 0.031650379 6 26.47450801 19.50841809 27.24773126 26.45065155 19.48133058 27.2192703 0.404136198 0.304530481 0.577825976 0.023856461 0.027087509 0.028460957 7 27.21814799 19.96035984 27.88220827 27.19451124 19.93430927 27.85354374 0.404136198 0.304530481 0.577825976 0.023636755 0.026050576 0.02866453 8 27.88940635 20.31018598 28.44764624 27.87457921 20.29377989 28.43109794 0.404136198 0.304530481 0.577825976 0.014827139 0.016406092 0.0165483 9 28.48601845 20.60910367 28.89216548 28.47179514 20.59521179 28.87680388 0.404136198 0.304530481 0.577825976 0.014223309 0.013891884 0.015361596 10 29.0886451 20.88821919 29.35572407 29.07768941 20.87695932 29.34472081 0.404136198 0.304530481 0.577825976 0.010955691 0.011259878 0.011003255 }\empaspline

\pgfplotstableread{ k x1 x2 x3 x4 x5 x6 x7 x8 x9 1 18.21376464 15.08225972 19.26981639 18.14018907 14.97738242 19.17183983 0.466663562 0.30740805 0.682441868 2 21.39172538 16.8779238 23.21987308 21.32001502 16.79694204 23.12958975 0.466663562 0.30740805 0.682441868 3 23.21036393 18.09507539 24.6618819 23.13834286 18.01638289 24.57154117 0.466663562 0.30740805 0.682441868 4 24.60064279 19.09498298 25.87211299 24.52912277 19.01716806 25.78375924 0.466663562 0.30740805 0.682441868 5 25.77610958 19.68258797 26.58531642 25.7068148 19.60723891 26.49836989 0.466663562 0.30740805 0.682441868 6 26.8378879 20.15756436 27.44367613 26.77045494 20.0850586 27.36156266 0.466663562 0.30740805 0.682441868 7 27.75956697 20.68667443 28.41417755 27.70003269 20.62507016 28.34869038 0.466663562 0.30740805 0.682441868 8 28.60887668 21.04634712 28.93592615 28.55213651 20.98544788 28.87342997 0.466663562 0.30740805 0.682441868 9 29.37462613 21.40000969 29.43111884 29.31848225 21.33905417 29.36910247 0.466663562 0.30740805 0.682441868 10 30.11243206 21.9021766 30.1198585 30.05793753 21.84589042 30.06123271 0.466663562 0.30740805 0.682441868 }\empbspline

\pgfplotstableread{ k x1 x2 x3 x4 x5 x6 x7 x8 x9 x10 x11 x12 1 13.10776141 9.670879077 12.33359277 12.69000145 9.123067886 11.87889585 0.864456072 0.604413358 1.080307192 0.417759955 0.547811191 0.454696914 2 15.01725349 11.70886333 13.86303032 14.59759032 11.17794811 13.39901714 0.864456072 0.604413358 1.080307192 0.419663172 0.530915224 0.464013174 3 17.81232534 14.6544419 17.69184302 17.42751736 14.1291686 17.21712697 0.864456072 0.604413358 1.080307192 0.384807981 0.525273295 0.474716052 4 18.60071732 15.08747128 18.68605936 18.21071773 14.70728754 18.25149473 0.864456072 0.604413358 1.080307192 0.389999586 0.380183732 0.434564631 5 20.44039553 16.53566919 21.32885478 20.16753269 16.20164243 21.02411215 0.864456072 0.604413358 1.080307192 0.272862839 0.334026752 0.30474263 6 20.87068851 16.99828725 21.72060871 20.74406607 16.8352111 21.55560507 0.864456072 0.604413358 1.080307192 0.126622439 0.163076146 0.165003644 7 21.90544273 18.04922017 23.28705198 21.69665669 17.79859404 22.98446656 0.864456072 0.604413358 1.080307192 0.208786038 0.250626122 0.302585418 8 22.72418108 18.75859548 24.22033442 22.5712626 18.57500396 24.00162082 0.864456072 0.604413358 1.080307192 0.152918472 0.18359152 0.218713602 9 23.04429504 19.11544664 24.58025596 22.8526714 18.89543952 24.27495593 0.864456072 0.604413358 1.080307192 0.191623645 0.220007119 0.305300034 10 23.44119958 19.64001011 25.17929301 23.28802816 19.43136666 24.93887761 0.864456072 0.604413358 1.080307192 0.153171413 0.208643453 0.240415398 }\empcspline

\pgfplotstableread{ k x1 x2 x3 x7 x8 x9 1 3.730955782 1.457065701 3.445342152 0.841677528 0.514120961 1.016947396 2 9.803144174 6.509265921 9.005506164 0.841677528 0.514120961 1.016947396 3 11.3797279 7.958427723 10.11507774 0.841677528 0.514120961 1.016947396 4 15.31450372 11.87442339 14.35138896 0.841677528 0.514120961 1.016947396 5 16.14869052 12.70308811 15.42907235 0.841677528 0.514120961 1.016947396 6 16.54821647 13.03980857 15.86031005 0.841677528 0.514120961 1.016947396 7 17.96624755 14.09265656 17.85728909 0.841677528 0.514120961 1.016947396 8 18.28952 14.42483044 18.17414921 0.841677528 0.514120961 1.016947396 9 19.18600499 15.32249453 19.28049809 0.841677528 0.514120961 1.016947396 10 19.50205769 15.62131374 19.54298715 0.841677528 0.514120961 1.016947396 }\empregressedpcaspline

figure[figure omitted — 9,130 chars of source]
bibunit\def\spacingset#1{ {#1}} \spacingset{1.8} {4pt} {4pt} {4pt} {4pt}