EconBase
← Back to paper

Semiparametric Conditional Factor Models in Asset Pricing

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

121,242 characters · 0 sections · 100 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Semiparametric Conditional Factor Models in Asset Pricing

bibunit\pdfbookmark[1]{Title}{title} \begin{abstract} We introduce a simple and tractable methodology for estimating semiparametric conditional latent factor models. Our approach disentangles the roles of characteristics in capturing factor betas of asset returns from “alpha.” We construct factors by extracting principal components from Fama-MacBeth managed portfolios. Applying this methodology to the cross-section of U.S. individual stock returns, we find compelling evidence of substantial nonzero pricing errors, even though our factors demonstrate superior performance in standard asset pricing tests. Unexplained “arbitrage” portfolios earn high Sharpe ratios, which decline over time. Combining factors with these orthogonal portfolios produces out-of-sample Sharpe ratios exceeding 4. \end{abstract} \begin{center} Keywords: Characteristics, Fama-MacBeth regression, factor models, managed portfolios, nonlinearity, PCA, sieve estimation \end{center} \section{Introduction} A central question in empirical asset pricing is why different assets earn different average returns. While asset pricing theory attributes cross-sectional differences in returns to variations in risk exposures, considerable evidence suggests that mispricing\textemdash captured by the dependence of returns on asset characteristics\textemdash also plays a significant role, suggesting potential market inefficiencies. Much of the debate centers on multi-factor models that aim to link average returns to factor loadings, building on the influential framework of FamaFrench_Commonrisk_1993, who introduced a portfolio-sorting approach to constructing asset pricing factors. Following their seminal work, researchers have proposed hundreds of factors, leading to what Cochrane_Presidential_2011 memorably termed the “factor zoo,” a concept further explored by Harveyetal_FactorZoo_2016. While some factor models have an explicit justification based on economic theory, many implicitly rely on the idea that factors capture common variation in returns, thus appealing to arbitrage pricing theory and its extensions Ross_APT_1976,ChamberlainRothschild_FactorStuctures_1982,ConnorKorajczyk_Performance_1986,ConnorKorajczyk_RiskReturn_1988,Reisman/IAPT:1992. Since implementing the latter requires estimating the conditional covariance matrix of returns, which becomes impractical when the number of assets ($N$) exceeds the number of time periods ($T$), most studies rely on asset characteristics to proxy for (imperfectly measured) risk exposures, often employing the portfolio-sorting approach. However, this makes distinguishing between risk-based explanations and those rooted in mispricing virtually impossible, as exemplified by the “characteristics versus covariances” debate DanielTitman_Characteristics_1997. In our analysis we consider a canonical conditional factor model: \begin{align} y_{it} = \alpha(z_{it}) + \beta(z_{it})^{\prime}f_{t} + \varepsilon_{it}, i=1,\ldots,N,t=1,\ldots,T. \end{align} Here, $y_{it}$ is the excess return of asset $i$ at time period $t$, $z_{it}$ is an $M\times 1$ vector of pre-specified asset characteristics (which may include a constant term) that is observed at the beginning of time period $t$,\footnote{In asset pricing, $z_{i,t-1}$, the characteristics observed at time period $t-1$, is usually used in (ref). For notational simplicity, here we use $z_{it}$ rather than $z_{i,t-1}$.} $f_t$ is a $K\times 1$ vector of unobserved {\em latent} factors, $\beta(\cdot)$ is a $K\times 1$ vector of unknown factor loading functions, $\alpha(\cdot)$ is an unknown intercept function, $\varepsilon_{it}$ is the idiosyncratic component that is orthogonal to the common factors $f_t$.\footnote{While our main focus is on cross-sectional asset pricing, the model has other potential applications, which include modelling the implied volatility of options Parketal_FactorDynamics_2009 and describing consumer demand system Lewbel_DemandSystems_1991, among others.} The model describes a {\em conditional} factor model, in the sense that it captures time-variation in asset return exposures to the common factors (i.e., $\beta(z_{it})$) as well as the pricing errors (i.e., $\alpha(z_{it})$), which are both functions of characteristics. This model is well suited for resolving the “characteristics versus covariances” debate, since it potentially allows for distinguishing between the risk and mispricing explanations of the role of characteristics in predicting asset returns.\footnote{While useful, it might not be sufficient to resolve the debate, since distinguishing between the different explanations requires understanding the economic nature of the latent factors - e.g., see Kozaketal_Interpreting_2018.} Meanwhile, the model allows for pooling the information in a multitude of characteristics and summarizing the common variation using a small number of factors, thereby helping to “tame the factor zoo.” The challenge of using the model is threefold: first, the identities of the common factors $f_t$ are unknown since the factors are latent; second, the functional forms of the $alpha$ and $beta$ functions are also generally unknown; finally, the cross-sectional dimension $N$ is typically much larger than the sample time-series length $T$, which renders standard tools of factor analysis inapplicable, especially when conditional covariances are time-varying. We introduce a simple and tractable estimation method to recover both the latent factors and the functional parameters of the model, alongside formal inference procedures. First, we develop an easy-to-compute estimator for $\alpha(\cdot)$, $\beta(\cdot)$, and $f_t$ based on a sieve approximation to the nonparametric functions $\alpha(\cdot)$ and $\beta(\cdot)$. The estimation involves two steps: (i) regressing $y_{it}$ on sieve functions of $z_{it}$ for each $t$; and (ii) applying principal component analysis (PCA) to the estimated coefficients obtained in step (i). We refer to this approach as the regressed-PCA. The first step aligns with the cross-sectional regressions of FamaMacBeth_RiskReturn_1973, where the estimated coefficients at each point in time represent returns of “pure play” characteristic-managed portfolios. The second step is effectively a standard PCA on a relatively small set of characteristic-managed portfolios constructed via Fama-MacBeth regressions. Second, we develop a bootstrap inference framework to assess the significance of $\alpha(\cdot)$ as well as test the linearity of $\alpha(\cdot)$ and $\beta(\cdot)$. Third, we establish large-sample properties of the estimators under mild conditions, including consistency, rate of convergence, and asymptotic normality, as well as the validity of the proposed tests. Notably, the asymptotic results possess several advantages: (i) they do not require large $T$; (ii) they accommodate time-varying and potentially nonstationary $z_{it}$; (iii) they apply to unbalanced panels, which is particularly beneficial for securities with varying lifespans. Our Monte Carlo simulations demonstrate that the proposed estimators and tests exhibit satisfactory finite-sample performance and remain robust even when $T$ is small, provided $N$ is large. In addition to offering formal inference procedures and well-founded asymptotic properties, regressed-PCA presents several advantages over existing methods such as instrumented PCA (IPCA) Kellyetal_Characteristics_2019 and projected-PCA Fanetal_ProjectedPCA_2016,Kimetal_Arbitrage_2019. Specifically, regressed-PCA is computationally efficient and accommodates nonzero alphas, time-varying characteristics, unbalanced panels, and short samples, making it particularly well-suited for empirical asset pricing applications. We apply our new methodology to analyzing the cross-section of individual stock returns. Our analysis uses the same dataset as Kellyetal_Characteristics_2019, the study most closely aligned with ours in terms of empirical aims. However, our econometric approach and empirical findings differ significantly from theirs. First, unlike Kellyetal_Characteristics_2019, Kellyetal_IPCA_2017, our method does not aim to simultaneously maximize the “fit” of the factor model to individual asset returns in both the time-series and cross-section. Instead, we extract factors that capture the most time-series comovement within a set of portfolios, which, in turn, reflect the most cross-sectional variation in individual asset returns. Second, our approach allows the $\alpha(\cdot)$ and $\beta(\cdot)$ functions to be nonlinear. We test\textemdash and reject\textemdash the validity of linear specifications empirically, which reveals the strong evidence of nonlinearity in factor loadings and pricing errors. Third, our inference procedure also enables us to test the significance of pricing errors. Our empirical results reveal that the pricing errors associated with many characteristics are statistically significant, leading to the rejection of the risk-based model. Lastly, our methodology facilitates rolling sub-sample analyses to accommodate evolving factor dynamics, as it does not rely on large $T$. We find that both in-sample and out-of-sample goodness-of-fit measures for all factor models decline from 1970 until roughly 2000 but improve thereafter. This pattern aligns with the findings in Campbell2001have and Campbell2022idiosyncratic on the time-variation in the amount of idiosyncratic volatility in the U.S. stock market. We also document a significant decline in pricing errors in more recent years, particularly since 2000. This decline may reflect the growing prevalence of quantitative investing, which reduces mispricing by exploiting characteristic-related anomalies, as suggested by McLeanPontiff_DoesAcademic_2016 and Greenetal_Characteristics_2017. Based on these findings, we construct trading strategies. The pure-$alpha$ portfolios constructed based on nonzero pricing errors are associated with annualized Sharpe ratios typically above 3 (as is common in the literature, we refer to these as “arbitrage” portfolios, even though their returns are far from riskless). Meanwhile, the mean-variance efficient (MVE) portfolios constructed from the corresponding factors deliver substantially lower Sharpe ratios. This is different from the case in IPCA, where the Sharpe ratios of pure-alpha and MVE factor portfolios are comparable. Moreover, we approximate the stock market's MVE portfolio with the combined MVE portfolios of the pure-alpha portfolios and factors. Regressed-PCA consistently yields higher Sharpe ratios than IPCA, highlighting the advantages of our method. The higher Sharpe ratios in regressed-PCA primarily stem from the pure-\textit{alpha} portfolios, in contrast to the case in IPCA, where both the pure-\textit{alpha} portfolios and factors contribute comparably. Furthermore, we document that the nonlinear specifications consistently produce MVE factor and combined MVE portfolios with higher Sharpe ratios than the linear specification, underscoring the significance of incorporating nonlinearity. Our results indicate that low-dimensional factors are unlikely to span the conditional efficient frontier. At the same time they demonstrate that imposing a factor structure on the conditional covariance matrix of returns yields robust estimates of the stochastic discount factor, as evidenced by the high Sharpe ratios of the out-of-sample MVE portfolios that we obtain using our approach. In order to further validate our factors, we evaluate their performance in standard asset pricing tests. Our empirical results demonstrate that the regressed-PCA factors consistently outperform IPCA's factors in pricing a large set of testing portfolios, as evidenced by smaller pricing errors, $t$-statistics, and $GRS$ statistics. IPCA factors' inferior performance primarily arises from its much larger regression $R^2$'s, which indicates that these factors capture more time-series variation in returns well but less cross-sectional variation. Moreover, our factors from the nonlinear specifications also outperform FamaFrench_FiveFactor_2015's factors, which justifies the advantages of regressed-PCA over the traditional portfolio-sorting approach. Our paper contributes to several strands of the literature. A number of studies have estimated models similar to (ref) under the assumption that $z_{it}$ are time-invariant, at least within subsamples. These include ConnorLinton_Semparametric_2007, Connoretal_EfficientFFFactor_2012, Fanetal_ProjectedPCA_2016, Kimetal_Arbitrage_2019, LiGeLinton_Dynamic_2020, and Fanetal_StructuralDeep_2022. GagliardiniMa_Extracting_2019 and Guetal_Autoencoder_2021 explore conditional latent factor models that impose the absence of arbitrage, i.e., $\alpha(\cdot)=0$. There are numerous studies of conditional models with observed factors; see Gagliardinietal_Timevarying_2016 and Gagliardinietal_EstimationConditionalFactor_2019 for a comprehensive review. Another strand of literature studies time-varying factor models in which factor loadings evolve smoothly as functions of $t/T$ or aggregate variables;\footnote{The broader literature on conditional models with observable factors has extensively explored time-varying factor loadings that depend on aggregate variables rather than firm-specific characteristics. For example, FersonHarvey_conditioning_1999 use a linear specification, while Roussanov_Composition_2014 employs nonparametric kernel-based approaches. Building on our methodology, Chen_UnifiedFramework_2022 extends the estimation of conditional latent factor models to include heterogeneous \textit{alpha} and \textit{beta} functions, accommodating aggregate variables within $z_{it}$.} see, for example, Mottaetal_LocallyStationary_2011, SuWang_TimeVarying_2017, and PelgerXiong_State-varying_2019. The literature on the cross-section of asset returns is vast; here we focus on multi-factor models motivated by the arbitrage pricing theory. Empirical analysis that exploits the ability of characteristics to predict asset returns typically follows either the portfolio-sorting approach FamaFrench_Commonrisk_1993,FamaFrench_FiveFactor_2015,DanielTitman_Characteristics_1997 or the characteristic-based approach RosenbergMcKibben_ThePrediction_1973, Jacobs/Levy:88, Lewellen_Crosssection_2015, Greenetal_Characteristics_2017,Freybergeretal_Dissecting_2017, Kirby_FirmChar_2020, GiglioXiu_Asset_2019,KozakNagel_Whydospan_2022. The significance of nonlinear relationships in asset pricing has been underscored by several empirical studies Connoretal_EfficientFFFactor_2012, Kirby_FirmChar_2020 and more recently explored through machine learning methods Guetal_Autoencoder_2021, Chenetal_DeepLearning_2020. The remainder of the paper is organized as follows. Section (ref) introduces the estimation method\textemdash regressed-PCA\textemdash along with its key properties and advantages. Section (ref) interprets the method in the context of asset pricing. Section (ref) establishes large-sample properties of the estimators and develops bootstrap inference procedures. Section (ref) applies the new methodology to analyze the cross-section of individual stock returns in the U.S. market. Finally, Section (ref) briefly concludes. The Online Appendix includes estimators for the number of factors, assumptions, proofs of theoretical results, additional discussions, simulation results, and additional empirical findings. \section{Estimation Method} In this section, we introduce a method for estimating the model in (ref), which we term regressed principal component analysis or regressed-PCA, along with its key properties and advantages. To illustrate the underlying idea of our regressed-PCA approach, we begin by assuming that $\alpha(\cdot)$ is null and $\beta(\cdot)$ is linear, i.e., $\alpha(\cdot)=0$ and $\beta(z_{it}) = \Gamma^{\prime}z_{it}$ for some $M\times K$ matrix $\Gamma$. Let $Y_t \equiv (y_{1t},\ldots, y_{Nt})^{\prime}$, $Z_t \equiv (z_{1t},\ldots, z_{Nt})^{\prime}$, and $\varepsilon_t \equiv (\varepsilon_{1t},\ldots, \varepsilon_{Nt})^{\prime}$. The model in (ref) can then be written in matrix form as: \begin{align} Y_t = Z_t\Gamma f_{t} + \varepsilon_{t}. \end{align} A key challenge in applying PCA to estimate $\Gamma$ and $f_t$ is the presence of $Z_t$ in the first term on the right-hand side of (ref). To address this, we first regress $Y_t$ on $Z_t$, yielding: \begin{align} (Z_t^{\prime}Z_t)^{-1}Z_t^{\prime}Y_t = \Gamma f_{t} + (Z_t^{\prime}Z_t)^{-1}Z_t^{\prime}\varepsilon_{t}. \end{align} Heuristically, variation in the common component $Z_t\Gamma f_{t}$ over $t$ comes from two sources: $Z_t$ and $f_t$, and regressing $Y_t$ on $Z_t$ disentangles these sources by isolating $Z_t$ from the common component. Given the factor structure on the right-hand side of (ref), we can apply PCA to the series $\{(Z_t^{\prime}Z_t)^{-1}Z_t^{\prime}Y_t\}_{t\leq T}$ to obtain estimators for $\Gamma$ and $f_t$. Alternatively, the model in (ref) can be viewed as a panel data model with time-varying slope coefficients $\Gamma f_{t} $, which exhibit a factor structure. Essentially, regressed-PCA first estimates the time-varying slope coefficients by period-by-period cross-sectional regressions and then exploits the underlying factor structure by using PCA. \subsection{Regressed-PCA} Now, we consider the general case where $\alpha(\cdot)$ is nonzero and show how to estimate $\alpha(\cdot)$ and $\beta(\cdot) = (\beta_{1}(\cdot),\ldots,\beta_{K}(\cdot))^{\prime}$ nonparametrically. To avoid the curse of dimensionality when $z_{it}$ is multivariate, we assume $\alpha(\cdot)$ and $\beta_k(\cdot)$ are separable. Specifically, we assume there exist functions $\{\alpha_{m}(\cdot)\}_{m\leq M}$ and $\{\beta_{km}(\cdot)\}_{m\leq M}$ such that: \begin{align} \alpha(z_{it}) = \sum_{m=1}^{M}\alpha_{m}(z_{it,m}) \text{ and } \beta_{k}(z_{it}) = \sum_{m=1}^{M}\beta_{km}(z_{it,m}), \end{align} where $z_{it,m}$ is the $m$th entry of $z_{it}$. We adopt the sieve method to estimate $\alpha_{m}(\cdot)$ and $\beta_{km}(\cdot)$. Let $\{\phi_{j}(\cdot)\}_{j\geq 1}$ be a set of basis functions (e.g., B-splines, Fourier series, polynomials) that span a dense linear space of the functional space for $\alpha_{m}(\cdot)$ and $\beta_{km}(\cdot)$. Then, we can express: \begin{align} \alpha_{m}(z_{it,m}) &= \sum_{j=1}^{J}a_{m,j} \phi_{j}(z_{it,m}) + r_{m,J}(z_{it,m}),\\ \beta_{km}(z_{it,m}) &= \sum_{j=1}^{J}b_{km,j} \phi_{j}(z_{it,m}) + \delta_{km,J}(z_{it,m}). \end{align} Here, $\{a_{m,j}\}_{j\leq J}$ and $\{b_{km,j}\}_{j\leq J}$ are the sieve coefficients; $r_{m,J}(\cdot)$ and $\delta_{km,J}(\cdot)$ are “remaining functions” representing the approximation errors; $J$ denotes the sieve size.\footnote{For notational simplicity, we use the same basis functions in (ref) for different $m$'s and the same sieve size. Our results remain valid if different basis functions and different sieve sizes are used for each $m$.} The basic assumption for the sieve method is that $\sup_{z}|r_{m,J}(z)|\to 0$ and $\sup_{z}|\delta_{km,J}(z)|\to 0$ as $J\to\infty$. Let $\bar{\phi}(z_{it,m}) \equiv (\phi_{1}(z_{it,m}),\ldots,\phi_{J}(z_{it,m}))^{\prime}$, ${\phi}(z_{it}) \equiv (\bar{\phi}(z_{it,1})^{\prime},\ldots, \bar{\phi}(z_{it,M})^{\prime})^{\prime}$, and define the vectors $a\equiv (a_{1,1},\ldots,a_{1,J},\ldots, a_{M,1},\ldots,a_{M,J})^{\prime}$, $b_{k}\equiv (b_{k1,1},\ldots,b_{k1,J},\ldots, $ $b_{kM,1},\ldots,b_{kM,J})^{\prime}$, and the matrix $B\equiv(b_1,\cdots, b_{K})$. Let $r(z_{it}) \equiv \sum_{m=1}^{M}r_{m,J}(z_{it,m})$ and $\delta(z_{it}) \equiv (\sum_{m=1}^{M}\delta_{1m,J}(z_{it,m}),\ldots, \sum_{m=1}^{M}\delta_{Km,J}(z_{it,m}))^{\prime}$. Thus, we have: \begin{align} \alpha(z_{it}) = a^{\prime}\phi(z_{it}) + r(z_{it}) \text{ and } \beta(z_{it}) = B^{\prime}\phi(z_{it}) + \delta(z_{it}). \end{align} This shows that $\alpha(z_{it})$ and $\beta(z_{it})$ can be approximated by $a^{\prime}\phi(z_{it})$ and $B^{\prime}\phi(z_{it})$, respectively, and estimating $\alpha(\cdot)$ and $\beta(\cdot)$ reduces to estimating $a$ and $B$. Next, we adapt the regressed-PCA to estimate $a$, $B$, and $f_t$ using the sieve approximation in (ref). Let $\Phi(Z_t) \equiv ({\phi}(z_{1t}),\ldots, {\phi}(z_{Nt}))^{\prime}$, $R(Z_t)\equiv (r(z_{1t}),\ldots,r(z_{Nt}))^{\prime}$, and $\Delta(Z_{t}) \equiv (\delta(z_{1t}),\ldots, \delta(z_{Nt}))^{\prime}$. Using the sieve approximation, we write the model in (ref) in matrix form as: \begin{align} Y_t = \Phi(Z_t)a + \Phi(Z_t)Bf_{t} + R(Z_t) + \Delta(Z_{t})f_t + \varepsilon_{t}. \end{align} Under the basic sieve assumption, the term $ R(Z_t) + \Delta(Z_{t})f_t$ is negligible. The main challenge in applying PCA to estimate $a$, $B$, and $f_t$ lies in the presence of $\Phi(Z_t)$ in the first two terms on the right-hand side of (ref). To address this, we regress $Y_t$ on $\Phi(Z_t)$, yielding: \begin{align} \tilde{Y}_{t} = a+ Bf_{t} +(\Phi(Z_t)^{\prime}\Phi(Z_t))^{-1}\Phi(Z_t)^{\prime}(R(Z_t) + {\Delta}(Z_{t})f_t + {\varepsilon}_{t}), \end{align} where $\tilde{Y}_{t} = (\Phi(Z_t)^{\prime}\Phi(Z_t))^{-1}\Phi(Z_t)^{\prime}Y_t$. Thus, we estimate $a$, $B$, and $f_t$ as follows. First, since $\tilde{Y}_{t}\approx a +Bf_t$, we remove $a$ by subtracting $\bar{\tilde{Y}} = \sum_{t=1}^{T}\tilde{Y}_{t}/T$ from $\tilde{Y}_{t}$ and estimate $B$ by applying PCA to the demeaned series $\{\tilde{Y}_{t}-\bar{\tilde{Y}}\}_{t\leq T}$. Second, for identifying $a$ (and thus $\alpha(\cdot)$), we impose the condition $a^{\prime}B = 0$. Since $\bar{\tilde{Y}}\approx a +B\bar{f}$ (with $\bar{f}=\sum_{t=1}^{T}f_t/T$), we estimate $a$ as $a\approx [I_{JM} - B(B^{\prime}B)^{-1}B]\bar{\tilde{Y}}$. Finally, $f_t$ is estimated as $f_t\approx (B^{\prime}B)^{-1}B^{\prime}\tilde{Y}_{t}$. The formal estimators for $a$, $B$, $\alpha(\cdot)$, $\beta(\cdot)$, and $F=(f_1,\ldots,f_{T})^{\prime}$ are defined as follows. Let $\hat{a}$, $\hat{B}$, $\hat{\alpha}(\cdot)$, $\hat{\beta}(\cdot)$, and $\hat{F}$ denote the respective estimators. Let $\tilde{Y} \equiv (\tilde{Y}_1,\ldots, \tilde{Y}_T)$ and $M_T\equiv I_{T} - 1_T1_T^{\prime}/T$, where $1_{T}$ is a $T\times 1$ vector of ones. Using the normalization $B^{\prime}B=I_{K}$ and $F^{\prime}M_T F/T$ being diagonal with descending diagonal entries, the columns of $\hat{B}$ are the eigenvectors corresponding to the largest $K$ eigenvalues of $\tilde{Y}M_T\tilde{Y}^{\prime}/T$. We then have: $\hat{a} = (I_{JM}-\hat{B}\hat{B}^{\prime})\bar{\tilde{Y}}$, \begin{align} \hat{\alpha}(z) = \hat{a}^{\prime}\phi(z), \hat{\beta}(z) = \hat{B}^{\prime}\phi(z), \text{ and }\hat{F} = (\hat{f}_1,\ldots,\hat{f}_T)^{\prime} = \tilde{Y}^{\prime}\hat{B}. \end{align} We assume that the number of factors, $K$, is fixed and known. In Section (ref), we establish asymptotic properties of the estimators and develop inference methods. In Appendix (ref), we propose two consistent estimators for $K$, ensuring that our results extend to the case of an unknown $K$ by using a conditioning argument. \subsection{Key Properties} Our regressed-PCA method has several appealing properties and is straightforward to implement. First, as discussed in Section (ref), it accommodates time-varying $z_{it}$ and does not require a large time dimension $T$. This flexibility allows us to examine the evolving relationship between risk and return using both full-sample and sub-sample analyses. Second, the estimation procedure is well-suited for unbalanced panels, which is particularly relevant in cross-sectional asset pricing applications. The key step of regressed-PCA is to compute $\tilde{Y}_t$. Specifically, we can express $\tilde{Y}_t$ as: \begin{align} \tilde{Y}_t=\left(\sum_{i=1}^{N}\phi(z_{it})\phi(z_{it})^{\prime}\right)^{-1}\sum_{i=1}^{N}\phi(z_{it})y_{it}. \end{align} For unbalanced panels, $\tilde{Y}_t$ can still be calculated by summing only over the $i$'s for which both $z_{it}$ and $y_{it}$ are observed at time period $t$. This approach is equivalent to treating missing data as zeros, allowing us to proceed as if working with a balanced panel. The asymptotic results we derive remain valid as long as $\min_{t\leq T} N_t \to\infty$, where $N_t$ is the sample size at time period $t$. Third, our method continues to be effective even when the pricing errors and risk exposures are not fully explained by $z_{it}$. Let $e_{\alpha,it}$ and $e_{\beta,it}$ be the error terms in the pricing errors and the risk exposures, respectively, which are orthogonal to $z_{it}$. In this case, the model becomes: \begin{align} y_{it} = [\alpha(z_{it})+e_{\alpha,it}] + [\beta(z_{it})+e_{\beta,it}]^{\prime}f_{t} + \varepsilon_{it} = \alpha(z_{it}) + \beta(z_{it})^{\prime}f_{t} + \varepsilon^{\ast}_{it}, \end{align} where $\varepsilon^{\ast}_{it} = \varepsilon_{it} + e_{\alpha,it} + e_{\beta,it}^{\prime}f_{t}$. Since we are not interested in estimating $e_{\alpha,it}$ and $e_{\beta,it}$, our asymptotic results remain valid if we replace $\varepsilon_{it}$ in the original model with $\varepsilon^{\ast}_{it}$. Finally, efficiency of our estimation procedure could be improved by using generalized least squares in its first step. Our asymptotic results continue to hold if we replace $\Phi(Z_t)$ and $\varepsilon_t$ with their transformed counterparts $V_{t}^{-1/2}\Phi(Z_t)$ and $V_{t}^{-1/2}\varepsilon_t$, where $V_t$ is the conditional covariance matrix of $Y_t$ at each time period $t$. While $V_t$ is usually unknown, Hoberg_Welch_Optimized_2009 suggest some practical guidance to account for the cross-correlation and heteroskedasticity of the idiosyncratic noise $\varepsilon_{it}$. \subsection{Comparing Methods} How does our regressed-PCA compare with existing methods that have been proposed in the literature? What are the advantages of our regressed-PCA? First, the projected-PCA proposed by Fanetal_ProjectedPCA_2016 applies PCA to the series $\{\Phi(Z_t)(\Phi(Z_t)^{\prime}\Phi(Z_t))^{-1}\Phi(Z_t)^{\prime}Y_t\}_{t\leq T}$. In contrast, our regressed-PCA applies PCA to the series $\{(\Phi(Z_t)^{\prime}\Phi(Z_t))^{-1}\Phi(Z_t)^{\prime}Y_t\}_{t\leq T}$. The two methods are fundamentally different: regressed-PCA applies PCA to estimated coefficients, while projected-PCA focuses on fitted values. The regression step in regressed-PCA is designed to extract $Z_t$ from the common component for consistent estimation, whereas projected-PCA aims to remove noise in non-time-varying factor loadings to achieve efficiency in estimation. Consequently, projected-PCA may yield inconsistent estimates when $Z_t$ is time-varying.\footnote{For instance, as illustrated in (ref), we have $Z_t(Z_t^{\prime}Z_t)^{-1}Z_t^{\prime}Y_t \approx Z_t\Gamma f_{t}$, which does not conform to a factor structure unless $Z_t$ remains constant over $t$. Therefore, applying PCA to $\{Z_t(Z_t^{\prime}Z_t)^{-1}Z_t^{\prime}Y_t\}_{t=1}^{T}$ can result in failure to estimate $f_t$.} As noted by Fanetal_ProjectedPCASupp_2016 and further investigated by Chengetal_UniformPredictive_2020, ensuring the consistency of projected-PCA may necessitate imposing smoothness conditions on how $Z_t$ varies with $t$, which our regressed-PCA does not require. Additionally, projected-PCA often requires dropping certain observations to maintain a balanced panel, whereas regressed-PCA is applicable to unbalanced panels. While Kimetal_Arbitrage_2019 extend projected-PCA to accommodate nonzero $\alpha(\cdot)$, they do not develop an inference procedure. Similarly, Fanetal_StructuralDeep_2022 extend projected-PCA by employing deep neural networks and propose a local version to capture slowly changing alphas and betas. However, while neural networks can alleviate the curse of dimensionality for prediction tasks, they are not typically well-suited for inference\textemdash a key focus of our paper. Our regressed-PCA not only allows for nonzero $\alpha(\cdot)$ but also provides formal testing procedures, which are crucial for evaluating and comparing factor models. Moreover, by incorporating rapidly varying $Z_t$, our regressed-PCA is able to capture abrupt changes in both alphas and betas effectively. Second, consider the least squares estimation approach introduced by Parketal_FactorDynamics_2009, which is at the core of the IPCA of Kellyetal_Characteristics_2019. The least squares method minimizes the following objective function: \begin{align} \sum_{t=1}^{T}({Y}_t - \Phi(Z_t)a -\Phi(Z_t)Bf_{t})^{\prime} ({Y}_t - \Phi(Z_t)a -\Phi(Z_t) Bf_{t}), \end{align} while regressed-PCA minimizes:\footnote{The equality in (ref) follows because $\tilde{Y}_{t} = (\Phi(Z_t)^{\prime}\Phi(Z_t))^{-1}\Phi(Z_t)^{\prime}Y_t$. } \begin{align} &\sum_{t=1}^{T}(\tilde{Y}_t - a -Bf_{t})^{\prime} (\tilde{Y}_t - a - Bf_{t})\notag\\ =&\sum_{t=1}^{T}({Y}_t - \Phi(Z_t)a -\Phi(Z_t)Bf_{t})^{\prime} S_{t} ({Y}_t - \Phi(Z_t)a -\Phi(Z_t) Bf_{t}), \end{align} where $S_{t} = \Phi(Z_t)(\Phi(Z_t)'\Phi(Z_t))^{-1}(\Phi(Z_t)'\Phi(Z_t))^{-1}\Phi(Z_t)^{\prime}$. The objective functions of these two approaches differ, except when $\Phi(Z_t)'\Phi(Z_t)/N=I_{JM}$.\footnote{See Appendix (ref) for discussion.} Essentially, the least squares approach in (ref) maximizes in-sample $R^2$, while regressed-PCA in (ref) optimizes time-series comovement of $\tilde Y_{t}$, which captures the most cross-sectional variations of individual asset returns. A key challenge with the least squares approach is that its minimization problem is nonconvex and cannot be solved explicitly. While Parketal_FactorDynamics_2009 develop a numerical algorithm to find the estimators, and Kellyetal_Characteristics_2019 propose an alternating least squares procedure, both methods may require careful selection of initial values to ensure convergence to the correct solution. Furthermore, the asymptotic properties of these algorithms are not well understood. In addition to the asymptotic properties that we derive, our regressed-PCA provides estimators that can always be explicitly solved. Unlike IPCA, which relies on a long time series of returns, our regressed-PCA does not require a large time dimension $T$, enabling sub-sample analyses and capturing potential time variation in the coefficients ($a$ and $B$). Another advantage of our regressed-PCA approach lies in factor construction. As the number of factors, $K$, increases by one, regressed-PCA simply constructs an additional factor on top of existing ones through PCA. In contrast, IPCA requires reconstructing all factors by recomputing alternating least squares, making it sensitive to the number of factors. For example, the first factor under $K=2$ may differ significantly from the first factor under $K=1$, and the factor space under $K=2$ does not necessarily nest the factor space under $K=1$. Moreover, IPCA's factors can be correlated. In contrast, regressed-PCA produces factors that are uncorrelated by construction and remain stable, regardless of the number of factors chosen. Overall, in addition to its formal inference procedures and well-established asymptotic properties, our regressed-PCA offers significant computational simplicity. It accommodates nonzero alphas, time-varying characteristics, unbalanced panels, and short samples, making it particularly suitable for empirical asset pricing. Moreover, it provides stable and reliable factor construction. \section{Asset Pricing Interpretation} Regressed-PCA has deep roots in asset pricing. In a typical asset pricing application, $y_{it}$ represents the realized returns on asset $i$ at the end of time period $t$, while $z_{it,m}$ represents the $m$'th attribute or characteristic of asset $i$ that is known at the {\em beginning} of time period $t$ (or, alternatively, at the “end” of time period $t-1$). The regressed-PCA first estimates the time-varying slope coefficients by period-by-period cross-sectional regressions of returns on (functions of) characteristics, and then exploits the factor structure by using PCA. These period-by-period cross-sectional regressions are known as Fama-MacBeth regressions FamaMacBeth_RiskReturn_1973, which help transform a large unbalanced panel of noisy individual asset returns into a lower-dimensional balanced panel of portfolio returns that are largely free of idiosyncratic noise, $\tilde{Y}_{t}$. Furthermore, $\tilde{Y}_{t}$ can be interpreted as the time $t$ realization of returns on a set of $JM$ characteristic-managed portfolios, sometimes referred to as “optimized portfolios,” “characteristic pure plays,” or “cross-section factors” (e.g., as in Hoberg_Welch_Optimized_2009, BackKapadiaOstdiek_Testing_2015, and FamaFrench_CSTS_2020). We also refer to them as Fama-MacBeth managed portfolios. In particular, if the basis functions $\phi(z_{it})$ include a constant term (e.g., as the first element in $\phi(z_{it})$) and are standardized to have a zero mean in each cross-section, then the intercept in the Fama-MacBeth regressions (the first element in $\tilde{Y}_{t}$) represents a “level” return. This is essentially the equal-weighted average excess return across all individual assets, with weights summing up to unity or costing \$1, and it has no ex-ante loadings on any characteristics.\footnote{This is due to two properties. First, the intercept in $\tilde{Y}_{t}$ is equal to $\sum_{i=1}^{N} y_{it}/N$ when the nonconstant regressors have a zero mean. Second, the weights in $\tilde{Y}_{t}$, given by $W_t = (\Phi(Z_t)^{\prime}\Phi(Z_t))^{-1}\Phi(Z_t)^{\prime}$, satisfy the property that $W_t\Phi(Z_t) = I_{JM}$. The second property also clarifies the property of the weights for other portfolios in $\tilde{Y}_{t}$.} This is sometimes referred to as a “naively diversified” or $1/N$ portfolio. As shown by Fama_76, the period-by-period slope coefficients corresponding to the time-varying basis functions (all elements of $\tilde{Y}_{t}$ starting from the second) are excess returns on zero-cost portfolios, as long as a constant term is included in the basis functions. These portfolios have weights on individual assets that set the weighted average value of the relevant basis function to one and those of all the remaining basis functions to zeros;\footnote{See Footnote (ref).} see Hoberg_Welch_Optimized_2009 and Kirby_FirmChar_2020 for other attractive properties. In simpler terms, each portfolio has spread in only one basis function, or a pure play on a particular basis function. Moreover, adding one more basis function in the Fama-MacBeth regressions to a benchmark corresponds to introducing one additional zero-cost portfolio, which raises average exposure to that specific basis function by one unit while maintaining average exposures to other basis functions unchanged. Therefore, these portfolios can also be interpreted as “slope” returns with respect to the time-varying basis functions. Moreover, Fama_76 demonstrates that the portfolios in $\tilde{Y}_{t}$ exhibit minimum variance and often low correlations under OLS-like i.i.d. assumptions, rendering them maximally diversified.\footnote{See Appendix (ref) for discussion.} Due to their low noise and correlations, several studies have explored the advantages of these portfolios in asset pricing tests. For instance, Hoberg_Welch_Optimized_2009, BackKapadiaOstdiek_Testing_2015, and Kirby_FirmChar_2020 highlight their utility as test assets (dependent variables), while BackKapadiaOstdiek_Slope_2013 and FamaFrench_CSTS_2020 emphasize their effectiveness as pricing factors (independent variables), particularly when compared with the portfolio-sorting approach. However, these studies primarily focus on linear specifications (i.e., $\phi(z_{it})= z_{it}$ which includes a constant term) with a small number of characteristics, and their findings may not generalize to more complex cases. In the presence of a large number of characteristics or basis functions, the dimension of $\tilde{Y}_{t}$ can become substantial, and the portfolios in $\tilde{Y}_{t}$ may exhibit high correlations, even though each portfolio is maximally diversified. Understanding the factor structure of the portfolios in $\tilde{Y}_{t}$ as test assets is crucial to avoid spurious fit of misspecified asset pricing models that happen to be correlated with some of the latent factors, as emphasized by Lewellenetal_Skeptical_2010. A straightforward solution is to apply PCA to extract uncorrelated principal components from $\tilde{Y}_{t}$, which is the essence of our approach. \subsection{Comparison with Portfolio Sorting} The extracted factors from regressed-PCA in ((ref)) are linear combinations of characteristic-managed portfolios $\tilde Y_{t}$, and thus are themselves tradable portfolios. As a way of factor construction, regressed-PCA shares a strong connection with the portfolio-sorting approach. Sorting assets into portfolios can be equivalently framed as running cross-sectional regressions of returns on dummies that represent groups sorted by characteristics at each time period, whether the sorting is independent or dependent. Specifically, the return of each sorted portfolio (i.e., the average return of individual assets within a group) is the coefficient of the corresponding group dummy in the regression. Using our notation, the returns of sorted portfolios can be represented by $\tilde{Y}_{t}$ when the basis functions $\phi(z_{it})$ are group dummies.\footnote{See Appendix (ref) for discussion.} Thus, the key difference between the first steps of the regressed-PCA and portfolio-sorting approaches lies in their choice of basis functions. Sorting has been recognized as a nonparametric method for examining the relationship between average returns and characteristics, as highlighted by FamaFrench_Dissecting_2008, Cochrane_Presidential_2011, and CattaneoCrumpFarrellSchaumburg_2020_Characteristic. However, regressed-PCA offers several advantages over sorting. First, sorting quickly encounters the curse of dimensionality and rarely handles more than four characteristics simultaneously. When multiple characteristics are present, double sorting is typically used for each pair of characteristics, making it difficult to infer which characteristics uniquely affect average returns. Regressed-PCA addresses this limitation by employing a separable additive specification (see (ref)) and B-splines basis functions, which allow for a large number of characteristics and facilitate testing their significance. Second, sorting fails to fully exploit the variation in characteristics within each sorted group. In contrast, regressed-PCA takes advantage of the full variation in characteristics. Third, sorting struggles to effectively explore the nonlinear relationship between average returns and characteristics: sorting essentially uses step functions, which suffer from several well-known shortcomings, such as discontinuities at cutoffs, poor extrapolation, and unstable estimates that are highly sensitive to outlier assets HastieTibshiraniFriedman_StatisticalLearning_2011. These issues are mitigated by the use of B-splines, which our regressed-PCA incorporates. Although long-short factors (such as high-minus-low and small-minus-big factors) are straightforward to interpret, regressed-PCA offers several advantages over them. First, the long-short approach focuses only on portfolios in extreme groups and ignores those in the middle. As a result, it may fail when factor loadings exhibit non-monotonicity with respect to characteristics, such as a “tent” shape. Regressed-PCA, on the other hand, uses all portfolios formed by Fama-MacBeth sieve regressions, making it more adaptable to capturing underlying nonlinearities. Second, long-short factors do not effectively distinguish between the risk and mispricing explanations of the role of characteristics in predicting asset returns, a distinction at the heart of the “characteristics versus covariances” debate. Regressed-PCA is based on a latent factor model for individual asset returns, which is ideally suited to resolve this debate. Third, when multiple characteristics are present, long-short methods can quickly lead to the “\textit{factor zoo}” problem as distentangling the roles of correlated characteristics can be difficult in multi-way sorts with potentially too few securities in each bin. Regressed-PCA mitigates this issue by utilizing Fama-Macbeth managed portfolios, which are maximally diversified “pure plays” on characteristics, together with PCA, which has also been proven effective in reducing the dimensionality of sorted portfolios Kozaketal_Interpreting_2018,Kozaketal_Shrinking_2020,LettauPelger_Factorstimeseries_2020. \section{Econometric Analysis} In this section, we establish the asymptotic properties of our estimators, including consistency, the rate of convergence, and their asymptotic distribution. Additionally, we develop bootstrap inference procedures. We begin by defining some notation that will be used throughout the paper. For a symmetric matrix $A$, we denote its $k$th largest eigenvalue by $\lambda_{k}(A)$, and its smallest and largest eigenvalues by $\lambda_{\min}(A)$ and $\lambda_{\max}(A)$, respectively. The operator norm of a matrix $A$ is denoted by $\|A\|_2$, and its Frobenius norm by $\|A\|_{F}$. The vectorization of $A$ is written as $\mathrm{vec}(A)$. The Euclidian norm of a column vector $x$ is denoted by $\|x\|$. Finally, for matrices $A$ and $B$, we use $A\otimes B$ to denote their Kronecker product. \subsection{Asymptotic Properties} Before presenting formal theorems, we revisit (ref) to briefly illustrate why a large $T$ is not required and $Z_t$ can be nonstationary over $t$. Consider the case when $T\geq K+1$ and $M\geq K$. Since the columns of $\hat{B}$ and $\Gamma$ are the eigenvectors of $\tilde{Y}M_{T}\tilde{Y}^{\prime}$ and $\Gamma F^{\prime} M_{T} F \Gamma^{\prime}$, respectively, corresponding to the first $K$ largest eigenvalues, by the matrix perturbation theorem (see, for example, Yuetal_Useful_2014), the consistency of $\hat{B}$ to $\Gamma$ (up to a rotational transformation) can be established if we can show that: \begin{align} \|\tilde{Y}M_{T} - \Gamma F^{\prime} M_{T}\|_{F} = o_{p}(1) \text{ as } N\to\infty. \end{align} Since $\tilde{Y}= \Gamma F^{\prime} + ((Z_1^{\prime}Z_1)^{-1}Z_1^{\prime}\varepsilon_1,\ldots, (Z_T^{\prime}Z_T)^{-1}Z_T^{\prime}\varepsilon_T)$, (ref) simplifies to: \begin{align} \|((Z_{1}^{\prime}Z_1)^{-1}Z_{1}^{\prime}\varepsilon_1, \ldots, (Z_{T}^{\prime}Z_T)^{-1}Z_{T}^{\prime}\varepsilon_T)M_{T}\|_{F}= o_{p}(1) \text{ as } N\to\infty. \end{align} When $T$ is fixed, (ref) is equivalent to $(Z_{t}^{\prime}Z_t)^{-1}Z_{t}^{\prime}\varepsilon_t = o_{p}(1)$ for each $t$. Therefore, only regularity conditions on $Z_t$ and $\varepsilon_t$ for each $t$ are needed to apply the law of large numbers. This result also implies that $Z_t$ can vary over $t$ in a nonstationary fashion. Let $H \equiv (F^{\prime}M_T\hat{F})(\hat{F}^{\prime}M_T\hat{F})^{-1}$, which represents a rotational transformation matrix that governs the convergence limit of $\hat{B}$, $\hat{F}$, and $\hat{\beta}(\cdot)$. Define $\xi_{J}\equiv \sup_{z}\|\bar{\phi}(z)\|$, which scales as $O(\sqrt{J})$ for B-splines and Fourier series, and $O(J)$ for polynomials (see, for example, BelloniChernozhukovChetverikovKato_SeriesEstimator_2015). \begin{thm} Suppose Assumptions (ref)-(ref) hold. Let $\hat{a}$, $\hat{B},\hat{F}$, $\hat{\alpha}(\cdot)$, and $\hat{\beta}(\cdot)$ be given in (ref). Assume (i) $N\to\infty$; (ii) $T\geq K+1$ ($T$ may stay fixed or grow simultaneously with $N$); (iii) $J\to\infty$ with $J^{2}\xi^{2}_{J}\log J=o(N)$. Then \begin{align*} \|\hat{a} - a\|^{2}&=O_{p}\left(\frac{1}{J^{2\kappa}}+\frac{J}{N^2}+\frac{{J}}{{NT}}\right),\notag\\ \|\hat{B} - B H\|^{2}_{F}&=O_{p}\left(\frac{1}{J^{2\kappa}}+\frac{J}{N^2}+\frac{{J}}{{NT}}\right),\notag\\ \frac{1}{T}\|\hat{F}-F(H^{\prime})^{-1}\|_{F}^{2}&=O_{p}\left(\frac{1}{J^{2\kappa}}+\frac{1}{{N}}\right),\notag\\ \sup_{z}|\hat{\alpha}(z)-\alpha(z)|^{2}&=O_{p}\left(\frac{1}{J^{2\kappa-1}}+\frac{J^2}{N^2}+\frac{{J^2}}{{NT}}\right)\max_{j\leq J}\sup_{z}|\phi_{j}(z)|^{2},\notag\\ \sup_{z}\|\hat{\beta}(z)-H^{\prime}\beta(z)\|^{2}&=O_{p}\left(\frac{1}{J^{2\kappa-1}}+\frac{J^2}{N^2}+\frac{{J^2}}{{NT}}\right)\max_{j\leq J}\sup_{z}|\phi_{j}(z)|^{2}, \end{align*} where $\kappa>1/2$ is a constant representing the smoothness of $\alpha(\cdot)$ and $\beta(\cdot)$. \end{thm} All the assumptions are outlined in Appendix (ref). Theorem (ref) establishes that $a$ and $\alpha(\cdot)$ can be consistently estimated by $\hat{a}$ and $\hat\alpha(\cdot)$, while $B$, $F$, and $\beta(\cdot)$ can be consistently estimated by $\hat{B}$, $\hat{F}$, and $\hat\beta(\cdot)$, respectively, up to a rotational transformation. This consistency holds as long as $J\to\infty$ under both large $N$ and either fixed or large $T$. Notably, the large $J$ requirement differs from Fanetal_ProjectedPCA_2016. This distinction highlights the importance of controlling sieve approximation errors in $\alpha(\cdot)$ and $\beta(\cdot)$ to ensure consistent estimation of $F$. Appendix (ref) illustrates how misspecifications in $\alpha(\cdot)$ and $\beta(\cdot)$ can result in inconsistent estimation of $F$, motivating the need for a specification test, which is addressed in Section (ref). Additionally, Theorem (ref) shows that the estimators achieve fast convergence rates. In particular, $\hat{F}$ attains the optimal rate $1/N$ when $\alpha(\cdot)$ and $\beta(\cdot)$ are sufficiently smooth (i.e., $\kappa$ is sufficiently large). This implies that the nonparametric modelling of $\alpha(\cdot)$ and $\beta(\cdot)$ do not degrade the convergence rate for estimating $F$, as long as the smoothness conditions are satisfied. This result is crucial for developing the specification test for $\alpha(\cdot)$ and $\beta(\cdot)$ in Section (ref), and also important for utilizing $\hat{F}$ in subsequential asset pricing tests of Section (ref). Notably, if the functional forms of $\alpha(\cdot)$ and $\beta(\cdot)$ are known and correctly specified, sieve approximation errors can be avoided, and the asymptotic results hold for a fixed $J$. Theorem (ref) allows for weak dependence of the errors $\{\varepsilon_{it}\}_{i\leq N,t\leq T}$ over both $i$ and $t$, which is relevant in asset pricing. Let $\Omega\equiv\sum_{i=1}^{N}\sum_{t=1}^{T}\sum_{s=1}^{T}f_{t}^{\dag}f_{s}^{\dag\prime}Q_{t}^{-1}E[\phi(z_{it})\phi(z_{is})^{\prime}]$ $\times Q_{s}^{-1}E[\varepsilon_{it}\varepsilon_{is}]/NT$, where $f^{\dag}_t=(1,(f_t-\bar{f})^{\prime})^{\prime}$ and $Q_{t} = \sum_{i=1}^{N}E[\phi(z_{it})\phi(z_{it})^{\prime}]/N$. This defines a variance-covariance matrix, which appears in the asymptotic distributions of $\hat{a}$ and $\hat{B}$. \begin{thm} Suppose Assumptions (ref)-(ref) hold. Let $\hat{a}$ and $\hat{B}$ be given in (ref). Assume (i) $N\to\infty$; (ii) $T\geq K+1$; (iii) $J\to\infty$ with $J^{2}\xi^{2}_{J}\log J=o(N)$. Then there is a $JM\times (K+1)$ random matrix $\mathbb{N}$ with $\mathrm{vec}(\mathbb{N})\sim N(0,\Omega)$ such that: \begin{align*} \|\sqrt{{NT}}(\hat{a} - a)-\mathbb{G}_{a}\| = O_{p}\left(\frac{\sqrt{NT}}{J^{\kappa}}+\frac{\sqrt{TJ}}{\sqrt{N}}+\frac{J^{5/6}}{N^{1/6}}+\frac{\sqrt{J\xi_{J}}\log^{1/4}J}{N^{1/4}}\right) \end{align*} and \begin{align*} \|\sqrt{{NT}}(\hat{B} - B H)-\mathbb{G}_{B}\|_{F} = O_{p}\left(\frac{\sqrt{NT}}{J^{\kappa}}+\frac{\sqrt{TJ}}{\sqrt{N}}+\frac{J^{5/6}}{N^{1/6}}+\frac{\sqrt{J\xi_{J}}\log^{1/4}J}{N^{1/4}}\right), \end{align*} where $\kappa>1/2$ is a constant representing the smoothness of $\alpha(\cdot)$ and $\beta(\cdot)$, $\mathbb{G}_{a} =(I_{JM}-B\mathcal{H}\mathcal{H}^{\prime}B^{\prime})(\mathbb{N}_1-\mathbb{G}_{B}\mathcal{H}^{-1}\bar{f})-B\mathcal{H}\mathbb{G}_{B}^{\prime}a$, and $\mathbb{G}_{B} = \mathbb{N}_2 B^{\prime}B\mathcal{M}$. The matrices $\mathcal{H}$ and $\mathcal{M}$ are nonrandom, as given in Lemma (ref), while $\mathbb{N}_1$ and $\mathbb{N}_2$ are the first column and the last $K$ columns of $\mathbb{N}$, respectively. \end{thm} Theorem (ref) establishes a strong approximation, demonstrating that $(\sqrt{{NT}}(\hat{a} - a),\sqrt{{NT}}(\hat{B} - B H))$ can be well approximated by a normal random matrix $(\mathbb{G}_{a},\mathbb{G}_{B})$. Specifically, the difference between them converges in probability to zero under the conditions $T=o(N)$, $NT J^{-2\kappa}=o(1)$, and $J=o(\min\{N^{1/5},N/T\})$. Since the dimensions of $\sqrt{{NT}}(\hat{a} - a)$ and $\sqrt{{NT}}(\hat{B} - B H)$ grow with $N$, rendering the classical central limit theorem inapplicable, we employ Yurinskii’s coupling to establish this strong approximation. This approach accommodates weak temporal dependence in the errors $\{\varepsilon_{it}\}_{i\leq N,t\leq T}$. Furthermore, the result can be readily extended to allow for cluster-type dependence across $i$ in the errors $\{\varepsilon_{it}\}_{i\leq N,t\leq T}$; see the discussion following Assumption (ref). Notably, distributional results of this kind are not provided in Fanetal_ProjectedPCA_2016. \subsection{Weighted Bootstrap} We develop a weighted bootstrap approach to estimating the distribution of $(\mathbb{G}_{a},\mathbb{G}_{B})$. Let $\{w_i\}_{i\leq N}$ be a sequence of i.i.d. positive random variables, with $E[w_i]=1$ and $var(w_i)=\omega_0>0$. For instance, the $w_i$'s can be drawn from a standard exponential distribution, where $\omega_0=1$. To preserve the time dependence, we assign the same weight $w_i$ to all observations over $t$. Define $\Phi(Z_t)^{\ast}\equiv ({\phi}(z_{1t})w_1,\ldots, {\phi}(z_{Nt})w_N)^{\prime}$ and $\tilde{Y}_{t}^{\ast}\equiv (\Phi(Z_t)^{\ast\prime}\Phi(Z_t))^{-1}\Phi(Z_t)^{\ast\prime}Y_t$, which is the bootstrap version of $\tilde{Y}_t$. To define the bootstrap estimators of $a$ and $B$, let $\tilde{Y}^{\ast} \equiv (\tilde{Y}^{\ast}_1,\ldots, \tilde{Y}^{\ast}_T)$ and $\bar{\tilde{Y}}^{\ast}\equiv \sum_{t=1}^{T}\tilde{Y}^{\ast}_t/T$. The bootstrap estimators are given by: \begin{align} \hat{B}^{\ast} = \tilde{Y}^{\ast}M_T\hat{F}(\hat{F}^{\prime}M_T\hat{F})^{-1} \text{ and } \hat{a}^{\ast} = (I_{JM} - \hat{B}^{\ast}(\hat{B}^{\ast\prime}\hat{B}^{\ast})^{-1}\hat{B}^{\ast\prime})\bar{\tilde{Y}}^{\ast}, \end{align} which mimic the original estimators: $\hat{B} = \tilde{Y}M_T\hat{F}(\hat{F}^{\prime}M_T\hat{F})^{-1}$ and $\hat{a} = (I_{JM}-\hat{B}\hat{B}^{\prime})\bar{\tilde{Y}} =(I_{JM}-\hat{B}(\hat{B}^{\prime}\hat{B})^{-1}\hat{B}^{\prime})\bar{\tilde{Y}} $. We propose estimating the distribution of $(\mathbb{G}_{a},\mathbb{G}_{B})$ by the distribution of $(\sqrt{{NT/\omega_0}}$ $(\hat{a}^{\ast}-\hat{a}),\sqrt{{NT/\omega_0}}(\hat{B}^{\ast}-\hat{B}))$ conditional on the data.\footnote{A more natural bootstrap estimator for $B$ is given by the eigenvectors of $\tilde{Y}^{\ast}M_T\tilde{Y}^{\ast\prime}/T$ corresponding to its first $K$ largest eigenvalues. However, the approach generally fails due to rational transformation matrices; see Appendix (ref) for discussion.} The bootstrap procedure can be easily adapted for unbalanced panels. The key step is obtaining $\tilde{Y}^{\ast}_t$. For balanced panels, we write: \begin{align} \tilde{Y}^{\ast}_t= \left(\sum_{i=1}^{N}\phi(z_{it})\phi(z_{it})^{\prime}w_i\right)^{-1}\sum_{i=1}^{N}\phi(z_{it})y_{it}w_i. \end{align} In unbalanced panels, we adjust by taking the sums over the $i$'s for which both $z_{it}$ and $y_{it}$ are observed at time period $t$. This effectively replaces missing data with zeros, making the procedure identical to the balanced case. The asymptotic results established below remain valid as long as as $\min_{t\leq T} N_t \to\infty$, where $N_t$ represents the sample size at time period $t$. Moreover, the bootstrap can easily accommodate cluster-type dependence across $i$ by assigning the same weight within each cluster. \begin{thm} Suppose Assumptions (ref)-(ref) hold. Let $\hat{a}$, $\hat{B}$, $\hat{a}^{\ast}$, and $\hat{B}^{\ast} $ be given in (ref) and (ref). Assume (i) $N\to\infty$; (ii) $T\geq K+1$; (iii) $J\to\infty$ with $J^{2}\xi^{2}_{J}\log J=o(N)$. Then there is a $JM\times (K+1)$ random matrix $\mathbb{N}^{\ast}$ with $\mathrm{vec}(\mathbb{N}^{\ast})\sim N(0,\Omega)$ conditional on $\{Y_t, Z_t\}_{t\leq T}$ such that: \begin{align*} \|\sqrt{{NT/\omega_0}}(\hat{a}^{\ast} - \hat{a})-\mathbb{G}^{\ast}_{a}\| = O_{p^{\ast}}\left(\frac{\sqrt{NT}}{J^{\kappa}}+\frac{\sqrt{TJ}}{\sqrt{N}}+\frac{J^{5/6}}{N^{1/6}}+\frac{\sqrt{J\xi_{J}}\log^{1/4}J}{N^{1/4}}\right) \end{align*} and \begin{align*} \|\sqrt{{NT/\omega_0}}(\hat{B}^{\ast} - \hat{B})-\mathbb{G}_{B}^{\ast}\|_{F} = O_{p^{\ast}}\left(\frac{\sqrt{NT}}{J^{\kappa}}+\frac{\sqrt{TJ}}{\sqrt{N}}+\frac{J^{5/6}}{N^{1/6}}+\frac{\sqrt{J\xi_{J}}\log^{1/4}J}{N^{1/4}}\right), \end{align*} where $p^{\ast}$ is the probability measure with respect to $\{w_i\}_{i\leq N}$ conditional on $\{Y_t, Z_t\}_{t\leq T}$, $\kappa>1/2$ is a constant representing the smoothness of $\alpha(\cdot)$ and $\beta(\cdot)$, $\mathbb{G}^{\ast}_{a} = (I_{JM}-B\mathcal{H}\mathcal{H}^{\prime}B^{\prime})(\mathbb{N}^{\ast}_1-\mathbb{G}^{\ast}_{B}\mathcal{H}^{-1}\bar{f})-B\mathcal{H}\mathbb{G}_{B}^{\ast\prime}a$, and $\mathbb{G}^{\ast}_{B} = \mathbb{N}^{\ast}_2B^{\prime}B\mathcal{M}$. The matrices $\mathcal{H}$ and $\mathcal{M}$ are nonrandom, as given in Lemma (ref), while $\mathbb{N}^{\ast}_1$ and $\mathbb{N}^{\ast}_2$ are the first column and the last $K$ columns of $\mathbb{N}^{\ast}$, respectively. \end{thm} Theorem (ref) demonstrates that the distribution of $(\mathbb{G}_{a},\mathbb{G}_{B})$, which aligns with the distribution of $(\mathbb{G}^{\ast}_{a},\mathbb{G}^{\ast}_{B})$, can be approximated by the distribution of $(\sqrt{{NT/\omega_0}}(\hat{a}^{\ast} - \hat{a}),\sqrt{{NT/\omega_0}}(\hat{B}^{\ast}-\hat{B}))$ conditional on the data, under the conditions $T=o(N)$, $NTJ^{-2\kappa}$ $=o(1)$, and $J=o(\min\{N^{1/5},N/T\})$. Theorems (ref) and (ref) can then be directly applied to conduct significance tests. To test whether $\alpha(\cdot) = 0$, we compare $NT\hat{a}^{\prime}\hat{a}$ with the $1-\alpha$ quantile of $NT(\hat{a}^{\ast}-\hat{a})^{\prime}(\hat{a}^{\ast}-\hat{a})/\omega_0$ conditional on the data for $0<\alpha<1$. Similarly, we can test whether each component of $\phi(z_{it})$ is significant in $\alpha(z_{it})$, which is equivalent to testing whether the corresponding element of $a$ is zero. We can also test the joint significance of each component of $\phi(z_{it})$ in $\beta(z_{it})$, which corresponds to testing whether the associated row of $BH$ is zero. \subsection{Specification Test} To test for linearity of $\alpha(\cdot)$ and $\beta(\cdot)$, we consider the following hypotheses: \begin{align} &\mathrm{H}_0: \alpha(z_{it})=\gamma^{\prime} z_{it} \text{ and } \beta(z_{it})=\Gamma^{\prime} z_{it} \text{ for some } \gamma \text{ and } \Gamma \text{ versus }\notag\\ &\mathrm{H}_1: \inf_{\pi}E[|\alpha(z_{it})-\pi^{\prime} z_{it}|^{2}]>0 \text{ or } \inf_{\Pi}E[\|\beta(z_{it})-\Pi^{\prime} z_{it}\|^{2}]>0. \end{align} We develop a test by comparing the estimators under $\mathrm{H}_0$ and $\mathrm{H}_1$. The estimators of $\alpha(\cdot)$ and $\beta(\cdot)$ under $\mathrm{H}_1$ are given by $\hat{\alpha}(\cdot)$ and $\hat{\beta}(\cdot)$, as defined in (ref). Let $\vec{Y}_{t}\equiv (Z_{t}^{\prime}Z_{t})^{-1}Z_{t}^{\prime}Y_t$, $\vec{Y}\equiv (\vec{Y}_{1},\ldots, \vec{Y}_{T})$, and $\bar{\vec{Y}}\equiv \sum_{t=1}^{T}\vec{Y}_{t}/T$. The estimators of $\alpha(z_{it})$ and $\beta(z_{it})$ under $\mathrm{H}_0$ are given by $\hat{\gamma}^{\prime}z_{it}$ and $\hat{\Gamma}^{\prime}z_{it}$, where $\hat{\Gamma}=\vec{Y}M_T\hat{F}(\hat{F}^{\prime}M_T\hat{F})^{-1}$ and $\hat{\gamma}=\bar{\vec{Y}}-\hat{\Gamma}\sum_{t=1}^{T}\hat{f}_t/T$.\footnote{It is crucial to use the unrestricted estimator $\hat{F}$ in both $\hat{\Gamma}$ and $\hat{\gamma}$, rather than the restricted one under $\mathrm{H}_0$. This ensures that $\hat{\Gamma}^{\prime}z_{it}$ and $\hat{\beta}(z_{it})$ share a common rotational transformation matrix, justifying the validity of the test. It also avoids the full-rank requirement for $\Gamma$. Thanks to the optimal rate of $\hat{F}$ established in Theorem (ref), this does not cause an issue.} Our test statistic is given by: \begin{align} \mathcal{S} = \frac{1}{J}\sum_{i=1}^{N}\sum_{t=1}^{T}|\hat{\gamma}^{\prime}z_{it} - \hat{\alpha}(z_{it})|^{2}+ \frac{1}{J}\sum_{i=1}^{N}\sum_{t=1}^{T}\|\hat{\Gamma}^{\prime}z_{it} - \hat{\beta}(z_{it})\|^{2}. \end{align} To obtain critical values, we adopt a bootstrap method. Let $\vec{Y}^{\ast}_{t}\equiv (Z_{t}^{\ast\prime}Z_{t})^{-1}Z_{t}^{\ast\prime}Y_t$, $\vec{Y}^{\ast}\equiv (\vec{Y}^{\ast}_{1},\ldots, \vec{Y}^{\ast}_{T})$, and $\bar{\vec{Y}}^{\ast}\equiv \sum_{t=1}^{T}\vec{Y}^{\ast}_{t}/T$, where $Z^{\ast}_t = (z_{1t}w_1,$ $\ldots, z_{Nt}w_{N})^{\prime}$. It is shown in the proof of Theorem (ref) that under $\mathrm{H}_0$, $\mathcal{S} = \sum_{i=1}^{N}\sum_{t=1}^{T}|(\hat{\gamma}-\gamma)^{\prime}z_{it} - (\hat{a}-a)^{\prime}\phi(z_{it})|^2/J+\sum_{i=1}^{N}\sum_{t=1}^{T}\|(\hat{\Gamma}-\Gamma H)^{\prime}z_{it} - (\hat{B}-BH)^{\prime}\phi(z_{it})\|^2/J+o_{p}(J^{-1/2})$. Given this, we estimate the null distribution of $\mathcal{S}$ by the distribution of \begin{align} \mathcal{S}^{\ast} &=\frac{1}{J\omega_0}\sum_{i=1}^{N}\sum_{t=1}^{T}|(\hat{\gamma}^{\ast}-\hat{\gamma})^{\prime}z_{it} - (\hat{a}^{\ast}-\hat{a})^{\prime}\phi(z_{it})|^2\notag\\ &+\frac{1}{J\omega_0}\sum_{i=1}^{N}\sum_{t=1}^{T}\|(\hat{\Gamma}^{\ast}-\hat{\Gamma})^{\prime}z_{it} -(\hat{B}^{\ast}-\hat{B})^{\prime}\phi(z_{it})\|^2 \end{align} conditional on the data. Here, $\hat{\Gamma}^{\ast}\hspace{-0.05cm}=\hspace{-0.05cm}\vec{Y}^{\ast}\hspace{-0.05cm}M_T\hat{F}(\hat{F}^{\prime}\hspace{-0.05cm}M_T\hspace{-0.05cm}\hat{F})^{-1}$ and $\hat{\gamma}^{\ast}\hspace{-0.05cm}=\hspace{-0.05cm}\bar{\vec{Y}}^{\ast}\hspace{-0.05cm}-\hat{\Gamma}^{\ast}\hspace{-0.05cm}(\hat{B}^{\ast\prime}\hat{B}^{\ast})^{-1}\hspace{-0.05cm}\hat{B}^{\ast\prime}\bar{\tilde{Y}}^{\ast}$. For $0<\alpha<1$, let $c_{1-\alpha}$ be the $1-\alpha$ quantile of $\mathcal{S}^{\ast}$ conditional on the data. Thus, we construct the test as follows: reject $\mathrm{H}_0$ if $\mathcal{S}>c_{1-\alpha}$. \begin{thm} Suppose Assumptions (ref)-(ref) hold. Let $\mathcal{S}$ be given in (ref) and $c_{1-\alpha}$ be given after (ref) for $0<\alpha<1$. Assume (i) $N\to\infty$; (ii) $T\geq K+1$; (iii) $J\to\infty$ with $J^{2}\xi^{2}_{J}\log J=o(N)$. In addition, assume $T=o(N)$, $J=o(\min\{N^{1/5},N/T\})$, and $NTJ^{-2\kappa}=o(1)$, where $\kappa>1/2$ is a constant representing the smoothness of $\alpha(\cdot)$ and $\beta(\cdot)$. Then \begin{align*} P(\mathcal{S}>c_{1-\alpha})\to\alpha \text{ under } \mathrm{H}_0 \text{ and } P(\mathcal{S}>c_{1-\alpha})\to1 \text{ under } \mathrm{H}_1. \end{align*} \end{thm} \section{Empirical Analysis} Our empirical analysis is based on the model specified in (ref). Following standard practice in asset pricing, we depart from the notation used in previous sections and denote characteristics observed at time $t-1$ as $z_{i,t-1}$ instead of $z_{it}$. We begin by estimating the model and evaluating its overall performance, and then assess the performance of the extracted factors. The primary objective of our analysis is to investigate whether pricing errors are associated with characteristics and to evaluate the performance of our factors in asset pricing tests. \subsection{Data and Methodology} We use the same dataset as Kellyetal_Characteristics_2019, which is originally from Freybergeretal_Dissecting_2017. The dataset contains monthly returns of $12,813$ individual stocks and $36$ time-varying characteristics, covering the sample period from July 1962 to May 2014. The data is in the form of an unbalanced panel, for which our method is applicable. For the detailed descriptions, refer to these papers. To ensure comparability, we use the same $36$ characteristics as those authors. Following the procedure in Kellyetal_Characteristics_2019, we transform the values of the characteristics into relative rankings within the range $[-0.5, 0.5]$. This transformation standardizes the contributions of characteristics to pricing errors and risk exposures such that the estimation only depends on the rankings of characteristics and is robust to extreme values, sharing the similar logic with the sorting procedure as in FamaFrench_Commonrisk_1993, FamaFrench_FiveFactor_2015. To meet the large $N$ requirement, we select a sample period during which at least $1,000$ individual stocks have observations for both returns and the $36$ characteristics. This results in a sample spanning from September 1968 to May 2014. Based on this dataset, we construct the market factor and five long-short factors following FamaFrench_FiveFactor_2015. These factors exhibit close means and standard deviations and show high correlations with the corresponding factors from Kenneth R. French’s website, as shown in Table (ref). We implement regressed-PCA estimation by selecting the basis functions to be either linear (i.e., $\phi(z_{it})=z_{it}$ including a constant term) or non-linear (via linear B-splines of $z_{it}$).\footnote{{Our econometric theory accommodates a variety of basis functions, such as Fourier series, polynomials, splines, and wavelets. Following Guetal_EmpiricalAsset_2020 and Freybergeretal_Dissecting_2017, we employ splines due to their flexibility, which arises from increasing the number of knots. Unlike polynomials that require a higher degree for flexibility, splines generally produce more stable estimates HastieTibshiraniFriedman_StatisticalLearning_2011.}} Using $\phi(z_{it})=z_{it}$ leads to linear specifications of $\alpha(\cdot)$ and $\beta(\cdot)$, while setting $\phi(z_{it})$ as linear B-splines of $z_{it}$ results in nonlinear specifications of $\alpha(\cdot)$ and $\beta(\cdot)$, where $\alpha(\cdot)$ and $\beta(\cdot)$ are continuous, piecewise linear functions.\footnote{The one dimensional linear B-splines $\{\psi_{j}(z)\}^{J}_{j=1}$ are defined over a set of consecutive, equidistant knots: $\{z_{1},...,z_{J+1}\}$. For $j< J$, $\psi_{j}(z)=(z-z_{j})/(z_{j+1}-z_{j})$ on $(z_{j}, z_{j+1}]$, $\psi_{j}(z)=(z_{j+2}-z)/(z_{j+2}-z_{j+1})$ on $(z_{j+1}, z_{j+2}]$, and 0 elsewhere. For $j=J$, $\psi_{j}(z)=(z-z_{j})/(z_{j+1}-z_{j})$ on $(z_{j}, z_{j+1}]$ and 0 elsewhere.} For ease of comparison, we maintain the same parameter dimension across different specifications. Specifically, we consider 18 characteristics with one internal knot and 12 characteristics with two internal knots in the linear B-splines specifications. The most significant 18/12 characteristics are selected based on the linear specification, which are collected in Table (ref). To implement the weighted bootstrap, we let the bootstrap weights $w_{i}$'s be i.i.d. random variables following the standard exponential distribution. For testing $\alpha(\cdot) = 0$ and linearity of $\alpha(\cdot)$ and $\beta(\cdot)$, we set the number of bootstrap draws to 499. In order to evaluate the performance of the models estimated via regressed-PCA, we compute several measures of fit and prediction. First, we calculate Fama-MacBeth cross-sectional regression $R^{2}$, denoted as $R^2_{\tilde Y}$, which captures the variation in individual stock returns explained by the Fama-MacBeth managed portfolios $\tilde{Y}_t$ constructed from $\phi(z_{i,t-1})$. Next, we report the variation in these managed portfolios explained by the extracted factors $\hat{f}_t$, denoted as $R^{2}_{K}$. We then consider the following three types of $R^2$ measures that directly speak to the ability of the factor models to explain the cross-section of individual stock returns. The first measure (denoted as $R^2$) is total $R^2$ as used in Kellyetal_Characteristics_2019. The second measure ($R^2_{T,N}$) calculates the cross-sectional average of time series $R^2$ across all stocks, which reflects the ability of the factors to capture common variation in stock returns. The third measure ($R^2_{N,T}$) computes the time-series average of cross-sectional goodness-of-fit measures, approximating the Fama-MacBeth cross-sectional regression $R^2$. This measure is particularly relevant for evaluating the model's capacity to explain the cross-section of average returns. The measures are defined as follows:\footnote{The differences among the three $R^2$'s are provided in Appendix (ref).} \begin{align} R^2 & = 1-\frac{\sum_{i, t}[y_{it}- \hat{\alpha}(z_{i,t-1}) - \hat{\beta}(z_{i,t-1})^{\prime}\hat{f}_t]^2}{\sum_{i, t} y_{it}^{2}}, \\ R^2_{T,N} & = 1 - \frac{1}{N} \sum_{i} \frac{\sum_{t}[y_{it}- \hat{\alpha}(z_{i,t-1}) - \hat{\beta}(z_{i,t-1})^{\prime}\hat{f}_t]^2}{\sum_{t}y_{it}^{2}},\\ R^2_{N,T} & = 1 - \frac{1}{T} \sum_{t} \frac{\sum_{i}[y_{it}- \hat{\alpha}(z_{i,t-1}) - \hat{\beta}(z_{i,t-1})^{\prime}\hat{f}_t]^2}{\sum_{i}y_{it}^{2}}. \end{align} We also report a version of these goodness-of-fit measures that zero in on the role of factors in explaining the time-series as well as the cross-section of stock returns, by excluding the conditional alphas $\hat{\alpha}(z_{i,t-1})$; see (ref)-(ref) for the corresponding formulas. Finally, we assess the out-of-sample prediction and fit using expanding-window estimations. For $t\geq 120$, we use the data up to time period $t-1$ to implement the regressed-PCA and obtain estimates such as $\hat{a}_{t-1}$, $\hat{B}_{t-1}$, $\hat{\alpha}_{t-1}(\cdot)$, $\hat{\beta}_{t-1}(\cdot)$, and $\hat{F}_{t-1}\equiv (\hat{f}^{(t-1)}_{1},\ldots, \hat{f}^{(t-1)}_{t-1})^{\prime}$. Using these, we compute the out-of-sample prediction of $y_{it}$ as $\hat{\alpha}_{t-1}(z_{i,t-1})+\hat{\beta}_{t-1}(z_{i,t-1})^{\prime}\hat{\lambda}_{t}$, where $\hat{\lambda}_{t} = \sum_{s\leq t-1}\hat{f}^{(t-1)}_s/(t-1)$, which is the average of factor estimates through time period $t-1$. The out-of-sample predictive $R^2$ is: \begin{align} R^2_O & = 1-\frac{\sum_{i, t\geq 120}[y_{it}- \hat{\alpha}_{t-1}(z_{i,t-1}) - \hat{\beta}_{t-1}(z_{i,t-1})^{\prime}\hat{\lambda}_t]^2}{\sum_{i, t\geq 120} y_{it}^{2}}; \end{align} see (ref)-(ref) for another two versions. We calculate the out-of-sample realized factor returns at time period \(t\) as: $\hat{f}_{t-1, t} = \hat B^{\prime}_{t-1}\tilde Y_{t} = \hat B^{\prime}_{t-1}(\Phi(Z_{t-1})^{\prime} \Phi(Z_{t-1}))^{-1}\Phi^{\prime}(Z_{t-1})Y_{t}$. Although the resulting factor returns are only known ex post at time period $t$, they represent returns on portfolios that are constructed ex ante, using weights based on estimates obtained at time period $t-1$. Using these, we can access how much of the cross-sectional variation of individual stock returns can be explained by the pre-estimated $\hat{\beta}_{t-1}(z_{i,t-1})$. We then define the out-of-sample fit $R^2$ as: \begin{align} R^2_{f,O} & = 1-\frac{\sum_{i, t\geq 120}[y_{it}- \hat{\beta}_{t-1}(z_{i,t-1})^{\prime}{\hat f_{t-1,t}}]^2}{\sum_{i, t\geq 120} y_{it}^{2}} ; \end{align} see (ref)-(ref) for another two versions. \subsection{Empirical Results} \subsubsection{Model Estimation} The main findings presented in Tables (ref)-(ref) can be summarized as follows. First, all of our measures of fit indicate that a low-dimensional factor model is unlikely to explain the time-series\textemdash or the cross-section\textemdash of individual stock returns (rather than the managed portfolios). In all specifications at least 5 or 6 factors are required for most of the in-sample R-squared to exceed $10\%$. That said, the total in-sample $R^2$'s in our model are smaller than those of Kellyetal_Characteristics_2019. This is not surprising, since the objective of their IPCA estimation is maximizing the total in-sample $R^2$, as discussed in (ref). We extract factors that capture the most time-series comovement within a set of portfolios, which, in turn, reflect the most cross-sectional variation in individual asset returns. In contrast, our out-of-sample $R_{O}^{2}$'s are 0.54% for the linear specification and 0.59% and 0.57% for the two nonlinear specifications,\footnote{Our out-of-sample predictive $R^2$'s are invariant to the number of factors, because $\hat{\alpha}_{t-1}(z_{i,t-1})+ \hat{\beta}_{t-1}(z_{i,t-1})^{\prime}\hat{\lambda}_t = \phi(z_{i,t-1})^{\prime}[\hat{a}_{t-1} + \hat{B}_{t-1}\hat{F}^{\prime}_{t-1}1_{t-1}/(t-1)] = \phi(z_{i,t-1})^{\prime}\sum_{s=1}^{t-1}\tilde{Y}_t/{t-1}$, which does not depend on $K$.} which are comparable to the 0.60% in Kellyetal_Characteristics_2019's linear specification with six factors. Similarly, the out-sample-sample fits are close to those of Kellyetal_Characteristics_2019. With six factors, the out-sample-sample $R^{2}_{f,O}$'s are 15.38%, 16.20%, and 16.16% for the three specifications, which are comparable to the 17.80% in Kellyetal_Characteristics_2019. Moreover, all in-sample and out-of-sample fits improve as the number of factors increases, as factor loadings soak up more of the variation in managed portfolio returns that is otherwise attributed to the alphas. In addition, compared to the linear specification, the nonlinear specifications based on linear B-splines significantly improve both in-sample and out-of-sample fits in most cases. They also slightly enhance the out-of-sample predictive $R^2$'s. Both the improved fits and robustness highlight the advantage of the nonlinear specifications. Second, the results of testing the linearity of $\alpha(\cdot)$ and $\beta(\cdot)$ explain why we observe lower $R^2$'s in the linear specification. Linearity is strongly rejected at the $1\%$ level in all factor models estimated with one to ten factors (Tables (ref)-(ref) report the $p$-values concisely to save space, all of which indicate rejection at the $1\%$ level). Moreover, we also find robust evidence rejecting the null hypothesis $\alpha(\cdot) = 0$ across all cases. Additional empirical results are provided in Appendix (ref) (see Tables (ref)-(ref)). Before analyzing the contribution of each characteristic to pricing errors and risk exposures, we first determine the signs of the extracted factors. Under the normalization $B^{\prime}B= I_K$ and $F^{\prime}M_TF/T$ being diagonal with descending diagonal entries, the signs of the factors remain undetermined. To address this, we set the sample means of the factors to be positive, ensuring the unconditional risk premium on each factor is positive. To interpret the factors, we examine their correlations with the market factor and five long-short factors from Kenneth R. French’s website as discussed in Table (ref), and conduct projection regressions, as detailed in Appendix (ref) (see Tables (ref)-(ref)). We find the substantial correlations between these factors and our factors, with both sets explaining significant variations in each other. Figures (ref) and (ref) illustrate the contribution of each characteristic to pricing errors and risk exposures under the linear specification. Figure (ref) reveals that the $95\%$ confidence intervals of characteristic coefficients in pricing errors remain relatively stable as the number of factors increases from one to six. Notably, 22 out of 36 characteristics remain significant at the $5\%$ level for $K=6$, in stark contrast to the results reported by Kellyetal_Characteristics_2019. Figure (ref) displays the characteristic coefficients in risk exposures for the first six factors. In contrast to Kellyetal_Characteristics_2019's finding that 13 out 36 characteristics are significant in driving risk exposures, we find 24 significant characteristics for $K=6$. Specifically, the coefficient of “market cap" in the first factor is negative and large in magnitude; the fourth and sixth factors exhibit substantial positive loadings on “market beta" (i.e., “beta"); the second and fifth factors display significant positive loadings on “book-to-market ratio” (i.e.,“bm"). These findings align with the traditional views of asset pricing anomalies as discussed in FamaFrench_Commonrisk_1993, FamaFrench_FiveFactor_2015. More results for the two nonlinear specifications are provided in Appendix (ref) (see Figures (ref)-(ref)). Taking advantage of the fact that our regressed-PCA method does not require a large time dimension ($T$), we perform subsample analyses using five-year intervals starting in January 1970. Figure (ref) presents the key results. The left panel shows that average pricing errors (measured by $\|a\|^2$) under the nonlinear specifications are significantly smaller than those under the linear specification. For the linear case, average pricing errors are the highest during the initial subsample period (1970-1974), decline over time, rise again during the equity market “boom” of the 1990s, peak in the early 2000s, and then drop sharply. Under the nonlinear specifications, the patterns differ somewhat, with average pricing errors spiking around 1990-1994 and subsequently decreasing to levels comparable to those in the linear case by the end of the sample period. This decline may reflect the growing prevalence of quantitative investing, which reduces mispricing by exploiting characteristic-related anomalies, as suggested by McLeanPontiff_DoesAcademic_2016 and Greenetal_Characteristics_2017. The right panel of Figure (ref) illustrates the proportion of time-series and cross-sectional variation in stock returns explained by common factors (measured by $R^{2}_{f,T,N}$ and $R^{2}_{f,N,T}$), which appears similar across model specifications. Notably, all the reported $R^2$ measures decline from 1970, reach a trough in the mid-1990s, and then steadily rise until the sample ends in 2014. This observation aligns with the empirical findings: Campbell2001have document a noticeable increase in firm-level volatility between 1962 and 1997; while extending this analysis to 2021, Campbell2022idiosyncratic find that idiosyncratic volatility declined after peaking in 1999-2000. Similar trends are evident in the out-of-sample fit measures shown in Figure (ref). \subsubsection{Trading Strategies} We construct trading strategies based on our estimated models. While constructing the MVE portfolio on individual stocks is usually infeasible due to the challenge in estimating a high-dimensional covariance matrix, the model in (ref) enables us to devise trading strategies by leveraging the \textit{alpha} (i.e., mispricing) and \textit{beta} (i.e., risk) roles of characteristics. Specifically, the initial step of the regressed-PCA method (i.e., Fama-MacBeth regressions) reduces a large number of individual stock returns into a smaller set of characteristic-managed portfolios. These portfolios exhibit a classical factor structure as outlined in (ref), facilitating the construction of a pure-$alpha$ strategy and an MVE factor portfolio. By (ref) and Theorem (ref), $\hat{a}^{\prime}\tilde{Y}_{t}\xrightarrow{p} \|a\|^{2}$ for each $t$ as $N\to\infty$, implying that $\hat{a}^{\prime}\tilde{Y}_{t}$ represents a portfolio with positive returns (if $a\neq0$) and no risk asymptotically. Using the expanding-window procedure in Section (ref), we construct a pure-$alpha$ “arbitrage” portfolio for $t\geq 120$ as follows: \begin{align} R_{\alpha,t} = \hat{a}_{t-1}^{\prime}\tilde{Y}_{t} = \hat{a}_{t-1}^{\prime} (\Phi(Z_{t-1})^{\prime} \Phi(Z_{t-1}))^{-1}\Phi(Z_{t-1})^{\prime} Y_{t}. \end{align} This portfolio can be equivalently constructed from individual stocks by using weights $\Phi(Z_{t-1}) (\Phi(Z_{t-1})^{\prime} \Phi(Z_{t-1}))^{-1}\hat{a}_{t-1}$. Since $R_{\alpha,t}$ relies on estimates obtained at time period $t-1$, it is tradable (ex ante), while $\hat{a}^{\prime}\tilde{Y}_{t}$ is based on full-sample estimates. We consider $R_{\alpha,t}$, the out-of-sample version of $\hat{a}^{\prime}\tilde{Y}_{t}$, and refer to their Sharpe ratios as the out-of-sample and in-sample Sharpe ratios of the pure-$alpha$ portfolio, respectively. Similarly, we construct the out-of-sample version of $\hat{f}_t = \hat{B}^{\prime}\tilde{Y}_t$ as $\hat{f}_{t-1, t} = \hat B^{\prime}_{t-1}\tilde Y_{t}$, which has been introduced in (ref). This enables the construction of an MVE factor portfolio based on the same expanding-window procedure: for $t\geq 120$, \begin{align} R_{\beta,t} = \hat{\mu}_{t-1}^{\prime} \hat{\Sigma}^{-1}_{t-1} \hat{f}_{t-1, t} = \hat{\mu}_{t-1}^{\prime} \hat{\Sigma}_{t-1}\hat B^{\prime}_{t-1}(\Phi(Z_{t-1})^{\prime} \Phi(Z_{t-1}))^{-1}\Phi^{\prime}(Z_{t-1})Y_{t}, \end{align} where $\hat{\mu}_{t-1}$ and $\hat{\Sigma}_{t-1}$ are estimates of the (conditional) mean and covariance matrix of $f_t$ at time period $t-1$.\footnote{Specifically, $\hat{\mu}_{t-1} = \sum_{s\leq t-1}\hat{f}^{(t-1)}_s/(t-1)$ and $\hat{\Sigma}_{t-1} =\sum_{s\leq t-1}(\hat{f}^{(t-1)}_s - \hat{\mu}_{t-1})(\hat{f}^{(t-1)}_s - \hat{\mu}_{t-1})^{\prime} /(t-2)$.} Weights for $R_{\beta,t}$ can also be derived from and applied on individual stocks, making it tradable. The in-sample counterpart of $R_{\beta,t}$ is $\hat{\mu}^{\prime} \hat{\Sigma}^{-1} \hat{f}_{t}$, where $\hat{\mu}$ and $\hat{\Sigma}$ are the full-sample estimates of the (unconditional) mean and covariance matrix of $f_t$.\footnote{Specifically, $\hat{\mu} = \sum_{s\leq T}\hat{f}_s/T$ and $\hat{\Sigma} =\sum_{s\leq T}(\hat{f}_s - \hat{\mu})(\hat{f}_s - \hat{\mu})^{\prime} /(T-1)$.} Their Sharpe ratios are referred to as the out-of-sample and in-sample Sharpe ratios of the MVE factor portfolio, respectively. We further combine $R_{\alpha,t}$ and $\hat{f}_{t-1, t}$ to form a set of $K+1$ factor portfolios, constructing an MVE portfolio following the procedure outlined for $R_{\beta,t}$. The in-sample counterpart is derived from $\hat{a}^{\prime}\tilde{Y}_{t}$ and $\hat{f}_t$. The resulting Sharpe ratios are referred to as the out-of-sample and in-sample Sharpe ratios of the combined MVE portfolio, respectively. By imposing a factor structure on the conditional covariance of individual asset returns as in (ref), the combined MVE portfolio provides an approximation to the stock market's MVE portfolio. Tables (ref) and (ref) present the annualized in-sample and out-of-sample Sharpe ratios for the pure-$alpha$, MVE factor, and combined MVE portfolios. In all subsequent tables, “Regressed-PCA” denotes the results under linear specifications of $\alpha(\cdot)$ and $\beta(\cdot)$ with 36 characteristics, while “Regressed-PCA S1” and “Regressed-PCA S2” correspond to nonlinear specifications with 18 and 12 characteristics, respectively. “IPCA” represents the results based on the linear specification using IPCA estimations with the same 36 characteristics. Two remarks are essential for understanding Tables (ref) and (ref). First, the in-sample Sharpe ratio of our MVE factor portfolio increases with $K$, whereas that of IPCA's MVE factor portfolio may not. This reflects the stability of our factor construction approach compared to IPCA, in particular the inherent orthogonality of regressed PCA factors (see Section (ref)). We report the out-of sample Sharpe ratio of each incremental regressed PCA factor in Table (ref) along with the pure-$alpha$ and MVE portfolios, but we do not report individual Sharpe ratios for IPCA factors as they are not incremental, since all of the factors are estimated jointly for each $K$. Second, for our regressed-PCA approach, the squared in-sample Sharpe ratio of the combined MVE portfolio equals the sum of those of the pure-\textit{alpha} and MVE factor portfolios as the two are orthogonal by construction.\footnote{This follows from $\sum_{t=1}^{T}\hat{a}^{\prime}(\tilde{Y}_{t}-\bar{\tilde{Y}})\hat{f}_t/T = \hat{B}^{\prime}\tilde{Y}M_T \tilde{Y}^{\prime}\hat{a}/T = \hat{\Lambda}\hat{B}^{\prime}\hat{a} = 0$, where $\hat{\Lambda}$ is a diagonal matrix with diagonal entries being the large $K$ eigenvalues of $\tilde{Y}M_T \tilde{Y}^{\prime}/T$.} As a result, the in-sample Sharpe ratios of the pure-\textit{alpha} and MVE factor portfolios are identical to their respective contributions to the combined MVE portfolio and are omitted in Table (ref). Similarly, the out-of-sample Sharpe ratio of our MVE factor portfolio equals its contribution to the combined MVE portfolio, which is also omitted in Table (ref). The orthogonality property does not hold for IPCA, highlighting another advantage of regressed-PCA. The main findings are summarized as follows. First, under the linear specification, we compare the Sharpe ratios of the pure-\textit{alpha} and MVE factor portfolios constructed based on regressed-PCA and IPCA, separately. As $K$ increases, the Sharpe ratio of our pure-\textit{alpha} portfolio remains high (in-sample: 3.89 to 4.50; out-of-sample: 3.18 to 3.84), while that of the MVE factor portfolio is comparatively low (in-sample: 0.65 to 0.89; out-of-sample: 0.44 to 0.72). The observed increase in the Sharpe ratio of the pure-\textit{alpha} portfolio as $K$ grows suggests that factors play a crucial role in hedging common variation in stock returns, reducing the volatility of the pure-\textit{alpha} portfolio at a rate exceeding the decline in alphas. The findings also align with the testing evidence of nonzero pricing errors in the linear specification found in Section (ref), reinforcing our conclusions. While the Sharpe ratios of IPCA's pure-\textit{alpha} and MVE factor portfolios are more comparable (in-sample: 1.61 to 3.13 vs. 1.07 to 2.85; out-of-sample: 1.31 to 2.84 vs. 0.92 to 1.72), the former is higher than the latter. Table (ref) also shows that the higher Sharpe ratio of our pure-\textit{alpha} portfolio compared with IPCA's is due to its low volatility, consistent with the idea of no risk asymptotically. The lower volatility of our pure-\textit{alpha} portfolio underscores the superior hedging properties of our factors. Notably, regressed-PCA consistently yields higher in-sample and out-of-sample Sharpe ratios for the combined MVE portfolio than IPCA across all $K$. This indicates that the combined MVE portfolio constructed with regressed-PCA better approximates the stock market's MVE portfolio compared to IPCA. In particular, the high Sharpe ratios of the combined MVE portfolio primarily arise from the high Sharpe ratios of the pure-\textit{alpha} portfolio. The better approximation is also attributed to the stability and orthogonality properties of our regressed-PCA as discussed above. Second, the nonlinear specifications yield higher in-sample and out-of-sample Sharpe ratios for the MVE factor and combined MVE portfolios than the linear specification across all values of $K$ (except $K=1$ in Regressed-PCA S2 of Table (ref)). As $K$ increases, the Sharpe ratio of the pure-$alpha$ portfolio under the nonlinear specifications falls below that under the linear specification, indicating that nonlinear models yield smaller magnitudes of pricing errors. Meanwhile, the nonlinear specifications also yield better factors, reflected in higher Sharpe ratios of the MVE factor portfolio (in-sample: 0.61 to 3.96; out-of-sample: 0.51 to 3.33) than the linear specification (in-sample: 0.65 to 0.89; out-of-sample: 0.44 to 0.72). More importantly, the Sharpe ratios of the combined MVE portfolio from the nonlinear specifications are substantially higher than those from the linear specification, implying the potential nonlinearity of stochastic discount factor in the U.S. stock market. Similarly, the high Sharpe ratios of our combined MVE portfolio are primarily derived by the pure-\textit{alpha} portfolio for $K\leq 5$, the Sharpe ratios of the MVE factor portfolio catch up and become comparable for larger $K$. Nevertheless, the Sharpe ratios of our pure-\textit{alpha} portfolio remain consistently high (in-sample: 3.47 to 4.80; out-of-sample: 3.09 to 4.26) compared with those of the MVE factor portfolio (in-sample: 0.61 to 3.96; out-of-sample: 0.51 to 3.33), providing strong evidence of nonzero pricing errors. This finding corroborates the evidence of nonzero pricing errors presented in Section (ref). All these findings align with the strong evidence of nonlinearity found in Section (ref) and underscore the advantage of the flexible nonlinear specifications. Lastly, we perform a subsample analysis of the pure-\textit{alpha} strategy, using five-year intervals starting in January 1970. Within each subsample, we construct the pure-\textit{alpha} “arbitrage” portfolio defined in (ref), employing expanding window estimation from the second year onward. The key results are presented in Figure (ref). Notably, the left panel of Figure (ref) displays the decline in Sharpe ratios of the portfolio, and the right panel indicates that the decline is primarily driven by a reduction in the portfolio’s average returns, rather than an increase in its standard deviations. This is consistent with the findings in Figure (ref), where we observe a significant decline in pricing errors since 2000. In summary, we provide further evidence supporting the findings of nonlinearity and nonzero pricing errors, as well as their significant decline over time in Section (ref). This section also reinforces the advantages of regressed-PCA. It is important to note that while nonzero pricing errors are observed, they do not necessarily imply market inefficiency unless the possibility of model misspecification is ruled out. \subsubsection{Asset Pricing Tests} We now evaluate the performance of our factors in asset pricing tests, considering a broad class of testing portfolios. Specifically, we examine three groups of Fama-MacBeth managed portfolios: Regressed-PCA, Regressed-PCA S1, and Regressed-PCA S2, as well as IPCA's managed portfolios. Additionally, we include two groups of single sorted portfolios based on 55 characteristics from Haddadetal_FactorTiming_2020 and our 36 characteristics, along with several groups of double sorted portfolios following FamaFrench_CSTS_2020. Table (ref) reports the bilateral correlations and standard deviations of these portfolios. The Fama-MacBeth managed portfolios exhibit lower bilateral correlations than others, and smaller standard deviations compared to the sorted portfolios. This supports the optimality of the Fama-MacBeth managed portfolios, which are maximally diversified, as discussed in Section (ref). We compare our factors with several existing sets: IPCA's factors, the five factors from FamaFrench_FiveFactor_2015 (denoted as FF5), and the factors constructed following Kozaketal_Interpreting_2018 (denoted as KNS). The comparison statistics from time series regressions are reported following the analysis in FamaFrench_CSTS_2020. The results for five groups of testing portfolios with $K=5$ are presented in Tables (ref) and (ref), with additional results provided in Appendix (ref) (see Tables (ref)-(ref)). The main findings are summarized as follows. First, for the Fama-MacBeth managed portfolios in Group I, both our factors and IPCA's factors outperform FF5 and KNS's factors in terms of average absolute intercepts ($A|a|$). They also achieve larger average regression $R^2$'s ($AR^2$), leading to smaller average standard errors ($As(a)$). The resulting average absolute $t$-statistics ($A|t(a)|$) and $GRS$ statistics ($GRS$) are comparable across all factors. Notably, our factors under the nonlinear specifications do not improve performance, as the testing portfolios are derived from the linear specification. Second, our factors consistently outperform IPCA's factors in pricing the sorted portfolios in Groups II, III, V, and VI, as evidenced by smaller average absolute intercepts, $t$-statistics, and $GRS$ statistics. The higher average absolute $t$-statistics and $GRS$ statistics for IPCA's factors arise from their larger average regression $R^2$'s (or smaller average residual standard deviations ($As(e)$) or standard errors). To investigate this, we project IPCA's factors onto our factors under the linear specification (without a constant term) and treat the resulting residuals as new factors (denoted as IPCA$\setminus$Regressed-PCA). These residuals yield even larger average absolute intercepts and substantial average regression $R^2$'s, indicating that IPCA's factors capture more time-series variation in returns but less cross-sectional variation, likely due to overfitting idiosyncratic noise rather than extracting true signals. This overfitting is much less pronounced in Group I, as the Fama-MacBeth managed portfolios are maximally diversified, unlike the sorted portfolios, as discussed in Section (ref) and shown in Table (ref). Third, our factors under the nonlinear specifications significantly reduce average absolute intercepts and $t$-statistics for sorted portfolios. This improvement stems from the nonparametric nature of sorting, as discussed in Section (ref). However, due to the high correlations among sorted portfolios, as shown in Table (ref), no noticeable improvement in $GRS$ statistics is observed. Moreover, factors under the nonlinear specifications with 12 characteristics outperform FF5, yielding smaller average absolute $t$-statistics and $GRS$ statistics, with comparable average absolute intercepts. The higher $t$-statistics and $GRS$ statistics for FF5 also arise from their larger regression $R^2$'s, suggesting the presence of unpriced components similar to IPCA's factors Danieletal_CrossSection_2020,KozakNagel_Whydospan_2022. Moreover, our factors also outperform KNS's factors, achieving smaller average absolute intercepts and $t$-statistics. Lastly, for IPCA's managed portfolios in Group IV, IPCA's factors exhibit inferior performance compared to other factors, with larger average absolute intercepts, $t$-statistics, and $GRS$ statistics. This finding is surprising given that IPCA factors are derived from its managed portfolios. As the testing portfolios are derived under the linear specification, our factors under the linear specification outperform those under the nonlinear specifications. The performance of our factors under the linear specification is comparable to FF5 and KNS's factors. In addition, the performance comparisons based on relative metrics (e.g., $Aa^2/V\overline{r}$ or $A\lambda^2/V\overline{r}$) align with those based on absolute metrics ($A|a|$), reinforcing the robustness of our results. In summary, our factors demonstrate superior performance compared to IPCA's factors and long-short factors in asset pricing tests, with robustness across a wide range of testing portfolios. \section{Conclusion} In this paper, we considered semiparametric conditional latent factor models to address the “characteristics versus covariances” debate and the “\textit{factor zoo}” problem in cross-sectional asset pricing. We proposed a simple and tractable sieve estimation approach combined with a weighted-bootstrap procedure for conducting inference on the \textit{alpha} and \textit{beta} functions. We established large-sample properties of the estimators and validity of the tests under large $N$, even when $T$ is small. In addition to offering formal inference procedures and well-founded asymptotic properties, our approach presents several advantages over existing methods such as IPCA and projected-PCA. Specifically, it is computationally efficient and accommodates nonzero alphas, time-varying characteristics, unbalanced panels, and short samples, making it particularly suitable for empirical asset pricing applications. These results enable the estimation of conditional factor structures for a large set of individual assets by incorporating numerous characteristics, accounting for nonlinearity without requiring pre-specified factors. Moreover, our approach disentangles the role of risk from the purely predictive power of return characteristics that is unrelated to common risk exposures. We applied this method to analyze the cross-sectional differences in individual stock returns in the U.S. market. The findings provide robust evidence of large nonzero pricing errors and nonlinearity in both \textit{alpha} and \textit{beta} functions, leading to the formation of “arbitrage” portfolios with exceptionally high Sharpe ratios (exceeding 3). Additionally, we documented a significant decline in pricing errors since 2000. Our method delivers stable and reliable factor construction without the risk of overfitting, yielding out-of-sample mean-variance efficient portfolios with Sharpe ratios in excess of 4. We also demonstrated that our factors outperform existing alternatives in explaining the cross-section of U.S. stock returns. {3.3pt} \begin{table}[!htbp] \begin{threeparttable} \caption{Results under linear specifications of $\alpha(\cdot)$ and $\beta(\cdot)$ with 36 characteristics\tnote{\dag}} \begin{tabular}{clccccccccccc} \hline\hline &\multicolumn{11}{c}{Unrestricted ($\alpha(\cdot) \neq 0$)}&\\ \cline{2-12} &$K$&$R^{2}_{K}$&$R^{2}$&$R^{2}_{T,N}$&$R^{2}_{N,T}$&$R^{2}_{f}$&$R^{2}_{f,T,N}$&$R^{2}_{f,N,T}$&$R^{2}_{f,O}$&$R^{2}_{f,T,N,O}$&$R^{2}_{f,N,T,O}$\\ \cline{2-12} &1&26.55& 2.54&1.37&0.36& 2.07&0.59&0.11& 6.23 &3.79&5.65\\ &2&36.42&4.52&2.43&1.76&4.08&1.75&1.37&13.59&10.63&11.28\\ &3&45.03&5.70&3.70&2.70&5.24&2.95&2.31&14.09&11.10&11.67\\ &4&52.55&11.69&8.55&9.27&11.28&7.92&8.69&14.74&12.15&12.11\\ &5&58.65&11.90&8.73&9.48&11.49&7.99&8.90&15.17&12.90&12.42\\ &6&64.20&13.90&10.30&11.80&13.53&9.79&11.24&15.38&13.19&12.63\\ &7&69.15&15.59&12.23&13.76&15.23&11.71&13.23&15.62&13.32&12.87\\ &8&72.84&15.93&12.59&13.98&15.56&12.00&13.44&15.90&13.58&13.12\\ &9&76.26&16.08&12.67&14.19&15.72&12.15&13.64&16.13&13.83&13.33 \\ &10&79.15&16.23&12.82&14.35&15.87&12.34&13.80&16.29&14.06&13.47\\ \cline{2-12} &$K$&$R^{2}_{\tilde Y}$&$R^{2}_{O}$&$R^{2}_{T,N,O}$&$R^{2}_{N,T,O}$&$p_{\alpha}$&$p_{\text{lin}}$\\ \cline{2-12} &1-10&20.89& 0.54&0.64&0.21&$<1\%$&$<1\%$&\\ \hline\hline \end{tabular} \begin{tablenotes} • $K$: the number of factors specified; $R^{2}_{\tilde Y}$: Fama-MacBeth cross-sectional regression $R^2$ ($\%$); $R^2_{K}$: the variation of the Fama-MacBeth managed portfolios $\tilde{Y}_t$ captured by the extracted factors $\hat{f}_t$ ($\%$); $R^{2}$, $R^{2}_{T,N}$, $R^{2}_{N,T}$: various in-sample $R^2$'s ($\%$), see (ref)-(ref); $R^{2}_{f}$, $R^{2}_{f,T,N}$, $R^{2}_{f,N,T}$: various in-sample $R^2$'s without $\alpha(\cdot)$ ($\%$), see (ref)-(ref); $R^{2}_{f,O}$, $R^{2}_{f,T,N,O}$, $R^{2}_{f,N,T,O}$: various out-of-sample fit $R^2$'s ($\%$), see (ref)-(ref); $R^{2}_O$, $R^{2}_{T,N,O}$, $R^{2}_{N,T,O}$: various out-of-sample predictive $R^2$'s ($\%$), see (ref)-(ref); $p_{\alpha}$ and $p_{\text{lin}}$: the $p$-values of \textit{alpha} test ($\alpha(\cdot)=0$) and model specification test (joint linearity of $\alpha(\cdot)$ and $\beta(\cdot)$), respectively. \end{tablenotes} \end{threeparttable} \end{table} {3.3pt} \begin{table}[!htbp] \begin{threeparttable} \caption{Results under nonlinear specifications of $\alpha(\cdot)$ and $\beta(\cdot)$ with 18 characteristics\tnote{\dag}} \begin{tabular}{clccccccccccc} \hline\hline &\multicolumn{11}{c}{Unrestricted ($\alpha(\cdot) \neq 0$)}&\\ \cline{2-12} &$K$&$R^{2}_{K}$&$R^{2}$&$R^{2}_{T,N}$&$R^{2}_{N,T}$&$R^{2}_{f}$&$R^{2}_{f,T,N}$&$R^{2}_{f,N,T}$&$R^{2}_{f,O}$&$R^{2}_{f,T,N,O}$&$R^{2}_{f,N,T,O}$\\ \cline{2-12} &1 & 41.61& 5.94 &3.47 & 3.60 & 5.52 & 2.99 & 3.11 &11.27&7.81&8.93 & \\ &2 & 59.05& 9.56 & 6.17 & 6.91 & 9.18 & 5.67 & 6.33 &14.04&11.31&11.29 &\\ &3 & 64.47 & 10.42 & 6.78 & 7.96 & 10.03 & 6.27 & 7.38 &14.64&11.93&11.95 & \\ &4 & 68.99& 13.83 & 10.26 & 11.52 & 13.40 & 9.80 & 10.90 &15.44&12.98&12.54&\\ &5 &72.33& 14.32 & 10.73 & 11.98 & 13.91 & 10.29 & 11.38 &15.78&13.43&12.89 & \\ &6 & 75.35& 14.71 & 10.97 & 12.40 & 14.29 & 10.55 & 11.86 &16.20&14.16 &13.18 &\\ &7 &77.63& 15.28 & 11.78 & 12.99 & 14.84 & 11.27 & 12.42 &16.45&14.34&13.37 & \\ &8 &80.83& 15.44 & 11.98 & 13.16 & 15.10 & 11.59 & 12.73 &16.59&14.50&13.52 & \\ &9 &82.88& 15.84 & 12.33 & 13.49 & 15.48 & 11.87 & 13.05 &16.86&14.69&13.81 &\\ &10 &85.61 &16.39 & 12.89 & 13.93 & 15.71 & 11.80 & 13.14 &16.98&14.72 &13.86 & \\ \cline{2-12} &$K$&$R^{2}_{\tilde Y}$&$R^{2}_{O}$&$R^{2}_{T,N,O}$&$R^{2}_{N,T,O}$&$p_{\alpha}$&$p_{\text{lin}}$\\ \cline{2-12} &1-10&21.11& 0.59&0.64&0.28&$<1\%$&$<1\%$&\\ \hline\hline \end{tabular} \begin{tablenotes} • $K$: the number of factors specified; $R^{2}_{\tilde Y}$: Fama-MacBeth cross-sectional regression $R^2$ ($\%$); $R^2_{K}$: the variation of the Fama-MacBeth managed portfolios $\tilde{Y}_t$ captured by the extracted factors $\hat{f}_t$ ($\%$); $R^{2}$, $R^{2}_{T,N}$, $R^{2}_{N,T}$: various in-sample $R^2$'s ($\%$), see (ref)-(ref); $R^{2}_{f}$, $R^{2}_{f,T,N}$, $R^{2}_{f,N,T}$: various in-sample $R^2$'s without $\alpha(\cdot)$ ($\%$), see (ref)-(ref); $R^{2}_{f,O}$, $R^{2}_{f,T,N,O}$, $R^{2}_{f,N,T,O}$: various out-of-sample fit $R^2$'s ($\%$), see (ref)-(ref); $R^{2}_O$, $R^{2}_{T,N,O}$, $R^{2}_{N,T,O}$: various out-of-sample predictive $R^2$'s ($\%$), see (ref)-(ref); $p_{\alpha}$ and $p_{\text{lin}}$: the $p$-values of \textit{alpha} test ($\alpha(\cdot)=0$) and model specification test (joint linearity of $\alpha(\cdot)$ and $\beta(\cdot)$), respectively. \end{tablenotes} \end{threeparttable} \end{table} {3.3pt} \begin{table}[!htbp] \begin{threeparttable} \caption{Results under nonlinear specifications of $\alpha(\cdot)$ and $\beta(\cdot)$ with 12 characteristics\tnote{\dag}} \begin{tabular}{clccccccccccccc} \hline\hline &\multicolumn{11}{c}{Unrestricted ($\alpha(\cdot) \neq 0$)}&\\ \cline{2-12} &$K$&$R^{2}_{K}$&$R^{2}$&$R^{2}_{T,N}$&$R^{2}_{N,T}$&$R^{2}_{f}$&$R^{2}_{f,T,N}$&$R^{2}_{f,N,T}$&$R^{2}_{f,O}$&$R^{2}_{f,T,N,O}$&$R^{2}_{f,N,T,O}$\\ \cline{2-12} &1 & 42.78 & 5.57 & 2.98 & 3.32 & 5.19 & 2.54 & 2.83 &11.08&7.57&8.77 \\ &2 & 61.36 & 9.56 & 5.97 & 6.87 & 9.18 & 5.51 & 6.26& 13.85 &11.12& 10.99 \\ &3 & 67.77 & 10.59 & 6.65 & 7.88 & 10.20 & 6.15 & 7.29 &14.66&12.25 &11.83 \\ &4 & 72.86 & 13.62 & 10.09 & 11.35 & 13.17 & 9.64 & 10.67 &15.39&13.53&12.53 \\ &5 & 76.92 & 14.14 & 10.43 & 12.01 & 13.73 & 10.01 & 11.48&15.82&13.94 &12.90 \\ &6 & 80.63 & 14.94 & 11.45 & 12.75 & 14.42 & 10.51 & 12.05 &16.16&14.20 &13.22\\ &7 & 84.29 & 15.17 & 11.59 & 12.94 & 14.76 & 10.77 & 12.45 &16.57&14.59&13.57\\ &8 & 87.42 & 15.45 & 11.87 & 13.23 & 15.26 & 11.47 & 12.98&16.94&14.83 &13.88 \\ &9 & 89.11 & 16.33 & 12.68 & 13.94 & 16.16 & 12.31 & 13.72 &17.12&15.00 &14.09 \\ &10 & 90.72 & 16.54 & 12.91 & 14.17 & 16.38 & 12.54 & 13.95 &17.30&15.19&14.29\\ \cline{2-12} &$K$&$R^{2}_{\tilde Y}$&$R^{2}_{O}$&$R^{2}_{T,N,O}$&$R^{2}_{N,T,O}$&$p_{\alpha}$&$p_{\text{lin}}$\\ \cline{2-12} &1-10 &20.72& 0.57 & 0.57 & 0.27 &$<1\%$&$<1\%$&\\ \hline\hline \end{tabular} \begin{tablenotes} • $K$: the number of factors specified; $R^{2}_{\tilde Y}$: Fama-MacBeth cross-sectional regression $R^2$ ($\%$); $R^2_{K}$: the variation of the Fama-MacBeth managed portfolios $\tilde{Y}_t$ captured by the extracted factors $\hat{f}_t$ ($\%$); $R^{2}$, $R^{2}_{T,N}$, $R^{2}_{N,T}$: various in-sample $R^2$'s ($\%$), see (ref)-(ref); $R^{2}_{f}$, $R^{2}_{f,T,N}$, $R^{2}_{f,N,T}$: various in-sample $R^2$'s without $\alpha(\cdot)$ ($\%$), see (ref)-(ref); $R^{2}_{f,O}$, $R^{2}_{f,T,N,O}$, $R^{2}_{f,N,T,O}$: various out-of-sample fit $R^2$'s ($\%$), see (ref)-(ref); $R^{2}_O$, $R^{2}_{T,N,O}$, $R^{2}_{N,T,O}$: various out-of-sample predictive $R^2$'s ($\%$), see (ref)-(ref); $p_{\alpha}$ and $p_{\text{lin}}$: the $p$-values of \textit{alpha} test ($\alpha(\cdot)=0$) and model specification test (joint linearity of $\alpha(\cdot)$ and $\beta(\cdot)$), respectively. \end{tablenotes} \end{threeparttable} \end{table} \begin{figure}[!htbp] \caption{$95\%$ confidence intervals for coefficients in $\alpha(\cdot)$ under linear specifications of $\alpha(\cdot)$ and $\beta(\cdot)$ with 36 characteristics} \end{figure} \begin{figure}[!htbp] \caption{Estimates of coefficients in $\beta(\cdot)$ under linear specifications of $\alpha(\cdot)$ and $\beta(\cdot)$ with 36 characteristics (blue: significant at the $5\%$ level; red: insignificant)} \end{figure} \begin{landscape} \begin{figure}[!htbp] \caption{$95\%$ confidence intervals for $\|a\|^2$, $R^{2}_{f}$, $R^{2}_{f,T,N}$, and $R^{2}_{f,N,T}$ ((ref)-(ref)) with $K=10$: subsample analysis} \end{figure} \end{landscape} \begin{landscape} {5pt} \begin{table}[!htbp] \begin{threeparttable} \caption{In-sample Sharpe ratios \tnote{\dag}} \begin{tabular}{ccccc|ccc|ccc|cccccc} \hline\hline & &\multicolumn{3}{c}{Regressed-PCA} &\multicolumn{3}{c}{Regressed-PCA S1} & \multicolumn{3}{c}{Regressed-PCA S2}& \multicolumn{5}{c}{IPCA}&\\ \cline{2-16} & $K$ & $SR_{\alpha}$ &$SR_{f}$& $SR_{M}$ & $SR_{\alpha}$ &$SR_{f}$& $SR_{M}$& $SR_{\alpha}$ &$SR_{f}$& $SR_{M}$&$SR_{\alpha}$ &$SR_{f}$& $SR_{M}$ & $SR_{M,\alpha}$& $SR_{M,f}$& \\ \cline{2-16} & 1 & 3.89 & 0.65 & 3.94 & 4.31 & 0.61 & 4.35 & 3.67 & 0.66 & 3.73 & 1.61 & 1.07 & 2.10 & 1.61 & 1.07& \\ & 2 & 4.00 & 0.67 & 4.06 & 4.76 & 0.68 & 4.81 & 4.31 & 0.79 & 4.39 & 1.98 & 1.36 & 2.22 & 1.98 & 1.36& \\ & 3 & 4.02 & 0.67 & 4.08 & 4.78 & 0.73 & 4.84 & 4.32 & 0.80 & 4.39 & 2.64 & 1.07 & 2.85 & 2.64 & 1.05& \\ & 4 & 4.07 & 0.68 & 4.13& 4.80 & 0.87 & 4.88 & 4.32 & 0.89 & 4.41 & 3.13 & 1.07 & 3.32 & 3.13 & 1.03& \\ & 5 & 4.10 & 0.69 & 4.15 & 4.75 & 1.38 & 4.95 & 4.28 & 1.24 & 4.45 & 3.00 & 1.11 & 3.21 & 3.00 & 1.07& \\ & 6 & 4.19 & 0.72 & 4.25 & 4.71 & 1.63 & 4.98 & 3.90 & 2.25 & 4.51 & 2.57 & 1.99 & 3.20 & 2.57 & 1.97& \\ & 7 & 4.37 & 0.80 & 4.44 & 4.69 & 1.69 & 4.99 & 3.47 & 3.07 & 4.63 & 2.74 & 2.09 & 3.41 & 2.74 & 2.08& \\ & 8 & 4.48 & 0.88 & 4.56& 4.16 & 2.95 & 5.10 & 3.72 & 3.74 & 5.28 & 2.50 & 2.85 & 3.69 & 2.50 & 2.84& \\ & 9 & 4.49 & 0.89 & 4.58& 4.07 & 3.11 & 5.12 & 3.74 & 3.75 & 5.29 & 2.41 & 2.77 & 3.56 & 2.41 & 2.76& \\ & 10 & 4.50 & 0.89 & 4.58& 3.78 & 3.96 & 5.48 & 3.78 & 3.76 & 5.34 & 2.40 & 2.85 & 3.62 & 2.40 & 2.84& \\ \hline\hline \end{tabular} \begin{tablenotes} • $K$: the number of factors specified; $SR_{\alpha}$: annualized Sharpe ratios of $\hat{a}^{\prime}\tilde{Y}_t$; $SR_{f}$: annualized Sharpe ratios of $\hat{\mu}^{\prime}\hat{\Sigma}\hat{f}_t$; $SR_{M}$: annualized Sharpe ratios of the combined MVE portfolios on $\hat{a}^{\prime}\tilde{Y}_t$ and $\hat{f}_t$; $SR_{M,\alpha}$: annualized Sharpe ratios of the component from $\hat{a}^{\prime}\tilde{Y}_t$ in the combined MVE portfolios; $SR_{M,f}$: annualized Sharpe ratios of the component from $\hat{f}_t$ in the combined MVE portfolios. \end{tablenotes} \end{threeparttable} \end{table} \end{landscape} \begin{landscape} {6pt} \begin{table}[!htbp] \begin{threeparttable} \caption{Out-of-sample Sharpe ratios \tnote{\dag}} \begin{tabular}{ccccccccc|ccccccccc} \hline\hline & &\multicolumn{7}{c}{Regressed-PCA} & \multicolumn{7}{c}{IPCA}&\\ \cline{2-16} & $K$ & Mean&Std&$SR_{\alpha}$ & $SR_{f,K}$ & $SR_{f}$ & $SR_{M}$ & $SR_{M,\alpha}$ & Mean&Std&$SR_{\alpha}$ & $SR_{f}$ & $SR_{M}$ & $SR_{M,\alpha}$& $SR_{M,f}$ \\ \cline{2-16} & 1 & 1.72&0.54& 3.18 & 0.61&0.61 & 3.25&3.24 & 2.70&2.07&1.31&1.26&1.80&1.33&1.26\\ & 2 & 1.74&0.52&3.36&-0.12&0.55 & 3.39 &3.38 & 2.38&1.41&1.69&1.44 &2.20&1.87&1.54 \\ & 3 & 1.77&0.50&3.56&-0.34&0.46 & 3.55 &3.56& 2.01 &0.93&2.16&1.14 &2.33&2.20&1.17 \\ & 4 & 1.77&0.47&3.74&0.02&0.44 & 3.67&3.68&1.85&0.68&2.70&0.92&2.71&2.74&0.93 \\ & 5 &1.70&0.44&3.84&0.42&0.53 &3.81 &3.78&1.75&0.62 &2.84&0.98&2.79 &2.84&0.99 \\ & 6 &1.68&0.44&3.78&0.23&0.57 & 3.80&3.76 &1.39 &0.52&2.68 &1.34&2.70& 2.66&1.31 \\ & 7 & 1.63&0.44&3.73&0.60&0.68& 3.76&3.69&1.33& 0.52&2.57& 1.42&2.61& 2.52& 1.34 \\ & 8 &1.61&0.42&3.79&0.24&0.72 & 3.80&3.74&1.22& 0.50&2.44&1.49& 2.67&2.44& 1.43\\ & 9 &1.61&0.42&3.80&-0.06&0.69 & 3.83&3.76 & 1.23& 0.49&2.51&1.53&2.73&2.52&1.47\\ & 10 &1.60&0.42&3.82&0.11&0.67& 3.87&3.79 & 1.19&0.48& 2.49&1.72&2.74 &2.45& 1.66 \\ \cline{2-16} & &\multicolumn{7}{c}{Regressed-PCA S1} & \multicolumn{7}{c}{Regressed-PCA S2}&\\ \cline{2-16} & $K$ & Mean&Std&$SR_{\alpha}$ & $SR_{f,K}$ & $SR_{f}$ & $SR_{M}$ & $SR_{M,\alpha}$ & Mean&Std&$SR_{\alpha}$ & $SR_{f,K}$ & $SR_{f}$ & $SR_{M}$ & $SR_{M,\alpha}$ \\ \cline{2-16} & 1 &2.46 & 0.69 & 3.54 &0.51& 0.51 & 3.57&3.52 & 3.29 & 0.99 & 3.33 &0.54& 0.54 & 3.27&3.24 \\ & 2 & 2.39& 0.57 & 4.22 &0.18 &0.53 & 4.16 &4.13 & 3.01 & 0.80 & 3.78&0.47& 0.70 & 3.77&3.72 \\ & 3 & 2.36 & 0.57 & 4.17 &0.45& 0.64 & 4.13 &4.08& 2.94 & 0.80 & 3.69 &0.51& 0.78 & 3.75 &3.64 \\ & 4 & 2.19 & 0.53 & 4.12& 0.85& 1.04 & 4.21 &4.08 &2.97 & 0.78 & 3.81&-0.18& 0.59 &3.84& 3.77 \\ & 5 & 2.19 & 0.51 & 4.26 &-0.03& 0.93 & 4.29 &4.21 & 2.98 & 0.76 & 3.91& -0.04& 0.56 & 3.90&3.87\\ & 6 & 1.95 & 0.49 & 3.96 &1.23& 1.62 & 4.34 &4.10 & 1.51 & 0.41 & 3.73 &2.47& 2.55& 4.08&3.74 \\ & 7 & 1.90 & 0.48 & 3.93 &0.45& 1.66 & 4.36 &4.11 & 1.01 & 0.33 & 3.09 &1.87& 3.20 & 4.02&3.26\\ & 8 &1.73 & 0.47 & 3.66& 1.11& 1.99 & 4.37 &3.92 & 0.73 & 0.22 & 3.36 &1.25& 3.29 & 4.41&3.60 \\ & 9 & 1.31 & 0.40 & 3.26 &1.81& 2.80& 4.30 &3.49 & 0.71 & 0.20 & 3.62 &0.28& 3.24 & 4.54&3.74\\ & 10 & 0.88 & 0.28 & 3.14& 1.72& 3.33 & 4.47 &3.50 & 0.69 & 0.18 & 3.90 &0.23& 3.19 & 4.64&3.90 \\ \hline\hline \end{tabular} \begin{tablenotes} • $K$: the number of factors specified; Mean: annualized means of the pure-\textit{alpha} portfolios $R_{\alpha,t}$ in (ref) ($\%$); Std: annualized standard deviations of $R_{\alpha,t}$ ($\%$); $SR_{\alpha}$: annualized Sharpe ratios of $R_{\alpha,t}$; $SR_{f,K}$: annualized Sharpe ratios of the $K$th component in $\hat{f}_{t-1,t}$; $SR_{f}$: annualized Sharpe ratios of the MVE factor portfolios $R_{\beta,t}$ in (ref); $SR_{M}$: annualized Sharpe ratios of the combined MVE portfolios on $R_{\alpha,t}$ and $\hat{f}_{t-1,t}$; $SR_{M,\alpha}$: annualized Sharpe ratios of the component from $R_{\alpha,t}$ in the combined MVE portfolios; $SR_{M,f}$: annualized Sharpe ratios of the component from $\hat{f}_{t-1,t}$ in the combined MVE portfolios. \end{tablenotes} \end{threeparttable} \end{table} \end{landscape} \begin{landscape} \begin{figure}[!htbp] \caption{Annualized realized excess returns and Sharpe ratios of the pure-\textit{alpha} portfolio with $K=10$: subsample analysis} \end{figure} \end{landscape} \begin{landscape} {4.2pt} \begin{table}[!htbp] \begin{threeparttable} \caption{Comparing asset pricing tests: $K=5$\tnote{\dag}} \begin{tabular}{clcccccccccccccc} \hline\hline &Testing portfolios/Factors & $A|a|$ & $A|t(a)|$ & $Aa^2/V\overline{r}$ & $A\lambda^2/V\overline{r}$ & $As(a)$ & $As(e)$ & $Sh^2(a)$ & $Sh^2(f)$ & $AR^2$ & $GRS$ & $p(GRS)$&\\ \cline{2-13} &\multicolumn{12}{l}{Group I: Regressed-PCA's 36 managed portfolios}\\ \cline{2-13} & Regressed-PCA & 0.40 & 3.40 & 0.64 & 0.61 & 0.13 & 2.93 & 2.57 & 0.04 & 28.93 & 34.89 & 0.00 \\ & Regressed-PCA S1 & 0.40 & 3.10 & 0.56 & 0.50 & 0.17 & 3.64 & 2.48 & 0.16 & 17.97 & 30.23 & 0.00 \\ & Regressed-PCA S2 & 0.43 & 3.28 & 0.59 & 0.53 & 0.17 & 3.73 & 2.71 & 0.13 & 16.08 & 33.86 & 0.00 \\ & IPCA & 0.40 & 3.15 & 0.64 & 0.58 & 0.17 & 3.73 & 3.59 & 0.10 & 17.76 & 45.92 & 0.00 \\ & IPCA$\setminus$Regressed-PCA & 0.50 & 3.34 & 1.03 & 0.96 & 0.17 & 3.94 & 3.49 & 0.06 & 12.76 & 46.40 & 0.00 \\ & FF5 & 0.49 & 3.11 & 0.99 & 0.92 & 0.18 & 4.03 & 2.62 & 0.12 & 8.25 & 33.16 & 0.00 \\ & KNS & 0.50 & 3.34 & 1.07 & 1.01 & 0.17 & 3.99 & 2.56 & 0.04 & 9.92 & 34.89 & 0.00 \\ \cline{2-13} &\multicolumn{12}{l}{Group II: 100 sorted portfolios (double sort on Size and BM, OP, INV, and MOM)}\\ \cline{2-13} & Regressed-PCA & 0.85 & 5.04 & 14.14 & 13.60 & 0.17 & 3.97 & 1.00 & 0.04 & 50.81 & 4.26 & 0.00 \\ & Regressed-PCA S1 & 0.57 & 3.21 & 6.71 & 6.13 & 0.18 & 3.89 & 1.27 & 0.16 & 52.33 & 4.85 & 0.00 \\ & Regressed-PCA S2 & 0.44 & 2.42 & 4.18 & 3.59 & 0.18 & 4.00 & 1.35 & 0.13 & 49.67 & 5.32 & 0.00 \\ & IPCA & 0.90 & 11.07 & 16.08 & 15.93 & 0.09 & 1.96 & 4.56 & 0.10 & 86.75 & 18.38 & 0.00 \\ & IPCA$\setminus$Regressed-PCA & 1.09 & 5.62 & 23.32 & 22.56 & 0.20 & 4.61 & 1.87 & 0.06 & 36.19 & 7.81 & 0.00 \\ & FF5 & 0.41 & 4.81 & 3.55 & 3.39 & 0.09 & 2.05 & 1.76 & 0.12 & 86.91 & 6.98 & 0.00 \\ & KNS & 0.96 & 6.25 & 17.25 & 16.82 & 0.16 & 3.55 & 0.92 & 0.04 & 59.83 & 3.96 & 0.00 \\ \cline{2-13} &\multicolumn{12}{l}{Group III: 110 sorted portfolios (double sort on Size and Beta, Accruals, NI, and Variance)}\\ \cline{2-13} & Regressed-PCA & 0.85 & 5.08 & 16.06 & 15.43 & 0.17 & 3.98 & 1.25 & 0.04 & 50.36 & 4.73 & 0.00 \\ & Regressed-PCA S1 & 0.55 & 3.09 & 7.42 & 6.74 & 0.18 & 3.94 & 1.55 & 0.16 & 50.83 & 5.29 & 0.00 \\ & Regressed-PCA S2 & 0.42 & 2.32 & 4.62 & 3.92 & 0.18 & 4.03 & 1.61 & 0.13 & 48.28 & 5.64 & 0.00 \\ & IPCA & 0.88 & 10.63 & 17.68 & 17.50 & 0.09 & 2.00 & 4.56 & 0.10 & 85.72 & 16.31 & 0.00 \\ & IPCA$\setminus$Regressed-PCA & 1.09 & 5.70 & 26.13 & 25.21 & 0.21 & 4.66 & 2.25 & 0.06 & 35.61 & 8.36 & 0.00 \\ & FF5 & 0.43 & 4.88 & 4.22 & 4.03 & 0.09 & 2.08 & 2.09 & 0.12 & 86.06 & 7.37 & 0.00 \\ & KNS & 0.95 & 6.15 & 19.41 & 18.90 & 0.16 & 3.59 & 1.11 & 0.04 & 58.32 & 4.22 & 0.00 \\ \hline\hline \end{tabular} \begin{tablenotes} • $A|a|$: average absolute intercept; $A|t(a)|$: average absolute $t$-statistic for the intercepts; $Aa^2/V\overline{r}$: average squared intercept over the cross-section variance of $\overline{r}$, average returns of the testing portfolios; $A\lambda^2/V\overline{r}$: average difference between each squared intercept and its squared standard error divided by the variance of $\overline{r}$; $As(a)$: average standard error of the intercepts; $As(e)$: average residual standard deviation; $Sh^2(a)$: maximized squared Sharpe ratio for the intercepts; $Sh^2(f)$: maximized squared Sharpe ratio for the factors; $AR^2$: average regression $R^2$ (%); $GRS$: $GRS$ statistic of Gibbonsetal_efficiency_1989; $p(GRS)$: $p$-value of $GRS$. \end{tablenotes} \end{threeparttable} \end{table} \end{landscape} \begin{landscape} {4.2pt} \begin{table}[!htbp] \begin{threeparttable} \caption{Comparing asset pricing tests: $K=5$ (continued)\tnote{\dag}} \begin{tabular}{clcccccccccccccc} \hline\hline &Testing portfolios/Factors & $A|a|$ & $A|t(a)|$ & $Aa^2/V\overline{r}$ & $A\lambda^2/V\overline{r}$ & $As(a)$ & $As(e)$ & $Sh^2(a)$ & $Sh^2(f)$ & $AR^2$ & $GRS$ & $p(GRS)$&\\ \cline{2-13} &\multicolumn{12}{l}{Group IV: IPCA's 36 managed portfolios}\\ \cline{2-13} & Regressed-PCA & 0.05 & 3.44 & 0.87 & 0.82 & 0.01 & 0.33 & 1.50 & 0.04 & 25.22 & 20.30 & 0.00 \\ & Regressed-PCA S1 & 0.06 & 5.12 & 1.41 & 1.38 & 0.01 & 0.27 & 1.85 & 0.16 & 47.24 & 22.50 & 0.00 \\ & Regressed-PCA S2 & 0.06 & 5.30 & 1.32 & 1.29 & 0.01 & 0.25 & 1.83 & 0.13 & 51.44 & 22.90 & 0.00 \\ & IPCA & 0.07 & 6.72 & 1.53 & 1.51 & 0.01 & 0.21 & 3.06 & 0.10 & 64.08 & 39.14 & 0.00 \\ & IPCA$\setminus$Regressed-PCA & 0.06 & 4.53 & 1.22 & 1.17 & 0.01 & 0.31 & 2.20 & 0.06 & 39.29 & 29.28 & 0.00 \\ & FF5 & 0.04 & 3.32 & 0.71 & 0.67 & 0.01 & 0.27 & 1.42 & 0.12 & 50.00 & 17.96 & 0.00 \\ & KNS & 0.04 & 3.87 & 0.76 & 0.73 & 0.01 & 0.26 & 1.47 & 0.04 & 52.97 & 20.03 & 0.00 \\ \cline{2-13} &\multicolumn{12}{l}{Group V: P1&10 of sorted portfolios (single sort on 55 characteristics in Kozaketal_Interpreting_2018, 110 portfolios)}\\ \cline{2-13} & Regressed-PCA & 0.73 & 4.05 & 7.24 & 6.82 & 0.19 & 4.27 & 1.36 & 0.04 & 43.38 & 5.16 & 0.00 \\ & Regressed-PCA S1 & 0.47 & 2.61 & 3.78 & 3.37 & 0.19 & 4.01 & 1.42 & 0.16 & 49.01 & 4.82 & 0.00 \\ & Regressed-PCA S2 & 0.37 & 1.98 & 2.51 & 2.10 & 0.19 & 4.13 & 1.44 & 0.13 & 45.95 & 5.02 & 0.00 \\ & IPCA & 0.73 & 6.43 & 7.71 & 7.53 & 0.12 & 2.66 & 3.15 & 0.10 & 77.14 & 11.26 & 0.00 \\ & IPCA$\setminus$Regressed-PCA & 0.90 & 4.68 & 10.96 & 10.45 & 0.21 & 4.67 & 1.85 & 0.06 & 34.10 & 6.87 & 0.00 \\ & FF5 & 0.40 & 4.50 & 2.40 & 2.25 & 0.11 & 2.34 & 2.54 & 0.12 & 82.37 & 8.97 & 0.00 \\ & KNS & 0.92 & 5.77 & 10.54 & 10.24 & 0.16 & 3.66 & 1.31 & 0.04 & 56.83 & 5.01 & 0.00 \\ \cline{2-13} &\multicolumn{12}{l}{Group VI: P1&10 of sorted portfolios (single sort on 36 characteristics, 72 portfolios)}\\ \cline{2-13} & Regressed-PCA & 0.69 & 3.55 & 5.72 & 5.28 & 0.21 & 4.71 & 0.81 & 0.04 & 41.77 & 5.13 & 0.00 \\ & Regressed-PCA S1 & 0.49 & 2.52 & 3.25 & 2.83 & 0.20 & 4.36 & 0.96 & 0.16 & 48.78 & 5.40 & 0.00 \\ & Regressed-PCA S2 & 0.39 & 1.89 & 2.35 & 1.91 & 0.21 & 4.50 & 1.02 & 0.13 & 45.53 & 5.94 & 0.00 \\ & IPCA & 0.70 & 5.43 & 6.11 & 5.90 & 0.14 & 3.07 & 1.84 & 0.10 & 74.60 & 10.92 & 0.00 \\ & IPCA$\setminus$Regressed-PCA & 0.88 & 4.20 & 9.17 & 8.62 & 0.23 & 5.13 & 1.34 & 0.06 & 33.26 & 8.29 & 0.00 \\ & FF5 & 0.42 & 4.38 & 2.48 & 2.29 & 0.12 & 2.73 & 2.48 & 0.12 & 80.46 & 14.53 & 0.00 \\ & KNS & 0.91 & 5.61 & 9.35 & 9.07 & 0.17 & 3.79 & 0.88 & 0.04 & 59.94 & 5.54 & 0.00 \\ \hline\hline \end{tabular} \begin{tablenotes} • $A|a|$: average absolute intercept; $A|t(a)|$: average absolute $t$-statistic for the intercepts; $Aa^2/V\overline{r}$: average squared intercept over the cross-section variance of $\overline{r}$, average returns of the testing portfolios; $A\lambda^2/V\overline{r}$: average difference between each squared intercept and its squared standard error divided by the variance of $\overline{r}$; $As(a)$: average standard error of the intercepts; $As(e)$: average residual standard deviation; $Sh^2(a)$: maximized squared Sharpe ratio for the intercepts; $Sh^2(f)$: maximized squared Sharpe ratio for the factors; $AR^2$: average regression $R^2$ (%); $GRS$: $GRS$ statistic of Gibbonsetal_efficiency_1989; $p(GRS)$: $p$-value of $GRS$. \end{tablenotes} \end{threeparttable} \end{table} \end{landscape} \addcontentsline{toc}{section}{References} \putbib
bibunit