Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
162,166 characters · 29 sections · 101 citation commands
Identifying and exploiting alpha in linear asset pricing models with strong, semi-strong, and latent factors
JEL Classifications: { C38, G10}
Key Words: F{ actor strength, pricing errors, risk premia, missing factors, pooled Lasso, mean-variance and phi-portfolios.} \thispagestyle{empty}
\pagenumbering{arabic} \onehalfspacing {0pt} {0pt} {0pt} {0pt}
There are two approaches to the estimation of risk premia and testing of market efficiency, often referred as the beta and the SDF (stochastic discount factor) methods, see Jagannathan2002. This paper adopts the beta method, and following the literature uses a linear factor pricing model (LFPM) to explain the time series of excess returns on individual securities, $r_{it}=R_{it}-r_{t}^{f}$, where $R_{it}$ is the return and $r_{t}^{f}$ the risk free rate for $i=1,2,...,n;$ $t=1,2,....,T,$ by a set of observed tradable risk factors. We use individual securities rather than portfolios since, as we will show, if risk factors are not strong, large $n$ is required for accurate estimation. Ang2020 discuss the general issues in the choice between portfolios and individual stocks. Pesaran2023 discuss both the use of portfolios and the relationship between the SDF and LFPM approaches.
The LFPM\ explains the excess return on each security, $r_{it},$ by an intercept, $\mathit{\alpha}_{i}$, labelled alpha, and a $K\times1$ vector of traded risk factors, $\mathbf{f}_{t}=(f_{1t},f_{2t},...,f_{Kt})^{\prime}$ with loadings, $\boldsymbol{\beta}_{i}$:
Under the Arbitrage Pricing Theory (APT) due to ROSS1976341 , we have
where $\eta_{i}$ is a firm-specific idiosyncratic pricing error. Ross allowed $\eta_{i}$ to be non-zero for some $i$ but required that these errors are bounded in the sense that $\sum_{i=1}^{n}\eta_{i}^{2}<C<\infty$. We relax this condition and impose weaker restrictions on $\eta_{i}$'s as discussed in sub-section (ref). But the paper's main contribution lies in the fact that even if there are no idiosyncratic pricing errors, it is still possible to have non-zero $\mathit{\alpha}_{i}$. To see this, taking unconditional expectations of ((ref)) and using ((ref)) we note that \[ E(r_{it})=\mathit{\alpha}_{i}+\boldsymbol{\beta}_{i}^{\prime}\mathbf{\mu }=c+\boldsymbol{\beta}_{i}^{\prime}\boldsymbol{\lambda}+\eta_{i}. \] where $E(\mathbf{f}_{t})=\boldsymbol{\mu}$, which in turn yields
We denote $\boldsymbol{\lambda}-\boldsymbol{\mu}$ by $\boldsymbol{\phi}$, and consider the implications of a non-zero $\boldsymbol{\phi}$ for portfolio optimization. We show that when $\boldsymbol{\phi\neq0}$, one can construct what we call "phi-portfolios" that exploit the systematic components of $\mathit{\alpha}_{i}$, as captured by the non-zero elements of $\boldsymbol{\phi}\ $\ in ((ref)).\footnote{We are grateful to one of the anonymous reviewers for suggesting the idea of phi-portfolios, as a way of motivating and interpreting the role of $\boldsymbol{\phi}$ in asset pricing models.} In contrast, the idiosyncratic pricing errors, $\eta_{i}$, cannot be exploited, as it is likely to involve insider trading. But our proposed phi-portfolios can be implemented and applies irrespective of whether $\eta_{i}=0$ or not.
The extent to which alpha can be exploited is discussed in sub-section ((ref)) where it is shown that $\sum_{i=1}^{n}\left( \mathit{\alpha }_{i}-\mathit{\bar{\alpha}}\right) ^{2}$ depends on the magnitude of $\boldsymbol{\phi}^{\prime}\boldsymbol{\phi}$ and the strength of the risk factors. Since this alpha can be exploited by the construction of the phi-portfolios, this is not consistent with a no arbitrage condition. While it is true that $E(r_{it})=\boldsymbol{\beta}_{i}^{\prime}\boldsymbol{\lambda}$, estimating $\boldsymbol{\lambda}$\ does not tell us whether $\boldsymbol{\phi }\neq0,$\ which can be exploited to obtain Sharpe ratios that are larger than the Sharpe of the tangency portfolio. The claim that tangency portfolio has the highest Sharpe ratio rests on the implicit assumption that $\boldsymbol{\phi=0}$, even if there are no idiosyncratic pricing errors ($\eta_{i}=0$ for all $i$).
The focus of much of the literature has been on testing $\mathit{\alpha} _{i}=0$, or on estimating the risk premium $\boldsymbol{\lambda}$. But given that estimating a non-zero $\boldsymbol{\phi}$ \ provides a way to identify and exploit the alpha in a linear factor pricing model for large $n,$ $\boldsymbol{\phi}$ is an interesting parameter in itself, which will be the focus of this paper. Furthermore, the tangency condition used to construct mean-variance (MV) portfolios has the implicit assumption that $\mathit{\alpha }_{i}=0$, and the MV portfolio implied by the tangency condition is not efficient unless $\boldsymbol{\phi=0}$. If $\boldsymbol{\phi\neq0}$ there is no guarantee that the tangency portfolio will achieve a maximal Sharpe ratio, and is likely to be dominated by the phi-portfolio. Specifically, abstracting from model and parameter uncertainties, we show that when $\boldsymbol{\phi \neq0}$ the Sharpe of the tangency portfolio is bounded in the number of securities, whilst the Sharpe of the phi portfolio rises in $n$ and there exists a sufficiently large $n$ that ensures that the Sharpe of the phi-portfolio will exceed that of the tangency portfolio.
More specifically, first we introduce our proposed phi-portfolio in terms of the factor loadings $\mathbf{B}_{n}=\left( \boldsymbol{\beta}_{1} ,\boldsymbol{\beta}_{2},...,\boldsymbol{\beta}_{n}\right) ^{\prime}$ and $\boldsymbol{\phi}$, and compare its limiting properties (as $n\rightarrow \infty$) with the standard MV portfolio. We show that for known factor loadings\ and non-zero $\boldsymbol{\phi}$ there exist phi-portfolios with strictly positive returns, given by $\boldsymbol{\phi}^{\prime} \boldsymbol{\phi}>0$, that are fully diversified (their variance tends to zero with $n$). The rate at which the variance of phi-portfolio returns tends to zero will depend on the strength of the traded risk factors compared to the strength of the missing (latent) factors, highlighting the importance of factor strengths in portfolio analysis. Since the phi-portfolios have Sharpe ratios that tend to infinity with $n$, they dominate MV portfolios whose Sharpe ratio is bounded in $n$. Note that unlike MV\ portfolios, the phi-portfolios do not require knowledge of the inverse of the covariance matrix of returns, which is particularly difficult to estimate accurately when $n$ is large. As is usual, these theoretical results apply to population values where there are no restrictions on trading many individual securities.
Next we consider estimation of and inference about $\boldsymbol{\phi}$ using a large number of individual securities, taking account of firm-specific pricing errors, traded factors that are not strong, missing latent factors, and panel data sets where the time dimension, $T$, is small relative to $n$. Factor strength plays a central role in our analysis of $\boldsymbol{\phi}$. We use a measure of factor strength, $\alpha_{k},$ developed in bailey2016exponent, bailey2021measurement , which is defined in terms of the proportion of non-zero factor loadings, $\beta_{ik}$.\footnote{Alpha is used both for the LFPM intercepts and the measure of factor strength, because these are established usages, but it will be clear from context and subscripts which is being referred to.} A factor is strong if this proportion is very close to unity, it is semi-strong if $1>\alpha_{k}>1/2$, and it is weak if $\alpha_{k}<1/2$. Use of this measure allows us to be precise about the degree of pervasiveness and show how the strengths of the observed factors, the missing factors and the pricing errors, each influence estimation and inference about $\boldsymbol{\phi}$.
In practice, exploitation of a non-zero $\boldsymbol{\phi}$ requires $n$ to be large and rebalancing such long-short portfolios for so many securities may incur high transactions costs or not be feasible. In addition, model uncertainty, estimation uncertainty, time variation in both $\boldsymbol{\beta }_{i}$ and in conditional volatility pose additional difficulties in implementing a strategy to exploit the potential returns revealed by $\boldsymbol{\phi}$. In developing the theory we will abstract from such practical difficulties, but in the empirical section we illustrate some of these issues with a comparison of the performance of $\boldsymbol{\phi}$ based portfolios relative to MV portfolios which would face similar difficulties.
We estimate $\boldsymbol{\phi}=(\phi_{1},\phi_{2},...,\phi_{K})^{\prime}$, using a two-step estimator. In the first step we estimate the intercepts, $\hat{\alpha}_{i}$, and the factor loadings, $\boldsymbol{\hat{\beta}}_{i}$, from least squares regressions of excess returns on an intercept and risk factors. In the second step, $\boldsymbol{\phi}$ is estimated from the cross section regression of $\hat{\alpha}_{i}$ on $\boldsymbol{\hat{\beta}}_{i}$. As with the two-step estimator of $\boldsymbol{\lambda}$, such a two-step estimator of $\phi_{k}$ will also be biased, and requires bias-correction. Following shanken1992estimation , we consider a bias-corrected version of the two-step estimator of $\boldsymbol{\phi}$, which we denote by $\boldsymbol{\tilde{\phi}}_{nT}$. We develop the asymptotic distribution of $\boldsymbol{\tilde{\phi}}_{nT}$ under quite general set of assumptions regarding the idiosyncratic pricing errors, error cross-sectional dependence, and the presence of missing (latent) factors. The paper also investigates the implications of factor strengths for the precision with which $\boldsymbol{\phi}$ can be estimated. The LFPM, following cham1983arbi , assumes that all the observed factors are strong and the eigenvalues of the covariance matrix of the errors are bounded.
In developing the arbitrage pricing theory, APT, ROSS1976341 , whose concerns were primarily theoretical, assumed the factors had mean zero: $\boldsymbol{\mu}=0,$ so $\boldsymbol{\phi}=\boldsymbol{\lambda}$, is the risk premium. For traded factors under market efficiency, where $\boldsymbol{\phi}=\mathbf{0}$, the risk premium is the factor mean $\boldsymbol{\mu}=\boldsymbol{\lambda}$. Were one interested in estimating $\boldsymbol{\lambda}$ there may be statistical reasons to estimate $\phi_{k}$ and $\mu_{k}$ separately, then summing them to obtain an estimate of $\lambda_{k}$. The factor mean, $\mu_{k}=E(f_{kt})$, can be estimated consistently at the regular $\sqrt{T}$ rate directly using time series data on the risk factors, $f_{kt}$, for $t=1,2,...,T,$ and does not require knowledge of the factor loadings or $n$. In contrast, estimation of $\phi_{k}$ requires panel data to estimate the factor loadings and hence both $n$ and $T$ dimensions are important. In some cases it may be beneficial to use different time series dimensions, $T_{\mu}$ and $T_{\phi},$ to estimate $\mu_{k}$ and $\phi_{k}$, respectively. For example, when factor loadings are subject to breaks it is advisable to use a relatively short sample, and when some factors are not sufficiently strong a large $n$ is required. Thus one could use large $T$ to estimate $\mu_{k}$ and small $T$ to estimate $\phi_{k}$ and $\lambda_{k}$ can be simply estimated by adding the estimates of $\mu_{k}$ and $\phi_{k}.$ We do not pursue the use of different $T$, but if the same $T$ is used to estimate the mean of factor $k,$ say $\hat{\mu}_{k,T},$ and the bias-corrected $\tilde{\phi}_{k,nT}$ then their sum is the same as the shanken1992estimation bias-corrected risk premium for factor $k,$ $\tilde{\lambda}_{k,nT}.$ This decomposition can be used to obtain the rate at which $\tilde{\lambda}_{k,nT}$ converges to its true value, $\lambda_{0,k}=\phi_{0,k}+\mu_{0,k}$.
The main theoretical results of the paper are set out around five theorems under a number of key assumptions, with proofs provided in the mathematical appendix. The small sample properties of $\boldsymbol{\tilde{\phi}}_{nT}$ are investigated using extensive Monte Carlo experiments, allowing for a mixture of strong and semi-strong observed factors, latent factors, pricing errors, GARCH effects and non-Gaussian errors. Small sample results are in line with our theoretical findings, and confirm that the bias-corrected estimator, $\boldsymbol{\tilde{\phi}}_{nT}$, has the correct size and good power properties for samples with time series dimensions of $T=120$ and $T=240$. They also show, in accord with the theory, that the precision with which $\phi_{k}$ is estimated falls with $\alpha_{k}$, the strength of the $k^{th}$ factor.
Our theoretical derivations and Monte Carlo simulations assume that the list of relevant observed factors is known. But in practice the relevant factors must be selected. Extending the theory to the high dimensional case where factors are selected, rather than given, is beyond the scope of the present paper. Since the rate of convergence of $\boldsymbol{\tilde{\phi}}_{nT}$ to its true value is given by $\sqrt{T}n^{(\alpha_{k}+\alpha_{\min}-1)/2}$, and the Monte Carlo confirms the crucial role of factor strengths in estimation and inference on $\boldsymbol{\phi,}$ it seems sensible to select factors on the basis of their factor strength. Weak factors whose strength is around $1/2$ can be ignored and absorbed into the error term.
The above selection procedure is applied in a high dimensional setting with both a large number of securities ($n$ from $1,090$ to $1,175)$ and a large number of potential risk factors ($m$ from $177$ to $189)$, taken from the chen2022open , which can all be traded. We used monthly data over the period $1996m1-2022m12$ and considered balanced panels obtained by including all existing stocks in a given month for which there are $T$ observations. We considered $T=120$ and $240$ months, and focus on the latter which we found to be more reliable, given the large number of securities being considered. Various procedures could be used to select risk factors for a given security, $i$. We used Lasso for this purpose and then selected a subset of these factors that were chosen by a sufficiently large number of securities in the sample, and whose estimated strength were above the given threshold value of $\alpha_{k}>0.75$. We refer to this selection procedure as pooled Lasso. Using this procedure with $T=240$ we ended up with $7$ risk factors for the sample ending in $2015$, declining to $4$ in $2021$. Interestingly, the three Fama-French factors were always included in the set of factors selected by pooled Lasso.
Accordingly, we considered three linear asset pricing models for our portfolio analysis: the pooled Lasso selected at the start of our evaluation sample, denoted by PL7, the Fama-French three factors model, FF3, and the Fama-French five factors model, FF5, which includes two factors not in PL7. The three models are estimated using rolling samples of size $T=240$, starting with a sample ending in $2015m12$ and finishing with a sample ending in $2022m11$. The hypothesis that $\boldsymbol{\phi}=\mathbf{0}$ was rejected for all $84$ rolling samples and all three models, albeit less strongly over the post Covid-19 period. The test results suggested possible unexploited return opportunities, and to investigate this possibility further, we constructed phi-portfolios and compared their Sharpe ratios with the ones based on standard MV portfolios over the full sample evaluation sample, $2016m1-2022m12$, and sample ending $2019m12$, that excludes the Covid-19 period. We find that in five out of the six cases (3 models 2 samples) the phi-portfolio has a higher SR than the corresponding MV portfolio. The exception is the FF5 pre Covid-19. This illustrates that if $\boldsymbol{\phi }\neq0,$ it is possible to construct a portfolio that outperforms the mean variance portfolio. In both the pre Covid-19 sample and the full sample the highest SR was obtained by the PL7 phi-portfolios, which also outperformed the S&P500. The SRs for the sample ending in $2022$ were substantially lower than the sample ending in $2019$, consistent with a falling value of the probability that $\boldsymbol{\phi}$ was non-zero.
Related literature: On estimation of risk premia, following Fama and MacBeth and shanken1992estimation , estimation of risk premia is further examined by shanken2007estimating , kan2013pricing , and BAI201531 . The survey paper by jagannathan2010analysis provide further references.
Testing for market efficiency dates back to jensen1968performance who proposes testing a$_{i}=0$ for each $i$ separately. gibbons1989test provide a joint test for the case where the errors are Gaussian and $n<T$. gagliardini2016time develop two-pass regressions of individual stock returns, allowing time-varying risk premia, and propose a standardised Wald test. raponi2019testing propose a test of pricing error in cross section regression for fixed number of time series observations. They use a bias-corrected estimator of shanken1992estimation to standardise their test statistic. ma2020testing employ polynomial spline techniques to allow for time variations in factor loadings when testing for alphas. feng2022high propose a max-of-square type test of the intercepts instead of the average used in the literature, and recommend using a combination of the two testing procedures. he2022testing propose two statistics, a Wald type statistic which require $n$ and $T$ to be of the same order of magnitude and a standardised t-ratio. kleibergen2009tests considers testing in the case where the loadings are small. pesaran2012testing, pesaran2023testing consider testing that the intercepts in the LFPM are zero when $n$ is large relative to $T$ and there may be non Gaussian errors and weakly cross-correlated errors.
A large number of risk factors have been considered in the empirical literature. We use the fama1993common three factors in our Monte Carlo design. In our empirical application we use the five factors proposed by fama2015five and the large set of factors provided by chen2022open . harvey2019census document over 400 factors published in top finance journals. dello2022missing argue that despite the hundreds of systematic risk factors considered in the literature, there is still a sizable pricing error and that this can be explained by asset specific risk that reflects market frictions and behavioral biases. There is a large Bayesian literature, including chib2020comparing , and hwang2022bayesian on selecting factor models. The issue of factor selection is also addressed by fama2018choosing .
Strong and weak factors in asset returns are considered by ANATOLYEV2022103 , connor2022semi , and giglio2023test. beaulieu2020arbitrage discuss the lack of identification of risk premia when many of the loadings are zero. There has also been concern about the consequences of omitted factors. giglio2021asset discuss the problem and try to deal with it using a three-pass method which is valid even when not all factors in the model are specified or observed using principal components of the test assets. onatski2012asymptotics and lettau2020estimating, lettau2020factors provide extensive discussions of weak factor and latent factors, respectively. More recent contributions include BaiNg2023Weak and uematsu2023inference.
There is also a large literature on portfolio construction. Herskovic2019 discuss low cost methods of hedging risk factors. Preite2024 derive an SDF in which there is compensation for unsystematic risk within the framework of the APT. Korsaye2021 introduce model-free smart SDFs which give rise to non-parametric SDF bounds for testing asset pricing models.\ \ Daniel2020 discuss the common practice of creating characteristic portfolios by sorting on characteristics associated with average returns and show that these portfolios capture not only the priced risk associated with the characteristic but also unpriced risk. Quaini2023 propose an estimator of tradable factor risk premia.
Paper's outline: The rest of the paper is organized as follows: Section (ref) provides the framework for estimation of $\boldsymbol{\phi}.$ Section (ref) sets out the assumptions and states the main theorems. Theorem (ref) shows that the standard Fama-MacBeth estimator is valid only when there are no pricing errors and $n/T\rightarrow0$. Theorem (ref) shows that the Shanken bias-corrected estimator of $\lambda_{k}$ continues to be consistent for a fixed $T$ as $n\rightarrow\infty$, even in presence of weak pricing errors and weak missing common factors. Theorem (ref) provides conditions under which the bias-corrected estimator, $\boldsymbol{\tilde{\phi}}_{nT}$, is consistent for $\boldsymbol{\phi}_{0}$, and derives the asymptotic distribution of $\boldsymbol{\tilde{\phi}}_{nT},$ assuming the observed factors are strong $(\alpha_{k}=1$ for all $k$). Theorem (ref) extends the results to cases where one or more risk factors are semi-strong, and establishes the rate at which $\boldsymbol{\tilde{\phi}}_{nT}$ converges to its true value, assuming the idiosyncratic pricing errors are sufficiently weak, as discussed after the theorem. For example, for a factor with strength $\alpha_{k}$ we show that $\tilde{\phi}_{k,nT}-\phi_{0,k}=O_{p}\left( T^{-1/2}n^{-(\alpha_{k} +\alpha_{\min}-1)/2}\right) $, and as a consequence \[ \tilde{\lambda}_{k,nT}-\lambda_{0,k}=O_{p}\left( T^{-1/2}n^{-(\alpha _{k}+\alpha_{\min}-1)/2}\right) +O_{p}\left( T^{-1/2}\right) , \] where $\lambda_{0,k}$ is the true value of the risk premia associated to factor $f_{kt}$, and $\alpha_{\min}$ is the strength of the least strong factor included. This consistency condition is weaker than the one derived by giglio2023test who also assume $\eta_{i}=0$, for all $i$. Finally, Theorem (ref) gives conditions for consistent estimation of the asymptotic variance of $\boldsymbol{\tilde{\phi}}_{nT},$ using a suitable threshold estimator of the covariance matrix. Section (ref) presents the Monte Carlo (MC) design, its calibration and a summary of the main findings. Section (ref)\ discusses the problem of factor selection from a large number of potential factors. Section (ref) gives the empirical application using monthly data on a large number of individual US stocks and risk factors over the period 1996-2021. It selects factors, estimates $\boldsymbol{\phi}\mathbf{,}$ and compares the performance of MV and phi-portfolios. Section (ref) provides some concluding remarks.
Detailed mathematical proofs are provided in a mathematical appendix. Further information on data sources, MC calibration plus some supplementary material for the empirical application are provided in the online supplement A. To save space all MC results are given in the online supplement B.
Let $R_{it}$ denote the holding period return on traded security $i$, which can be bought long or short without transaction costs, $r_{t}^{f}$ \ is the risk free rate, and $r_{it}=R_{it}-r_{t}^{f}$ is the excess return. We start with the linear factor pricing model (LFPM)
for $i=1,2,...,n$ and $t=1,2,...,T$, where $r_{it}$ is explained in terms of the $K\times1$ vector of factors $\mathbf{f}_{t}=(f_{1t},f_{2t},...,f_{Kt} )^{\prime}$. The intercept $\mathit{\alpha}_{i}$ and the $K\times1$ vector of factor loadings, $\boldsymbol{\beta}_{i}=(\beta_{i1},\beta_{i2},...,\beta _{iK})^{\prime}$, are unknown. The idiosyncratic errors, $u_{it}$ have zero means and are assumed to be serially uncorrelated. The factors, $\mathbf{f} _{t}$, are assumed to be covariance stationary with the constant mean $\boldsymbol{\mu}=E\left( \mathbf{f}_{t}\right) $, and $\boldsymbol{\Sigma }_{f}=E\left[ \left( \mathbf{f}_{t}-\boldsymbol{\mu}\right) \left( \mathbf{f}_{t}-\boldsymbol{\mu}\right) ^{\prime}\right] $.
Under the Arbitrage Pricing Theory (APT) due to Ross (1976) the pricing errors, $\eta_{i}$ for $i=1,2,...,n$ defined by
are bounded such that
where $c$ is zero-beta expected excess return, $\boldsymbol{\lambda}$ is the $K\times1$ vector of risk premia. Under APT, we have
and reduces to the standard beta representation of an unconditional asset pricing model, if $c+\eta_{i}=0$. In this case, $\boldsymbol{\beta} _{i}=-Cov(r_{it},m_{t})/var(m_{t})$, where $m_{t}$ is the stochastic discount factor (SDF), which satisfies the moment condition $E_{t}(r_{i,t+1}m_{t+1} )=0$. See Pesaran2023 for a discussion of the link. For empirical assessment, we consider the more general unconditional pricing model ((ref)), and relate to the linear factor pricing model. To this end, taking unconditional expectations of ((ref)) and using the APT\ condition we have
which in turn yields
where
Under the APT condition, ((ref)), $\boldsymbol{\phi}$ can be identified from the cross section regression of $\alpha_{i}$ on $\boldsymbol{\beta}_{i}$. We relax the APT\ condition and derive restrictions on the degree to which idiosyncratic pricing errors, $\eta_{i}$, can be pervasive to achieve identification.
The focus of the literature has been on testing for alpha, $\mathit{\alpha }_{i}=0$, and the estimation of the risk premia, $\boldsymbol{\lambda}$, using panel data on excess returns, $\left\{ r_{it},1,2,...,n\text{; } t=1,2,...,T\right\} ,$ and $\mathbf{F}$, the $T\times K$ matrix of observations on the factors. It is clear that $\boldsymbol{\phi}$ plays an important role in tests for alpha in LFPM, and a non-zero $\boldsymbol{\phi}$ implies non-zero alphas, which in turn implies exploitable excess profitable opportunities. More specifically, we show that for known values of $\boldsymbol{\beta}_{i},$ $i=1,2,...,n\,$\ and when $\boldsymbol{\phi}$ is non-zero, there exists phi-based portfolios with non-zero means that are fully diversified (their variance tends to zero with $n$), namely have Sharpe ratios that tend to infinity with $n$, and hence dominate the MV portfolios. For the MV portfolios to be efficient it is necessary that $\boldsymbol{\phi=0}$.
Substitute the expression for $\mathit{\alpha}_{i}$ given by ((ref)) in ((ref)) to obtain
and write the $n$ return equations more compactly as
where $\mathbf{r}_{\circ t}=(r_{1t},r_{2t},....,r_{nt})^{\prime},$ $\boldsymbol{\tau}_{n}$ is an $n$-dimensional vector of ones, $\mathbf{B} _{n}=(\boldsymbol{\beta}_{\circ1},\boldsymbol{\beta}_{\circ2} ,...,\boldsymbol{\beta}_{\circ K}),$ $\boldsymbol{\beta}_{\circ k}=(\beta _{1k},\beta_{2k},...,\beta_{nk})^{\prime},$ $\mathbf{u}_{\circ t} =(u_{1t},u_{2t},....,u_{nt})^{\prime}$, $\boldsymbol{\eta}_{n}=\left( \eta_{1},\eta_{2},...,\eta_{n}\right) ^{\prime}$, and $\mathbf{V} _{u}=E\left( \mathbf{u}_{\circ t}\mathbf{u}_{\circ t}^{\prime}\right) $. Suppose that the factors, $\mathbf{f}_{t},$ are traded, and $\boldsymbol{\phi }^{\prime}\boldsymbol{\phi}>0$. Consider the $n\times1$ vector of phi-portfolio weights,$\mathbf{w}_{\phi}$, given by
where $\mathbf{M}_{n}=\mathbf{I}_{n}-n^{-1}\boldsymbol{\tau}_{n} \boldsymbol{\tau}_{n}^{\prime}$. Finally, consider the long-short hedged portfolio return
and using ((ref)) note that
For a given vector of pricing errors, $\boldsymbol{\eta}_{n}$, $E\left( \rho_{t,\phi}\right) =\boldsymbol{\phi}^{\prime}\boldsymbol{\phi} +\mathbf{w}_{\phi}^{\prime}\boldsymbol{\eta}_{n}$, and $Var\left( \rho_{t,\phi}\right) =\mathbf{w}_{\phi}^{\prime}\mathbf{V}_{u}\mathbf{w} _{\phi}\mathbf{,}$ where $\mathbf{V}_{u}=E(\mathbf{u}_{\circ t}\mathbf{u} _{\circ t}^{\prime})$. Using ((ref)) it follows that \[ \left\Vert \mathbf{w}_{\phi}^{\prime}\boldsymbol{\eta}_{n}\right\Vert ^{2} \leq\left\Vert \boldsymbol{\eta}_{n}\right\Vert ^{2}\left\Vert \mathbf{w} _{\phi}\right\Vert ^{2}\leq\left\Vert \boldsymbol{\eta}_{n}\right\Vert ^{2}\left\Vert \boldsymbol{\phi}\right\Vert ^{2}\lambda_{max}\left[ \left( \mathbf{B}_{n}^{\prime}\mathbf{M}_{n}\mathbf{B}_{n}\right) ^{-1}\right] =\frac{\left\Vert \boldsymbol{\eta}_{n}\right\Vert ^{2}\left\Vert \boldsymbol{\phi}\right\Vert ^{2}}{\lambda_{min}\left( \mathbf{B}_{n} ^{\prime}\mathbf{M}_{n}\mathbf{B}_{n}\right) }. \] Similarly,
Hence, $E\left( \boldsymbol{\rho}_{t,\phi}\right) \rightarrow \boldsymbol{\phi}^{\prime}\boldsymbol{\phi}$, and $Var\left( \rho_{t,\phi }\right) \rightarrow0$, as $n\rightarrow\infty$, if
These conditions are met if $\lambda_{min}\left( \mathbf{B}_{n}^{\prime }\mathbf{M}_{n}\mathbf{B}_{n}\right) \rightarrow\infty$, $\lambda _{max}\left( \mathbf{V}_{u}\right) <C$, and the APT condition ((ref) ) holds. The first two conditions follow if the LFPM given by ((ref)) is an approximate factor model as assumed by cham1983arbi , namely when the factors are strong and the errors, $u_{it}$, are weakly cross correlated.\footnote{The diversification conditions in ((ref)) are met more generally, and accommodate the presence of semi-strong traded factors in $f_{t}$, and less restrictive conditions on $\mathbf{V}_{u}$.and the pricing errors $\eta_{i}.$ See Remark (ref) below.}
Therefore, knowledge of $\boldsymbol{\phi}^{\prime}\boldsymbol{\phi}$ and its statistical significance can play an important role in portfolio analysis. To illustrate this point suppose that $c=0$ and $\boldsymbol{\eta}_{n} =\mathbf{0}$, but $\boldsymbol{\phi}^{\prime}\boldsymbol{\phi}>0$, and consider the hedged return $\rho_{t,\phi}=\boldsymbol{\phi}^{\prime}\left[ \left( \mathbf{B}_{n}^{\prime}\mathbf{V}_{u}^{-1}\mathbf{B}_{n}\right) ^{-1}\mathbf{B}_{n}\mathbf{V}_{u}^{-1}\mathbf{r}_{\circ t}-\mathbf{f} _{t}\right] $. Then under LFPM we have\footnote{When $c\neq0$, it is not possible to eliminate $c$ (which is unpriced), and at the same time exploit the error covariance matrix $\mathbf{V}_{u}$.}
and its squared Sharpe ratio is given by
Since, \[ \boldsymbol{\phi}^{\prime}\left( \mathbf{B}_{n}^{\prime}\mathbf{V}_{u} ^{-1}\mathbf{B}_{n}\right) ^{-1}\boldsymbol{\phi}\leq\left( \boldsymbol{\phi }^{\prime}\boldsymbol{\phi}\right) \lambda_{\max}\left[ \left( \mathbf{B}_{n}^{\prime}\mathbf{V}_{u}^{-1}\mathbf{B}_{n}\right) ^{-1}\right] =\frac{\boldsymbol{\phi}^{\prime}\boldsymbol{\phi}}{\lambda_{min}\left( \mathbf{B}_{n}^{\prime}\mathbf{V}_{u}^{-1}\mathbf{B}_{n}\right) }, \] then
As a result, if $\lambda_{min}\left( \mathbf{B}_{n}^{\prime}\mathbf{V} _{u}^{-1}\mathbf{B}_{n}\right) \rightarrow\infty$ as $n\rightarrow\infty$, $SR_{\phi}^{2}$ increases in $n$ without bounds if $\boldsymbol{\phi}^{\prime }\boldsymbol{\phi}>0$. In contrast, it is easily seen that the Sharpe ratio of the standard mean-variance (MV) portfolio, defined by $SR_{MV}^{2} =\boldsymbol{\mu}_{R}^{\prime}\mathbf{V}_{R}^{-1}\boldsymbol{\mu}_{R}$ is bounded in $n$, and in consequence $SR_{\phi}^{2}$ will eventually dominate $SR_{MV}^{2}$ if $\boldsymbol{\phi}^{\prime}\boldsymbol{\phi}>0$. To see this, note that under the LFPM given by ((ref)) with $c=0$ and $\boldsymbol{\eta }_{n}=\mathbf{0}$, the MV portfolio is given by $\rho_{MV,t}=\boldsymbol{\mu }_{R}^{\prime}\mathbf{V}_{R}^{-1}\mathbf{r}_{\circ t}$, where $\boldsymbol{\mu }_{R}=\mathbf{B}_{n}\boldsymbol{\lambda}$ and $\mathbf{V}_{R}=\mathbf{B} _{n}\mathbf{\Sigma}_{f}\mathbf{B}_{n}^{\prime}+\mathbf{V}_{u}$. Also, since $\mathbf{B}_{n}\mathbf{\Sigma}_{f}\mathbf{B}_{n}^{\prime}$ is rank deficient then \[ \mathbf{V}_{R}^{-1}=\mathbf{V}_{u}^{-1}\mathbf{-V}_{u}^{-1}\mathbf{B}\left( \mathbf{\Sigma}_{f}^{-1}+\mathbf{B}_{n}^{\prime}\mathbf{V}_{u}^{-1} \mathbf{B}_{n}\right) ^{-1}\mathbf{B}_{n}^{\prime}\mathbf{V}_{u}^{-1}, \] and it follows that
Hence, the Sharpe ratio of the MV portfolio continues to be bounded in $n$ even if $\boldsymbol{\phi\neq0}$. Whether $\boldsymbol{\phi=0}$ or not affects the magnitude of $SR_{MV}^{2}$ but does not alter the fact that $SR_{MV}^{2}$ will be bounded if $\lambda_{min}\left( \mathbf{B}_{n}^{\prime}\mathbf{V} _{u}^{-1}\mathbf{B}_{n}\right) \rightarrow\infty$ as $n\rightarrow\infty$. Furthermore, it is not possible to improve over the MV portfolio by using only the factor loadings. The optimal portfolio formed using the $K$ beta-based portfolios, { $\boldsymbol{\rho}_{B,t}=$ }$\left( \mathbf{B} _{n}^{\prime}\mathbf{V}_{u}^{-1}\mathbf{B}_{n}\right) ^{-1}\mathbf{B} _{n}^{\prime}\mathbf{V}_{u}^{-1}${ $\mathbf{r}_{\circ t}$, }has the same Sharpe ratio as the mean-variance portfolio. See Lemma (ref) for a proof.
In short, to make sure that the mean-variance portfolio is efficient we must have $\boldsymbol{\phi}^{\prime}\boldsymbol{\phi}=0$, and it is of special interest to estimate $\boldsymbol{\phi}$ and develop reliable procedures for testing its statistical significance, using a large number of securities. In practice, the number of tradeable securities, $n$, might not be sufficiently large, and there are important specification and estimation uncertainties, and the Sharpe ratio of the phi-based portfolio, $\boldsymbol{\rho}_{t,\phi}$, is likely to be bounded even if $\boldsymbol{\phi}^{\prime}\boldsymbol{\phi}>0$. We turn to these issues in the empirical application provided in\ Section (ref).
It will prove convenient to write ((ref)) in matrix notation by stacking the excess returns by $t=1,2,...,T$, for each security $i$
where $\mathbf{r}_{i\circ}=(r_{i1},r_{i2},...,r_{iT})^{\prime}$, $\mathbf{F=(f}_{1},\mathbf{f}_{2},...,\mathbf{f}_{T})^{\prime}$, $\mathbf{u}_{i\circ}=\left( u_{i1},u_{i2},...,u_{iT}\right) ^{\prime}$, and $\boldsymbol{\tau}_{T}$ is a $T\times1$ vector of ones. Similarly, stacking the excess returns by $i$ for each $t$ we have ((ref)) which we rewrite as
where $\boldsymbol{\alpha}_{n}=(\alpha_{1},\alpha_{2},...,\alpha_{n})^{\prime }=c\mathbf{\tau}_{n}+\mathbf{B}_{n}\boldsymbol{\phi}+\mathbf{\eta}_{n}$.
The risk premia are usually estimated using a two-pass procedure suggested by fama1973risk . The first-pass runs time series regressions of excess returns, $r_{it}$, on the $K$ observed factors to give estimates of the factor loadings, $\boldsymbol{\beta}_{i}:$
The second-pass runs a cross section regression of\ average returns, $\bar {r}_{i\circ}=T^{-1}\sum_{t=1}^{T}r_{it}$ on the estimated factor loadings, to obtain the FM estimator of $\boldsymbol{\lambda}$, namely
where $\mathbf{\hat{B}}_{nT}=(\boldsymbol{\hat{\beta}}_{1T},\boldsymbol{\hat {\beta}}_{2T},...,\boldsymbol{\hat{\beta}}_{nT})^{\prime},$ $\mathbf{\bar{r} }_{n\circ}=(\bar{r}_{1T},\bar{r}_{2T},...,\bar{r}_{nT})^{\prime},$ $\mathbf{M}_{T}=\mathbf{I}_{T}-T^{-1}\boldsymbol{\tau}_{T}\boldsymbol{\tau }_{T}^{\prime}$, $\boldsymbol{\tau}_{T}$ is a $T$-dimensional vector of ones, $\mathbf{M}_{n}=\mathbf{I}_{n}-n^{-1}\boldsymbol{\tau}_{n}\boldsymbol{\tau }_{n}^{\prime}$, and $\boldsymbol{\tau}_{n}$ is an n-dimensional vector of ones.
As is well known, when $T$ is finite FM's two-pass estimator is biased due the errors in estimation of factor loadings that do not vanish. The small $T$ bias of the two-pass estimator of $\boldsymbol{\lambda}$ has been a source of concern in the empirical literature. Under standard regularity conditions and as $n\rightarrow\infty$, we have
where $\mathbf{d}_{fT}=\mathbf{\hat{\mu}}_{T}-\mathbf{\mu}_{0},$ $\boldsymbol{\lambda}_{0}$ and $\mathbf{\mu}_{0}$ are the true values of $\boldsymbol{\lambda}$ and $\mathbf{\mu}$, respectively, $\mathbf{\Sigma }_{\beta\beta}=\lim_{n\rightarrow\infty}\left( n^{-1}\mathbf{B}_{n}^{\prime }\mathbf{M}_{n}\mathbf{B}_{n}\right) $, and $\overline{\sigma}^{2} =\lim_{n\rightarrow\infty}n^{-1}\sum_{i=1}^{n}\sigma_{i}^{2}>0$. Following shanken1992estimation , $\overline{\sigma}_{n}^{2}$ can be consistently estimated (for a fixed $T$) by
where
and as before $\hat{\alpha}_{iT}$ and $\boldsymbol{\hat{\beta}}_{iT}$ are the OLS estimators of $\mathit{\alpha}_{i}$ and $\boldsymbol{\beta}_{i}$. Using these results the bias-corrected version of the two-pass estimator is given by\footnote{See also shanken2007estimating , kan2013pricing , and BAI201531 , and the survey paper by jagannathan2010analysis for further references.}
where
When all the risk factors are strong, under certain regularity conditions, there exists a fixed $T_{0}$ such that for all $T>T_{0}$, then
where $\boldsymbol{\mu}_{0}$ indicates the true value of the factor mean. Shanken refers to $\boldsymbol{\lambda}_{T}^{\ast}$ as "ex-post" risk premia to be distinguished from $\boldsymbol{\lambda}_{0}$, referred to as "ex ante" risk premia. See also Section 3.7 of jagannathan2010analysis .
In this paper we exploit Shanken's bias correction procedure by applying it to $\boldsymbol{\phi}=$ $\boldsymbol{\lambda}-\boldsymbol{\mu}$ which we identify directly using ((ref)) from the regression of $\mathit{\alpha}_{i}$ on $\boldsymbol{\beta}_{i}$ for $i=1,2,...,n$, assuming the idiosyncratic pricing errors, $\eta_{i}$, are sufficiently weak relative to the strengths of the risk factors in a sense which will be made precise below.
In view of ((ref)), the estimation of $\boldsymbol{\phi}$ can be carried out following a two-step procedure whereby in the first step $\mathit{\alpha }_{i}$ and $\boldsymbol{\beta}_{i}$ are estimated from the least squares regressions of $r_{it}$ on an intercept and $\mathbf{f}_{t}$, and these are then used in a second step regression to estimate $\boldsymbol{\phi}$, namely,
where $\boldsymbol{\hat{\alpha}}_{nT}=(\hat{\alpha}_{1T},\hat{\alpha} _{1T},...,\hat{\alpha}_{nT})^{\prime}=\mathbf{\bar{r}}_{nT}-\mathbf{\hat{B} }_{nT}\boldsymbol{\hat{\mu}}_{T},$ and as before $\mathbf{\hat{B}} _{nT}=(\boldsymbol{\hat{\beta}}_{1T},\boldsymbol{\hat{\beta}}_{2T} ,...,\boldsymbol{\hat{\beta}}_{nT})^{\prime}$. This estimator is consistent for $\boldsymbol{\phi}_{0}$ so long as $n$ and $T\rightarrow\infty$, and bias-corrections are necessary to ensure the large $n$ consistency of the estimator when $T$ is fixed. A Shanken type bias-corrected estimator of $\boldsymbol{\phi}_{0}$ is given by
where $\mathbf{H}_{nT\ }$ and $\widehat{\bar{\sigma}}_{nT}^{2}$ are given by ((ref)) and ((ref)), respectively. It is also easily established that
and for a fixed $T$ and as $n\rightarrow\infty$, we have \[ p\lim_{n\rightarrow\infty}\boldsymbol{\tilde{\phi}}_{nT}=p\lim_{n\rightarrow \infty}\boldsymbol{\tilde{\lambda}}_{nT}-\boldsymbol{\hat{\mu}}_{T}. \] Hence, upon using ((ref))
and there exists a fixed $T_{0}$ such that for all $T>T_{0},$ $\boldsymbol{\tilde{\phi}}_{nT}$ converges to $\boldsymbol{\phi}_{0}$ as $n\rightarrow\infty$. \ Also using ((ref)) and ((ref)), and noting that $\boldsymbol{\lambda}_{0}-\boldsymbol{\mu}_{0}=\boldsymbol{\phi }_{0}$, interestingly we have \[ \boldsymbol{\tilde{\lambda}}_{nT}-\boldsymbol{\lambda}_{T}^{\ast }=\boldsymbol{\tilde{\phi}}_{nT}+\boldsymbol{\hat{\mu}}_{T} -\boldsymbol{\lambda}_{T}^{\ast}=\boldsymbol{\tilde{\phi}}_{nT} -\boldsymbol{\phi}_{0}. \] So inference using the Shanken bias-corrected estimator of $\boldsymbol{\lambda}$ around $\boldsymbol{\lambda}_{T}^{\ast}$, is the same as making inference using $\boldsymbol{\tilde{\phi}}_{nT}$ around $\boldsymbol{\phi}_{0}$.
The asymptotic distribution of $\boldsymbol{\tilde{\phi}}_{nT}$ depends on both $n$ and $T$. Assuming the observed factors are strong and under certain regularity conditions, to be introduced below, we have
where \[ \mathbf{V}_{\xi}=\left( 1+\boldsymbol{\lambda}_{0}^{\prime}\mathbf{\Sigma }_{f}^{-1}\boldsymbol{\lambda}_{0}\right) p\lim_{n\rightarrow\infty}\left[ n^{-1}\mathbf{B}_{n}^{\prime}\mathbf{M}_{n}\mathbf{V}_{u}\mathbf{M} _{n}\mathbf{B}_{n\ }\right] \emph{.} \] The variance of $\boldsymbol{\tilde{\phi}}_{nT}$ is consistently estimated by
where $\mathbf{H}_{nT\ }$ is given by ((ref)),
and
$\boldsymbol{\tilde{\lambda}}_{nT}$ is defined by ((ref)), and $\mathbf{\tilde{V}}_{u}\mathbf{\ }$is a suitable estimator of $\mathbf{V} _{u}=E\left( \mathbf{u}_{\circ t}\mathbf{u}_{\circ t}^{\prime}\right) $. How to estimate $\mathbf{V}_{u}$ and the conditions under which $Var\left( \boldsymbol{\tilde{\phi}}_{nT}\right) $ is consistently estimated by $\widehat{Var\left( \boldsymbol{\tilde{\phi}}_{nT}\right) }$ is discussed in sub-section (ref).
Note that even when $\mathbf{V}_{u}=\sigma^{2}\mathbf{I}_{T}$ the variance of $\boldsymbol{\tilde{\phi}}_{nT}$ does not reduce to $\sigma^{2} \boldsymbol{\Sigma}_{\beta\beta}^{-1}$, the standard least squares formula used for the case of known factor loadings. When the loadings are estimated the scaling term $\left( 1+\boldsymbol{\lambda}_{0}^{\prime} \boldsymbol{\Sigma}_{f}^{-1}\boldsymbol{\lambda}_{0}\right) $ is required and its neglect can lead to serious over-rejection even if $n/T\rightarrow0$ as $n$ and $T\rightarrow\infty$.
In this paper we deviate from the standard literature and allow the observed and latent factors to have different degrees of strength, depending on how pervasively they impact the security returns. bailey2021measurement define the strength of factor, $f_{kt}$, in terms of the number of its non-zero factor loadings. For a factor to be strong almost all of its $n$ loadings must differ from zero. Given our focus on estimation of risk premia, we adopt the following definition which directly relates to the covariance of $\boldsymbol{\beta}_{i}.$ See also chudik2011weak .
In the above definition $\ominus_{p}\left( n^{\alpha_{k}}\right) $ denotes the rate at which additional securities add to the factor's strength and $\alpha_{k}$ can be viewed as a logarithmic expansion rate in terms of $n$ and relates to the proportion of non-zero factor loadings. In the literature it is commonly assumed that the covariance matrix of factor loadings defined by
is positive definite, where $\boldsymbol{\bar{\beta}}_{n}=n^{-1}\sum_{i=1} ^{n}\boldsymbol{\beta}_{i}=(\bar{\beta}_{1},\bar{\beta}_{2},...,\bar{\beta }_{k})^{\prime}$. For $\mathbf{\Sigma}_{\beta\beta}$ to be positive definite matrix it is necessary that all the $K$ risk factors under consideration are strong in the sense that
In terms of our definition of factor strength, $\mathbf{\Sigma}_{\beta\beta} $\ will be positive definite if all the observed factors are strong, namely if $\alpha_{k}=1$ for $k=1,2,...,K$. However, such an assumption is quite restrictive and is unlikely to be satisfied for many risk factors being considered in the literature. bailey2021measurement show that, apart from the market factor, only a handful of 144 factors in the literature considered by feng2020taming come close to being strong. giglio2023test consider the estimation of PCA-based risk premia in presence of weak factors. However, their definition of factor strength involves both $n~$and $T$, and is best viewed as a consistency condition rather than factor strength as such. See the discussion following Theorem (ref). Our notion of factor strength, $\alpha_{k}$, is in line with the recent literature. See, for example, BaiNg2023Weak and uematsu2023inference.
We now turn to the structure of the errors, $u_{it}$, in the returns equations, and consider two possible sources of error cross-sectional dependence: a missing or latent factor and production networks. The issue of missing factors has been investigated in the recent literature by giglio2021asset and ANATOLYEV2022103 . The issue of production networks has been investigated in the recent literature by herskovic2018networks , who derives two risk factors based on the changes in network concentration and network sparsity, and gofman2020production , who focus on the vertical dimension of production by modelling a supply chain, in terms of supplier-customer links. They find that the further away a firm is from final consumers the higher its return. They use this to create a factor TMB (top minus bottom). Both sources of cross-sectional error dependence could be important, since network dependence cannot be represented using latent factor models. See Section 3 of chudik2011weak.
To allow for both forms of error cross-sectional dependence we consider the following decomposition of $u_{it}$
where $g_{t}$ is the missing (latent) factor and $v_{it}$ is weakly cross-correlated in the sense of approximate factor models due to Chamberlain1983 and cham1983arbi. Here we allow for a single missing factor to simplify the exposition, but note that increasing the number of missing factors has little impact on our analysis, so long as the number of missing factors is fixed. Using the normalization $E(g_{t}^{2})=1$, and assuming that $\gamma_{i}g_{t}$ and $v_{it}$ are independently distributed then $E(u_{it}u_{jt})=\sigma_{ij}=\gamma_{i}\gamma_{j}+\sigma_{v,ij}$, and as shown in Lemma (ref), $n^{-1}\sum_{i=1}^{n}\sum_{j=1}^{n}\left\vert \sigma_{ij}\right\vert =O(1)$ so long as the strength of $g_{t},$ $\alpha_{\gamma}<1/2$ and $\lambda_{\max}(\mathbf{V}_{v})<\infty.$ This is despite the fact that $\lambda_{\max}(\mathbf{V}_{u})=O(n^{\alpha_{\gamma}})$, where $\mathbf{V}_{v}=E\left( \mathbf{v}_{i}\mathbf{v}_{i}^{\prime}\right) $ and $\mathbf{V}_{u}=E\left( \mathbf{u}_{i}\mathbf{u}_{i}^{\prime}\right) $.\footnote{Note that Chamberlain's approximate factor model specification requires $\lambda_{\max}(\mathbf{V}_{u})=O(1)$ and is violated if $\alpha_{\gamma}>0$.} In the Monte Carlo experiments, we consider the possibility of missing factors, as well as weak spatial and network cross-dependence that satisfy conditions of approximate factor models.
The APT condition ((ref)), given by (18)\ in Theorem II of ROSS1976341 , ensures that under APT the idiosyncratic pricing errors are sparse. In this paper we relax the\ Ross's condition to
where the exponent $\alpha_{\eta}$ measures the degrees of pervasiveness of pricing errors. Deviations from APT are measured in terms of $\alpha_{\eta}$ ($0\leq\alpha_{\eta}<1$). We investigate the robustness of our proposed estimator of $\boldsymbol{\phi}$ to $\alpha_{\eta}$. This extension is important for tests of market efficiency where the null of interest is $H_{0}:$ $\mathit{\alpha}_{i}=c$ for all $i$ in ((ref)). We note that under the alternative hypothesis $H_{1}:$ $\mathit{\alpha}_{i}=c+\boldsymbol{\beta }_{i}^{\prime}\boldsymbol{\phi}+\eta_{i}$, therefore it is desirable to develop a test of $\boldsymbol{\phi}=\boldsymbol{0}$ which is robust to a wider class of pricing errors than those entertained originally by Ross, where $\alpha_{\eta}=0$.
Under the alternative hypothesis the power of testing $H_{0}$ will depend on the rate at which $\sum_{i=1}^{n}\left( \mathit{\alpha}_{i}-\mathit{\bar {\alpha}}\right) ^{2}$ rises with $n$. This in turn depends on the degree of pervasiveness of the idiosyncratic pricing errors, $\alpha_{\eta}$, the strength of the factors, $\alpha_{j}$, for $j=1,2,...,K$ and the magnitude of $\boldsymbol{\phi}^{\prime}\boldsymbol{\phi}\,$. Using $\mathit{\alpha} _{i}=c+\boldsymbol{\beta}_{i}^{\prime}\boldsymbol{\phi}+\eta_{i}$ \[ \sum_{i=1}^{n}\left( \mathit{\alpha}_{i}-\mathit{\bar{\alpha}}\right) ^{2}=\boldsymbol{\phi}^{\prime}\left( \sum_{i=1}^{n}\left( \boldsymbol{\beta }_{i}-\boldsymbol{\bar{\beta}}_{n}\right) \left( \boldsymbol{\beta} _{i}-\boldsymbol{\bar{\beta}}_{n}\right) ^{\prime}\right) \boldsymbol{\phi }+\sum_{i=1}^{n}\left( \eta_{i}-\bar{\eta}\right) ^{2}+2\boldsymbol{\phi }^{\prime}\sum_{i=1}^{n}\left( \boldsymbol{\beta}_{i}-\boldsymbol{\bar{\beta }}_{n}\right) \eta_{i}, \] and in the absence of idiosyncratic pricing errors
When the risk factors are all strong $\lambda_{\min}\left( n^{-1}\sum _{i=1}^{n}\left( \boldsymbol{\beta}_{i}-\boldsymbol{\bar{\beta}}_{n}\right) \left( \boldsymbol{\beta}_{i}-\boldsymbol{\bar{\beta}}_{n}\right) ^{\prime }\right) >0$, then $\sum_{i=1}^{n}\left( \mathit{\alpha}_{i}-\mathit{\bar {\alpha}}\right) ^{2}=\ominus(n)$ if and only if $\boldsymbol{\phi} \neq\boldsymbol{0}$. Namely, the extent to which alpha can be exploited will depend on the magnitudes of $\phi_{j}$ and the strength of the risk factors.
We make the following standard assumptions about $\mathbf{f}_{t},g_{t}$, $v_{it}$, $\boldsymbol{\beta}_{i}$, $\eta_{i}$, and $\gamma_{i}$ (the drivers of asset returns):
As we shall see, to estimate and conduct inference on the risk premia associated with the observed factors, $f_{kt}$, we require $\alpha_{k} >\alpha_{\gamma}<1/2$, where $\alpha_{\gamma}$ denotes the strength of the latent factor, $g_{t},$ and similarly defined by $\sum_{i=1}^{n}\gamma_{i} ^{2}=\ominus(n^{\alpha_{\gamma}})$. Namely, the latent factor must be sufficiently weak so that ignoring it will be inconsequential, and observed factors sufficiently strong so that they can be distinguished from the weak latent factor.
The main theoretical results of the paper are set out around five theorems. Theorem (ref) considers the Fama-MacBeth two-step estimator and derives its limiting property as $n$ and $T\rightarrow\infty$. To eliminate the bias of Fama-MacBeth estimator we require $n/T\rightarrow0$, and to eliminate the effects of pricing errors we need $Tn^{\alpha_{\eta} }/n\rightarrow0$, which results in a contradiction. Thus the Fama-MacBeth estimator is valid only when there are no pricing errors ($\eta_{i}=0$ for all $i$) and when $n/T\rightarrow0$. Theorem (ref) provides a proof that the estimator of $\bar{\sigma}_{n}^{2}$ (denoted by $\widehat{\bar{\sigma}} _{nT}^{2}$) proposed by shanken1992estimation continues to be unbiased for a fixed $T$ as $n\rightarrow\infty$, even under the general setting of the current paper that allows for missing factors as well as pricing errors. Theorem (ref) also establishes that $\widehat{\bar{\sigma}}_{nT}^{2}-\bar{\sigma}_{n}^{2}\rightarrow O_{p}(n^{-1/2}T^{-1/2})$, which is essential for establishing the results for the bias-corrected estimator of $\boldsymbol{\phi}_{0}$, namely $\boldsymbol{\tilde{\phi}}_{nT}$ given by ((ref)), summarized in Theorem (ref). This theorem provides conditions under which $\boldsymbol{\tilde{\phi}}_{nT}$ is a consistent estimator of $\boldsymbol{\phi}_{0}$, and derives its asymptotic distribution assuming the observed factors are strong, again allowing for pricing errors, a missing factor, and other forms of weak error cross-sectional dependence. Theorem (ref) extends the results of Theorem (ref) to the case where one or more of the observed risk factors are semi-strong and shows how factor strength impacts the precision with which the elements of $\boldsymbol{\phi }_{0}$ are estimated. Finally, Theorem (ref) presents the conditions under which the asymptotic variance of $\boldsymbol{\tilde{\phi}}_{nT}$ can be consistently estimated.
The proof is provided in Section (ref) of the mathematical appendix.
To derive the asymptotic distribution of $\boldsymbol{\hat{\lambda}} _{nT}-\boldsymbol{\lambda}_{0}$ it is required that both $n$ and $T\rightarrow\infty$, jointly. Also, noting that \[ \boldsymbol{\hat{\lambda}}_{nT}-\boldsymbol{\lambda}_{0}=\left( \boldsymbol{\hat{\mu}}_{T}-\boldsymbol{\mu}_{0}\right) +\left( \boldsymbol{\hat{\phi}}_{nT}-\boldsymbol{\phi}_{0}\right) , \] it is clear that increasing $n$ is not relevant for the distribution of $\boldsymbol{\hat{\mu}}_{T}-\boldsymbol{\mu}_{0}$, but joint $n$ and $T$ asymptotics are required when investigating the distribution of $\boldsymbol{\hat{\phi}}_{nT}-\boldsymbol{\phi}_{0}$. Focussing on the latter, and using result ((ref))\ in the Appendix, we have
Where $\boldsymbol{U}_{nT}=\left( \mathbf{u}_{1\circ},\mathbf{u}_{2\circ },...,\mathbf{u}_{n\circ}\right) ^{\prime},$ $\mathbf{u}_{i\circ}=\left( u_{i1},u_{i2},...,u_{iT}\right) ^{\prime},$ $\mathbf{G}_{T}=$ $\mathbf{M} _{T}\mathbf{F}\left( \mathbf{F}^{\prime}\mathbf{M}_{T}\mathbf{F}\right) ^{-1}$, $\overline{\mathbf{u}}_{n\circ}=(\overline{u}_{1\circ},\overline {u}_{2\circ},...,\overline{u}_{n\circ})^{\prime},$ and $\overline{u}_{i\circ }=T^{-1}\sum_{t=1}^{T}u_{it}.$ Consider first the terms that include the pricing errors, $\boldsymbol{\eta}_{n}$, and using the results in Lemma (ref) note that \[ n^{-1/2}T^{1/2}\mathbf{B}_{n}^{\prime}\mathbf{M}_{n}\boldsymbol{\eta} _{n}=O_{p}\left( T^{1/2}n^{-1/2+\alpha_{\eta}}\right) ,\text{ } n^{-1/2}T^{1/2}\mathbf{G}_{T}^{\prime}\mathbf{U}_{nT}^{\prime}\mathbf{M} _{n}\boldsymbol{\eta}_{n}=O_{p}\left( n^{-1/2+\frac{\alpha_{\eta} +\alpha_{\gamma}}{2}}\right) . \] It is clear that the effects of pricing errors on the distribution of $\boldsymbol{\hat{\phi}}_{nT}$ vanish only if $T^{1/2}n^{-1/2+\alpha_{\eta} }\rightarrow0$, and $\alpha_{\eta}+\alpha_{\gamma}<1$. Also
Finally, for the first two terms involving $\mathbf{B}_{n}$ and $\mathbf{U} _{nT}$ we have
It is clear that the small $T$ bias of the asymptotic distribution of the two-step estimator, given by $\sqrt{\frac{n}{T}}\bar{\sigma}_{n}^{2}\left( \frac{\mathbf{F}^{\prime}\mathbf{M}_{T}\mathbf{F}}{T}\right) ^{-1} \boldsymbol{\lambda}_{T}^{\ast}$, does not vanish unless, $n/T\rightarrow0$. At the same time for the pricing errors to have no impact on the distribution of the two-step estimator we must have $Tn^{a_{\eta}}/n\rightarrow0$. Both conditions cannot be met simultaneously. It is possible to derive the asymptotic distribution of $\boldsymbol{\hat{\phi}}_{nT}$, and hence that of $\boldsymbol{\hat{\lambda}}_{nT}$, when $n/T\rightarrow0$ and $\mathbf{\eta =0}$, but these are quite restrictive conditions, and to avoid them we follow shanken1992estimation and instead consider a bias-corrected version of $\boldsymbol{\hat{\phi}}_{nT}$, namely $\boldsymbol{\tilde{\phi}}_{nT}$ given by ((ref)). As noted earlier $\boldsymbol{\tilde{\phi}} _{nT}=\boldsymbol{\tilde{\lambda}}_{nT}-\boldsymbol{\hat{\mu}}_{T}$, where $\boldsymbol{\tilde{\lambda}}_{nT}$ is the bias-corrected version of $\boldsymbol{\hat{\lambda}}_{nT}$ originally proposed by Shanken.
To investigate the asymptotic properties of $\boldsymbol{\tilde{\phi}}_{nT}$ we first need to establish conditions under which $\widehat{\bar{\sigma}} _{nT}^{2}$, defined by ((ref)), is a consistent estimator of $\bar{\sigma}_{n}^{2}=n^{-1}\sum_{i=1}^{n}\sigma_{i}^{2}$, which enters the bias-corrected estimator. The proof of consistency in the literature does not allow for missing factors or pricing errors and only considers the case where $T$ is fixed as $n\rightarrow\infty$. For derivation of asymptotic distribution of $\boldsymbol{\tilde{\phi}}_{nT}$ we also need to consider the limiting properties of $\widehat{\bar{\sigma}}_{nT}^{2}$ under joint $n$ and $T$ asymptotics. The following theorem provides the required results for $\widehat{\bar{\sigma}}_{nT}^{2}$ as an estimator of $\bar{\sigma}_{n}^{2}$.
For a proof see sub-section (ref) in the Appendix.
Result ((ref)) shows that $\widehat{\bar{\sigma}}_{nT}^{2}$ continues to be a consistent estimator of $\bar{\sigma}^{2}=\lim_{n\rightarrow\infty }\bar{\sigma}_{n}^{2}$ for a fixed $T$ as $n\rightarrow\infty$, even in the presence of pricing errors and a missing common factor. This result also holds when one or more of the factors are semi-strong.
Equipped with the above result we are now in a position to present the theorem that sets out the asymptotic distribution of $\boldsymbol{\tilde{\phi}}_{nT}$.
For a proof see sub-section (ref) in the Appendix.
Result ((ref)) establishes the finite $T$ consistency of $\boldsymbol{\tilde{\phi}}_{nT}$ for $\boldsymbol{\phi}_{0}$ so long as $\alpha_{\gamma}<1/2$ and $\alpha_{\eta}<1$, thus extending the Shanken result to a much more general setting. To the best of our knowledge the asymptotic distribution in ((ref)) is new and shows that the asymptotic covariance matrix of $\boldsymbol{\tilde{\phi}}_{nT}$ includes the term $\boldsymbol{\lambda}_{T}^{\ast\prime}(T^{-1}\mathbf{F}^{\prime}\mathbf{M} _{T}\mathbf{F)}^{-1}\boldsymbol{\lambda}_{T}^{\ast}$, that arises from the first stage estimation of the factor loadings, and must be included in the analysis for valid inference. It is also clear that this additional term does not vanish with $T\rightarrow\infty$, and tends to $\boldsymbol{\lambda} _{0}^{\prime}\mathbf{\Sigma}_{f}^{-1}\boldsymbol{\lambda}_{0}\geq\left( \boldsymbol{\lambda}_{0}^{\prime}\boldsymbol{\lambda}_{0}\right) \lambda_{\max}\left( \mathbf{\Sigma}_{f}^{-1}\right) =\left( \boldsymbol{\lambda}_{0}^{\prime}\boldsymbol{\lambda}_{0}\right) \lambda_{\min}\left( \mathbf{\Sigma}_{f}\right) >0$, which is strictly non-zero unless $\boldsymbol{\lambda}_{0}=\mathbf{0}$. Shanken type bias correction addresses the mean of the asymptotic distribution of $\boldsymbol{\tilde{\phi}}_{nT}$, but not its covariance.
The $O_{p}\left( T^{-1/2}\right) $ term in ((ref)) arises from the sampling errors involved in the estimation of the factor loadings and $\bar{\sigma}_{n}^{2}$, and tends to zero at the regular $\sqrt{T}$ rate. But $n$ has to be sufficiently large to eliminate the effects of pricing errors on identification of $\boldsymbol{\phi}_{0}$, as dictated by condition $\sqrt{\frac{T}{n}}n^{\alpha_{\eta}}\rightarrow0,$ as $n$ and $T\rightarrow \infty$.\footnote{The condition $\sqrt{\frac{T}{n}}n^{\alpha_{\eta} }\rightarrow0$ can be weakened somewhat to $\sqrt{\frac{T}{n}}n^{\alpha_{\eta }/2}\rightarrow0$ if we also assume that $\beta_{ik}-\bar{\beta}_{k}$ are independently distributed over $i$, but will still require $n$ to be larger than $T$.} The requirement that $T$ need not be too large relative to $n$ for estimation of $\boldsymbol{\phi}_{0}$ is consistent with separating the estimation of $\boldsymbol{\phi}_{0}$ from that of $\boldsymbol{\mu}_{0}$, allowing the use a relatively small $T$ and a large $n$ to estimate $\boldsymbol{\phi}_{0}$ and a relatively large $T$ when estimating $\boldsymbol{\mu}_{0}.$
We now turn to an intermediate case where one or more of the observed factors are semi-strong, in the sense that their factor strength, $\alpha_{k}$ lies between $1/2$ and $1$. The case of weak risk factors is already covered in the proceeding analysis, and such factors can be included in the error term, $u_{it}$, with little consequence for the estimation of risk premia of the remaining factors that are strong or semi-strong. Weak factors do not have any explanatory power and can be dropped from the analysis.
When one or more of the observed factors is semi-strong $\mathbf{\Sigma }_{\beta\beta}$ is no longer positive definite and Theorem (ref) does not apply, but it is possible to adapt the proofs to establish the limiting properties of $\tilde{\phi}_{k,nT}\boldsymbol{\ }$(the $k^{th}$ element of $\boldsymbol{\tilde{\phi}}_{nT}$)\ for different values of \thinspace $\alpha_{k}\,$.
To this end, analogously to $\boldsymbol{\tilde{\phi}}_{nT}$, we introduce the following estimator of $\boldsymbol{\phi}_{0}$
where
It is now easily seen that
where
$\mathbf{D}_{\alpha}$ is defined by ((ref)), and $\boldsymbol{\alpha }=\mathbf{(}\alpha_{1},\alpha_{2},...,\alpha_{K})^{\prime}$. It is easily established that numerically $\boldsymbol{\tilde{\phi}}_{nT}\left( \boldsymbol{\alpha}\right) $ is identical to $\boldsymbol{\tilde{\phi}}_{nT} $, and its introduction is primarily for the purpose of establishing the limiting properties of $\tilde{\phi}_{k,nT}-\phi_{0,k}$ that do depend on $\alpha_{k}$. Note that \[ \mathbf{H}_{nT}\left( \boldsymbol{\alpha}\right) =n\mathbf{D}_{\alpha} ^{-1}\mathbf{H}_{nT}\mathbf{D}_{\alpha}^{-1},\text{ and }\mathbf{q} _{nT}\left( \boldsymbol{\alpha}\right) =n\mathbf{D}_{\alpha}^{-1} \mathbf{s}_{nT} \] where $\mathbf{s}_{nT}$ and $\mathbf{H}_{nT}$ are already defined by ((ref)) and ((ref)). Using these in ((ref)) we have \[ \mathbf{D}_{\alpha}\left( \boldsymbol{\tilde{\phi}}_{nT}\left( \boldsymbol{\alpha}\right) -\boldsymbol{\phi}_{0}\right) =\left( n\mathbf{D}_{\alpha}^{-1}\mathbf{H}_{nT}\mathbf{D}_{\alpha}^{-1}\right) ^{-1}n\mathbf{D}_{\alpha}^{-1}\mathbf{s}_{nT}=\mathbf{D}_{\alpha} \mathbf{H}_{nT}^{-1}\mathbf{s}_{nT}, \] and it follows that $\boldsymbol{\tilde{\phi}}_{nT}\left( \boldsymbol{\alpha }\right) -\boldsymbol{\phi}_{0}=\mathbf{H}_{nT}^{-1}\mathbf{s}_{nT} =\boldsymbol{\tilde{\phi}}_{nT}\left( \boldsymbol{\tau}_{K}\right) =\boldsymbol{\tilde{\phi}}_{nT}$. See ((ref)).
The convergence results for $\boldsymbol{\tilde{\phi}}_{nT}\left( \boldsymbol{\alpha}\right) $ are set out in the following theorem.
See sub-section (ref) of the Appendix for a proof.
The result in ((ref)) establishes the consistency of $\tilde{\phi }_{k,nT}\left( \boldsymbol{\alpha}\right) =\tilde{\phi}_{k,nT}$ even if $f_{kt}$ is semi-strong so long as $n\rightarrow\infty$, and $\alpha_{\min }>1/2,$ $\alpha_{k}+\alpha_{\min}>\alpha_{\eta}+\alpha_{\gamma}$, and $\alpha_{k}+\alpha_{\min}>2\alpha_{\eta}$. Clearly, these results reduce to the case of strong factors where $\alpha_{\min}=\alpha_{k}=1$. Turning to the asymptotic distribution of $\tilde{\phi}_{k,nT}\left( \boldsymbol{\alpha }\right) $, again only convergence rates are affected, and instead of the regular rate of $\sqrt{nT}$, we have $\ \sqrt{T}n^{(\alpha_{k}+\alpha_{\min }-1)/2}$, and using ((ref)) we have
The conditions needed for eliminating the effects of the pricing errors are the same as before and are given by $\alpha_{\eta}+\alpha_{\gamma}<1$ and $\sqrt{T}n^{-1/2+\alpha_{\eta}}\rightarrow0$. The asymptotic distribution is unaffected except for the slower rate of convergence alluded to above. It is also of interest to note that adding semi-strong factors can adversely affect the convergence rate of the strong factor with $\alpha_{k}=1$. As an example suppose the asset pricing model contains two factors, one strong, $\alpha _{1}=1$ and one semi strong with $\alpha_{2}<1$.$\,\ $Then the convergence rate of $\tilde{\phi}_{1,nT}\left( \boldsymbol{\alpha}\right) -\phi_{0,1}$ is given by$\sqrt{T}$ $n^{(1+\alpha_{2}-1)/2}$ which is slower than the rate we would have obtained for $\tilde{\phi}_{1,nT}\left( \boldsymbol{\alpha }\right) -\phi_{0,1}$ if both factors were strong ($\alpha_{min}=\alpha _{k}=1)$, namely the regular rate of $\sqrt{nT}$.
Furthermore, when conditions $\alpha_{\eta}+\alpha_{\gamma}<1$ and $\sqrt {T}n^{-1/2+\alpha_{\eta}}\rightarrow0$ are met we have
and $\phi_{0,k}$ is consistently estimated if $T^{-1/2}n^{-(\alpha_{k} +\alpha_{\min}-1)/2}\rightarrow0$. Also, using ((ref)), an estimator of the risk premia, $\lambda_{k}$, is given by $\tilde{\lambda}_{k,nT}\left( \boldsymbol{\alpha}\right) =\tilde{\phi}_{k,nT}\left( \boldsymbol{\alpha }\right) +\hat{\mu}_{k,T}$, where $\hat{\mu}_{k,T}=T^{-1}\sum_{t=1}^{T} f_{kt}$. Hence \[ \tilde{\lambda}_{k,nT}\left( \boldsymbol{\alpha}\right) -\lambda _{0,k}=\left[ \tilde{\phi}_{k,nT}\left( \boldsymbol{\alpha}\right) -\phi_{0,k}\right] +\left( \hat{\mu}_{k,T}-\mu_{0,k}\right) . \] Under Assumption (ref) $\hat{\mu}_{k,T}-\mu_{0,k}=O_{p}\left( T^{-1/2}\right) $, and using ((ref)) it then follows that
and $\tilde{\lambda}_{k,nT}\left( \boldsymbol{\alpha}\right) $ is a consistent estimator of $\lambda_{0,k}$ if $T\rightarrow\infty$ as well as $T^{-1/2}n^{-(\alpha_{k}+\alpha_{\min}-1)/2}\rightarrow0$. More specifically, suppose $T=\ominus\left( n^{d}\right) $ for some $d>0$, where $\ominus \left( \cdot\right) $ denotes $T$ and $n^{d}$ are of the same order of magnitude. Then for any $d>0$, the condition for consistency of $\tilde {\lambda}_{k,nT}\left( \boldsymbol{\alpha}\right) $ is given by $\alpha _{k}+\alpha_{min}+d>1$ and $d>0$. In the case where all risk factors have the same strength, $\alpha$, the consistency condition reduces to $\alpha>\left( 1-d\right) /2$, which is weaker than the one derived by giglio2023test , namely $n/(\left\Vert \beta\right\Vert ^{2}T)\rightarrow0$, where, in terms of our notation, $\left\Vert \beta\right\Vert ^{2}=\ominus\left( n^{\alpha }\right) $. This latter condition will be met if $\alpha>1-d$. In practice where $T$ is small relative to $n$, the accuracy of $\tilde{\lambda} _{k,nT}\left( \boldsymbol{\alpha}\right) $ as an estimator $\lambda_{0,k}$ does depend on $\alpha$, and our weaker condition on $\alpha>\left( 1-d\right) /2$ is advantageous.
To carry out inference on $\boldsymbol{\phi}_{0}$, or any of its elements individually, we require a consistent estimator of $Var\left( \boldsymbol{\tilde{\phi}}_{nT}\right) $. Using ((ref)) and ((ref)) we first note that $\mathbf{\Sigma}_{\beta\beta}$ is consistently estimated by $\mathbf{H}_{nT\ }$ given by ((ref)). Therefore, it is sufficient to find a suitable estimator of $\mathbf{V}_{u}=(\sigma_{ij})$ such that $\mathbf{V}_{\xi}$ given by ((ref)) is consistently estimated. Under suitable sparsity restrictions $\mathbf{V}_{u}$ can be consistently estimated using the various thresholding procedures advanced in the statistical literature by bickel2008covariance, bickel2008regularized , cai2011adaptive , and BPS2019multiple . fan2011high, fan2013large also show that the adaptive threshold technique of Cai and Liu applies equally to the residuals from an approximate factor model. Here we consider the threshold estimator proposed by BPS which does not require cross-validation and is shown to have desirable small sample properties. It is given by $\mathbf{\tilde{V}}_{u}=\left( \tilde{\sigma}_{ij}\right) $
where
and $c_{p}(n,d)=\Phi^{-1}\left( 1-\frac{p}{2n^{d}}\right) ,$ is a normal critical value function, $p$ is the the nominal size of testing of $\sigma_{ij}=0$, ($i\neq j$) and $d$ is chosen to take account of the $n(n-1)/2$ multiple tests being carried out. Monte Carlo experiments carried out by BPS suggest setting $d=2$. The variance estimator given by ((ref)) does not require a knowledge of the factor strength and applies to risk factors of differing degrees.
Under Assumptions (ref), (ref), and (ref), $\left\Vert \mathbf{V}_{u}\right\Vert =O\left( n^{\alpha_{\gamma}}\right) $, and using results in fan2011high, fan2013large we have
Consider the following estimator of $\mathbf{V}_{\xi}$ \[ \mathbf{\hat{V}}_{\xi,nT}=\left( 1+\hat{s}_{nT}\right) \left( n^{-1}\mathbf{\hat{B}}_{nT}^{\prime}\mathbf{M}_{n}\mathbf{\tilde{V}} _{u}\mathbf{M}_{n}\mathbf{\hat{B}}_{nT}\right) \text{.} \] where $\hat{s}_{nT}=\boldsymbol{\tilde{\lambda}}_{nT}^{^{\prime}}\left( T^{-1}\mathbf{F}^{\prime}\mathbf{M}_{T}\mathbf{F}\right) ^{-1} \boldsymbol{\tilde{\lambda}}_{nT}$. Under Assumption (ref), $T^{-1}\mathbf{F}^{\prime}\mathbf{M}_{T}\mathbf{F\rightarrow}_{p} \mathbf{\Sigma}_{f}$ $\ $and using the results above we have $\boldsymbol{\tilde{\lambda}}_{nT}=\boldsymbol{\tilde{\phi}}_{nT} +\boldsymbol{\hat{\mu}}_{T}\rightarrow_{p}\boldsymbol{\phi}_{0} +\boldsymbol{\mu}_{0}=\boldsymbol{\lambda}_{0}$. Hence, $\hat{s} _{nT}\rightarrow_{p}\boldsymbol{\lambda}_{0}^{\prime}\mathbf{\Sigma}_{f} ^{-1}\boldsymbol{\lambda}_{0}$ as $n,T\rightarrow\infty$, jointly, and it is sufficient to show that
The following theorem provides a formal statement of the conditions under which $\mathbf{\hat{V}}_{\xi,nT}$ is a consistent estimator of $\mathbf{V} _{\xi}$.
For a proof see sub-section (ref) in the Appendix.
This theorem shows that consistent estimation of $Var\left( \boldsymbol{\tilde{\phi}}_{nT}\right) $ can be achieved by using a suitable threshold estimator of $\mathbf{V}_{u}$, so long as the strength of the missing factor, $\alpha_{\gamma}$, is sufficiently weak in the sense that $n^{\alpha_{\gamma}}\sqrt{\ln(n)/T}\rightarrow0$ as $n,T\rightarrow\infty$.
This section presents Monte Carlo simulations to investigate the small sample properties of estimators and tests for $\boldsymbol{\phi}_{0}$. In the empirical application of the next section the factors are selected from a large list. But here we assume $K=3$ and mimic the 3 Fama-French factors, namely the market return minus the risk free rate, MKT, the value factor (high minus low book to market portfolios, HML) and the size factor (small minus big portfolios, SMB). These are denoted by $f_{kt},$ $k=M,H,S$.\footnote{Data on factors and the risk free rate are downloaded from Kenneth French's data library: https://mba.tuck.dartmouth.edu/pages/faculty/ken.french/data_library.html} For further details see Section (ref) of the online supplement A.
To calibrate the loadings, $\beta_{ik}$, we used excess returns on a large number securities observed over the shorter sample covering the 20 years $2002m1$ $-2021m12$ ($T=240$). Monthly returns for NYSE and NASDAQ stocks code 10 and 11 from CRSP were downloaded from Wharton Research Data Services and converted to excess returns over the risk free rate, taken from Kenneth French's webpages, in percent per month. Only stocks with available data for the full sample were included, yielding a balanced panel, and to avoid outliers influencing the results, stocks with a kurtosis greater than 16 were excluded. There were $1289$ stocks before exclusion on the basis of kurtosis and $1175$ after. The summary statistics giving mean, median, standard deviation of the estimates of $\beta_{ik}$ and their histograms are provided in the online supplement A.
For factor strength, we considered a range of DGPs. Given the evidence that most factors, other than the market factor, are not strong, we focus on the case where there is one strong factor, namely the market factor with $\alpha_{M}=1,$ plus two semi-strong factors, with the value factor, $HML$, being quite strong with $\alpha_{H}=0.85,$ and the size factor, $SML$, being only moderately strong with $\alpha_{S}=0.65$. These estimates are also informed by the results provided in bailey2021measurement who propose methods for estimation of factor strength. For a given factor strength, $\alpha_{k}$, the associated loadings, $\beta_{ik}$, are generated as $\boldsymbol{\beta}_{k}=(\beta_{1k},\beta_{2k},,...,\beta_{\alpha_{k} },0,0,...0)$ where $n_{\alpha_{k}}=\lfloor n^{\alpha_{k}}\rfloor$ the integer part of $n^{\alpha_{k}}$, with non-zero and zero values of $\boldsymbol{\beta }_{k}$ given by
where $\lfloor n^{\alpha_{k}}\rfloor$ denotes the integer part of $n^{\alpha_{k}}$. \ Since the security returns are randomly generated, it does not matter how zero and non-zero values of $\beta_{ik}$ are distributed across $i$. Also, the zero loadings can also be replaced by an exponentially decaying sequence without any implications for the simulation results.\footnote{See also footnote 5 of \citet *[p.942]{bailey2016exponent}.} We also set
which match the mean and standard deviation of the estimates of $\beta_{ik}$. See above.
The pricing errors in ((ref)) can be considered as firm-specific characteristics and are set as $\boldsymbol{\eta}_{n}=\left( \eta_{1} ,\eta_{2},...,\eta_{n_{\eta}},0,0,...,0\right) ^{\prime}.$ The non-zero loadings of $\boldsymbol{\eta}_{n}$ for $i\leq n_{\eta}=\lfloor n^{\alpha _{\eta}}\rfloor$\ are drawn from $IIDU(0.7,0.9)$, and $\eta_{i}=0$ for $i=n_{\eta}+1,n_{\eta}+2,....,n.$ We consider $\alpha_{\eta}=(0,0.3).$ When $\alpha_{\eta}=0$ we have $\eta_{i}=0$ ,\ for all $i.$ As in the case of factor loadings the non-zero values of $\boldsymbol{\eta }_{n}$ must be randomly allocated to different groups.
The return equation errors, $u_{it}$, are generated following ((ref)) as a combination of a missing factor, $g_{t}\sim IIDN\left( 0,1\right) $ plus an idiosyncratic error, $v_{jt}.$ The loadings $\boldsymbol{\gamma}=\left( \gamma_{1},\gamma_{2},...,\gamma_{n_{\gamma}},0,0,...,0\right) ^{\prime}$ of the missing factor are set as
where $\alpha_{\gamma}$ is the strength of the missing factor $g_{t}$. We consider $\alpha_{\gamma}=1/4$ and $1/2$.
For the idiosyncratic errors, $v_{it}$, we consider spatial as well as a block diagonal specification, with the spatial specification including a diagonal specification as the special case. Under the spatial specification the idiosyncratic errors are generated as the first order spatial autoregressive model $v_{it}=\rho_{\varepsilon}\sum_{j=1}^{n}w_{ij}v_{jt}+\kappa \varepsilon_{it},$ which can be written in matrix notation as $\mathbf{v} _{t}=\rho_{\varepsilon}\mathbf{Wv}_{t}+\kappa\boldsymbol{\varepsilon}_{t},$ and solved for as $\mathbf{v}_{t}=\kappa\left( \mathbf{I}_{n}-\rho _{\varepsilon}\mathbf{W}\right) ^{-1}\boldsymbol{\varepsilon}_{t}$. Adding the missing factor now yields
The spatial coefficient $\rho_{\varepsilon}$ is such that $|\rho_{\varepsilon }|<1$, $\mathbf{W=(}w_{ij})$ with $w_{ii}=0$, and $\sum_{j=1}^{n}w_{ij}=1$. The diagonal case is obtained by setting $\rho_{\varepsilon}=0$, with $\rho_{\varepsilon}=0.5$ characterizing the SAR specification. The weight matrix $\mathbf{W}=(w_{ij})$ is set to follow the familiar rook pattern where all its elements are set to zero except for $w_{i+1,i}=w_{j-1,j}=0.5$ for $i=1,2,...,n-2$ and $j=3,4...,n$, with $w_{1,2}=w_{n,n-1}=1$.
Under the block error covariance specification, $\mathbf{v}_{t}$ is generated as $\mathbf{v}_{t}=\kappa\mathbf{\hat{S}}\boldsymbol{\varepsilon}_{t},$ where $\mathbf{\hat{S}}$ is a block diagonal matrix with its $b^{th}$ block given by $\mathbf{\hat{S}}_{b}$ for $b=1,2,...,B$, and $\boldsymbol{\varepsilon} _{t}=(\mathbf{\varepsilon}_{1t}^{\prime},\mathbf{\varepsilon}_{2t}^{\prime },...,\mathbf{\varepsilon}_{Bt}^{\prime})^{\prime}$, and $\boldsymbol{\varepsilon}_{bt}=(\varepsilon_{b,1t},\varepsilon_{b,2t} ,...,\varepsilon_{b,n_{b},t})^{\prime}$. $\mathbf{\hat{S}}$ is set as a Cholesky factor of the correlation matrix of $\mathbf{u}_{t}$. Denoting this correlation matrix by $\mathbf{\hat{R}}_{u},$ \[ \mathbf{\hat{R}}_{u}=\left[ Diag(\mathbf{\hat{V}}_{Bu})\right] ^{-1/2}\mathbf{\hat{V}}_{Bu}\left[ Diag(\mathbf{\hat{V}}_{Bu})\right] ^{-1/2}=Diag(\mathbf{\hat{R}}_{bu},\text{ }b=1,2,...,B), \] where $\mathbf{\hat{V}}_{Bu}$ is the threshold estimator of $\mathbf{V}_{u}$ subject to the additional restriction that $\mathbf{V}_{u}$ is block diagonal. For each block $\mathbf{\hat{R}}_{bu}$ we set the number of distinct non-zero elements of this block equal to the integer part of $[n_{b}(n_{b}-1)/2]\times q_{b}$ where $q_{b}$ is the proportion of non-zero distinct elements in block $b$ of our calibrated sample and computed by the calibration over the sample $2001m10-2021m9$. The non-zero elements are drawn randomly from $IIDU(0,0.5)$. Similarly, adding the missing factor, we have
The block diagonal structure is intended to capture possible within industry correlations not picked up by observed or weak missing factors, with each block representing an industry or sector. To calibrate the block structure estimates of the pair-wise correlations between the residuals of the return regressions using the Fama-French three factors of the $T=240$ sample ending in $2021$ were obtained. Then all the statistically insignificant correlations were set to zero, allowing for the multiple testing nature of the tests. For the majority of securities (668 out of the 1168), the pair-wise return correlations were not statistically significant. The securities with a relatively large number of non-zero correlations were either in the banking or energy related industries. Considering stocks by 2-digit SIC classifications, a division into $B=14$ contiguous groups ranging in size from $33$ to $145$ stocks, seemed sensible. More detail on the process is given in Section (ref) of the online supplement A.
The primitive errors, $\varepsilon_{it}$ for $i=1,2,...,n$ in ((ref)) and ((ref))\ are generated as $\varepsilon_{it}=\sqrt{\sigma_{ii}} \varpi_{it}$, where $\varpi_{it}\sim IIDN\left( 0,1\right) $, and $\varepsilon_{it}=\sqrt{\sigma_{ii}}\left[ \sqrt{\frac{v-2}{v}}\varpi _{it}\right] ,$ where $\varpi_{it}\sim IID$ $t(v)$, with $t(v)$ denotes a standard $t$ distributed variate with $v=5$ degrees of freedom. Also $\sigma_{ii}\thicksim IID$ \ $0.5(1+\chi_{1}^{2})$ $,\ for$\ $i=1,2,...,n_{b}$ and $b=1,2,...,B$. In this way, it is ensured that $Var(\varepsilon _{b,it})=\sigma_{b,ii}$, and on average $E\left[ Var(\varepsilon _{b,it})\right] =E(\sigma_{ii})=1$, under both Gaussian and t-distributed errors. Note that $Var\left( \nu_{b,it}\right) =v/(v-2)$. All the experiments are designed to give an $R^{2}$ of about $0.3,$ similar to that obtained in the empirical applications. For further details see sub-section (ref) of the online supplement A.
In total, we consider 12 experimental designs: six designs with Gaussian errors and six with $t(5)$ distributed errors. We considered designs with GARCH effects, with and without pricing errors, $\eta_{i}$, and with and without the missing factor, $g_{t}$. We also considered designs with spatial patterns in the idiosyncratic errors, $v_{it}$. All experiments are implemented using $R=2,000$ replications. Details of of the 12 experiments are summarized in Table S-1 of the online supplement B (MC results).
Subsection (ref) considered consistent estimation of the variance of $\tilde{\phi}_{nT}$ using $\mathbf{\tilde{V}}_{u}$ a threshold estimator for $\mathbf{V}_{u},$ given by equation ((ref))$.$ For comparison purposes we also considered two other estimators of $\mathbf{V}_{u}.$ These were the sample covariance matrix $\mathbf{\hat{V}}_{u}=\sum_{t=1} ^{T}\mathbf{\hat{u}}_{t}\mathbf{\hat{u}}_{t}^{^{\prime}}/T$ and a diagonal covariance matrix, where the off-diagonal elements of $\mathbf{\hat{V}}_{u},$ $\hat{\sigma}_{ij},$ are set to zero. Thus we have three designs for the return error covariance matrix, $\mathbf{V}_{u},$ and three different estimators of it. A comparison of the results for the different covariance matrices is available on request. The diagonal estimator, as to be expected, performed poorly when the true covariance matrix was not diagonal, particularly for the spatial error covariance matrix and when the strength of the missing factor was close to $1/2$. For these designs the sample and threshold estimators of the covariance matrix generally performed similarly and given that there is a theoretical justification for the threshold estimator and there are structures of the error covariance matrix for which the sample estimator is unlikely to perform well we report the results using the threshold estimator in the simulations below.
We focus on a comparison of two-step (defined by ((ref)) and the bias-corrected (BC) estimator (defined by ((ref))), and report bias, root mean square error (RMSE) and size for testing $H_{0j}\,:\phi_{0k}=0,$ $k=M,H,S$ at the five per cent nominal level, for all $n=100,500,1,000,3,000$ and $T=60,$ $120,$ $240$ combinations. The results for all 12 experiments are summarized in Tables S-A-E1 to S-A-E12 in the online supplement B. In terms of bias and RMSE the two-step estimator does much better than the bias-corrected (BC) estimator when $T=60$ and $n=100$, but this gap closes quickly as $n$ is increased. In fact for $T=60$ and $n=3,000,$ the bias and RMSE of the BC estimator (at $0.0010$ and $0.0607)$ are much less than those of the two-step estimator (at -$0.0080$ and $0.1489)$. This pattern continues to hold when $T=120$ and $240$. Bias correction can cause the RMSE to "blow up" for small samples, such as $n=100$ and $T=60$, but for $n=500$ and above the bias-corrected estimator always has a smaller RMSE than the two-step estimator. As discussed in the theoretical section, having a large $n$ is important for the properties of the estimators.
But most importantly, the two-step estimator is subject to substantial size distortions, particularly when $T$ is small relative to $n$. As predicted by the theory, the degree of over-rejection of the tests based on FM estimator falls with $T,$ but increases with $n$. For example, the two-step test sizes rise from $11.1\%$ when $T=60$ and $n=100$ to $60.9\%$ when $T=60$ and $n=3,000$. Increasing $T$ reduces the size distortion of the two-step estimator but test sizes are still substantially above the $5\%$ nominal value when $n$ is large. The strong tendency of the tests based on the two-step\ estimator to over-reject could be an important contributory factor leading to false discovery of a large number of apparently significant factors in the literature. In contrast, sizes of the tests based on the BC estimator, using the variance estimator given by ((ref)), are all close to its nominal value, irrespective of the factor strength or sample size combinations. We only note some elevated test sizes in the case of the experimental design 12, and when we consider the semi-strong factors. The highest test size of $7.85$ per cent is obtained for the least strong factor, $f_{st}$, when $n=3000$ and $T=60$. See Table S-A-E12 of the online supplement B.
We also experimented with raising $\alpha_{\eta}$ from $0.3$ to $0.5,$ making the pricing errors much more pervasive and $\rho_{\varepsilon}$ from $0.5$ to $0.85,$ introducing more spatial correlation$.$ This increased the rejection rate in experiments 9 and 10.
Plots of the empirical power functions for testing the null hypothesis $H_{0j}\,:\phi_{0k}=0,$ $k=M,H,S$, are also provided in the online supplement B for the 12 experiments (Figures S-A-E1 to S-A-E12). All power functions have the familiar bell curve shape and tend to unity as $n$ and/or $T$ are increased, showing the test has satisfactory power, particularly for $n$ sufficiently large even when $T=60$. Again the power functions are quite similar across the 12 different experiments and show similar patterns for strong and semi-strong factors. However, this similarity hides the fact that the test of $\phi_{k}=0$ for the strong factor is much more powerful than corresponding tests for the semi-strong factors, with the test power declining as factor strength is reduced.
These differences in the effects of factor strength on the power of the test of $\phi_{k}=0$ are in line with our theoretical results, and are also reflected in the rate at which the RMSE of the estimators of $\phi_{M}$, $\phi_{H}$ and $\phi_{S}$ fall with $n$. For example, using results in Tables S-B-E10-12 in the online supplement B for the bias-corrected estimator in the case of design 12 with $T=240,$ we note that the ratio of RMSE of $n=3,000$ to $n=100$ is $17\%$ for the strong factor $\left( \alpha_{M}=1\right) $, $23\%$ for the first semi-strong factor with $\alpha_{H}=0.85$, and $36\%$ for the second semi-strong factor with $\alpha_{S}=0.65.$ As strength falls one needs larger cross section samples of securities to attain the same level of precision. In the case of the two-step estimator there was the same pattern, but the fall in the RMSE with $n$ was much slower. The ratio for $T=240$ of RMSE of $n=3,000$ to $n=100$ is $25\%$ for $\phi_{M}$, rather than $17\%$ (for the bias-corrected estimator); $48\%$ rather than $23\%$ for $\phi_{H}$.
So far, we have assumed that the DGP is correctly specified with two semi-strong and no observed weak factors. Here we consider the implications of incorrectly excluding semi-strong factors or correctly including weak factors on the small sample properties of the bias-corrected\ estimator of $\phi_{M},$ the coefficient of the strong factor. Using the same DGP (which includes one strong factor and two semi-strong factors), we carried out additional MC experiments (designs 1-12) where we also estimated $\phi_{M}$ without the semi-strong factors being included in the regressions. Comparative results, with and without the semi-strong factors, are summarized in Tables S-C-E1-3 to S-C-E10-12 of the online supplement B. We find that incorrectly excluding semi-strong factors can be quite costly, both in terms of bias and RMSE as well as size distortions. In terms of RMSE it was almost always better to estimate the model with the semi strong factors included. The exception was for the case of $T=60,$ $n=100,$ where including the semi-strong factors caused the RMSE to blow up. Size distortions resulting from the exclusion of the semi-strong factors tended to be more pronounced for large $n$ and $T$ samples. These conclusions were not sensitive to the choice of the experimental design. For these experiments the lesson seems to be that it is important to have $n$ large and include relevant semi-strong factors provided that they are sufficiently strong.
When the DGP includes one strong factor ($\alpha_{M}=1$) and two weak factors ($\alpha_{H}=$ $\alpha_{S}=0.5)$, in terms of the bias and RMSE for $\phi_{M}$ it is unambiguously better to exclude the weak factors from the regression, even though they are in the DGP. Weak factors are best treated as missing and absorbed in the error term.
The conclusions from the Monte Carlo simulations are that the bias corrected estimator of $\boldsymbol{\phi}_{0}\boldsymbol{,}$ generally works well. Although, it can generate a large RMSE for small $n$ and $T$, this can be solved by increasing $n.$ This performance is robust to non-Gaussian errors, GARCH effects, missing weak factors, pricing errors and weak cross-sectional dependence. Test sizes are generally correct and the power good. The rate at which RMSEs decline with $n$ depends on the strength of the underlying factors. Semi-strong factors need much larger values of $n$ for precise estimation. Tests of the joint significance of $\phi_{M}=\phi_{H}=\phi_{S}=0,$ not reported here, also performed well, as might be expected given the good power performance of the separate induced tests which are reported. Including weak factors could be harmful, but there are potential advantages of adding semi-strong factors, although the issue of how best to select such factors is an open question to which we now turn.
Our theoretical derivations and Monte Carlo simulations both assume that the number of risk factors included in the return regressions is fixed and the factors are known. For the empirical application we face the additional challenge of selecting a small number of relevant risk factors from a possibly large number of potential factors, $m$. This problem has been the subject of a number of recent studies. harvey2016and propose a multiple testing approach aimed at controlling the false discovery rate in the process of factor selection, and in a more recent paper harvey2020false suggest using a double-bootstrap method to calibrate the t-statistic used for controlling the desired level of the false discovery rate. giglio2021asset suggest applying the double-selection Lasso procedure by belloni2014inference to second pass regressions. None of these methods distinguish between strong, semi-strong or weak factors in their selection process, whereas the theory and simulations presented above indicate the importance of factor strength for estimation and inference.
Here we propose an alternative selection procedure where we first estimate the strength of all the $m$ factors under consideration, and then select factors with strength above a given threshold, the value of which is informed by the convergence results of Theorem (ref). This theorem showed that if factor $f_{kt}$ has strength $\alpha_{k},$ then for a given $T$ the BC estimator of $\phi_{k}$ converges to its true value, $\phi_{0k}$, at the rate of $n^{(\alpha_{k}+\alpha_{\min}-1)/2}$, where $\alpha_{\min}=\min_{i}(\alpha _{i})$. As is recognized in non-parametric estimation literature, if the rate of convergence is less than $1/3,$ the gain in precision with $n$ is so slow that the estimator may not be that useful.\footnote{The manski1985semiparametric maximum score estimator for a binary response model has $n^{1/3}$ convergence and this is regarded as very slow and there are suggested modifications such as horowitz1992smoothed to increase the rate of convergence to $n^{2/5}$.} To achieve rate of $n^{1/3}$ we need to set the threshold value of $\alpha_{k}$, denoted by $\underline{\alpha}$, such that $\underline{\alpha}+\alpha_{\min}>1+2/3$. The smallest value of such a threshold is obtained when $\alpha_{\min }=\underline{\alpha}$ or if $\underline{\alpha}>1/2+1/3.$ Given the threshold the main issue is how to estimate factor strength. This problem is already addressed in \citet*[BKP]{bailey2021measurement} when $m$ is fixed. In this setting they base their estimation on the statistical significance of $f_{kt}$ in the first stage time-series regressions of excess returns on all the factors under consideration, whilst allowing for the $n$ multiple testing problem which their approach entails. When $m\,$(the number of factors) is also large the first stage regressions will also be subject to the multiple testing problem and penalized regression techniques such as Lasso or the one covariate at a time (OCMT) selection procedure technique proposed by chudik2018one could be used. Irrespective of selection technique used at the level of individual security returns, we end up with $n$ different subsets of the $m$ factors under consideration. Factor strengths can then be estimated similarly to BKP\ from their selection frequencies across the $n$ securities.
To be more specific, denote the set of $m$ factors under consideration by $\mathcal{S}$ and denote the set of selected factors for security $i$ by $\hat{S}_{i}$ and their numbers by $\hat{m}_{i}=|\hat{S}_{i}|$. Clearly $\hat{S}_{i}\subseteq\mathcal{S}$, and $\hat{m}_{i}\leq m$, for $i=1,2,...,n$. Then compute the proportion of stocks in which the $k^{th}$ factor is selected, $\hat{\pi}_{k},$ for $k\in\{1,2,...,m\}$ based on $\hat{S}_{i},$ $i=1,2,...,n$ by $\hat{\pi}_{k}=\frac{1}{n}\sum_{i=1}^{n}\mathcal{I}\{k\in \hat{S}_{i}\}$. Then the strength of the $k^{th}$ factor is measured by
The transformation from $\hat{\pi}_{k}$ to $\hat{\alpha}_{k}$ is explained and justified in BKP, where it is shown that considering strength aids interpretation because it is not dependent on $n$. It is beyond the scope of the present paper to provide theoretical justification for the proposed factor selection procedure, but using extensive Monte Carlo experiments yoo2022factor has shown that the proposed method has desirable small sample properties whether Lasso or OCMT is used for factor selection at the level of individual security returns.
This section uses the results above in the explanation of monthly returns for a large number, $n,$ of U.S. securities, by a large active set of $m$ potential risk factors. We first briefly describe the sources and characteristics of the data for the stock returns and factors, which cover different sub-samples over the period $1996m1-2022m12.$ We then consider the selection of a subset of $K$ factors from the active set. Finally we test $\boldsymbol{\phi}_{0}=0$, and construct and evaluate phi-portfolios and corresponding mean-variance portfolios for alternative models.
Monthly returns (inclusive of dividends) for NYSE and NASDAQ stocks from CRSP with codes $10$ and $11$ were downloaded from Wharton Research Data Services. They were converted to excess returns by subtracting the risk free rate, which was taken from Kenneth French's data base. To obtain balanced panels of stock returns and factors, only variables for which there was data for the full sample under consideration were used. Excess returns are measured in percent per month. To avoid outliers influencing the results, stocks with a kurtosis greater than 16 were excluded. To examine factor selection, four samples were considered, each had $20$ years of data, $T=240$, ending in $2015m12$, $2017m12$, $2019m12$, $2021m12$. Filtering out the stocks with kurtosis larger than $16$ removed about $100$ of the roughly $1200$ stocks. The number of stocks ($n$) considered for each of the four $T=240$ samples are given in panel A of Table (ref). Further detail is given in Section (ref) of the online supplement A.\footnote{Summary statistics for the excess returns across the different samples are given in Table (ref) of the online supplement A.} For analysis of the phi-portfolios, the factors selected in the sample ending in $2015m12$ were used to construct portfolios up to $2022m12.$
For factor selection we used a sample of $T=240$ observations\footnote{Some results for $T=120$ are included in the online supplement.}. The set of factors considered combine the $5$ Fama-French factors with the $207$ factors from the chen2022open , Open Source Asset Pricing webpages, both downloaded July 6 2022. Only factors with data for the full sample were considered so the return regressions constitute a balanced panel. The number of factors in each of the four $20$ year samples ending in the years $2015$, $2017$, $2019$, $2021$ is also given in panel A of Table (ref), and range between $187$ to $199$. Summary statistics for the factors in the active set are given in (ref) of the online supplement A.
To implement the factor selection procedure set out in Section (ref), Lasso is used to carry out selection in the return regressions for individual securities and we refer to the factor selection procedure as pooled Lasso (PL). As is well known, Lasso does not work well with too many highly correlated regressors, therefore, factors with an absolute correlation with the market factor greater than $0.70$ were dropped. This still left between $177$ and $190$ risk factors in the active set $\mathcal{S}$, depending on the sample period (see panel A of Table (ref)).
Specifically, Lasso was applied $n$ times to the regressions of excess returns, $r_{it}=R_{it}-r_{t}^{f}$, for $i=1,2,...,n$, on the $177$ to $190$ factors in the active set, $\mathcal{S}$, to select the sub-set $\hat{S}_{i}$ for each $i$ over the four 20-year samples, separately. Following the literature, the tuning parameters in the Lasso algorithm were set by ten-fold cross-validation.\footnote{The post-Lasso and one covariate multiple testing (OCMT) approach of \citet *[CKP]{chudik2018one}. were also investigated, but Lasso seemed to work reasonably well. The details of the Lasso procedure used are given in Section 2.2 of the online supplement of CKP paper.} Interestingly, the market factor was selected by Lasso for almost all the securities, thus confirming the pervasive nature of the market factor. No other factor came close to being selected for all the securities. Lasso tended to choose a lot of non-market factors and every factor got chosen in at least one return regression. The mean number of non-market factors chosen by Lasso fell from $11.6$ in the $2015$ sample to $9.9$ in the $2021$ sample. The median was lower, falling from $10$ to $8$. There was a long right tail because Lasso tended to choose a very large number of non-market factors for some securities, ranging from a maximum of $48$ in the 2021 sample to $54$ in the $2017$ sample.
Apart from the market factor, there are no systematic patterns for the rest of selected factors in the return equations for the individual securities. Following the theory set out in Section (ref), the $K$ factors used to estimate $\boldsymbol{\phi}$ are chosen on the basis of their factor strength.\footnote{The idea of using factor strength could also be viewed as a kind of averaging of the factors selected in individual return regressions.} A minimum threshold of $0.7$ was used. This is below the threshold value of $\underline{\alpha}=1/2+1/3$ required to achieve the convergence rate of $n^{1/3}$, and is intended to capture borderline semi-strong factors. We also consider the values of $0.75,$and $0.80$ that are quite close to $\underline{\alpha}$. The number of selected factors for different choices of factor strength threshold is given in panel B of Table (ref), for the four different samples.
Using the threshold of $0.70$, $17$ factors (inclusive of the market factor) were selected for the sample ending in $2015$, with the number of selected factors declining to $15,13$ and $11,$ for the samples ending $2017,$ $2019$ and $2021$, respectively. At the other extreme, setting the threshold at $0.80$, the number of selected factors dropped to $4$ for the samples ending in $2015$, $2017$ and $2021$, and $2$ for the sample ending in $2021$. Since $17$ factors seemed too large and $2$ factors too small, the threshold value was set at the intermediate value of $0.75$. We considered always conditioning on the market factor, but since Lasso almost always selected it, this was unnecessary.
\floatstyle{plaintop} \restylefloat{table}
\onehalfspacing The list of factors selected by pooled Lasso for the four samples are given in Table (ref). The three Fama-French factors, Market, HML and SMB, are all selected in all four periods.\footnote{We also considered selecting the risk factors using the generalized one covariate at a time (OCMT) method proposed by sharifvaghefi2022variable . Using GOCMT the Fama-French three factors were again amongst the five strongest factors selected. The use of GOCMT\ for factor selection is also investigated by yoo2022factor , using Monte Carlo and empirical applications.} Of the Fama-French three, only the market factor is strong, with estimated strength in excess of $0.98$ across the four periods.\footnote{These results also support the choice of 3 Fama-French factors and their strength used in our Monte Carlo simulations.} The other factor which is selected across all the four periods is "short selling". This is proposed by dechow2001short who argue that short-sellers target firms that are priced high relative to fundamentals. It measures the extent to which investors are shorting the market as reflected in Compustat data. Two additional factors are selected in periods ending in $2019$ and earlier. One is "Beta Tail Risk" proposed by kelly2014tail which estimates a time-varying tail exponent from the cross section of returns. The other is "Cash Based Operating Profitability" (CBOP) suggested by ball2016accruals . This is operating profit less accruals, with working capital and R&D adjustments. For periods ending in 2017 and 2015 the "Sin Stock" indicator proposed by hong2009price is also selected. It takes the value of unity if the stock in question is involved in producing alcohol, tobacco, and gaming. They find that such stocks are held less by norm-constrained institutions such as pension plans.
The factor strengths are relatively stable across the periods, with many of the estimates close to the threshold value of $0.75$. Apart from the market factor only SMB, Short Selling and Beta Tail Risk factors have strengths in excess of $0.85$ when averaged across the four periods. From the large number of factors in the active set we have ended up with relatively few factors that are reasonably strong and for which $\boldsymbol{\phi}_{0}$ can be estimated reasonably accurately.\footnote{We do not report estimates of individual $\phi_{k}.$ Because of correlations between the loadings, the sign, size and significance of the coefficients are difficult to interpret and for phi-portfolio construction, discussed below, what matters is $\boldsymbol{\phi }^{\prime}\boldsymbol{\phi}$ which determines the return on the portfolio.}
\singlespacing
\onehalfspacing
The strengths of the selected factors are also closely related to the average measures of fit often used in the literature. Here we consider both $Ave\bar{R}^{2}=n^{-1}\sum_{i=1}^{n}\bar{R}_{i}^{2}$, a simple average of the fit of the individual return regressions adjusted for degrees of freedom, $\bar{R}_{i}^{2}=1-(T-K-1)^{-1}\sum_{t=1}^{T}\hat{u}_{it}^{2}/T^{-1}\sum _{t=1}^{T}\left( r_{it}-\bar{r}_{i\circ}\right) ^{2}$, and the adjusted pooled $R^{2}$ defined by $\overline{PR}^{2}=1-\widehat{\bar{\sigma}}_{nT} ^{2}/s_{r,nT}^{2}$, where $\widehat{\bar{\sigma}}_{nT}^{2}$ is the bias-corrected estimator of $\bar{\sigma}_{n}^{2}$ defined by ((ref)) and $=$ $\left( nT\right) ^{-1}\sum_{i=1}^{n}\sum_{t=1}^{T}\left( r_{it}-\bar{r}_{i\circ}\right) ^{2}$. Both of these measures behave very similarly, but the pooled version is less sensitive to outliers. As shown in Appendix (ref), for sufficiently large $n$ and $T$, $\overline{PR}^{2}$ is dominated by the contribution of the most strong factor(s). Since the only strong factor selected is the market factor in Table (ref)\ we report the $Ave\bar{R}^{2}$ and $\overline {PR}^{2}$ in the case of return regressions which just include the market factor and those which include all other factors with strength in excess of $0.75$. First, we note that the $\overline{PR}^{2}$ values are generally lower than the $Ave\bar{R}^{2}$. Second, the additional factors do add to the fit, but their relative contributions vary considerably across sample sizes and periods. In general, the marginal contribution of non-market factors tend to be smaller when $T$ is larger, which is consistent with the theory for adjusted pooled $R^{2}$ set out in Section (ref) of the online supplement A.
\singlespacing
\onehalfspacing
In principle, if $\boldsymbol{\phi}\neq0$ there are potentially exploitable excess returns. In practice, to construct an effective phi-portfolio a large number of securities is required and rebalancing such long-short portfolios for so many securities may not be feasible or may incur high transactions costs. In addition, model uncertainty, estimation uncertainty, time variation in both $\boldsymbol{\beta}_{i}$ and in conditional volatility pose additional difficulties in implementing a strategy to exploit the potential returns revealed by $\boldsymbol{\phi}$. We will abstract from such practical difficulties to provide some indication of the performance of phi-portfolios relative to alternatives which would face similar difficulties.
As our preferred asset pricing model, we consider the seven factors selected by pooled Lasso using the sample ending in $2015m12,$ which we label as PL7.\footnote{The selection of the PL7 model was reported in the earlier version of the paper submitted for publication and was not informed by the performance the phi-portfolio that we report in this version of the paper.} Recall that the PL7 includes the 3 Fama-French factors (Mkt., SMB, and HML) plus the four risk factors, Short Selling, CBOP, Beta Tail Risk, and Sin Stocks. But given uncertainties that surround the problem of model selection we also considered the two popular FF factor models, namely FF3 and FF5. The latter augments FF3 with RMW (robust minus weak operating profitability) and CMA (conservative minus aggressive investment portfolios). The three models are estimated using twenty-year rolling windows covering the $84$ months from $2015m12$ to $2022m11$, so that we can generate out of sample return forecasts for the months $2016m1-2022m12$.\footnote{Due to entry and exit of securities the number of securities included in our analysis varied across the rolling sample periods. We started with $n=1,090$ securities for the first rolling sample ending in $2015m12$, with the number of available securities with $240$ months of data falling to $953$ by $2017$, $838$ by $2019$, $767$ by $2021$, and $736$ by $2022$.} For all $3\times84$ model-sample interactions we computed the following Wald test statistics \[ W_{t\left\vert T\right. }^{2}=\boldsymbol{\tilde{\phi}}_{t\left\vert T\right. }^{\prime}\left[ \widehat{Var\left( \boldsymbol{\tilde{\phi} }_{t\left\vert T\right. }\right) }\right] ^{-1}\boldsymbol{\tilde{\phi} }_{t\left\vert T\right. }, \] for testing the null hypothesis $\boldsymbol{\phi}=0,$ where $\boldsymbol{\tilde{\phi}}_{t\left\vert T\right. }$ and $\widehat{Var\left( \boldsymbol{\tilde{\phi}}_{t\left\vert T\right. }\right) }$ denote the rolling versions of ((ref)) and ((ref)).\footnote{The formulae for the rolling estimates are provided in the sub-section (ref) of the online appendix.} For the PL7 the range of the test statistic was from 141.8 to 25.8, as compared to the 5 per cent $\chi^{2}(7)$ critical value of 14.07. Thus the hypothesis that $\boldsymbol{\phi}=0$ is strongly rejected in all the 84 rolling sample for PL7. This is also true for the FF5 and FF3 models where the test statistic ranged from 162.7 to 25.0, and from 67.9 to 25.8, compared to the 5 per cent critical values of 11.07 and 7.8, respectively.
The rolling values of the Wald statistics for testing $\boldsymbol{\phi}=0$ for the three models are shown in Figure 1. The horizontal line (pink) represents the critical value of the $\chi_{7}^{2}$ distribution at the 5 per cent level. The time profiles of these test statistics clearly show that $\boldsymbol{\phi}=0$ is rejected for all rolling samples and for all three models. But there is also a clear downward trend showing that the evidence against $\boldsymbol{\phi}=0$ has been getting weaker over time, irrespective of the choice of asset pricing model.
Having established that most likely $\boldsymbol{\phi\neq0}$ for the asset pricing models we have considered, we now turn to the performance of phi-portfolios based on these models. Using the recursive version of the phi-portfolio given by ((ref)), we consider the following phi-portfolio returns
\[ \hat{\rho}_{t+1,\phi}=\boldsymbol{\tilde{\phi}}_{t\left\vert T\right. }\left[ \left( \mathbf{\hat{B}}_{t\left\vert T\right. }^{\prime} \mathbf{M}_{n}\mathbf{\hat{B}}_{t\left\vert T\right. }\right) ^{-1} \mathbf{\hat{B}}_{t\left\vert T\right. }\mathbf{M}_{n}\mathbf{r}_{\circ ,t+1}-\mathbf{f}_{t+1}\right] , \] for $t=2016m1,2016m2,...,2022m12$, using the rolling estimates $\mathbf{\hat {B}}_{t\left\vert T\right. }=\left( \mathbf{\hat{\beta}}_{1t\left\vert T\right. },\mathbf{\hat{\beta}}_{2t\left\vert T\right. }....,\mathbf{\hat {\beta}}_{nt\left\vert T\right. }\right) ^{\prime}$ with $T=240$, for each of the three factor models, FF3, FF5 and PL7. We compare the annualised Sharpe ratios of phi-portfolios with the ones based on associated MV\ portfolios, given by $\rho_{t+1,MV}=\boldsymbol{\mu}_{R}^{\prime}\mathbf{V}_{R} ^{-1}\mathbf{r}_{\circ,t+1}$.\footnote{Given our focus on the Sharpe ratios, we have set the scaling of the MV portfolio to unity.} Although in principle, MV portfolios can be constructed without a reference to a particular factor model, reliable estimation of $\boldsymbol{\mu}_{R}$ and $\mathbf{V}_{R}^{-1}$ are challenging when $n$ is relatively large. For example, the rolling sample covariance matrix estimator of $\mathbf{V}_{R}$, given by $\mathbf{\mathring {V}}_{R,t\left\vert T\right. }=T^{-1}\sum_{\tau=t-T+1}^{t}\left( \mathbf{r}_{\circ,\tau}-\mathbf{\bar{r}}_{\circ,t\left\vert T\right. }\right) \left( \mathbf{r}_{\circ,\tau}-\mathbf{\bar{r}}_{\circ,t\left\vert T\right. }\right) ^{\prime}$, with $\mathbf{\bar{r}}_{\circ,t\left\vert T\right. }=T^{-1}\sum_{\tau=t-T+1}^{t}\mathbf{r}_{\circ,\tau}$, will be singular when $n>T$, and can be very poorly estimated if $T$ is not sufficiently large relative to $n$. There is a vast literature on consistent estimation of high dimensional covariance matrices like $\mathbf{V}_{R}.$ fan2011high use observed factors while fan2013large use principal components to filter out the effects of strong factors, in both cases assuming $\mathbf{V}_{u}$ is sparse, and then using a threshold method to estimate it. Shrinkage estimators of $\mathbf{V}_{R}$ are also proposed in the literature with a recent survey provided by LedoitWolfsurvey2022. However, the shrinkage estimators require $n$ and $T$ to be of the same order of magnitude and do not work well when $n$ is much larger than $T$, as in the present application. We follow fan2011high and base our estimation of $\boldsymbol{\mu}_{R}$ and $\mathbf{V}_{R}$ on the same factor model used to construct the phi-portfolio. For a given factor model, characterized by $c\mathbf{,B}$, and $\mathbf{F}$, we compute the MV portfolio returns as
\[ \hat{\rho}_{t+1.MV}=\mathbf{\hat{\mu}}_{R,t\left\vert T\right. }^{\prime }\mathbf{\hat{V}}_{R,t\left\vert T\right. }^{-1}\mathbf{r}_{\circ,t+1}, \] where $\mathbf{\hat{\mu}}_{R,t\left\vert T\right. }=\hat{c}_{t\left\vert T\right. }\mathbf{\tau}_{n}+\mathbf{\hat{B}}_{t\left\vert T\right. }\boldsymbol{\tilde{\lambda}}_{t\left\vert T\right. }$, $\boldsymbol{\tilde {\lambda}}_{t\left\vert T\right. }=\boldsymbol{\tilde{\phi}}_{t\left\vert T\right. }+\boldsymbol{\hat{\mu}}_{t\left\vert T\right. },$ \[ \mathbf{\hat{V}}_{R,t\left\vert T\right. }=\mathbf{\hat{B}}_{t\left\vert T\right. }^{\prime}\left( \frac{\mathbf{F}_{t\left\vert T\right. }^{\prime }\mathbf{M}_{T}\mathbf{F}_{t\left\vert T\right. }}{T}\right) ^{-1} \mathbf{\hat{B}}_{t\left\vert T\right. }\mathbf{+\mathbf{\tilde{V}} }_{u,t\left\vert T\right. }\mathbf{,} \] $\boldsymbol{\hat{\mu}}_{t\left\vert T\right. }=\boldsymbol{\bar{f} }_{t\left\vert T\right. }=T^{-1}\sum_{\tau=t-T+1}^{t}\mathbf{f}_{\tau}$, and $\mathbf{\mathbf{\tilde{V}}}_{u,t\left\vert T\right. }$ is the rolling estimate of $\mathbf{V}_{u}$. The algorithms used to compute the recursive estimates for the MV portfolio can be found in the sub-section (ref) in the online supplement A.
Table (ref) presents annualised SR of the phi-portfolios, for the FF3, FF5, and PL7 models, and their corresponding MV portfolios for two samples, both beginning in 2016m1, one ending in 2019m12, pre Covid-19, and one ending in 2022m12. \ In five of the six SR ratios reported in this table, the phi-portfolio has a higher SR than the corresponding MV portfolio. The exception is the SR associated to the FF5 model for the pre Covid-19 sample. This illustrates that if $\boldsymbol{\phi}\neq\mathbf{0},$ it is possible to construct portfolios that outperform the mean-variance portfolio. Amongst the 3 models considered, the phi-portfolio based on the PL7 model performed best during the pre Covid-19 and the full sample, even beating the S&P 500. The Sharpe ratio of the phi-portfolio based on the Pl7 model was 1.95 compared to 0.94 for the S&P 500 during the pre Covid-19 period, and fell sharply to 0.65 for the full sample as compared to 0.58 for the S&P 500. The SR for the MV portfolio using the same model were 0.87 and 0.42 for the two samples, respectively.
The sharp decline in the SRs as we add the post Covid-19 years is in line with the strong downward trend in the Wald statistics for the test of $\boldsymbol{\phi}=\mathbf{0}$ shown in Figure 1. As is well known SRs have large standard errors, and in the case of our application that are around 1, so none of the SRs are significantly different from zero, with the possible except of the largest SR of 1.95 for phi-portfolio based on PL7 model for the pre Covid-19 period. We also note that the reported Sharpe ratios do not allow for transaction costs and the fact that shorting might not be feasible for all the securities.
In this paper we have highlighted the importance of decomposing the risk premia, $\boldsymbol{\lambda}$, into the the factor mean, $\boldsymbol{\mu}$, and $\boldsymbol{\phi}$, and writing the alpha of security $i$, $\mathit{\alpha}_{i}$, in terms of $\boldsymbol{\phi}$ and the idiosyncratic pricing errors. We have shown that when $\boldsymbol{\phi\neq0}$, it is possible to construct a portfolio, denoted as phi-portfolio, that dominates the associated mean-variance portfolio when the number securities, $n$, is sufficiently large and the risk factors are sufficiently strong. Given the pivotal role played by $\boldsymbol{\phi}$ for estimating the risk premia, for formation of large $n$ portfolios, and for tests of market efficiency, we have focussed on estimation of $\boldsymbol{\phi}$, and its asymptotic distribution under quite a general setting that allows for missing factors and idiosyncratic pricing errors. Since factor means, $\boldsymbol{\mu}$, can be estimated at the regular rate of $T^{-1/2}$ from time series data, it is relatively straightforward to develop a mixed strategy for estimation of $\boldsymbol{\lambda}$ by adding a time series estimate of $\boldsymbol{\mu}$ to the bias-corrected estimator of $\boldsymbol{\phi}$. If we use the same time series sample, such an estimator reduces to the Shanken bias-corrected estimator of $\boldsymbol{\lambda}$. But in practice, given the concern over the instability of the factor loadings $\beta_{ik}$ over time, one could use relatively long time series, say $T_{\mu}$, when estimating $\boldsymbol{\mu} $, and a shorter time series, say $T_{\phi}<T_{\mu}$, when estimating $\boldsymbol{\phi}$. The distributional and small sample properties of such a mixed estimator of risk premia is a topic for further research.
Our theoretical and Monte Carlo results further highlight the important role played by factor strengths in estimation and inference on $\boldsymbol{\phi}$, and hence on $\boldsymbol{\lambda}$. For a fixed $T_{\phi}$, factors with strength below $2/3$ lead to estimates of $\boldsymbol{\phi}$ \ with convergence rate of $n^{-1/3}$ or worse, and their use in asset pricing models can be justified only when $n$ is very large. We have also shown that weak factors, with strength below $1/2$ are best treated as missing and absorbed in the error term. We have shown that estimation of $\boldsymbol{\phi}$ for strong or semi-strong factors is robust to weak missing factors, and the explicit inclusion of weak factors in the empirical analysis is likely to have adverse spill over effects on the estimates of $\boldsymbol{\phi\ }$\ for strong and semi-strong factors. In view of these results we have proposed a factor selection procedure where only factors with strength above $1/2+1/3$ are included in the asset pricing model. Developing a formal statistical theory for the proposed selection is another topic for future research.
The paper also provides an empirical application to a large number of U.S. securities with risk factors selected from a large number of potential risk factors according to their strength, and use a pooled Lasso approach to select $7$ risk factors out of over $180$. We find strong statistical evidence against $\boldsymbol{\phi=0}$ for the selected model as well as for the two popular Fama-French models (FF3 and FF5). Using rolling estimates of $\boldsymbol{\phi}$ we also construct phi portfolios with better Sharpe ratios as compared to associated mean-variance portfolios. But we also warn that these portfolio comparisons are preliminary and need to be further investigated by allowing for transaction costs, and the feasibility of the long-short trading strategies that are involved.
\singlespacing