EconBase
← Back to paper

Identifying and exploiting alpha in linear asset pricing models with strong, semi-strong, and latent factors

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

162,166 characters · 29 sections · 101 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Identifying and exploiting alpha in linear asset pricing models with strong, semi-strong, and latent factors

abstractThe risk premia of traded factors are the sum of factor means and a parameter vector we denote by $\boldsymbol{\phi}$ $\ $which is identified from the cross-section regression of $\alpha_{i}$ on the vector of factor loadings, $\boldsymbol{\beta}_{i}$. If $\boldsymbol{\phi}$ is non-zero, then $\alpha_{i}$ are non-zero and one can construct "phi-portfolios" which exploit the systematic components of non-zero alpha. We show that for known values of $\boldsymbol{\beta}_{i}$ and when $\boldsymbol{\phi}\ \ $is non-zero there exist phi-portfolios that dominate mean-variance portfolios. The paper then proposes a two-step bias corrected estimator of $\boldsymbol{\phi}$ \ and derives its asymptotic distribution allowing for idiosyncratic pricing errors, weak missing factors, and weak error cross-sectional dependence. Small sample results from extensive Monte Carlo experiments show that the proposed estimator has the correct size with good power properties. The paper also provides an empirical application to a large number of U.S. securities with risk factors selected from a large number of potential risk factors according to their strength and constructs phi-portfolios and compares their Sharpe ratios to mean variance and S&P portfolios.

JEL Classifications: { C38, G10}

Key Words: F{ actor strength, pricing errors, risk premia, missing factors, pooled Lasso, mean-variance and phi-portfolios.} \thispagestyle{empty}

\pagenumbering{arabic} \onehalfspacing {0pt} {0pt} {0pt} {0pt}

Introduction

There are two approaches to the estimation of risk premia and testing of market efficiency, often referred as the beta and the SDF (stochastic discount factor) methods, see Jagannathan2002. This paper adopts the beta method, and following the literature uses a linear factor pricing model (LFPM) to explain the time series of excess returns on individual securities, $r_{it}=R_{it}-r_{t}^{f}$, where $R_{it}$ is the return and $r_{t}^{f}$ the risk free rate for $i=1,2,...,n;$ $t=1,2,....,T,$ by a set of observed tradable risk factors. We use individual securities rather than portfolios since, as we will show, if risk factors are not strong, large $n$ is required for accurate estimation. Ang2020 discuss the general issues in the choice between portfolios and individual stocks. Pesaran2023 discuss both the use of portfolios and the relationship between the SDF and LFPM approaches.

The LFPM\ explains the excess return on each security, $r_{it},$ by an intercept, $\mathit{\alpha}_{i}$, labelled alpha, and a $K\times1$ vector of traded risk factors, $\mathbf{f}_{t}=(f_{1t},f_{2t},...,f_{Kt})^{\prime}$ with loadings, $\boldsymbol{\beta}_{i}$:

equation[equation omitted — 110 chars of source]

Under the Arbitrage Pricing Theory (APT) due to ROSS1976341 , we have

equation[equation omitted — 117 chars of source]

where $\eta_{i}$ is a firm-specific idiosyncratic pricing error. Ross allowed $\eta_{i}$ to be non-zero for some $i$ but required that these errors are bounded in the sense that $\sum_{i=1}^{n}\eta_{i}^{2}<C<\infty$. We relax this condition and impose weaker restrictions on $\eta_{i}$'s as discussed in sub-section (ref). But the paper's main contribution lies in the fact that even if there are no idiosyncratic pricing errors, it is still possible to have non-zero $\mathit{\alpha}_{i}$. To see this, taking unconditional expectations of ((ref)) and using ((ref)) we note that \[ E(r_{it})=\mathit{\alpha}_{i}+\boldsymbol{\beta}_{i}^{\prime}\mathbf{\mu }=c+\boldsymbol{\beta}_{i}^{\prime}\boldsymbol{\lambda}+\eta_{i}. \] where $E(\mathbf{f}_{t})=\boldsymbol{\mu}$, which in turn yields

equation[equation omitted — 170 chars of source]

We denote $\boldsymbol{\lambda}-\boldsymbol{\mu}$ by $\boldsymbol{\phi}$, and consider the implications of a non-zero $\boldsymbol{\phi}$ for portfolio optimization. We show that when $\boldsymbol{\phi\neq0}$, one can construct what we call "phi-portfolios" that exploit the systematic components of $\mathit{\alpha}_{i}$, as captured by the non-zero elements of $\boldsymbol{\phi}\ $\ in ((ref)).\footnote{We are grateful to one of the anonymous reviewers for suggesting the idea of phi-portfolios, as a way of motivating and interpreting the role of $\boldsymbol{\phi}$ in asset pricing models.} In contrast, the idiosyncratic pricing errors, $\eta_{i}$, cannot be exploited, as it is likely to involve insider trading. But our proposed phi-portfolios can be implemented and applies irrespective of whether $\eta_{i}=0$ or not.

The extent to which alpha can be exploited is discussed in sub-section ((ref)) where it is shown that $\sum_{i=1}^{n}\left( \mathit{\alpha }_{i}-\mathit{\bar{\alpha}}\right) ^{2}$ depends on the magnitude of $\boldsymbol{\phi}^{\prime}\boldsymbol{\phi}$ and the strength of the risk factors. Since this alpha can be exploited by the construction of the phi-portfolios, this is not consistent with a no arbitrage condition. While it is true that $E(r_{it})=\boldsymbol{\beta}_{i}^{\prime}\boldsymbol{\lambda}$, estimating $\boldsymbol{\lambda}$\ does not tell us whether $\boldsymbol{\phi }\neq0,$\ which can be exploited to obtain Sharpe ratios that are larger than the Sharpe of the tangency portfolio. The claim that tangency portfolio has the highest Sharpe ratio rests on the implicit assumption that $\boldsymbol{\phi=0}$, even if there are no idiosyncratic pricing errors ($\eta_{i}=0$ for all $i$).

The focus of much of the literature has been on testing $\mathit{\alpha} _{i}=0$, or on estimating the risk premium $\boldsymbol{\lambda}$. But given that estimating a non-zero $\boldsymbol{\phi}$ \ provides a way to identify and exploit the alpha in a linear factor pricing model for large $n,$ $\boldsymbol{\phi}$ is an interesting parameter in itself, which will be the focus of this paper. Furthermore, the tangency condition used to construct mean-variance (MV) portfolios has the implicit assumption that $\mathit{\alpha }_{i}=0$, and the MV portfolio implied by the tangency condition is not efficient unless $\boldsymbol{\phi=0}$. If $\boldsymbol{\phi\neq0}$ there is no guarantee that the tangency portfolio will achieve a maximal Sharpe ratio, and is likely to be dominated by the phi-portfolio. Specifically, abstracting from model and parameter uncertainties, we show that when $\boldsymbol{\phi \neq0}$ the Sharpe of the tangency portfolio is bounded in the number of securities, whilst the Sharpe of the phi portfolio rises in $n$ and there exists a sufficiently large $n$ that ensures that the Sharpe of the phi-portfolio will exceed that of the tangency portfolio.

More specifically, first we introduce our proposed phi-portfolio in terms of the factor loadings $\mathbf{B}_{n}=\left( \boldsymbol{\beta}_{1} ,\boldsymbol{\beta}_{2},...,\boldsymbol{\beta}_{n}\right) ^{\prime}$ and $\boldsymbol{\phi}$, and compare its limiting properties (as $n\rightarrow \infty$) with the standard MV portfolio. We show that for known factor loadings\ and non-zero $\boldsymbol{\phi}$ there exist phi-portfolios with strictly positive returns, given by $\boldsymbol{\phi}^{\prime} \boldsymbol{\phi}>0$, that are fully diversified (their variance tends to zero with $n$). The rate at which the variance of phi-portfolio returns tends to zero will depend on the strength of the traded risk factors compared to the strength of the missing (latent) factors, highlighting the importance of factor strengths in portfolio analysis. Since the phi-portfolios have Sharpe ratios that tend to infinity with $n$, they dominate MV portfolios whose Sharpe ratio is bounded in $n$. Note that unlike MV\ portfolios, the phi-portfolios do not require knowledge of the inverse of the covariance matrix of returns, which is particularly difficult to estimate accurately when $n$ is large. As is usual, these theoretical results apply to population values where there are no restrictions on trading many individual securities.

Next we consider estimation of and inference about $\boldsymbol{\phi}$ using a large number of individual securities, taking account of firm-specific pricing errors, traded factors that are not strong, missing latent factors, and panel data sets where the time dimension, $T$, is small relative to $n$. Factor strength plays a central role in our analysis of $\boldsymbol{\phi}$. We use a measure of factor strength, $\alpha_{k},$ developed in bailey2016exponent, bailey2021measurement , which is defined in terms of the proportion of non-zero factor loadings, $\beta_{ik}$.\footnote{Alpha is used both for the LFPM intercepts and the measure of factor strength, because these are established usages, but it will be clear from context and subscripts which is being referred to.} A factor is strong if this proportion is very close to unity, it is semi-strong if $1>\alpha_{k}>1/2$, and it is weak if $\alpha_{k}<1/2$. Use of this measure allows us to be precise about the degree of pervasiveness and show how the strengths of the observed factors, the missing factors and the pricing errors, each influence estimation and inference about $\boldsymbol{\phi}$.

In practice, exploitation of a non-zero $\boldsymbol{\phi}$ requires $n$ to be large and rebalancing such long-short portfolios for so many securities may incur high transactions costs or not be feasible. In addition, model uncertainty, estimation uncertainty, time variation in both $\boldsymbol{\beta }_{i}$ and in conditional volatility pose additional difficulties in implementing a strategy to exploit the potential returns revealed by $\boldsymbol{\phi}$. In developing the theory we will abstract from such practical difficulties, but in the empirical section we illustrate some of these issues with a comparison of the performance of $\boldsymbol{\phi}$ based portfolios relative to MV portfolios which would face similar difficulties.

We estimate $\boldsymbol{\phi}=(\phi_{1},\phi_{2},...,\phi_{K})^{\prime}$, using a two-step estimator. In the first step we estimate the intercepts, $\hat{\alpha}_{i}$, and the factor loadings, $\boldsymbol{\hat{\beta}}_{i}$, from least squares regressions of excess returns on an intercept and risk factors. In the second step, $\boldsymbol{\phi}$ is estimated from the cross section regression of $\hat{\alpha}_{i}$ on $\boldsymbol{\hat{\beta}}_{i}$. As with the two-step estimator of $\boldsymbol{\lambda}$, such a two-step estimator of $\phi_{k}$ will also be biased, and requires bias-correction. Following shanken1992estimation , we consider a bias-corrected version of the two-step estimator of $\boldsymbol{\phi}$, which we denote by $\boldsymbol{\tilde{\phi}}_{nT}$. We develop the asymptotic distribution of $\boldsymbol{\tilde{\phi}}_{nT}$ under quite general set of assumptions regarding the idiosyncratic pricing errors, error cross-sectional dependence, and the presence of missing (latent) factors. The paper also investigates the implications of factor strengths for the precision with which $\boldsymbol{\phi}$ can be estimated. The LFPM, following cham1983arbi , assumes that all the observed factors are strong and the eigenvalues of the covariance matrix of the errors are bounded.

In developing the arbitrage pricing theory, APT, ROSS1976341 , whose concerns were primarily theoretical, assumed the factors had mean zero: $\boldsymbol{\mu}=0,$ so $\boldsymbol{\phi}=\boldsymbol{\lambda}$, is the risk premium. For traded factors under market efficiency, where $\boldsymbol{\phi}=\mathbf{0}$, the risk premium is the factor mean $\boldsymbol{\mu}=\boldsymbol{\lambda}$. Were one interested in estimating $\boldsymbol{\lambda}$ there may be statistical reasons to estimate $\phi_{k}$ and $\mu_{k}$ separately, then summing them to obtain an estimate of $\lambda_{k}$. The factor mean, $\mu_{k}=E(f_{kt})$, can be estimated consistently at the regular $\sqrt{T}$ rate directly using time series data on the risk factors, $f_{kt}$, for $t=1,2,...,T,$ and does not require knowledge of the factor loadings or $n$. In contrast, estimation of $\phi_{k}$ requires panel data to estimate the factor loadings and hence both $n$ and $T$ dimensions are important. In some cases it may be beneficial to use different time series dimensions, $T_{\mu}$ and $T_{\phi},$ to estimate $\mu_{k}$ and $\phi_{k}$, respectively. For example, when factor loadings are subject to breaks it is advisable to use a relatively short sample, and when some factors are not sufficiently strong a large $n$ is required. Thus one could use large $T$ to estimate $\mu_{k}$ and small $T$ to estimate $\phi_{k}$ and $\lambda_{k}$ can be simply estimated by adding the estimates of $\mu_{k}$ and $\phi_{k}.$ We do not pursue the use of different $T$, but if the same $T$ is used to estimate the mean of factor $k,$ say $\hat{\mu}_{k,T},$ and the bias-corrected $\tilde{\phi}_{k,nT}$ then their sum is the same as the shanken1992estimation bias-corrected risk premium for factor $k,$ $\tilde{\lambda}_{k,nT}.$ This decomposition can be used to obtain the rate at which $\tilde{\lambda}_{k,nT}$ converges to its true value, $\lambda_{0,k}=\phi_{0,k}+\mu_{0,k}$.

The main theoretical results of the paper are set out around five theorems under a number of key assumptions, with proofs provided in the mathematical appendix. The small sample properties of $\boldsymbol{\tilde{\phi}}_{nT}$ are investigated using extensive Monte Carlo experiments, allowing for a mixture of strong and semi-strong observed factors, latent factors, pricing errors, GARCH effects and non-Gaussian errors. Small sample results are in line with our theoretical findings, and confirm that the bias-corrected estimator, $\boldsymbol{\tilde{\phi}}_{nT}$, has the correct size and good power properties for samples with time series dimensions of $T=120$ and $T=240$. They also show, in accord with the theory, that the precision with which $\phi_{k}$ is estimated falls with $\alpha_{k}$, the strength of the $k^{th}$ factor.

Our theoretical derivations and Monte Carlo simulations assume that the list of relevant observed factors is known. But in practice the relevant factors must be selected. Extending the theory to the high dimensional case where factors are selected, rather than given, is beyond the scope of the present paper. Since the rate of convergence of $\boldsymbol{\tilde{\phi}}_{nT}$ to its true value is given by $\sqrt{T}n^{(\alpha_{k}+\alpha_{\min}-1)/2}$, and the Monte Carlo confirms the crucial role of factor strengths in estimation and inference on $\boldsymbol{\phi,}$ it seems sensible to select factors on the basis of their factor strength. Weak factors whose strength is around $1/2$ can be ignored and absorbed into the error term.

The above selection procedure is applied in a high dimensional setting with both a large number of securities ($n$ from $1,090$ to $1,175)$ and a large number of potential risk factors ($m$ from $177$ to $189)$, taken from the chen2022open , which can all be traded. We used monthly data over the period $1996m1-2022m12$ and considered balanced panels obtained by including all existing stocks in a given month for which there are $T$ observations. We considered $T=120$ and $240$ months, and focus on the latter which we found to be more reliable, given the large number of securities being considered. Various procedures could be used to select risk factors for a given security, $i$. We used Lasso for this purpose and then selected a subset of these factors that were chosen by a sufficiently large number of securities in the sample, and whose estimated strength were above the given threshold value of $\alpha_{k}>0.75$. We refer to this selection procedure as pooled Lasso. Using this procedure with $T=240$ we ended up with $7$ risk factors for the sample ending in $2015$, declining to $4$ in $2021$. Interestingly, the three Fama-French factors were always included in the set of factors selected by pooled Lasso.

Accordingly, we considered three linear asset pricing models for our portfolio analysis: the pooled Lasso selected at the start of our evaluation sample, denoted by PL7, the Fama-French three factors model, FF3, and the Fama-French five factors model, FF5, which includes two factors not in PL7. The three models are estimated using rolling samples of size $T=240$, starting with a sample ending in $2015m12$ and finishing with a sample ending in $2022m11$. The hypothesis that $\boldsymbol{\phi}=\mathbf{0}$ was rejected for all $84$ rolling samples and all three models, albeit less strongly over the post Covid-19 period. The test results suggested possible unexploited return opportunities, and to investigate this possibility further, we constructed phi-portfolios and compared their Sharpe ratios with the ones based on standard MV portfolios over the full sample evaluation sample, $2016m1-2022m12$, and sample ending $2019m12$, that excludes the Covid-19 period. We find that in five out of the six cases (3 models 2 samples) the phi-portfolio has a higher SR than the corresponding MV portfolio. The exception is the FF5 pre Covid-19. This illustrates that if $\boldsymbol{\phi }\neq0,$ it is possible to construct a portfolio that outperforms the mean variance portfolio. In both the pre Covid-19 sample and the full sample the highest SR was obtained by the PL7 phi-portfolios, which also outperformed the S&P500. The SRs for the sample ending in $2022$ were substantially lower than the sample ending in $2019$, consistent with a falling value of the probability that $\boldsymbol{\phi}$ was non-zero.

Related literature: On estimation of risk premia, following Fama and MacBeth and shanken1992estimation , estimation of risk premia is further examined by shanken2007estimating , kan2013pricing , and BAI201531 . The survey paper by jagannathan2010analysis provide further references.

Testing for market efficiency dates back to jensen1968performance who proposes testing a$_{i}=0$ for each $i$ separately. gibbons1989test provide a joint test for the case where the errors are Gaussian and $n<T$. gagliardini2016time develop two-pass regressions of individual stock returns, allowing time-varying risk premia, and propose a standardised Wald test. raponi2019testing propose a test of pricing error in cross section regression for fixed number of time series observations. They use a bias-corrected estimator of shanken1992estimation to standardise their test statistic. ma2020testing employ polynomial spline techniques to allow for time variations in factor loadings when testing for alphas. feng2022high propose a max-of-square type test of the intercepts instead of the average used in the literature, and recommend using a combination of the two testing procedures. he2022testing propose two statistics, a Wald type statistic which require $n$ and $T$ to be of the same order of magnitude and a standardised t-ratio. kleibergen2009tests considers testing in the case where the loadings are small. pesaran2012testing, pesaran2023testing consider testing that the intercepts in the LFPM are zero when $n$ is large relative to $T$ and there may be non Gaussian errors and weakly cross-correlated errors.

A large number of risk factors have been considered in the empirical literature. We use the fama1993common three factors in our Monte Carlo design. In our empirical application we use the five factors proposed by fama2015five and the large set of factors provided by chen2022open . harvey2019census document over 400 factors published in top finance journals. dello2022missing argue that despite the hundreds of systematic risk factors considered in the literature, there is still a sizable pricing error and that this can be explained by asset specific risk that reflects market frictions and behavioral biases. There is a large Bayesian literature, including chib2020comparing , and hwang2022bayesian on selecting factor models. The issue of factor selection is also addressed by fama2018choosing .

Strong and weak factors in asset returns are considered by ANATOLYEV2022103 , connor2022semi , and giglio2023test. beaulieu2020arbitrage discuss the lack of identification of risk premia when many of the loadings are zero. There has also been concern about the consequences of omitted factors. giglio2021asset discuss the problem and try to deal with it using a three-pass method which is valid even when not all factors in the model are specified or observed using principal components of the test assets. onatski2012asymptotics and lettau2020estimating, lettau2020factors provide extensive discussions of weak factor and latent factors, respectively. More recent contributions include BaiNg2023Weak and uematsu2023inference.

There is also a large literature on portfolio construction. Herskovic2019 discuss low cost methods of hedging risk factors. Preite2024 derive an SDF in which there is compensation for unsystematic risk within the framework of the APT. Korsaye2021 introduce model-free smart SDFs which give rise to non-parametric SDF bounds for testing asset pricing models.\ \ Daniel2020 discuss the common practice of creating characteristic portfolios by sorting on characteristics associated with average returns and show that these portfolios capture not only the priced risk associated with the characteristic but also unpriced risk. Quaini2023 propose an estimator of tradable factor risk premia.

Paper's outline: The rest of the paper is organized as follows: Section (ref) provides the framework for estimation of $\boldsymbol{\phi}.$ Section (ref) sets out the assumptions and states the main theorems. Theorem (ref) shows that the standard Fama-MacBeth estimator is valid only when there are no pricing errors and $n/T\rightarrow0$. Theorem (ref) shows that the Shanken bias-corrected estimator of $\lambda_{k}$ continues to be consistent for a fixed $T$ as $n\rightarrow\infty$, even in presence of weak pricing errors and weak missing common factors. Theorem (ref) provides conditions under which the bias-corrected estimator, $\boldsymbol{\tilde{\phi}}_{nT}$, is consistent for $\boldsymbol{\phi}_{0}$, and derives the asymptotic distribution of $\boldsymbol{\tilde{\phi}}_{nT},$ assuming the observed factors are strong $(\alpha_{k}=1$ for all $k$). Theorem (ref) extends the results to cases where one or more risk factors are semi-strong, and establishes the rate at which $\boldsymbol{\tilde{\phi}}_{nT}$ converges to its true value, assuming the idiosyncratic pricing errors are sufficiently weak, as discussed after the theorem. For example, for a factor with strength $\alpha_{k}$ we show that $\tilde{\phi}_{k,nT}-\phi_{0,k}=O_{p}\left( T^{-1/2}n^{-(\alpha_{k} +\alpha_{\min}-1)/2}\right) $, and as a consequence \[ \tilde{\lambda}_{k,nT}-\lambda_{0,k}=O_{p}\left( T^{-1/2}n^{-(\alpha _{k}+\alpha_{\min}-1)/2}\right) +O_{p}\left( T^{-1/2}\right) , \] where $\lambda_{0,k}$ is the true value of the risk premia associated to factor $f_{kt}$, and $\alpha_{\min}$ is the strength of the least strong factor included. This consistency condition is weaker than the one derived by giglio2023test who also assume $\eta_{i}=0$, for all $i$. Finally, Theorem (ref) gives conditions for consistent estimation of the asymptotic variance of $\boldsymbol{\tilde{\phi}}_{nT},$ using a suitable threshold estimator of the covariance matrix. Section (ref) presents the Monte Carlo (MC) design, its calibration and a summary of the main findings. Section (ref)\ discusses the problem of factor selection from a large number of potential factors. Section (ref) gives the empirical application using monthly data on a large number of individual US stocks and risk factors over the period 1996-2021. It selects factors, estimates $\boldsymbol{\phi}\mathbf{,}$ and compares the performance of MV and phi-portfolios. Section (ref) provides some concluding remarks.

Detailed mathematical proofs are provided in a mathematical appendix. Further information on data sources, MC calibration plus some supplementary material for the empirical application are provided in the online supplement A. To save space all MC results are given in the online supplement B.

Identification and estimation of $\boldsymbol{\phi}$

Let $R_{it}$ denote the holding period return on traded security $i$, which can be bought long or short without transaction costs, $r_{t}^{f}$ \ is the risk free rate, and $r_{it}=R_{it}-r_{t}^{f}$ is the excess return. We start with the linear factor pricing model (LFPM)

equation[equation omitted — 111 chars of source]

for $i=1,2,...,n$ and $t=1,2,...,T$, where $r_{it}$ is explained in terms of the $K\times1$ vector of factors $\mathbf{f}_{t}=(f_{1t},f_{2t},...,f_{Kt} )^{\prime}$. The intercept $\mathit{\alpha}_{i}$ and the $K\times1$ vector of factor loadings, $\boldsymbol{\beta}_{i}=(\beta_{i1},\beta_{i2},...,\beta _{iK})^{\prime}$, are unknown. The idiosyncratic errors, $u_{it}$ have zero means and are assumed to be serially uncorrelated. The factors, $\mathbf{f} _{t}$, are assumed to be covariance stationary with the constant mean $\boldsymbol{\mu}=E\left( \mathbf{f}_{t}\right) $, and $\boldsymbol{\Sigma }_{f}=E\left[ \left( \mathbf{f}_{t}-\boldsymbol{\mu}\right) \left( \mathbf{f}_{t}-\boldsymbol{\mu}\right) ^{\prime}\right] $.

Under the Arbitrage Pricing Theory (APT) due to Ross (1976) the pricing errors, $\eta_{i}$ for $i=1,2,...,n$ defined by

equation[equation omitted — 130 chars of source]

are bounded such that

equation[equation omitted — 68 chars of source]

where $c$ is zero-beta expected excess return, $\boldsymbol{\lambda}$ is the $K\times1$ vector of risk premia. Under APT, we have

equation[equation omitted — 122 chars of source]

and reduces to the standard beta representation of an unconditional asset pricing model, if $c+\eta_{i}=0$. In this case, $\boldsymbol{\beta} _{i}=-Cov(r_{it},m_{t})/var(m_{t})$, where $m_{t}$ is the stochastic discount factor (SDF), which satisfies the moment condition $E_{t}(r_{i,t+1}m_{t+1} )=0$. See Pesaran2023 for a discussion of the link. For empirical assessment, we consider the more general unconditional pricing model ((ref)), and relate to the linear factor pricing model. To this end, taking unconditional expectations of ((ref)) and using the APT\ condition we have

equation[equation omitted — 172 chars of source]

which in turn yields

equation[equation omitted — 132 chars of source]

where

equation[equation omitted — 85 chars of source]

Under the APT condition, ((ref)), $\boldsymbol{\phi}$ can be identified from the cross section regression of $\alpha_{i}$ on $\boldsymbol{\beta}_{i}$. We relax the APT\ condition and derive restrictions on the degree to which idiosyncratic pricing errors, $\eta_{i}$, can be pervasive to achieve identification.

The focus of the literature has been on testing for alpha, $\mathit{\alpha }_{i}=0$, and the estimation of the risk premia, $\boldsymbol{\lambda}$, using panel data on excess returns, $\left\{ r_{it},1,2,...,n\text{; } t=1,2,...,T\right\} ,$ and $\mathbf{F}$, the $T\times K$ matrix of observations on the factors. It is clear that $\boldsymbol{\phi}$ plays an important role in tests for alpha in LFPM, and a non-zero $\boldsymbol{\phi}$ implies non-zero alphas, which in turn implies exploitable excess profitable opportunities. More specifically, we show that for known values of $\boldsymbol{\beta}_{i},$ $i=1,2,...,n\,$\ and when $\boldsymbol{\phi}$ is non-zero, there exists phi-based portfolios with non-zero means that are fully diversified (their variance tends to zero with $n$), namely have Sharpe ratios that tend to infinity with $n$, and hence dominate the MV portfolios. For the MV portfolios to be efficient it is necessary that $\boldsymbol{\phi=0}$.

Why $\boldsymbol{\phi}$ matters: introduction of the phi-portfolios

Substitute the expression for $\mathit{\alpha}_{i}$ given by ((ref)) in ((ref)) to obtain

equation[equation omitted — 173 chars of source]

and write the $n$ return equations more compactly as

equation[equation omitted — 174 chars of source]

where $\mathbf{r}_{\circ t}=(r_{1t},r_{2t},....,r_{nt})^{\prime},$ $\boldsymbol{\tau}_{n}$ is an $n$-dimensional vector of ones, $\mathbf{B} _{n}=(\boldsymbol{\beta}_{\circ1},\boldsymbol{\beta}_{\circ2} ,...,\boldsymbol{\beta}_{\circ K}),$ $\boldsymbol{\beta}_{\circ k}=(\beta _{1k},\beta_{2k},...,\beta_{nk})^{\prime},$ $\mathbf{u}_{\circ t} =(u_{1t},u_{2t},....,u_{nt})^{\prime}$, $\boldsymbol{\eta}_{n}=\left( \eta_{1},\eta_{2},...,\eta_{n}\right) ^{\prime}$, and $\mathbf{V} _{u}=E\left( \mathbf{u}_{\circ t}\mathbf{u}_{\circ t}^{\prime}\right) $. Suppose that the factors, $\mathbf{f}_{t},$ are traded, and $\boldsymbol{\phi }^{\prime}\boldsymbol{\phi}>0$. Consider the $n\times1$ vector of phi-portfolio weights,$\mathbf{w}_{\phi}$, given by

equation[equation omitted — 167 chars of source]

where $\mathbf{M}_{n}=\mathbf{I}_{n}-n^{-1}\boldsymbol{\tau}_{n} \boldsymbol{\tau}_{n}^{\prime}$. Finally, consider the long-short hedged portfolio return

equation[equation omitted — 325 chars of source]

and using ((ref)) note that

equation[equation omitted — 196 chars of source]

For a given vector of pricing errors, $\boldsymbol{\eta}_{n}$, $E\left( \rho_{t,\phi}\right) =\boldsymbol{\phi}^{\prime}\boldsymbol{\phi} +\mathbf{w}_{\phi}^{\prime}\boldsymbol{\eta}_{n}$, and $Var\left( \rho_{t,\phi}\right) =\mathbf{w}_{\phi}^{\prime}\mathbf{V}_{u}\mathbf{w} _{\phi}\mathbf{,}$ where $\mathbf{V}_{u}=E(\mathbf{u}_{\circ t}\mathbf{u} _{\circ t}^{\prime})$. Using ((ref)) it follows that \[ \left\Vert \mathbf{w}_{\phi}^{\prime}\boldsymbol{\eta}_{n}\right\Vert ^{2} \leq\left\Vert \boldsymbol{\eta}_{n}\right\Vert ^{2}\left\Vert \mathbf{w} _{\phi}\right\Vert ^{2}\leq\left\Vert \boldsymbol{\eta}_{n}\right\Vert ^{2}\left\Vert \boldsymbol{\phi}\right\Vert ^{2}\lambda_{max}\left[ \left( \mathbf{B}_{n}^{\prime}\mathbf{M}_{n}\mathbf{B}_{n}\right) ^{-1}\right] =\frac{\left\Vert \boldsymbol{\eta}_{n}\right\Vert ^{2}\left\Vert \boldsymbol{\phi}\right\Vert ^{2}}{\lambda_{min}\left( \mathbf{B}_{n} ^{\prime}\mathbf{M}_{n}\mathbf{B}_{n}\right) }. \] Similarly,

equation[equation omitted — 337 chars of source]

Hence, $E\left( \boldsymbol{\rho}_{t,\phi}\right) \rightarrow \boldsymbol{\phi}^{\prime}\boldsymbol{\phi}$, and $Var\left( \rho_{t,\phi }\right) \rightarrow0$, as $n\rightarrow\infty$, if

equation[equation omitted — 344 chars of source]

These conditions are met if $\lambda_{min}\left( \mathbf{B}_{n}^{\prime }\mathbf{M}_{n}\mathbf{B}_{n}\right) \rightarrow\infty$, $\lambda _{max}\left( \mathbf{V}_{u}\right) <C$, and the APT condition ((ref) ) holds. The first two conditions follow if the LFPM given by ((ref)) is an approximate factor model as assumed by cham1983arbi , namely when the factors are strong and the errors, $u_{it}$, are weakly cross correlated.\footnote{The diversification conditions in ((ref)) are met more generally, and accommodate the presence of semi-strong traded factors in $f_{t}$, and less restrictive conditions on $\mathbf{V}_{u}$.and the pricing errors $\eta_{i}.$ See Remark (ref) below.}

Therefore, knowledge of $\boldsymbol{\phi}^{\prime}\boldsymbol{\phi}$ and its statistical significance can play an important role in portfolio analysis. To illustrate this point suppose that $c=0$ and $\boldsymbol{\eta}_{n} =\mathbf{0}$, but $\boldsymbol{\phi}^{\prime}\boldsymbol{\phi}>0$, and consider the hedged return $\rho_{t,\phi}=\boldsymbol{\phi}^{\prime}\left[ \left( \mathbf{B}_{n}^{\prime}\mathbf{V}_{u}^{-1}\mathbf{B}_{n}\right) ^{-1}\mathbf{B}_{n}\mathbf{V}_{u}^{-1}\mathbf{r}_{\circ t}-\mathbf{f} _{t}\right] $. Then under LFPM we have\footnote{When $c\neq0$, it is not possible to eliminate $c$ (which is unpriced), and at the same time exploit the error covariance matrix $\mathbf{V}_{u}$.}

align*[align* omitted — 534 chars of source]

and its squared Sharpe ratio is given by

equation[equation omitted — 246 chars of source]

Since, \[ \boldsymbol{\phi}^{\prime}\left( \mathbf{B}_{n}^{\prime}\mathbf{V}_{u} ^{-1}\mathbf{B}_{n}\right) ^{-1}\boldsymbol{\phi}\leq\left( \boldsymbol{\phi }^{\prime}\boldsymbol{\phi}\right) \lambda_{\max}\left[ \left( \mathbf{B}_{n}^{\prime}\mathbf{V}_{u}^{-1}\mathbf{B}_{n}\right) ^{-1}\right] =\frac{\boldsymbol{\phi}^{\prime}\boldsymbol{\phi}}{\lambda_{min}\left( \mathbf{B}_{n}^{\prime}\mathbf{V}_{u}^{-1}\mathbf{B}_{n}\right) }, \] then

equation[equation omitted — 206 chars of source]

As a result, if $\lambda_{min}\left( \mathbf{B}_{n}^{\prime}\mathbf{V} _{u}^{-1}\mathbf{B}_{n}\right) \rightarrow\infty$ as $n\rightarrow\infty$, $SR_{\phi}^{2}$ increases in $n$ without bounds if $\boldsymbol{\phi}^{\prime }\boldsymbol{\phi}>0$. In contrast, it is easily seen that the Sharpe ratio of the standard mean-variance (MV) portfolio, defined by $SR_{MV}^{2} =\boldsymbol{\mu}_{R}^{\prime}\mathbf{V}_{R}^{-1}\boldsymbol{\mu}_{R}$ is bounded in $n$, and in consequence $SR_{\phi}^{2}$ will eventually dominate $SR_{MV}^{2}$ if $\boldsymbol{\phi}^{\prime}\boldsymbol{\phi}>0$. To see this, note that under the LFPM given by ((ref)) with $c=0$ and $\boldsymbol{\eta }_{n}=\mathbf{0}$, the MV portfolio is given by $\rho_{MV,t}=\boldsymbol{\mu }_{R}^{\prime}\mathbf{V}_{R}^{-1}\mathbf{r}_{\circ t}$, where $\boldsymbol{\mu }_{R}=\mathbf{B}_{n}\boldsymbol{\lambda}$ and $\mathbf{V}_{R}=\mathbf{B} _{n}\mathbf{\Sigma}_{f}\mathbf{B}_{n}^{\prime}+\mathbf{V}_{u}$. Also, since $\mathbf{B}_{n}\mathbf{\Sigma}_{f}\mathbf{B}_{n}^{\prime}$ is rank deficient then \[ \mathbf{V}_{R}^{-1}=\mathbf{V}_{u}^{-1}\mathbf{-V}_{u}^{-1}\mathbf{B}\left( \mathbf{\Sigma}_{f}^{-1}+\mathbf{B}_{n}^{\prime}\mathbf{V}_{u}^{-1} \mathbf{B}_{n}\right) ^{-1}\mathbf{B}_{n}^{\prime}\mathbf{V}_{u}^{-1}, \] and it follows that

align[align omitted — 531 chars of source]

Hence, the Sharpe ratio of the MV portfolio continues to be bounded in $n$ even if $\boldsymbol{\phi\neq0}$. Whether $\boldsymbol{\phi=0}$ or not affects the magnitude of $SR_{MV}^{2}$ but does not alter the fact that $SR_{MV}^{2}$ will be bounded if $\lambda_{min}\left( \mathbf{B}_{n}^{\prime}\mathbf{V} _{u}^{-1}\mathbf{B}_{n}\right) \rightarrow\infty$ as $n\rightarrow\infty$. Furthermore, it is not possible to improve over the MV portfolio by using only the factor loadings. The optimal portfolio formed using the $K$ beta-based portfolios, { $\boldsymbol{\rho}_{B,t}=$ }$\left( \mathbf{B} _{n}^{\prime}\mathbf{V}_{u}^{-1}\mathbf{B}_{n}\right) ^{-1}\mathbf{B} _{n}^{\prime}\mathbf{V}_{u}^{-1}${ $\mathbf{r}_{\circ t}$, }has the same Sharpe ratio as the mean-variance portfolio. See Lemma (ref) for a proof.

In short, to make sure that the mean-variance portfolio is efficient we must have $\boldsymbol{\phi}^{\prime}\boldsymbol{\phi}=0$, and it is of special interest to estimate $\boldsymbol{\phi}$ and develop reliable procedures for testing its statistical significance, using a large number of securities. In practice, the number of tradeable securities, $n$, might not be sufficiently large, and there are important specification and estimation uncertainties, and the Sharpe ratio of the phi-based portfolio, $\boldsymbol{\rho}_{t,\phi}$, is likely to be bounded even if $\boldsymbol{\phi}^{\prime}\boldsymbol{\phi}>0$. We turn to these issues in the empirical application provided in\ Section (ref).

Fama-MacBeth and Shanken estimators of risk premia

It will prove convenient to write ((ref)) in matrix notation by stacking the excess returns by $t=1,2,...,T$, for each security $i$

equation[equation omitted — 167 chars of source]

where $\mathbf{r}_{i\circ}=(r_{i1},r_{i2},...,r_{iT})^{\prime}$, $\mathbf{F=(f}_{1},\mathbf{f}_{2},...,\mathbf{f}_{T})^{\prime}$, $\mathbf{u}_{i\circ}=\left( u_{i1},u_{i2},...,u_{iT}\right) ^{\prime}$, and $\boldsymbol{\tau}_{T}$ is a $T\times1$ vector of ones. Similarly, stacking the excess returns by $i$ for each $t$ we have ((ref)) which we rewrite as

equation[equation omitted — 152 chars of source]

where $\boldsymbol{\alpha}_{n}=(\alpha_{1},\alpha_{2},...,\alpha_{n})^{\prime }=c\mathbf{\tau}_{n}+\mathbf{B}_{n}\boldsymbol{\phi}+\mathbf{\eta}_{n}$.

The risk premia are usually estimated using a two-pass procedure suggested by fama1973risk . The first-pass runs time series regressions of excess returns, $r_{it}$, on the $K$ observed factors to give estimates of the factor loadings, $\boldsymbol{\beta}_{i}:$

equation[equation omitted — 191 chars of source]

The second-pass runs a cross section regression of\ average returns, $\bar {r}_{i\circ}=T^{-1}\sum_{t=1}^{T}r_{it}$ on the estimated factor loadings, to obtain the FM estimator of $\boldsymbol{\lambda}$, namely

equation[equation omitted — 225 chars of source]

where $\mathbf{\hat{B}}_{nT}=(\boldsymbol{\hat{\beta}}_{1T},\boldsymbol{\hat {\beta}}_{2T},...,\boldsymbol{\hat{\beta}}_{nT})^{\prime},$ $\mathbf{\bar{r} }_{n\circ}=(\bar{r}_{1T},\bar{r}_{2T},...,\bar{r}_{nT})^{\prime},$ $\mathbf{M}_{T}=\mathbf{I}_{T}-T^{-1}\boldsymbol{\tau}_{T}\boldsymbol{\tau }_{T}^{\prime}$, $\boldsymbol{\tau}_{T}$ is a $T$-dimensional vector of ones, $\mathbf{M}_{n}=\mathbf{I}_{n}-n^{-1}\boldsymbol{\tau}_{n}\boldsymbol{\tau }_{n}^{\prime}$, and $\boldsymbol{\tau}_{n}$ is an n-dimensional vector of ones.

As is well known, when $T$ is finite FM's two-pass estimator is biased due the errors in estimation of factor loadings that do not vanish. The small $T$ bias of the two-pass estimator of $\boldsymbol{\lambda}$ has been a source of concern in the empirical literature. Under standard regularity conditions and as $n\rightarrow\infty$, we have

equation[equation omitted — 448 chars of source]

where $\mathbf{d}_{fT}=\mathbf{\hat{\mu}}_{T}-\mathbf{\mu}_{0},$ $\boldsymbol{\lambda}_{0}$ and $\mathbf{\mu}_{0}$ are the true values of $\boldsymbol{\lambda}$ and $\mathbf{\mu}$, respectively, $\mathbf{\Sigma }_{\beta\beta}=\lim_{n\rightarrow\infty}\left( n^{-1}\mathbf{B}_{n}^{\prime }\mathbf{M}_{n}\mathbf{B}_{n}\right) $, and $\overline{\sigma}^{2} =\lim_{n\rightarrow\infty}n^{-1}\sum_{i=1}^{n}\sigma_{i}^{2}>0$. Following shanken1992estimation , $\overline{\sigma}_{n}^{2}$ can be consistently estimated (for a fixed $T$) by

equation[equation omitted — 126 chars of source]

where

equation[equation omitted — 120 chars of source]

and as before $\hat{\alpha}_{iT}$ and $\boldsymbol{\hat{\beta}}_{iT}$ are the OLS estimators of $\mathit{\alpha}_{i}$ and $\boldsymbol{\beta}_{i}$. Using these results the bias-corrected version of the two-pass estimator is given by\footnote{See also shanken2007estimating , kan2013pricing , and BAI201531 , and the survey paper by jagannathan2010analysis for further references.}

equation[equation omitted — 189 chars of source]

where

equation[equation omitted — 236 chars of source]

When all the risk factors are strong, under certain regularity conditions, there exists a fixed $T_{0}$ such that for all $T>T_{0}$, then

equation[equation omitted — 229 chars of source]

where $\boldsymbol{\mu}_{0}$ indicates the true value of the factor mean. Shanken refers to $\boldsymbol{\lambda}_{T}^{\ast}$ as "ex-post" risk premia to be distinguished from $\boldsymbol{\lambda}_{0}$, referred to as "ex ante" risk premia. See also Section 3.7 of jagannathan2010analysis .

In this paper we exploit Shanken's bias correction procedure by applying it to $\boldsymbol{\phi}=$ $\boldsymbol{\lambda}-\boldsymbol{\mu}$ which we identify directly using ((ref)) from the regression of $\mathit{\alpha}_{i}$ on $\boldsymbol{\beta}_{i}$ for $i=1,2,...,n$, assuming the idiosyncratic pricing errors, $\eta_{i}$, are sufficiently weak relative to the strengths of the risk factors in a sense which will be made precise below.

Estimation of $\boldsymbol{\phi}$

In view of ((ref)), the estimation of $\boldsymbol{\phi}$ can be carried out following a two-step procedure whereby in the first step $\mathit{\alpha }_{i}$ and $\boldsymbol{\beta}_{i}$ are estimated from the least squares regressions of $r_{it}$ on an intercept and $\mathbf{f}_{t}$, and these are then used in a second step regression to estimate $\boldsymbol{\phi}$, namely,

equation[equation omitted — 233 chars of source]

where $\boldsymbol{\hat{\alpha}}_{nT}=(\hat{\alpha}_{1T},\hat{\alpha} _{1T},...,\hat{\alpha}_{nT})^{\prime}=\mathbf{\bar{r}}_{nT}-\mathbf{\hat{B} }_{nT}\boldsymbol{\hat{\mu}}_{T},$ and as before $\mathbf{\hat{B}} _{nT}=(\boldsymbol{\hat{\beta}}_{1T},\boldsymbol{\hat{\beta}}_{2T} ,...,\boldsymbol{\hat{\beta}}_{nT})^{\prime}$. This estimator is consistent for $\boldsymbol{\phi}_{0}$ so long as $n$ and $T\rightarrow\infty$, and bias-corrections are necessary to ensure the large $n$ consistency of the estimator when $T$ is fixed. A Shanken type bias-corrected estimator of $\boldsymbol{\phi}_{0}$ is given by

equation[equation omitted — 316 chars of source]

where $\mathbf{H}_{nT\ }$ and $\widehat{\bar{\sigma}}_{nT}^{2}$ are given by ((ref)) and ((ref)), respectively. It is also easily established that

equation[equation omitted — 127 chars of source]

and for a fixed $T$ and as $n\rightarrow\infty$, we have \[ p\lim_{n\rightarrow\infty}\boldsymbol{\tilde{\phi}}_{nT}=p\lim_{n\rightarrow \infty}\boldsymbol{\tilde{\lambda}}_{nT}-\boldsymbol{\hat{\mu}}_{T}. \] Hence, upon using ((ref))

equation[equation omitted — 286 chars of source]

and there exists a fixed $T_{0}$ such that for all $T>T_{0},$ $\boldsymbol{\tilde{\phi}}_{nT}$ converges to $\boldsymbol{\phi}_{0}$ as $n\rightarrow\infty$. \ Also using ((ref)) and ((ref)), and noting that $\boldsymbol{\lambda}_{0}-\boldsymbol{\mu}_{0}=\boldsymbol{\phi }_{0}$, interestingly we have \[ \boldsymbol{\tilde{\lambda}}_{nT}-\boldsymbol{\lambda}_{T}^{\ast }=\boldsymbol{\tilde{\phi}}_{nT}+\boldsymbol{\hat{\mu}}_{T} -\boldsymbol{\lambda}_{T}^{\ast}=\boldsymbol{\tilde{\phi}}_{nT} -\boldsymbol{\phi}_{0}. \] So inference using the Shanken bias-corrected estimator of $\boldsymbol{\lambda}$ around $\boldsymbol{\lambda}_{T}^{\ast}$, is the same as making inference using $\boldsymbol{\tilde{\phi}}_{nT}$ around $\boldsymbol{\phi}_{0}$.

The asymptotic distribution of $\boldsymbol{\tilde{\phi}}_{nT}$ depends on both $n$ and $T$. Assuming the observed factors are strong and under certain regularity conditions, to be introduced below, we have

equation[equation omitted — 229 chars of source]

where \[ \mathbf{V}_{\xi}=\left( 1+\boldsymbol{\lambda}_{0}^{\prime}\mathbf{\Sigma }_{f}^{-1}\boldsymbol{\lambda}_{0}\right) p\lim_{n\rightarrow\infty}\left[ n^{-1}\mathbf{B}_{n}^{\prime}\mathbf{M}_{n}\mathbf{V}_{u}\mathbf{M} _{n}\mathbf{B}_{n\ }\right] \emph{.} \] The variance of $\boldsymbol{\tilde{\phi}}_{nT}$ is consistently estimated by

equation[equation omitted — 179 chars of source]

where $\mathbf{H}_{nT\ }$ is given by ((ref)),

equation[equation omitted — 208 chars of source]

and

equation[equation omitted — 198 chars of source]

$\boldsymbol{\tilde{\lambda}}_{nT}$ is defined by ((ref)), and $\mathbf{\tilde{V}}_{u}\mathbf{\ }$is a suitable estimator of $\mathbf{V} _{u}=E\left( \mathbf{u}_{\circ t}\mathbf{u}_{\circ t}^{\prime}\right) $. How to estimate $\mathbf{V}_{u}$ and the conditions under which $Var\left( \boldsymbol{\tilde{\phi}}_{nT}\right) $ is consistently estimated by $\widehat{Var\left( \boldsymbol{\tilde{\phi}}_{nT}\right) }$ is discussed in sub-section (ref).

Note that even when $\mathbf{V}_{u}=\sigma^{2}\mathbf{I}_{T}$ the variance of $\boldsymbol{\tilde{\phi}}_{nT}$ does not reduce to $\sigma^{2} \boldsymbol{\Sigma}_{\beta\beta}^{-1}$, the standard least squares formula used for the case of known factor loadings. When the loadings are estimated the scaling term $\left( 1+\boldsymbol{\lambda}_{0}^{\prime} \boldsymbol{\Sigma}_{f}^{-1}\boldsymbol{\lambda}_{0}\right) $ is required and its neglect can lead to serious over-rejection even if $n/T\rightarrow0$ as $n$ and $T\rightarrow\infty$.

Factor strength

In this paper we deviate from the standard literature and allow the observed and latent factors to have different degrees of strength, depending on how pervasively they impact the security returns. bailey2021measurement define the strength of factor, $f_{kt}$, in terms of the number of its non-zero factor loadings. For a factor to be strong almost all of its $n$ loadings must differ from zero. Given our focus on estimation of risk premia, we adopt the following definition which directly relates to the covariance of $\boldsymbol{\beta}_{i}.$ See also chudik2011weak .

definition(Factor strengths) The strength of factor $f_{kt}$ is measured by its degree of pervasiveness as defined by the exponent $\alpha_{_{k}}$ in \begin{equation} \sum_{i=1}^{n}\left( \beta_{ik}-\bar{\beta}_{k}\right) ^{2}=\ominus (n^{\alpha_{k}}), \end{equation} and $0<\alpha_{k}\leq1$. We refer to $\left\{ \alpha_{k},\text{ }k=1,2,...,K\right\} $ as factor strengths. Factor $f_{kt}$ is said to be strong if $\alpha_{k}=1$, semi-strong if $1>\alpha_{k}>1/2$, and weak if $0\leq\alpha_{k}\leq1/2$. Condition ((ref)) applies irrespective of whether the loadings, $\beta_{ik}$, are viewed as deterministic or stochastic.\

In the above definition $\ominus_{p}\left( n^{\alpha_{k}}\right) $ denotes the rate at which additional securities add to the factor's strength and $\alpha_{k}$ can be viewed as a logarithmic expansion rate in terms of $n$ and relates to the proportion of non-zero factor loadings. In the literature it is commonly assumed that the covariance matrix of factor loadings defined by

equation[equation omitted — 277 chars of source]

is positive definite, where $\boldsymbol{\bar{\beta}}_{n}=n^{-1}\sum_{i=1} ^{n}\boldsymbol{\beta}_{i}=(\bar{\beta}_{1},\bar{\beta}_{2},...,\bar{\beta }_{k})^{\prime}$. For $\mathbf{\Sigma}_{\beta\beta}$ to be positive definite matrix it is necessary that all the $K$ risk factors under consideration are strong in the sense that

equation[equation omitted — 168 chars of source]

In terms of our definition of factor strength, $\mathbf{\Sigma}_{\beta\beta} $\ will be positive definite if all the observed factors are strong, namely if $\alpha_{k}=1$ for $k=1,2,...,K$. However, such an assumption is quite restrictive and is unlikely to be satisfied for many risk factors being considered in the literature. bailey2021measurement show that, apart from the market factor, only a handful of 144 factors in the literature considered by feng2020taming come close to being strong. giglio2023test consider the estimation of PCA-based risk premia in presence of weak factors. However, their definition of factor strength involves both $n~$and $T$, and is best viewed as a consistency condition rather than factor strength as such. See the discussion following Theorem (ref). Our notion of factor strength, $\alpha_{k}$, is in line with the recent literature. See, for example, BaiNg2023Weak and uematsu2023inference.

Missing factor

We now turn to the structure of the errors, $u_{it}$, in the returns equations, and consider two possible sources of error cross-sectional dependence: a missing or latent factor and production networks. The issue of missing factors has been investigated in the recent literature by giglio2021asset and ANATOLYEV2022103 . The issue of production networks has been investigated in the recent literature by herskovic2018networks , who derives two risk factors based on the changes in network concentration and network sparsity, and gofman2020production , who focus on the vertical dimension of production by modelling a supply chain, in terms of supplier-customer links. They find that the further away a firm is from final consumers the higher its return. They use this to create a factor TMB (top minus bottom). Both sources of cross-sectional error dependence could be important, since network dependence cannot be represented using latent factor models. See Section 3 of chudik2011weak.

To allow for both forms of error cross-sectional dependence we consider the following decomposition of $u_{it}$

equation[equation omitted — 58 chars of source]

where $g_{t}$ is the missing (latent) factor and $v_{it}$ is weakly cross-correlated in the sense of approximate factor models due to Chamberlain1983 and cham1983arbi. Here we allow for a single missing factor to simplify the exposition, but note that increasing the number of missing factors has little impact on our analysis, so long as the number of missing factors is fixed. Using the normalization $E(g_{t}^{2})=1$, and assuming that $\gamma_{i}g_{t}$ and $v_{it}$ are independently distributed then $E(u_{it}u_{jt})=\sigma_{ij}=\gamma_{i}\gamma_{j}+\sigma_{v,ij}$, and as shown in Lemma (ref), $n^{-1}\sum_{i=1}^{n}\sum_{j=1}^{n}\left\vert \sigma_{ij}\right\vert =O(1)$ so long as the strength of $g_{t},$ $\alpha_{\gamma}<1/2$ and $\lambda_{\max}(\mathbf{V}_{v})<\infty.$ This is despite the fact that $\lambda_{\max}(\mathbf{V}_{u})=O(n^{\alpha_{\gamma}})$, where $\mathbf{V}_{v}=E\left( \mathbf{v}_{i}\mathbf{v}_{i}^{\prime}\right) $ and $\mathbf{V}_{u}=E\left( \mathbf{u}_{i}\mathbf{u}_{i}^{\prime}\right) $.\footnote{Note that Chamberlain's approximate factor model specification requires $\lambda_{\max}(\mathbf{V}_{u})=O(1)$ and is violated if $\alpha_{\gamma}>0$.} In the Monte Carlo experiments, we consider the possibility of missing factors, as well as weak spatial and network cross-dependence that satisfy conditions of approximate factor models.

Pricing errors and market efficiency

The APT condition ((ref)), given by (18)\ in Theorem II of ROSS1976341 , ensures that under APT the idiosyncratic pricing errors are sparse. In this paper we relax the\ Ross's condition to

equation[equation omitted — 77 chars of source]

where the exponent $\alpha_{\eta}$ measures the degrees of pervasiveness of pricing errors. Deviations from APT are measured in terms of $\alpha_{\eta}$ ($0\leq\alpha_{\eta}<1$). We investigate the robustness of our proposed estimator of $\boldsymbol{\phi}$ to $\alpha_{\eta}$. This extension is important for tests of market efficiency where the null of interest is $H_{0}:$ $\mathit{\alpha}_{i}=c$ for all $i$ in ((ref)). We note that under the alternative hypothesis $H_{1}:$ $\mathit{\alpha}_{i}=c+\boldsymbol{\beta }_{i}^{\prime}\boldsymbol{\phi}+\eta_{i}$, therefore it is desirable to develop a test of $\boldsymbol{\phi}=\boldsymbol{0}$ which is robust to a wider class of pricing errors than those entertained originally by Ross, where $\alpha_{\eta}=0$.

Under the alternative hypothesis the power of testing $H_{0}$ will depend on the rate at which $\sum_{i=1}^{n}\left( \mathit{\alpha}_{i}-\mathit{\bar {\alpha}}\right) ^{2}$ rises with $n$. This in turn depends on the degree of pervasiveness of the idiosyncratic pricing errors, $\alpha_{\eta}$, the strength of the factors, $\alpha_{j}$, for $j=1,2,...,K$ and the magnitude of $\boldsymbol{\phi}^{\prime}\boldsymbol{\phi}\,$. Using $\mathit{\alpha} _{i}=c+\boldsymbol{\beta}_{i}^{\prime}\boldsymbol{\phi}+\eta_{i}$ \[ \sum_{i=1}^{n}\left( \mathit{\alpha}_{i}-\mathit{\bar{\alpha}}\right) ^{2}=\boldsymbol{\phi}^{\prime}\left( \sum_{i=1}^{n}\left( \boldsymbol{\beta }_{i}-\boldsymbol{\bar{\beta}}_{n}\right) \left( \boldsymbol{\beta} _{i}-\boldsymbol{\bar{\beta}}_{n}\right) ^{\prime}\right) \boldsymbol{\phi }+\sum_{i=1}^{n}\left( \eta_{i}-\bar{\eta}\right) ^{2}+2\boldsymbol{\phi }^{\prime}\sum_{i=1}^{n}\left( \boldsymbol{\beta}_{i}-\boldsymbol{\bar{\beta }}_{n}\right) \eta_{i}, \] and in the absence of idiosyncratic pricing errors

equation[equation omitted — 370 chars of source]

When the risk factors are all strong $\lambda_{\min}\left( n^{-1}\sum _{i=1}^{n}\left( \boldsymbol{\beta}_{i}-\boldsymbol{\bar{\beta}}_{n}\right) \left( \boldsymbol{\beta}_{i}-\boldsymbol{\bar{\beta}}_{n}\right) ^{\prime }\right) >0$, then $\sum_{i=1}^{n}\left( \mathit{\alpha}_{i}-\mathit{\bar {\alpha}}\right) ^{2}=\ominus(n)$ if and only if $\boldsymbol{\phi} \neq\boldsymbol{0}$. Namely, the extent to which alpha can be exploited will depend on the magnitudes of $\phi_{j}$ and the strength of the risk factors.

remarkHaving formalized the concepts of factor strength, missing factors, and the less restrictive APT condition given by ((ref)), it is now of interest to revisit the conditions under which the phi-portfolio fully diversifies. Consider ((ref)) and note that \[ \frac{\left\Vert \boldsymbol{\eta}_{n}\right\Vert ^{2}}{\lambda_{min}\left( \mathbf{B}_{n}^{\prime}\mathbf{M}_{n}\mathbf{B}_{n}\right) }=O\left[ n^{-\left( \alpha_{min}-\alpha_{\eta}\right) }\right] \text{, and } \frac{\lambda_{max}\left( \mathbf{V}_{u}\right) }{\lambda_{min}\left( \mathbf{B}_{n}^{\prime}\mathbf{M}_{n}\mathbf{B}_{n}\right) }=O\left[ n^{-\left( \alpha_{min}-\alpha_{\gamma}\right) }\right] . \] Using these results in ((ref))\ and ((ref)) and we have \[ E\left( \rho_{t,\phi}\right) =\boldsymbol{\phi}^{\prime}\boldsymbol{\phi }+O\left[ n^{-\left( \alpha_{min}-\alpha_{\eta}\right) }\right] \text{, and }Var\left( \rho_{t,\phi}\right) =O\left[ n^{-\left( \alpha _{min}-\alpha_{\gamma}\right) }\right] \] Therefore, for phi-portfolio to dominate the MV portfolio in addition to $\boldsymbol{\phi}^{\prime}\boldsymbol{\phi}>0$, it is also required that the strengths of the traded factors $\alpha_{k}$, $k=1,2,...,K$ are strictly larger than the strength of the missing factor, $\alpha_{\gamma},$ as well as the strength of the idiosyncratic pricing errors, $\alpha_{\eta}$.
remarkSimilarly, If we allow for idiosyncratic pricing errors, missing factors and non-strong risk factors the limiting expression for $\sum_{i=1}^{n}\left( \mathit{\alpha}_{i}-\mathit{\bar{\alpha}}\right) ^{2}$ becomes \begin{equation} \sum_{i=1}^{n}\left( \mathit{\alpha}_{i}-\mathit{\bar{\alpha}}\right) ^{2}=\ominus(\sum_{j=1}^{K}n^{\alpha_{j}}\phi_{j}^{2})+O\left( n^{\alpha _{\eta}}\right) +O\left( n^{\alpha_{\gamma}}\right) . \end{equation} In this more general case for alpha to be exploitable it is necessary that $\phi_{j}$ associated with the strongest factor is non-zero, and $\alpha _{\max}=\max_{j}(\alpha_{j})$ is larger than $\alpha_{\eta}$ and $\alpha_{\gamma}$.

Assumptions and theorems

We make the following standard assumptions about $\mathbf{f}_{t},g_{t}$, $v_{it}$, $\boldsymbol{\beta}_{i}$, $\eta_{i}$, and $\gamma_{i}$ (the drivers of asset returns):

assumption(Observed common factors) (a) The $K\times1$ vector of observed risk factors, $\mathbf{f}_{t}$, follows the general linear process \begin{equation} \mathbf{f}_{t}=\boldsymbol{\mu}+\sum_{\ell=0}^{\infty}\mathbf{\Psi}_{\ell }\mathbf{\zeta}_{t-\ell}, \end{equation} where $\left\Vert \boldsymbol{\mu}\right\Vert <C$, $\mathbf{\zeta} _{t}\thicksim IID(\mathbf{0},\mathbf{I}_{K})$, and $\mathbf{\Psi}_{\ell}$ are $K\times K$ exponentially decaying matrices such that $\left\Vert \mathbf{\Psi}_{\ell}\right\Vert <C\rho^{\ell}$ for some $C>0$ and $0<\rho<1$. (b) The $T\times K$ data matrix $\mathbf{F=(f}_{1},\mathbf{f}_{2} ,...,\mathbf{f}_{T})^{\prime}$ is full column rank and there exists $T_{0}$ such that for all $T>T_{0}$, $\boldsymbol{\hat{\Sigma}}_{f}=T^{-1} \mathbf{F}^{\prime}\mathbf{M}_{T}\mathbf{F}$ is a positive definite matrix, $\lambda_{\max}\left[ (T^{-1}\mathbf{F}^{\prime}\mathbf{M}_{T}\mathbf{F)} ^{-1}\right] <C$, $\boldsymbol{\hat{\Sigma}}_{f}\rightarrow_{p} \boldsymbol{\Sigma}_{f}=E\left( \mathbf{f}_{t}-\boldsymbol{\mu}_{0}\right) \left( \mathbf{f}_{t}-\boldsymbol{\mu}_{0}\right) ^{\prime}>\mathbf{0},$ where $\boldsymbol{\mu}_{0}$ is the true value of $\boldsymbol{\mu}$.
assumption(Observed factor loadings) (a) The factor loadings $\beta _{ik}$ for $i=1,2,...,n$ and $k=1,2,...,K$ are stochastically bounded such that $sup_{ik}E\left( \beta_{ik}^{2}\right) <C$, \begin{equation} \sum_{i=1}^{n}\left( \beta_{ik}-\bar{\beta}_{k}\right) ^{2}=\ominus _{p}(n^{\alpha_{k}}), for k=1,2,...,K. \end{equation} (b)The $n\times K$ matrix of factor loadings, $\mathbf{B}_{n} =(\boldsymbol{\beta}_{\circ1},\boldsymbol{\beta}_{\circ2} ,...,\boldsymbol{\beta}_{\circ K}),$ where $\boldsymbol{\beta}_{\circ k}=(\beta_{1k},\beta_{2k},...,\beta_{nk})^{\prime}$ satisfy \begin{equation} 0<c<\lambda_{min}\left( \mathbf{D}_{\alpha}^{-1}\mathbf{B}_{n}^{\prime }\mathbf{M}_{n}\mathbf{B}_{n}\mathbf{D}_{\alpha}^{-1}\right) <\lambda _{max}\left( \mathbf{D}_{\alpha}^{-1}\mathbf{B}_{n}^{\prime}\mathbf{M} _{n}\mathbf{B}_{n}\mathbf{D}_{\alpha}^{-1}\right) <C<\infty, \end{equation} for some small and large positive constants, $c$ and $C$, where $\mathbf{M} _{n}=\mathbf{I}_{n}-n^{-1}\boldsymbol{\tau}_{n}\boldsymbol{\tau}_{n}^{\prime} $, $\boldsymbol{\tau}_{n}=(1,1,...,1)^{\prime},$ and $\mathbf{D}_{\alpha} \,$\ is the $n\times n$ diagonal matrix \begin{equation} \mathbf{D}_{\alpha}=Diag(n^{\alpha_{1}/2},n^{\alpha_{2}/2},....,n^{\alpha _{K}/2}). \end{equation}
assumption(latent factor) (a) The latent factor, $g_{t}$, in ((ref)) is distributed independently of $\mathbf{f}_{t^{\prime}},$ for all $t$ and $t^{\prime}$, $g_{t}$ is serially independent with mean zero, $E(g_{t})=0,$ $E(g_{t}^{2})=1$, and a finite fourth order moment, $sup_{t}E(g_{t}^{4})<C$. (b) The loadings $\gamma_{i}$ are such that $sup_{i}\left\vert \gamma_{i}\right\vert <C$ and \begin{equation} \sum_{i=1}^{n}\left\vert \gamma_{i}\right\vert =O(n^{\alpha_{\gamma}}). \end{equation}
assumption(idiosyncratic errors) (a) The errors $\left\{ v_{it}\text{, }i=1,2,...,n;\text{ }t=1,2,...,T\right\} $ are distributed independently of the factors $f_{k,t^{\prime}}$, and $g_{t}$, for all $i,t,t^{\prime}$ and $k=1,2,...,K$, and their associated loadings $\beta_{ik},$ and $\gamma_{i}$. They are serially independent with $E(v_{it})=0$ and finite fourth order moments $E(v_{it}^{4})<\infty$, and covariances $E(v_{it}v_{jt})=\sigma _{v,ij}$, such that \begin{equation} \sup_{i}\sum_{j=1}^{n}\left\vert \sigma_{v,ij}\right\vert <\infty, and \sup_{i}\sum_{j=1}^{n}Cov(v_{it}^{2},v_{jt}^{2})<\infty, \end{equation} with $\lambda_{\min}\left( \mathbf{V}_{v}\right) >0$, where $\mathbf{V} _{v}=(\sigma_{v,ij})$. (b) The degree of cross-sectional dependence of $v_{it}$ is sufficiently weak so that \begin{equation} T^{-1/2}n^{-1/2}\sum_{t=1}^{T}\sum_{i=1}^{n}(\beta_{ik}-\bar{\beta}_{k} )v_{it}\rightarrow_{d}N(0,\omega_{k}^{2}), for k=1,2,...,K, \end{equation} where \begin{equation} \omega_{k}^{2}=p\lim_{n\rightarrow\infty}n^{-\alpha_{k}}\sum_{i=1}^{n} \sum_{j=1}^{n}(\beta_{ik}-\bar{\beta}_{k})(\beta_{jk}-\bar{\beta}_{k} )\sigma_{v,ij}. \end{equation}
assumption(Pricing errors) The pricing errors, $\eta_{i}$, for $i=1,2,...,n$ are individually bounded, $sup_{j}\left\vert \eta_{j}\right\vert <C$\thinspace, and are distributed independently of the factor loadings, $\beta_{jk}$, and $\gamma_{j}$ for all $i,j$ and $k=1,2,...,K$, as well as satisfying the condition \begin{equation} \sum_{i=1}^{n}\left\vert \eta_{i}\right\vert =O\left( n^{\alpha_{\eta} }\right) , \end{equation} with $\alpha_{\eta}<1/2$.
remarkUnder Assumption (ref) $E\left( \mathbf{f}_{t}\right) =\mathbf{\mu },$ and $Var(\mathbf{f}_{t})=\boldsymbol{\Sigma}_{f}=\sum_{\ell=0}^{\infty }\mathbf{\Psi}_{\ell}\mathbf{\Psi}_{\ell}^{\prime}$. Also since $\left\Vert \boldsymbol{\Sigma}_{f}\right\Vert \leq\sum_{\ell=0}^{\infty}\left\Vert \mathbf{\Psi}_{\ell}\right\Vert ^{2}$ it then follows from part (a) of Assumption (ref) that $\left\Vert \boldsymbol{\Sigma}_{f}\right\Vert <C$.
remarkUnder Assumption (ref) \begin{equation} \mathbf{D}_{\alpha}^{-1}\mathbf{B}_{n}^{\prime}\mathbf{M}_{n}\mathbf{B} _{n}\mathbf{D}_{\alpha}^{-1}\rightarrow_{p}\mathbf{\Sigma}_{\beta\beta }(\boldsymbol{\alpha}\mathbf{)}>0, \end{equation} where $\mathbf{\Sigma}_{\beta\beta}(\boldsymbol{\alpha}\mathbf{)}$ is a $k\times k$ symmetric positive definite matrix which is a function of $\boldsymbol{\alpha}=\mathbf{(}\alpha_{1},\alpha_{2},...,\alpha_{K})^{\prime} $. This follows from ((ref)) since for any non-zero $n\times1$ vector $\boldsymbol{\kappa}$, \[ \boldsymbol{\kappa}^{\prime}\mathbf{D}_{\alpha}^{-1}\mathbf{B}_{n}^{\prime }\mathbf{M}_{n}\mathbf{B}_{n}\mathbf{D}_{\alpha}^{-1}\boldsymbol{\kappa} \geq\left( \mathbf{\kappa}^{\prime}\mathbf{\kappa}\right) \lambda _{min}\left( \mathbf{D}_{\alpha}^{-1}\mathbf{B}_{n}^{\prime}\mathbf{M} _{n}\mathbf{B}_{n}\mathbf{D}_{\alpha}^{-1}\right) >0. \] In the standard case where the factors are all strong ($\alpha_{k}=1$ for all $k$), the above limit reduces to $n^{-1}\mathbf{B}_{n}^{\prime}\mathbf{M} _{n}\mathbf{B}_{n}\rightarrow_{p}\mathbf{\Sigma}_{\beta\beta}(\boldsymbol{\tau }_{K}\mathbf{)=\Sigma}_{\beta\beta}>0$.
remarkThe high level condition ((ref)) in Assumption (ref) is required for establishing the asymptotic normality of the estimator of $\boldsymbol{\phi}_{0}$, and is clearly met when $v_{it}$ and/or $\beta_{ik}$ are independently distributed. It is also possible to establish ((ref)) under weaker conditions assuming that $v_{it}$ and/or $\beta_{ik}$ satisfy some time-series type mixing conditions applied to cross section.
remarkThe exponent parameter, $\alpha_{\eta}$, of the pricing condition in ((ref)), can be viewed as the degree to which pricing errors are pervasive in large economies (as $n\rightarrow\infty$). Letting $\boldsymbol{\eta}_{n}=(\eta_{1},\eta_{2},...,\eta_{n})^{\prime}\,$\ we have \begin{equation} \sum_{i=1}^{n}\eta_{i}^{2}=\left\Vert \boldsymbol{\eta}_{n}\right\Vert ^{2}\leq\left\Vert \boldsymbol{\eta}_{n}\right\Vert _{\infty}\left\Vert \boldsymbol{\eta}_{n}\right\Vert _{1}=sup_{j}\left\vert \eta_{j}\right\vert \left( \sum_{i=1}^{n}\left\vert \eta_{i}\right\vert \right) , \end{equation} and under Assumption ((ref)) it also follows that \begin{equation} \sum_{i=1}^{n}\eta_{i}^{2}=O\left( n^{\alpha_{\eta}}\right) . \end{equation} Similarly \begin{equation} \sum_{i=1}^{n}\gamma_{i}^{2}=O(n^{\alpha_{\gamma}}). \end{equation}
remarkWhilst ((ref)) implies ((ref)), the reverse does not follow. By allowing for $\alpha_{\eta}>0$ we are relaxing the Ross's boundedness condition that requires setting $\alpha_{\eta}=0$.
remarkThe assumption that the observed and missing factors, $\mathbf{f}_{t}$ and $g_{t^{\prime}},$ are distributed independently is not restrictive and can be relaxed. For example, suppose that \[ g_{t}=\mu_{g}+\boldsymbol{\theta}^{\prime}\mathbf{f}_{t}+v_{gt}, \] where $\boldsymbol{f}_{t}$ and $v_{gt}$ are independently distributed. Then using ((ref)) we have \[ u_{it}=\gamma_{i}\mu_{g}+\gamma_{i}\left( \boldsymbol{\theta}^{\prime }\mathbf{f}_{t}\right) +\gamma_{i}v_{gt}+v_{it}, \] and the return equation ((ref)) can be written as \[ r_{it}=\left( \mathit{\alpha}_{i}+\gamma_{i}\mu_{g}\right) +\left( \boldsymbol{\beta}_{i}+\gamma_{i}\boldsymbol{\theta}\right) ^{\prime }\mathbf{f}_{t}+\gamma_{i}v_{gt}+v_{it}, \] with $v_{gt}$ now acting as the missing common factor, which, by construction, is distributed independently of $\mathbf{f}_{t}$.
remarkAssumptions (ref) and (ref) allow $u_{it}$ to be cross-sectionally weakly correlated, but require the errors to be serially uncorrelated. This requirement is not strong for asset pricing models, since realized returns are only mildly serially correlated and most likely such serial dependence will be captured by the serial correlation in the observed factors.

As we shall see, to estimate and conduct inference on the risk premia associated with the observed factors, $f_{kt}$, we require $\alpha_{k} >\alpha_{\gamma}<1/2$, where $\alpha_{\gamma}$ denotes the strength of the latent factor, $g_{t},$ and similarly defined by $\sum_{i=1}^{n}\gamma_{i} ^{2}=\ominus(n^{\alpha_{\gamma}})$. Namely, the latent factor must be sufficiently weak so that ignoring it will be inconsequential, and observed factors sufficiently strong so that they can be distinguished from the weak latent factor.

The main theoretical results of the paper are set out around five theorems. Theorem (ref) considers the Fama-MacBeth two-step estimator and derives its limiting property as $n$ and $T\rightarrow\infty$. To eliminate the bias of Fama-MacBeth estimator we require $n/T\rightarrow0$, and to eliminate the effects of pricing errors we need $Tn^{\alpha_{\eta} }/n\rightarrow0$, which results in a contradiction. Thus the Fama-MacBeth estimator is valid only when there are no pricing errors ($\eta_{i}=0$ for all $i$) and when $n/T\rightarrow0$. Theorem (ref) provides a proof that the estimator of $\bar{\sigma}_{n}^{2}$ (denoted by $\widehat{\bar{\sigma}} _{nT}^{2}$) proposed by shanken1992estimation continues to be unbiased for a fixed $T$ as $n\rightarrow\infty$, even under the general setting of the current paper that allows for missing factors as well as pricing errors. Theorem (ref) also establishes that $\widehat{\bar{\sigma}}_{nT}^{2}-\bar{\sigma}_{n}^{2}\rightarrow O_{p}(n^{-1/2}T^{-1/2})$, which is essential for establishing the results for the bias-corrected estimator of $\boldsymbol{\phi}_{0}$, namely $\boldsymbol{\tilde{\phi}}_{nT}$ given by ((ref)), summarized in Theorem (ref). This theorem provides conditions under which $\boldsymbol{\tilde{\phi}}_{nT}$ is a consistent estimator of $\boldsymbol{\phi}_{0}$, and derives its asymptotic distribution assuming the observed factors are strong, again allowing for pricing errors, a missing factor, and other forms of weak error cross-sectional dependence. Theorem (ref) extends the results of Theorem (ref) to the case where one or more of the observed risk factors are semi-strong and shows how factor strength impacts the precision with which the elements of $\boldsymbol{\phi }_{0}$ are estimated. Finally, Theorem (ref) presents the conditions under which the asymptotic variance of $\boldsymbol{\tilde{\phi}}_{nT}$ can be consistently estimated.

theorem(Small $T$ bias of Fama-MacBeth estimator of $\boldsymbol{\lambda}$) Consider the multi-factor linear return model ((ref)) with the missing factor $g_{t}$ in $u_{it}$ as defined by ((ref)) and the associated risk premia, $\boldsymbol{\lambda}$, defined by ((ref)). Suppose that Assumptions (ref), (ref), (ref), (ref) and (ref) hold and all observed factors are strong. Suppose further that the true value of the risk premia, $\boldsymbol{\lambda}_{0}$, is estimated by Fama-MacBeth two-pass estimator, $\boldsymbol{\hat{\lambda}}_{nT}$, defined by ((ref)). Then for any fixed $T>T_{0}$ such that $\lambda_{\min}\left( T^{-1}\mathbf{F}^{\prime }\mathbf{M}_{T}\mathbf{F}\right) >0,$ we have (as $n\rightarrow\infty$) \begin{equation} \boldsymbol{\hat{\lambda}}_{nT}-\boldsymbol{\lambda}_{0}=\left( \boldsymbol{\hat{\mu}}_{T}-\boldsymbol{\mu}_{0}\right) -\frac{\bar{\sigma }^{2}}{T}\left[ \boldsymbol{\Sigma}_{\beta\beta}+\bar{\sigma}^{2}\frac{1} {T}\left( \frac{\mathbf{F}^{\prime}\mathbf{M}_{T}\mathbf{F}}{T}\right) ^{-1}\right] ^{-1}\left( \frac{\mathbf{F}^{\prime}\mathbf{M}_{T}\mathbf{F} }{T}\right) ^{-1}\boldsymbol{\lambda}_{T}^{\ast}+o_{p}(1), \end{equation} where $\boldsymbol{\hat{\mu}}_{T}=T^{-1}\sum_{t=1}^{T}\mathbf{f}_{t}$, \[ \mathbf{\Sigma}_{\beta\beta}=\lim_{n\rightarrow\infty}\left( \frac {\mathbf{B}_{n}^{\prime}\mathbf{M}_{n}\mathbf{B}_{n}}{n}\right) \text{, and }\overline{\sigma}^{2}=\lim_{n\rightarrow\infty}n^{-1}\sum_{i=1}^{n}\sigma _{i}^{2}>0. \]

The proof is provided in Section (ref) of the mathematical appendix.

To derive the asymptotic distribution of $\boldsymbol{\hat{\lambda}} _{nT}-\boldsymbol{\lambda}_{0}$ it is required that both $n$ and $T\rightarrow\infty$, jointly. Also, noting that \[ \boldsymbol{\hat{\lambda}}_{nT}-\boldsymbol{\lambda}_{0}=\left( \boldsymbol{\hat{\mu}}_{T}-\boldsymbol{\mu}_{0}\right) +\left( \boldsymbol{\hat{\phi}}_{nT}-\boldsymbol{\phi}_{0}\right) , \] it is clear that increasing $n$ is not relevant for the distribution of $\boldsymbol{\hat{\mu}}_{T}-\boldsymbol{\mu}_{0}$, but joint $n$ and $T$ asymptotics are required when investigating the distribution of $\boldsymbol{\hat{\phi}}_{nT}-\boldsymbol{\phi}_{0}$. Focussing on the latter, and using result ((ref))\ in the Appendix, we have

align*[align* omitted — 592 chars of source]

Where $\boldsymbol{U}_{nT}=\left( \mathbf{u}_{1\circ},\mathbf{u}_{2\circ },...,\mathbf{u}_{n\circ}\right) ^{\prime},$ $\mathbf{u}_{i\circ}=\left( u_{i1},u_{i2},...,u_{iT}\right) ^{\prime},$ $\mathbf{G}_{T}=$ $\mathbf{M} _{T}\mathbf{F}\left( \mathbf{F}^{\prime}\mathbf{M}_{T}\mathbf{F}\right) ^{-1}$, $\overline{\mathbf{u}}_{n\circ}=(\overline{u}_{1\circ},\overline {u}_{2\circ},...,\overline{u}_{n\circ})^{\prime},$ and $\overline{u}_{i\circ }=T^{-1}\sum_{t=1}^{T}u_{it}.$ Consider first the terms that include the pricing errors, $\boldsymbol{\eta}_{n}$, and using the results in Lemma (ref) note that \[ n^{-1/2}T^{1/2}\mathbf{B}_{n}^{\prime}\mathbf{M}_{n}\boldsymbol{\eta} _{n}=O_{p}\left( T^{1/2}n^{-1/2+\alpha_{\eta}}\right) ,\text{ } n^{-1/2}T^{1/2}\mathbf{G}_{T}^{\prime}\mathbf{U}_{nT}^{\prime}\mathbf{M} _{n}\boldsymbol{\eta}_{n}=O_{p}\left( n^{-1/2+\frac{\alpha_{\eta} +\alpha_{\gamma}}{2}}\right) . \] It is clear that the effects of pricing errors on the distribution of $\boldsymbol{\hat{\phi}}_{nT}$ vanish only if $T^{1/2}n^{-1/2+\alpha_{\eta} }\rightarrow0$, and $\alpha_{\eta}+\alpha_{\gamma}<1$. Also

align*[align* omitted — 472 chars of source]

Finally, for the first two terms involving $\mathbf{B}_{n}$ and $\mathbf{U} _{nT}$ we have

equation[equation omitted — 210 chars of source]

It is clear that the small $T$ bias of the asymptotic distribution of the two-step estimator, given by $\sqrt{\frac{n}{T}}\bar{\sigma}_{n}^{2}\left( \frac{\mathbf{F}^{\prime}\mathbf{M}_{T}\mathbf{F}}{T}\right) ^{-1} \boldsymbol{\lambda}_{T}^{\ast}$, does not vanish unless, $n/T\rightarrow0$. At the same time for the pricing errors to have no impact on the distribution of the two-step estimator we must have $Tn^{a_{\eta}}/n\rightarrow0$. Both conditions cannot be met simultaneously. It is possible to derive the asymptotic distribution of $\boldsymbol{\hat{\phi}}_{nT}$, and hence that of $\boldsymbol{\hat{\lambda}}_{nT}$, when $n/T\rightarrow0$ and $\mathbf{\eta =0}$, but these are quite restrictive conditions, and to avoid them we follow shanken1992estimation and instead consider a bias-corrected version of $\boldsymbol{\hat{\phi}}_{nT}$, namely $\boldsymbol{\tilde{\phi}}_{nT}$ given by ((ref)). As noted earlier $\boldsymbol{\tilde{\phi}} _{nT}=\boldsymbol{\tilde{\lambda}}_{nT}-\boldsymbol{\hat{\mu}}_{T}$, where $\boldsymbol{\tilde{\lambda}}_{nT}$ is the bias-corrected version of $\boldsymbol{\hat{\lambda}}_{nT}$ originally proposed by Shanken.

To investigate the asymptotic properties of $\boldsymbol{\tilde{\phi}}_{nT}$ we first need to establish conditions under which $\widehat{\bar{\sigma}} _{nT}^{2}$, defined by ((ref)), is a consistent estimator of $\bar{\sigma}_{n}^{2}=n^{-1}\sum_{i=1}^{n}\sigma_{i}^{2}$, which enters the bias-corrected estimator. The proof of consistency in the literature does not allow for missing factors or pricing errors and only considers the case where $T$ is fixed as $n\rightarrow\infty$. For derivation of asymptotic distribution of $\boldsymbol{\tilde{\phi}}_{nT}$ we also need to consider the limiting properties of $\widehat{\bar{\sigma}}_{nT}^{2}$ under joint $n$ and $T$ asymptotics. The following theorem provides the required results for $\widehat{\bar{\sigma}}_{nT}^{2}$ as an estimator of $\bar{\sigma}_{n}^{2}$.

theoremConsider $\widehat{\bar{\sigma}}_{nT}^{2}$, the estimator of $\bar{\sigma}_{n}^{2}$ given by (see ((ref))), \begin{equation} \widehat{\bar{\sigma}}_{nT}^{2}=\frac{\sum_{t=1}^{T}\sum_{i=1}^{n}\hat{u} _{it}^{2}}{n(T-K-1)}, \end{equation} and suppose that Assumptions (ref), (ref), and (ref), are satisfied. Then for a fixed $T$ \begin{equation} \lim_{n\rightarrow\infty}E\left( \widehat{\bar{\sigma}}_{nT}^{2}\right) =\bar{\sigma}^{2}, \end{equation} where $\bar{\sigma}^{2}=\lim_{n\rightarrow\infty}\bar{\sigma}_{n}^{2}$, and $\bar{\sigma}_{n}^{2}=n^{-1}\sum_{i=1}^{n}\sigma_{i}^{2}$. Furthermore \begin{equation} \widehat{\bar{\sigma}}_{nT}^{2}-\bar{\sigma}_{n}^{2}=O_{p}\left( T^{-1/2}n^{-1/2}\right) . \end{equation}

For a proof see sub-section (ref) in the Appendix.

Result ((ref)) shows that $\widehat{\bar{\sigma}}_{nT}^{2}$ continues to be a consistent estimator of $\bar{\sigma}^{2}=\lim_{n\rightarrow\infty }\bar{\sigma}_{n}^{2}$ for a fixed $T$ as $n\rightarrow\infty$, even in the presence of pricing errors and a missing common factor. This result also holds when one or more of the factors are semi-strong.

Equipped with the above result we are now in a position to present the theorem that sets out the asymptotic distribution of $\boldsymbol{\tilde{\phi}}_{nT}$.

theoremConsider, $\boldsymbol{\tilde{\phi}}_{nT}$, the bias-corrected estimators of $\boldsymbol{\phi}_{0}$ given by ((ref)). Suppose Assumptions (ref), (ref), (ref), (ref) and (ref) hold, all the observed factors are strong, ($\alpha _{k}=1$, for $k=1,2,...,K$), and the strength of the missing factor, $\alpha_{\gamma}$ defined by ((ref)), satisfies $\alpha_{\gamma}$ $<1/2$. \begin{align} \boldsymbol{\tilde{\phi}}_{nT}-\boldsymbol{\phi}_{0} & =O_{p}\left( T^{-1/2}n^{^{-1/2}}\right) +O_{p}\left( T^{-1/2}n^{-1+\frac{\alpha_{\eta }+\alpha_{\gamma}}{2}}\right) \\ & +O_{p}\left( n^{-1+\alpha_{\eta}}\right) +O_{p}\left( T^{-1} n^{-1/2}\right) ,\nonumber \end{align} where $\alpha_{\eta}\,$denotes the degree of pervasiveness of the pricing errors defined by ((ref)). (a) When $T$ is fixed, $\alpha_{\gamma}<1/2$ and $\alpha_{\eta}<1$, then there exists $T_{0}$ such that for all $T>T_{0}$ \begin{equation} p\lim_{n\rightarrow\infty}\left( \boldsymbol{\tilde{\phi}}_{nT}\right) =\boldsymbol{\phi}_{0}. \end{equation} Also \begin{equation} \sqrt{nT}\left( \boldsymbol{\tilde{\phi}}_{nT}-\boldsymbol{\phi}_{0}\right) =\mathbf{\Sigma}_{\beta\beta}^{-1}\boldsymbol{\xi}_{nT}+O_{p}\left( n^{-\frac{1}{2}+\frac{\alpha_{\eta+\alpha_{\gamma}}}{2}}\right) +O_{p}\left( T^{1/2}n^{-1/2+\alpha_{\eta}}\right) +O_{p}\left( T^{-1/2}\right) , \end{equation} where $\mathbf{\Sigma}_{\beta\beta}=p\lim_{n\rightarrow\infty}\left( n^{-1}\mathbf{B}_{n}^{\prime}\mathbf{M}_{n}\mathbf{B}_{n\ }\right) ,$ \begin{equation} \boldsymbol{\xi}_{nT}=n^{-1/2}T^{-1/2}\mathbf{B}_{n}^{\prime}\mathbf{M} _{n}\mathbf{U}_{nT}\mathbf{a}_{T}, \end{equation} and $\mathbf{a}_{T}=\mathbf{\tau}_{T}-\mathbf{M}_{T}\mathbf{F}(T^{-1} \mathbf{F}^{\prime}\mathbf{M}_{T}\mathbf{F)}^{-1}\boldsymbol{\lambda} _{T}^{\ast}$. (b) If $\alpha_{\gamma}<1/2$, $\alpha_{\eta}<1/2,$ and $\sqrt{\frac{T}{n}}n^{\alpha_{\eta}}\rightarrow0$, as $n$ and $T\rightarrow \infty$ jointly, then \begin{equation} \sqrt{nT}\left( \boldsymbol{\tilde{\phi}}_{nT}-\boldsymbol{\phi}_{0}\right) \rightarrow_{d}N\left( \mathbf{0,\Sigma}_{\beta\beta}^{-1}\mathbf{V}_{\xi }\mathbf{\Sigma}_{\beta\beta}^{-1}\right) , \end{equation} where \begin{equation} \mathbf{V}_{\xi}=\left( 1+\boldsymbol{\lambda}_{0}^{\prime}\mathbf{\Sigma }_{f}^{-1}\boldsymbol{\lambda}_{0}\right) p\lim_{n\rightarrow\infty }\left( n^{-1}\mathbf{B}_{n}^{\prime}\mathbf{M}_{n}\mathbf{V}_{u} \mathbf{M}_{n}\mathbf{B}_{n\ }\right) . \end{equation}

For a proof see sub-section (ref) in the Appendix.

Result ((ref)) establishes the finite $T$ consistency of $\boldsymbol{\tilde{\phi}}_{nT}$ for $\boldsymbol{\phi}_{0}$ so long as $\alpha_{\gamma}<1/2$ and $\alpha_{\eta}<1$, thus extending the Shanken result to a much more general setting. To the best of our knowledge the asymptotic distribution in ((ref)) is new and shows that the asymptotic covariance matrix of $\boldsymbol{\tilde{\phi}}_{nT}$ includes the term $\boldsymbol{\lambda}_{T}^{\ast\prime}(T^{-1}\mathbf{F}^{\prime}\mathbf{M} _{T}\mathbf{F)}^{-1}\boldsymbol{\lambda}_{T}^{\ast}$, that arises from the first stage estimation of the factor loadings, and must be included in the analysis for valid inference. It is also clear that this additional term does not vanish with $T\rightarrow\infty$, and tends to $\boldsymbol{\lambda} _{0}^{\prime}\mathbf{\Sigma}_{f}^{-1}\boldsymbol{\lambda}_{0}\geq\left( \boldsymbol{\lambda}_{0}^{\prime}\boldsymbol{\lambda}_{0}\right) \lambda_{\max}\left( \mathbf{\Sigma}_{f}^{-1}\right) =\left( \boldsymbol{\lambda}_{0}^{\prime}\boldsymbol{\lambda}_{0}\right) \lambda_{\min}\left( \mathbf{\Sigma}_{f}\right) >0$, which is strictly non-zero unless $\boldsymbol{\lambda}_{0}=\mathbf{0}$. Shanken type bias correction addresses the mean of the asymptotic distribution of $\boldsymbol{\tilde{\phi}}_{nT}$, but not its covariance.

The $O_{p}\left( T^{-1/2}\right) $ term in ((ref)) arises from the sampling errors involved in the estimation of the factor loadings and $\bar{\sigma}_{n}^{2}$, and tends to zero at the regular $\sqrt{T}$ rate. But $n$ has to be sufficiently large to eliminate the effects of pricing errors on identification of $\boldsymbol{\phi}_{0}$, as dictated by condition $\sqrt{\frac{T}{n}}n^{\alpha_{\eta}}\rightarrow0,$ as $n$ and $T\rightarrow \infty$.\footnote{The condition $\sqrt{\frac{T}{n}}n^{\alpha_{\eta} }\rightarrow0$ can be weakened somewhat to $\sqrt{\frac{T}{n}}n^{\alpha_{\eta }/2}\rightarrow0$ if we also assume that $\beta_{ik}-\bar{\beta}_{k}$ are independently distributed over $i$, but will still require $n$ to be larger than $T$.} The requirement that $T$ need not be too large relative to $n$ for estimation of $\boldsymbol{\phi}_{0}$ is consistent with separating the estimation of $\boldsymbol{\phi}_{0}$ from that of $\boldsymbol{\mu}_{0}$, allowing the use a relatively small $T$ and a large $n$ to estimate $\boldsymbol{\phi}_{0}$ and a relatively large $T$ when estimating $\boldsymbol{\mu}_{0}.$

What if one or more of the risk factors are semi-strong?

We now turn to an intermediate case where one or more of the observed factors are semi-strong, in the sense that their factor strength, $\alpha_{k}$ lies between $1/2$ and $1$. The case of weak risk factors is already covered in the proceeding analysis, and such factors can be included in the error term, $u_{it}$, with little consequence for the estimation of risk premia of the remaining factors that are strong or semi-strong. Weak factors do not have any explanatory power and can be dropped from the analysis.

When one or more of the observed factors is semi-strong $\mathbf{\Sigma }_{\beta\beta}$ is no longer positive definite and Theorem (ref) does not apply, but it is possible to adapt the proofs to establish the limiting properties of $\tilde{\phi}_{k,nT}\boldsymbol{\ }$(the $k^{th}$ element of $\boldsymbol{\tilde{\phi}}_{nT}$)\ for different values of \thinspace $\alpha_{k}\,$.

To this end, analogously to $\boldsymbol{\tilde{\phi}}_{nT}$, we introduce the following estimator of $\boldsymbol{\phi}_{0}$

equation[equation omitted — 434 chars of source]

where

equation[equation omitted — 367 chars of source]

It is now easily seen that

equation[equation omitted — 265 chars of source]

where

equation[equation omitted — 423 chars of source]

$\mathbf{D}_{\alpha}$ is defined by ((ref)), and $\boldsymbol{\alpha }=\mathbf{(}\alpha_{1},\alpha_{2},...,\alpha_{K})^{\prime}$. It is easily established that numerically $\boldsymbol{\tilde{\phi}}_{nT}\left( \boldsymbol{\alpha}\right) $ is identical to $\boldsymbol{\tilde{\phi}}_{nT} $, and its introduction is primarily for the purpose of establishing the limiting properties of $\tilde{\phi}_{k,nT}-\phi_{0,k}$ that do depend on $\alpha_{k}$. Note that \[ \mathbf{H}_{nT}\left( \boldsymbol{\alpha}\right) =n\mathbf{D}_{\alpha} ^{-1}\mathbf{H}_{nT}\mathbf{D}_{\alpha}^{-1},\text{ and }\mathbf{q} _{nT}\left( \boldsymbol{\alpha}\right) =n\mathbf{D}_{\alpha}^{-1} \mathbf{s}_{nT} \] where $\mathbf{s}_{nT}$ and $\mathbf{H}_{nT}$ are already defined by ((ref)) and ((ref)). Using these in ((ref)) we have \[ \mathbf{D}_{\alpha}\left( \boldsymbol{\tilde{\phi}}_{nT}\left( \boldsymbol{\alpha}\right) -\boldsymbol{\phi}_{0}\right) =\left( n\mathbf{D}_{\alpha}^{-1}\mathbf{H}_{nT}\mathbf{D}_{\alpha}^{-1}\right) ^{-1}n\mathbf{D}_{\alpha}^{-1}\mathbf{s}_{nT}=\mathbf{D}_{\alpha} \mathbf{H}_{nT}^{-1}\mathbf{s}_{nT}, \] and it follows that $\boldsymbol{\tilde{\phi}}_{nT}\left( \boldsymbol{\alpha }\right) -\boldsymbol{\phi}_{0}=\mathbf{H}_{nT}^{-1}\mathbf{s}_{nT} =\boldsymbol{\tilde{\phi}}_{nT}\left( \boldsymbol{\tau}_{K}\right) =\boldsymbol{\tilde{\phi}}_{nT}$. See ((ref)).

The convergence results for $\boldsymbol{\tilde{\phi}}_{nT}\left( \boldsymbol{\alpha}\right) $ are set out in the following theorem.

theoremConsider, $\boldsymbol{\tilde{\phi}}_{nT}\left( \boldsymbol{\alpha}\right) $, the bias-corrected estimators of $\boldsymbol{\phi}_{0}$ given by ((ref)), and suppose Assumptions (ref), (ref), (ref), (ref) and (ref) hold, the strength of observed factors, $\boldsymbol{f} _{t}=(f_{1t},f_{2t},...,f_{Kt})^{\prime},$ is given by $\boldsymbol{\alpha }\mathbf{=(}\alpha_{1},\alpha_{2},...,\alpha_{K})^{\prime}$, and the strength of the missing factor, $g_{t}$, defined by ((ref)) is $\alpha_{\gamma}$. Let $\alpha_{\min}=\min_{k}(\alpha_{k})$ and suppose that $\alpha_{\gamma }<1/2$. Then \[ \mathbf{H}_{nT}\left( \boldsymbol{\alpha}\right) =\mathbf{D}_{\alpha} ^{-1}\mathbf{B}_{n}^{\prime}\mathbf{M}_{n}\mathbf{B}_{n}\mathbf{D}_{\alpha }^{-1}+O_{p}\left( T^{-1}n^{-\alpha_{\min}+1/2}\right) , \] where $\mathbf{H}_{nT}\left( \boldsymbol{\alpha}\right) $ is given by ((ref)), and by part (b) of Assumption (ref), $\mathbf{H} _{nT}\left( \boldsymbol{\alpha}\right) \rightarrow_{p}\mathbf{\Sigma} _{\beta\beta}(\boldsymbol{\alpha}\mathbf{)}>0$, for any fixed $T>T_{0}$ such that $\lambda_{\max}\left( \frac{\mathbf{F}^{\prime}\mathbf{M}_{T}\mathbf{F} }{T}\right) ^{-1}<C$ and $\alpha_{\min}>1/2>\alpha_{\gamma}$. Also \begin{align} \tilde{\phi}_{k,nT}\left( \boldsymbol{\alpha}\right) -\phi_{0,k} & =O_{p}\left( n^{-(\alpha_{k}+\alpha_{\min})/2+1/2}T^{-1/2}\right) +O_{p}\left( n^{\frac{-\left( \alpha_{k}+\alpha_{\min}\right) +\left( \alpha_{\eta+\alpha_{\gamma}}\right) }{2}}T^{-1/2}\right) \\ & +O_{p}\left( n^{-\left( \alpha_{k}+\alpha_{\min}\right) /2+\alpha_{\eta }}\right) +O_{p}\left( n^{-(\alpha_{k}+\alpha_{\min})/2+1/2}T^{-1}\right) .\nonumber \end{align}

See sub-section (ref) of the Appendix for a proof.

The result in ((ref)) establishes the consistency of $\tilde{\phi }_{k,nT}\left( \boldsymbol{\alpha}\right) =\tilde{\phi}_{k,nT}$ even if $f_{kt}$ is semi-strong so long as $n\rightarrow\infty$, and $\alpha_{\min }>1/2,$ $\alpha_{k}+\alpha_{\min}>\alpha_{\eta}+\alpha_{\gamma}$, and $\alpha_{k}+\alpha_{\min}>2\alpha_{\eta}$. Clearly, these results reduce to the case of strong factors where $\alpha_{\min}=\alpha_{k}=1$. Turning to the asymptotic distribution of $\tilde{\phi}_{k,nT}\left( \boldsymbol{\alpha }\right) $, again only convergence rates are affected, and instead of the regular rate of $\sqrt{nT}$, we have $\ \sqrt{T}n^{(\alpha_{k}+\alpha_{\min }-1)/2}$, and using ((ref)) we have

align[align omitted — 355 chars of source]

The conditions needed for eliminating the effects of the pricing errors are the same as before and are given by $\alpha_{\eta}+\alpha_{\gamma}<1$ and $\sqrt{T}n^{-1/2+\alpha_{\eta}}\rightarrow0$. The asymptotic distribution is unaffected except for the slower rate of convergence alluded to above. It is also of interest to note that adding semi-strong factors can adversely affect the convergence rate of the strong factor with $\alpha_{k}=1$. As an example suppose the asset pricing model contains two factors, one strong, $\alpha _{1}=1$ and one semi strong with $\alpha_{2}<1$.$\,\ $Then the convergence rate of $\tilde{\phi}_{1,nT}\left( \boldsymbol{\alpha}\right) -\phi_{0,1}$ is given by$\sqrt{T}$ $n^{(1+\alpha_{2}-1)/2}$ which is slower than the rate we would have obtained for $\tilde{\phi}_{1,nT}\left( \boldsymbol{\alpha }\right) -\phi_{0,1}$ if both factors were strong ($\alpha_{min}=\alpha _{k}=1)$, namely the regular rate of $\sqrt{nT}$.

Furthermore, when conditions $\alpha_{\eta}+\alpha_{\gamma}<1$ and $\sqrt {T}n^{-1/2+\alpha_{\eta}}\rightarrow0$ are met we have

equation[equation omitted — 167 chars of source]

and $\phi_{0,k}$ is consistently estimated if $T^{-1/2}n^{-(\alpha_{k} +\alpha_{\min}-1)/2}\rightarrow0$. Also, using ((ref)), an estimator of the risk premia, $\lambda_{k}$, is given by $\tilde{\lambda}_{k,nT}\left( \boldsymbol{\alpha}\right) =\tilde{\phi}_{k,nT}\left( \boldsymbol{\alpha }\right) +\hat{\mu}_{k,T}$, where $\hat{\mu}_{k,T}=T^{-1}\sum_{t=1}^{T} f_{kt}$. Hence \[ \tilde{\lambda}_{k,nT}\left( \boldsymbol{\alpha}\right) -\lambda _{0,k}=\left[ \tilde{\phi}_{k,nT}\left( \boldsymbol{\alpha}\right) -\phi_{0,k}\right] +\left( \hat{\mu}_{k,T}-\mu_{0,k}\right) . \] Under Assumption (ref) $\hat{\mu}_{k,T}-\mu_{0,k}=O_{p}\left( T^{-1/2}\right) $, and using ((ref)) it then follows that

equation[equation omitted — 205 chars of source]

and $\tilde{\lambda}_{k,nT}\left( \boldsymbol{\alpha}\right) $ is a consistent estimator of $\lambda_{0,k}$ if $T\rightarrow\infty$ as well as $T^{-1/2}n^{-(\alpha_{k}+\alpha_{\min}-1)/2}\rightarrow0$. More specifically, suppose $T=\ominus\left( n^{d}\right) $ for some $d>0$, where $\ominus \left( \cdot\right) $ denotes $T$ and $n^{d}$ are of the same order of magnitude. Then for any $d>0$, the condition for consistency of $\tilde {\lambda}_{k,nT}\left( \boldsymbol{\alpha}\right) $ is given by $\alpha _{k}+\alpha_{min}+d>1$ and $d>0$. In the case where all risk factors have the same strength, $\alpha$, the consistency condition reduces to $\alpha>\left( 1-d\right) /2$, which is weaker than the one derived by giglio2023test , namely $n/(\left\Vert \beta\right\Vert ^{2}T)\rightarrow0$, where, in terms of our notation, $\left\Vert \beta\right\Vert ^{2}=\ominus\left( n^{\alpha }\right) $. This latter condition will be met if $\alpha>1-d$. In practice where $T$ is small relative to $n$, the accuracy of $\tilde{\lambda} _{k,nT}\left( \boldsymbol{\alpha}\right) $ as an estimator $\lambda_{0,k}$ does depend on $\alpha$, and our weaker condition on $\alpha>\left( 1-d\right) /2$ is advantageous.

Consistent estimation of the variance of $\boldsymbol{\tilde{\phi }}_{nT}$

To carry out inference on $\boldsymbol{\phi}_{0}$, or any of its elements individually, we require a consistent estimator of $Var\left( \boldsymbol{\tilde{\phi}}_{nT}\right) $. Using ((ref)) and ((ref)) we first note that $\mathbf{\Sigma}_{\beta\beta}$ is consistently estimated by $\mathbf{H}_{nT\ }$ given by ((ref)). Therefore, it is sufficient to find a suitable estimator of $\mathbf{V}_{u}=(\sigma_{ij})$ such that $\mathbf{V}_{\xi}$ given by ((ref)) is consistently estimated. Under suitable sparsity restrictions $\mathbf{V}_{u}$ can be consistently estimated using the various thresholding procedures advanced in the statistical literature by bickel2008covariance, bickel2008regularized , cai2011adaptive , and BPS2019multiple . fan2011high, fan2013large also show that the adaptive threshold technique of Cai and Liu applies equally to the residuals from an approximate factor model. Here we consider the threshold estimator proposed by BPS which does not require cross-validation and is shown to have desirable small sample properties. It is given by $\mathbf{\tilde{V}}_{u}=\left( \tilde{\sigma}_{ij}\right) $

align[align omitted — 270 chars of source]

where

equation[equation omitted — 293 chars of source]

and $c_{p}(n,d)=\Phi^{-1}\left( 1-\frac{p}{2n^{d}}\right) ,$ is a normal critical value function, $p$ is the the nominal size of testing of $\sigma_{ij}=0$, ($i\neq j$) and $d$ is chosen to take account of the $n(n-1)/2$ multiple tests being carried out. Monte Carlo experiments carried out by BPS suggest setting $d=2$. The variance estimator given by ((ref)) does not require a knowledge of the factor strength and applies to risk factors of differing degrees.

Under Assumptions (ref), (ref), and (ref), $\left\Vert \mathbf{V}_{u}\right\Vert =O\left( n^{\alpha_{\gamma}}\right) $, and using results in fan2011high, fan2013large we have

equation[equation omitted — 156 chars of source]

Consider the following estimator of $\mathbf{V}_{\xi}$ \[ \mathbf{\hat{V}}_{\xi,nT}=\left( 1+\hat{s}_{nT}\right) \left( n^{-1}\mathbf{\hat{B}}_{nT}^{\prime}\mathbf{M}_{n}\mathbf{\tilde{V}} _{u}\mathbf{M}_{n}\mathbf{\hat{B}}_{nT}\right) \text{.} \] where $\hat{s}_{nT}=\boldsymbol{\tilde{\lambda}}_{nT}^{^{\prime}}\left( T^{-1}\mathbf{F}^{\prime}\mathbf{M}_{T}\mathbf{F}\right) ^{-1} \boldsymbol{\tilde{\lambda}}_{nT}$. Under Assumption (ref), $T^{-1}\mathbf{F}^{\prime}\mathbf{M}_{T}\mathbf{F\rightarrow}_{p} \mathbf{\Sigma}_{f}$ $\ $and using the results above we have $\boldsymbol{\tilde{\lambda}}_{nT}=\boldsymbol{\tilde{\phi}}_{nT} +\boldsymbol{\hat{\mu}}_{T}\rightarrow_{p}\boldsymbol{\phi}_{0} +\boldsymbol{\mu}_{0}=\boldsymbol{\lambda}_{0}$. Hence, $\hat{s} _{nT}\rightarrow_{p}\boldsymbol{\lambda}_{0}^{\prime}\mathbf{\Sigma}_{f} ^{-1}\boldsymbol{\lambda}_{0}$ as $n,T\rightarrow\infty$, jointly, and it is sufficient to show that

equation[equation omitted — 262 chars of source]

The following theorem provides a formal statement of the conditions under which $\mathbf{\hat{V}}_{\xi,nT}$ is a consistent estimator of $\mathbf{V} _{\xi}$.

theoremSuppose Assumptions (ref), (ref), (ref), (ref) and (ref) hold, and all the observed factors are strong, ($\alpha_{k}=1$, for $k=1,2,...,K$), and the strength of the missing factor, $g_{t}$, defined by ((ref)), $\alpha_{\gamma}$ $<1/2$. Then \begin{equation} \left\Vert \mathbf{\hat{V}}_{\xi,nT}-\mathbf{V}_{\xi}\right\Vert =O_{p}\left( n^{\alpha_{\gamma}}\sqrt{\frac{\ln(n)}{T}}\right) , \end{equation} where \begin{equation} \mathbf{\hat{V}}_{\xi,nT}=\left( 1+\hat{s}_{nT}\right) \left( n^{-1}\mathbf{\hat{B}}_{nT}^{\prime}\mathbf{M}_{n}\mathbf{\tilde{V}} _{u}\mathbf{M}_{n}\mathbf{\hat{B}}_{nT}\right) , \end{equation} $\mathbf{\tilde{V}}_{u}=\left( \tilde{\sigma}_{ij}\right) $, $\tilde{\sigma }_{ij}$ is the threshold estimator of $\sigma_{ij}$ given by ((ref) ), and \begin{equation} \mathbf{V}_{\xi}=\left( 1+\boldsymbol{\lambda}_{0}^{\prime}\mathbf{\Sigma }_{f}^{-1}\boldsymbol{\lambda}_{0}\right) p\lim_{n\rightarrow\infty }\left( n^{-1}\mathbf{B}_{n}^{\prime}\mathbf{M}_{n}\mathbf{V}_{u} \mathbf{M}_{n}\mathbf{B}_{n\ }\right) , \end{equation}

For a proof see sub-section (ref) in the Appendix.

This theorem shows that consistent estimation of $Var\left( \boldsymbol{\tilde{\phi}}_{nT}\right) $ can be achieved by using a suitable threshold estimator of $\mathbf{V}_{u}$, so long as the strength of the missing factor, $\alpha_{\gamma}$, is sufficiently weak in the sense that $n^{\alpha_{\gamma}}\sqrt{\ln(n)/T}\rightarrow0$ as $n,T\rightarrow\infty$.

Small sample properties of the estimators and tests for $\boldsymbol{\phi}$

Monte Carlo Design

This section presents Monte Carlo simulations to investigate the small sample properties of estimators and tests for $\boldsymbol{\phi}_{0}$. In the empirical application of the next section the factors are selected from a large list. But here we assume $K=3$ and mimic the 3 Fama-French factors, namely the market return minus the risk free rate, MKT, the value factor (high minus low book to market portfolios, HML) and the size factor (small minus big portfolios, SMB). These are denoted by $f_{kt},$ $k=M,H,S$.\footnote{Data on factors and the risk free rate are downloaded from Kenneth French's data library: https://mba.tuck.dartmouth.edu/pages/faculty/ken.french/data_library.html} For further details see Section (ref) of the online supplement A.

Loadings and factor strengths

To calibrate the loadings, $\beta_{ik}$, we used excess returns on a large number securities observed over the shorter sample covering the 20 years $2002m1$ $-2021m12$ ($T=240$). Monthly returns for NYSE and NASDAQ stocks code 10 and 11 from CRSP were downloaded from Wharton Research Data Services and converted to excess returns over the risk free rate, taken from Kenneth French's webpages, in percent per month. Only stocks with available data for the full sample were included, yielding a balanced panel, and to avoid outliers influencing the results, stocks with a kurtosis greater than 16 were excluded. There were $1289$ stocks before exclusion on the basis of kurtosis and $1175$ after. The summary statistics giving mean, median, standard deviation of the estimates of $\beta_{ik}$ and their histograms are provided in the online supplement A.

For factor strength, we considered a range of DGPs. Given the evidence that most factors, other than the market factor, are not strong, we focus on the case where there is one strong factor, namely the market factor with $\alpha_{M}=1,$ plus two semi-strong factors, with the value factor, $HML$, being quite strong with $\alpha_{H}=0.85,$ and the size factor, $SML$, being only moderately strong with $\alpha_{S}=0.65$. These estimates are also informed by the results provided in bailey2021measurement who propose methods for estimation of factor strength. For a given factor strength, $\alpha_{k}$, the associated loadings, $\beta_{ik}$, are generated as $\boldsymbol{\beta}_{k}=(\beta_{1k},\beta_{2k},,...,\beta_{\alpha_{k} },0,0,...0)$ where $n_{\alpha_{k}}=\lfloor n^{\alpha_{k}}\rfloor$ the integer part of $n^{\alpha_{k}}$, with non-zero and zero values of $\boldsymbol{\beta }_{k}$ given by

align*[align* omitted — 236 chars of source]

where $\lfloor n^{\alpha_{k}}\rfloor$ denotes the integer part of $n^{\alpha_{k}}$. \ Since the security returns are randomly generated, it does not matter how zero and non-zero values of $\beta_{ik}$ are distributed across $i$. Also, the zero loadings can also be replaced by an exponentially decaying sequence without any implications for the simulation results.\footnote{See also footnote 5 of \citet *[p.942]{bailey2016exponent}.} We also set

align*[align* omitted — 186 chars of source]

which match the mean and standard deviation of the estimates of $\beta_{ik}$. See above.

Generation of pricing errors

The pricing errors in ((ref)) can be considered as firm-specific characteristics and are set as $\boldsymbol{\eta}_{n}=\left( \eta_{1} ,\eta_{2},...,\eta_{n_{\eta}},0,0,...,0\right) ^{\prime}.$ The non-zero loadings of $\boldsymbol{\eta}_{n}$ for $i\leq n_{\eta}=\lfloor n^{\alpha _{\eta}}\rfloor$\ are drawn from $IIDU(0.7,0.9)$, and $\eta_{i}=0$ for $i=n_{\eta}+1,n_{\eta}+2,....,n.$ We consider $\alpha_{\eta}=(0,0.3).$ When $\alpha_{\eta}=0$ we have $\eta_{i}=0$ ,\ for all $i.$ As in the case of factor loadings the non-zero values of $\boldsymbol{\eta }_{n}$ must be randomly allocated to different groups.

Generation of return equation errors

The return equation errors, $u_{it}$, are generated following ((ref)) as a combination of a missing factor, $g_{t}\sim IIDN\left( 0,1\right) $ plus an idiosyncratic error, $v_{jt}.$ The loadings $\boldsymbol{\gamma}=\left( \gamma_{1},\gamma_{2},...,\gamma_{n_{\gamma}},0,0,...,0\right) ^{\prime}$ of the missing factor are set as

align*[align* omitted — 221 chars of source]

where $\alpha_{\gamma}$ is the strength of the missing factor $g_{t}$. We consider $\alpha_{\gamma}=1/4$ and $1/2$.

For the idiosyncratic errors, $v_{it}$, we consider spatial as well as a block diagonal specification, with the spatial specification including a diagonal specification as the special case. Under the spatial specification the idiosyncratic errors are generated as the first order spatial autoregressive model $v_{it}=\rho_{\varepsilon}\sum_{j=1}^{n}w_{ij}v_{jt}+\kappa \varepsilon_{it},$ which can be written in matrix notation as $\mathbf{v} _{t}=\rho_{\varepsilon}\mathbf{Wv}_{t}+\kappa\boldsymbol{\varepsilon}_{t},$ and solved for as $\mathbf{v}_{t}=\kappa\left( \mathbf{I}_{n}-\rho _{\varepsilon}\mathbf{W}\right) ^{-1}\boldsymbol{\varepsilon}_{t}$. Adding the missing factor now yields

equation[equation omitted — 178 chars of source]

The spatial coefficient $\rho_{\varepsilon}$ is such that $|\rho_{\varepsilon }|<1$, $\mathbf{W=(}w_{ij})$ with $w_{ii}=0$, and $\sum_{j=1}^{n}w_{ij}=1$. The diagonal case is obtained by setting $\rho_{\varepsilon}=0$, with $\rho_{\varepsilon}=0.5$ characterizing the SAR specification. The weight matrix $\mathbf{W}=(w_{ij})$ is set to follow the familiar rook pattern where all its elements are set to zero except for $w_{i+1,i}=w_{j-1,j}=0.5$ for $i=1,2,...,n-2$ and $j=3,4...,n$, with $w_{1,2}=w_{n,n-1}=1$.

Under the block error covariance specification, $\mathbf{v}_{t}$ is generated as $\mathbf{v}_{t}=\kappa\mathbf{\hat{S}}\boldsymbol{\varepsilon}_{t},$ where $\mathbf{\hat{S}}$ is a block diagonal matrix with its $b^{th}$ block given by $\mathbf{\hat{S}}_{b}$ for $b=1,2,...,B$, and $\boldsymbol{\varepsilon} _{t}=(\mathbf{\varepsilon}_{1t}^{\prime},\mathbf{\varepsilon}_{2t}^{\prime },...,\mathbf{\varepsilon}_{Bt}^{\prime})^{\prime}$, and $\boldsymbol{\varepsilon}_{bt}=(\varepsilon_{b,1t},\varepsilon_{b,2t} ,...,\varepsilon_{b,n_{b},t})^{\prime}$. $\mathbf{\hat{S}}$ is set as a Cholesky factor of the correlation matrix of $\mathbf{u}_{t}$. Denoting this correlation matrix by $\mathbf{\hat{R}}_{u},$ \[ \mathbf{\hat{R}}_{u}=\left[ Diag(\mathbf{\hat{V}}_{Bu})\right] ^{-1/2}\mathbf{\hat{V}}_{Bu}\left[ Diag(\mathbf{\hat{V}}_{Bu})\right] ^{-1/2}=Diag(\mathbf{\hat{R}}_{bu},\text{ }b=1,2,...,B), \] where $\mathbf{\hat{V}}_{Bu}$ is the threshold estimator of $\mathbf{V}_{u}$ subject to the additional restriction that $\mathbf{V}_{u}$ is block diagonal. For each block $\mathbf{\hat{R}}_{bu}$ we set the number of distinct non-zero elements of this block equal to the integer part of $[n_{b}(n_{b}-1)/2]\times q_{b}$ where $q_{b}$ is the proportion of non-zero distinct elements in block $b$ of our calibrated sample and computed by the calibration over the sample $2001m10-2021m9$. The non-zero elements are drawn randomly from $IIDU(0,0.5)$. Similarly, adding the missing factor, we have

equation[equation omitted — 123 chars of source]

The block diagonal structure is intended to capture possible within industry correlations not picked up by observed or weak missing factors, with each block representing an industry or sector. To calibrate the block structure estimates of the pair-wise correlations between the residuals of the return regressions using the Fama-French three factors of the $T=240$ sample ending in $2021$ were obtained. Then all the statistically insignificant correlations were set to zero, allowing for the multiple testing nature of the tests. For the majority of securities (668 out of the 1168), the pair-wise return correlations were not statistically significant. The securities with a relatively large number of non-zero correlations were either in the banking or energy related industries. Considering stocks by 2-digit SIC classifications, a division into $B=14$ contiguous groups ranging in size from $33$ to $145$ stocks, seemed sensible. More detail on the process is given in Section (ref) of the online supplement A.

The primitive errors, $\varepsilon_{it}$ for $i=1,2,...,n$ in ((ref)) and ((ref))\ are generated as $\varepsilon_{it}=\sqrt{\sigma_{ii}} \varpi_{it}$, where $\varpi_{it}\sim IIDN\left( 0,1\right) $, and $\varepsilon_{it}=\sqrt{\sigma_{ii}}\left[ \sqrt{\frac{v-2}{v}}\varpi _{it}\right] ,$ where $\varpi_{it}\sim IID$ $t(v)$, with $t(v)$ denotes a standard $t$ distributed variate with $v=5$ degrees of freedom. Also $\sigma_{ii}\thicksim IID$ \ $0.5(1+\chi_{1}^{2})$ $,\ for$\ $i=1,2,...,n_{b}$ and $b=1,2,...,B$. In this way, it is ensured that $Var(\varepsilon _{b,it})=\sigma_{b,ii}$, and on average $E\left[ Var(\varepsilon _{b,it})\right] =E(\sigma_{ii})=1$, under both Gaussian and t-distributed errors. Note that $Var\left( \nu_{b,it}\right) =v/(v-2)$. All the experiments are designed to give an $R^{2}$ of about $0.3,$ similar to that obtained in the empirical applications. For further details see sub-section (ref) of the online supplement A.

Experiments

In total, we consider 12 experimental designs: six designs with Gaussian errors and six with $t(5)$ distributed errors. We considered designs with GARCH effects, with and without pricing errors, $\eta_{i}$, and with and without the missing factor, $g_{t}$. We also considered designs with spatial patterns in the idiosyncratic errors, $v_{it}$. All experiments are implemented using $R=2,000$ replications. Details of of the 12 experiments are summarized in Table S-1 of the online supplement B (MC results).

Alternative estimators of $\mathbf{V}_{u}$

Subsection (ref) considered consistent estimation of the variance of $\tilde{\phi}_{nT}$ using $\mathbf{\tilde{V}}_{u}$ a threshold estimator for $\mathbf{V}_{u},$ given by equation ((ref))$.$ For comparison purposes we also considered two other estimators of $\mathbf{V}_{u}.$ These were the sample covariance matrix $\mathbf{\hat{V}}_{u}=\sum_{t=1} ^{T}\mathbf{\hat{u}}_{t}\mathbf{\hat{u}}_{t}^{^{\prime}}/T$ and a diagonal covariance matrix, where the off-diagonal elements of $\mathbf{\hat{V}}_{u},$ $\hat{\sigma}_{ij},$ are set to zero. Thus we have three designs for the return error covariance matrix, $\mathbf{V}_{u},$ and three different estimators of it. A comparison of the results for the different covariance matrices is available on request. The diagonal estimator, as to be expected, performed poorly when the true covariance matrix was not diagonal, particularly for the spatial error covariance matrix and when the strength of the missing factor was close to $1/2$. For these designs the sample and threshold estimators of the covariance matrix generally performed similarly and given that there is a theoretical justification for the threshold estimator and there are structures of the error covariance matrix for which the sample estimator is unlikely to perform well we report the results using the threshold estimator in the simulations below.

Monte Carlo results

We focus on a comparison of two-step (defined by ((ref)) and the bias-corrected (BC) estimator (defined by ((ref))), and report bias, root mean square error (RMSE) and size for testing $H_{0j}\,:\phi_{0k}=0,$ $k=M,H,S$ at the five per cent nominal level, for all $n=100,500,1,000,3,000$ and $T=60,$ $120,$ $240$ combinations. The results for all 12 experiments are summarized in Tables S-A-E1 to S-A-E12 in the online supplement B. In terms of bias and RMSE the two-step estimator does much better than the bias-corrected (BC) estimator when $T=60$ and $n=100$, but this gap closes quickly as $n$ is increased. In fact for $T=60$ and $n=3,000,$ the bias and RMSE of the BC estimator (at $0.0010$ and $0.0607)$ are much less than those of the two-step estimator (at -$0.0080$ and $0.1489)$. This pattern continues to hold when $T=120$ and $240$. Bias correction can cause the RMSE to "blow up" for small samples, such as $n=100$ and $T=60$, but for $n=500$ and above the bias-corrected estimator always has a smaller RMSE than the two-step estimator. As discussed in the theoretical section, having a large $n$ is important for the properties of the estimators.

But most importantly, the two-step estimator is subject to substantial size distortions, particularly when $T$ is small relative to $n$. As predicted by the theory, the degree of over-rejection of the tests based on FM estimator falls with $T,$ but increases with $n$. For example, the two-step test sizes rise from $11.1\%$ when $T=60$ and $n=100$ to $60.9\%$ when $T=60$ and $n=3,000$. Increasing $T$ reduces the size distortion of the two-step estimator but test sizes are still substantially above the $5\%$ nominal value when $n$ is large. The strong tendency of the tests based on the two-step\ estimator to over-reject could be an important contributory factor leading to false discovery of a large number of apparently significant factors in the literature. In contrast, sizes of the tests based on the BC estimator, using the variance estimator given by ((ref)), are all close to its nominal value, irrespective of the factor strength or sample size combinations. We only note some elevated test sizes in the case of the experimental design 12, and when we consider the semi-strong factors. The highest test size of $7.85$ per cent is obtained for the least strong factor, $f_{st}$, when $n=3000$ and $T=60$. See Table S-A-E12 of the online supplement B.

We also experimented with raising $\alpha_{\eta}$ from $0.3$ to $0.5,$ making the pricing errors much more pervasive and $\rho_{\varepsilon}$ from $0.5$ to $0.85,$ introducing more spatial correlation$.$ This increased the rejection rate in experiments 9 and 10.

Empirical power functions

Plots of the empirical power functions for testing the null hypothesis $H_{0j}\,:\phi_{0k}=0,$ $k=M,H,S$, are also provided in the online supplement B for the 12 experiments (Figures S-A-E1 to S-A-E12). All power functions have the familiar bell curve shape and tend to unity as $n$ and/or $T$ are increased, showing the test has satisfactory power, particularly for $n$ sufficiently large even when $T=60$. Again the power functions are quite similar across the 12 different experiments and show similar patterns for strong and semi-strong factors. However, this similarity hides the fact that the test of $\phi_{k}=0$ for the strong factor is much more powerful than corresponding tests for the semi-strong factors, with the test power declining as factor strength is reduced.

Differences in performance of strong and semi-strong factors

These differences in the effects of factor strength on the power of the test of $\phi_{k}=0$ are in line with our theoretical results, and are also reflected in the rate at which the RMSE of the estimators of $\phi_{M}$, $\phi_{H}$ and $\phi_{S}$ fall with $n$. For example, using results in Tables S-B-E10-12 in the online supplement B for the bias-corrected estimator in the case of design 12 with $T=240,$ we note that the ratio of RMSE of $n=3,000$ to $n=100$ is $17\%$ for the strong factor $\left( \alpha_{M}=1\right) $, $23\%$ for the first semi-strong factor with $\alpha_{H}=0.85$, and $36\%$ for the second semi-strong factor with $\alpha_{S}=0.65.$ As strength falls one needs larger cross section samples of securities to attain the same level of precision. In the case of the two-step estimator there was the same pattern, but the fall in the RMSE with $n$ was much slower. The ratio for $T=240$ of RMSE of $n=3,000$ to $n=100$ is $25\%$ for $\phi_{M}$, rather than $17\%$ (for the bias-corrected estimator); $48\%$ rather than $23\%$ for $\phi_{H}$.

Misspecification: Semi-strong versus weak factors

So far, we have assumed that the DGP is correctly specified with two semi-strong and no observed weak factors. Here we consider the implications of incorrectly excluding semi-strong factors or correctly including weak factors on the small sample properties of the bias-corrected\ estimator of $\phi_{M},$ the coefficient of the strong factor. Using the same DGP (which includes one strong factor and two semi-strong factors), we carried out additional MC experiments (designs 1-12) where we also estimated $\phi_{M}$ without the semi-strong factors being included in the regressions. Comparative results, with and without the semi-strong factors, are summarized in Tables S-C-E1-3 to S-C-E10-12 of the online supplement B. We find that incorrectly excluding semi-strong factors can be quite costly, both in terms of bias and RMSE as well as size distortions. In terms of RMSE it was almost always better to estimate the model with the semi strong factors included. The exception was for the case of $T=60,$ $n=100,$ where including the semi-strong factors caused the RMSE to blow up. Size distortions resulting from the exclusion of the semi-strong factors tended to be more pronounced for large $n$ and $T$ samples. These conclusions were not sensitive to the choice of the experimental design. For these experiments the lesson seems to be that it is important to have $n$ large and include relevant semi-strong factors provided that they are sufficiently strong.

When the DGP includes one strong factor ($\alpha_{M}=1$) and two weak factors ($\alpha_{H}=$ $\alpha_{S}=0.5)$, in terms of the bias and RMSE for $\phi_{M}$ it is unambiguously better to exclude the weak factors from the regression, even though they are in the DGP. Weak factors are best treated as missing and absorbed in the error term.

Main conclusions from MC experiments

The conclusions from the Monte Carlo simulations are that the bias corrected estimator of $\boldsymbol{\phi}_{0}\boldsymbol{,}$ generally works well. Although, it can generate a large RMSE for small $n$ and $T$, this can be solved by increasing $n.$ This performance is robust to non-Gaussian errors, GARCH effects, missing weak factors, pricing errors and weak cross-sectional dependence. Test sizes are generally correct and the power good. The rate at which RMSEs decline with $n$ depends on the strength of the underlying factors. Semi-strong factors need much larger values of $n$ for precise estimation. Tests of the joint significance of $\phi_{M}=\phi_{H}=\phi_{S}=0,$ not reported here, also performed well, as might be expected given the good power performance of the separate induced tests which are reported. Including weak factors could be harmful, but there are potential advantages of adding semi-strong factors, although the issue of how best to select such factors is an open question to which we now turn.

Factor selection when the number of securities and the number of factors are both large

Our theoretical derivations and Monte Carlo simulations both assume that the number of risk factors included in the return regressions is fixed and the factors are known. For the empirical application we face the additional challenge of selecting a small number of relevant risk factors from a possibly large number of potential factors, $m$. This problem has been the subject of a number of recent studies. harvey2016and propose a multiple testing approach aimed at controlling the false discovery rate in the process of factor selection, and in a more recent paper harvey2020false suggest using a double-bootstrap method to calibrate the t-statistic used for controlling the desired level of the false discovery rate. giglio2021asset suggest applying the double-selection Lasso procedure by belloni2014inference to second pass regressions. None of these methods distinguish between strong, semi-strong or weak factors in their selection process, whereas the theory and simulations presented above indicate the importance of factor strength for estimation and inference.

Here we propose an alternative selection procedure where we first estimate the strength of all the $m$ factors under consideration, and then select factors with strength above a given threshold, the value of which is informed by the convergence results of Theorem (ref). This theorem showed that if factor $f_{kt}$ has strength $\alpha_{k},$ then for a given $T$ the BC estimator of $\phi_{k}$ converges to its true value, $\phi_{0k}$, at the rate of $n^{(\alpha_{k}+\alpha_{\min}-1)/2}$, where $\alpha_{\min}=\min_{i}(\alpha _{i})$. As is recognized in non-parametric estimation literature, if the rate of convergence is less than $1/3,$ the gain in precision with $n$ is so slow that the estimator may not be that useful.\footnote{The manski1985semiparametric maximum score estimator for a binary response model has $n^{1/3}$ convergence and this is regarded as very slow and there are suggested modifications such as horowitz1992smoothed to increase the rate of convergence to $n^{2/5}$.} To achieve rate of $n^{1/3}$ we need to set the threshold value of $\alpha_{k}$, denoted by $\underline{\alpha}$, such that $\underline{\alpha}+\alpha_{\min}>1+2/3$. The smallest value of such a threshold is obtained when $\alpha_{\min }=\underline{\alpha}$ or if $\underline{\alpha}>1/2+1/3.$ Given the threshold the main issue is how to estimate factor strength. This problem is already addressed in \citet*[BKP]{bailey2021measurement} when $m$ is fixed. In this setting they base their estimation on the statistical significance of $f_{kt}$ in the first stage time-series regressions of excess returns on all the factors under consideration, whilst allowing for the $n$ multiple testing problem which their approach entails. When $m\,$(the number of factors) is also large the first stage regressions will also be subject to the multiple testing problem and penalized regression techniques such as Lasso or the one covariate at a time (OCMT) selection procedure technique proposed by chudik2018one could be used. Irrespective of selection technique used at the level of individual security returns, we end up with $n$ different subsets of the $m$ factors under consideration. Factor strengths can then be estimated similarly to BKP\ from their selection frequencies across the $n$ securities.

To be more specific, denote the set of $m$ factors under consideration by $\mathcal{S}$ and denote the set of selected factors for security $i$ by $\hat{S}_{i}$ and their numbers by $\hat{m}_{i}=|\hat{S}_{i}|$. Clearly $\hat{S}_{i}\subseteq\mathcal{S}$, and $\hat{m}_{i}\leq m$, for $i=1,2,...,n$. Then compute the proportion of stocks in which the $k^{th}$ factor is selected, $\hat{\pi}_{k},$ for $k\in\{1,2,...,m\}$ based on $\hat{S}_{i},$ $i=1,2,...,n$ by $\hat{\pi}_{k}=\frac{1}{n}\sum_{i=1}^{n}\mathcal{I}\{k\in \hat{S}_{i}\}$. Then the strength of the $k^{th}$ factor is measured by

equation[equation omitted — 182 chars of source]

The transformation from $\hat{\pi}_{k}$ to $\hat{\alpha}_{k}$ is explained and justified in BKP, where it is shown that considering strength aids interpretation because it is not dependent on $n$. It is beyond the scope of the present paper to provide theoretical justification for the proposed factor selection procedure, but using extensive Monte Carlo experiments yoo2022factor has shown that the proposed method has desirable small sample properties whether Lasso or OCMT is used for factor selection at the level of individual security returns.

An empirical application using a large number of U.S. securities and a large number of risk factors

This section uses the results above in the explanation of monthly returns for a large number, $n,$ of U.S. securities, by a large active set of $m$ potential risk factors. We first briefly describe the sources and characteristics of the data for the stock returns and factors, which cover different sub-samples over the period $1996m1-2022m12.$ We then consider the selection of a subset of $K$ factors from the active set. Finally we test $\boldsymbol{\phi}_{0}=0$, and construct and evaluate phi-portfolios and corresponding mean-variance portfolios for alternative models.

Monthly returns (inclusive of dividends) for NYSE and NASDAQ stocks from CRSP with codes $10$ and $11$ were downloaded from Wharton Research Data Services. They were converted to excess returns by subtracting the risk free rate, which was taken from Kenneth French's data base. To obtain balanced panels of stock returns and factors, only variables for which there was data for the full sample under consideration were used. Excess returns are measured in percent per month. To avoid outliers influencing the results, stocks with a kurtosis greater than 16 were excluded. To examine factor selection, four samples were considered, each had $20$ years of data, $T=240$, ending in $2015m12$, $2017m12$, $2019m12$, $2021m12$. Filtering out the stocks with kurtosis larger than $16$ removed about $100$ of the roughly $1200$ stocks. The number of stocks ($n$) considered for each of the four $T=240$ samples are given in panel A of Table (ref). Further detail is given in Section (ref) of the online supplement A.\footnote{Summary statistics for the excess returns across the different samples are given in Table (ref) of the online supplement A.} For analysis of the phi-portfolios, the factors selected in the sample ending in $2015m12$ were used to construct portfolios up to $2022m12.$

Factor selection

For factor selection we used a sample of $T=240$ observations\footnote{Some results for $T=120$ are included in the online supplement.}. The set of factors considered combine the $5$ Fama-French factors with the $207$ factors from the chen2022open , Open Source Asset Pricing webpages, both downloaded July 6 2022. Only factors with data for the full sample were considered so the return regressions constitute a balanced panel. The number of factors in each of the four $20$ year samples ending in the years $2015$, $2017$, $2019$, $2021$ is also given in panel A of Table (ref), and range between $187$ to $199$. Summary statistics for the factors in the active set are given in (ref) of the online supplement A.

To implement the factor selection procedure set out in Section (ref), Lasso is used to carry out selection in the return regressions for individual securities and we refer to the factor selection procedure as pooled Lasso (PL). As is well known, Lasso does not work well with too many highly correlated regressors, therefore, factors with an absolute correlation with the market factor greater than $0.70$ were dropped. This still left between $177$ and $190$ risk factors in the active set $\mathcal{S}$, depending on the sample period (see panel A of Table (ref)).

Specifically, Lasso was applied $n$ times to the regressions of excess returns, $r_{it}=R_{it}-r_{t}^{f}$, for $i=1,2,...,n$, on the $177$ to $190$ factors in the active set, $\mathcal{S}$, to select the sub-set $\hat{S}_{i}$ for each $i$ over the four 20-year samples, separately. Following the literature, the tuning parameters in the Lasso algorithm were set by ten-fold cross-validation.\footnote{The post-Lasso and one covariate multiple testing (OCMT) approach of \citet *[CKP]{chudik2018one}. were also investigated, but Lasso seemed to work reasonably well. The details of the Lasso procedure used are given in Section 2.2 of the online supplement of CKP paper.} Interestingly, the market factor was selected by Lasso for almost all the securities, thus confirming the pervasive nature of the market factor. No other factor came close to being selected for all the securities. Lasso tended to choose a lot of non-market factors and every factor got chosen in at least one return regression. The mean number of non-market factors chosen by Lasso fell from $11.6$ in the $2015$ sample to $9.9$ in the $2021$ sample. The median was lower, falling from $10$ to $8$. There was a long right tail because Lasso tended to choose a very large number of non-market factors for some securities, ranging from a maximum of $48$ in the 2021 sample to $54$ in the $2017$ sample.

Apart from the market factor, there are no systematic patterns for the rest of selected factors in the return equations for the individual securities. Following the theory set out in Section (ref), the $K$ factors used to estimate $\boldsymbol{\phi}$ are chosen on the basis of their factor strength.\footnote{The idea of using factor strength could also be viewed as a kind of averaging of the factors selected in individual return regressions.} A minimum threshold of $0.7$ was used. This is below the threshold value of $\underline{\alpha}=1/2+1/3$ required to achieve the convergence rate of $n^{1/3}$, and is intended to capture borderline semi-strong factors. We also consider the values of $0.75,$and $0.80$ that are quite close to $\underline{\alpha}$. The number of selected factors for different choices of factor strength threshold is given in panel B of Table (ref), for the four different samples.

Using the threshold of $0.70$, $17$ factors (inclusive of the market factor) were selected for the sample ending in $2015$, with the number of selected factors declining to $15,13$ and $11,$ for the samples ending $2017,$ $2019$ and $2021$, respectively. At the other extreme, setting the threshold at $0.80$, the number of selected factors dropped to $4$ for the samples ending in $2015$, $2017$ and $2021$, and $2$ for the sample ending in $2021$. Since $17$ factors seemed too large and $2$ factors too small, the threshold value was set at the intermediate value of $0.75$. We considered always conditioning on the market factor, but since Lasso almost always selected it, this was unnecessary.

\floatstyle{plaintop} \restylefloat{table}

table[table omitted — 1,316 chars of source]

\onehalfspacing The list of factors selected by pooled Lasso for the four samples are given in Table (ref). The three Fama-French factors, Market, HML and SMB, are all selected in all four periods.\footnote{We also considered selecting the risk factors using the generalized one covariate at a time (OCMT) method proposed by sharifvaghefi2022variable . Using GOCMT the Fama-French three factors were again amongst the five strongest factors selected. The use of GOCMT\ for factor selection is also investigated by yoo2022factor , using Monte Carlo and empirical applications.} Of the Fama-French three, only the market factor is strong, with estimated strength in excess of $0.98$ across the four periods.\footnote{These results also support the choice of 3 Fama-French factors and their strength used in our Monte Carlo simulations.} The other factor which is selected across all the four periods is "short selling". This is proposed by dechow2001short who argue that short-sellers target firms that are priced high relative to fundamentals. It measures the extent to which investors are shorting the market as reflected in Compustat data. Two additional factors are selected in periods ending in $2019$ and earlier. One is "Beta Tail Risk" proposed by kelly2014tail which estimates a time-varying tail exponent from the cross section of returns. The other is "Cash Based Operating Profitability" (CBOP) suggested by ball2016accruals . This is operating profit less accruals, with working capital and R&D adjustments. For periods ending in 2017 and 2015 the "Sin Stock" indicator proposed by hong2009price is also selected. It takes the value of unity if the stock in question is involved in producing alcohol, tobacco, and gaming. They find that such stocks are held less by norm-constrained institutions such as pension plans.

The factor strengths are relatively stable across the periods, with many of the estimates close to the threshold value of $0.75$. Apart from the market factor only SMB, Short Selling and Beta Tail Risk factors have strengths in excess of $0.85$ when averaged across the four periods. From the large number of factors in the active set we have ended up with relatively few factors that are reasonably strong and for which $\boldsymbol{\phi}_{0}$ can be estimated reasonably accurately.\footnote{We do not report estimates of individual $\phi_{k}.$ Because of correlations between the loadings, the sign, size and significance of the coefficients are difficult to interpret and for phi-portfolio construction, discussed below, what matters is $\boldsymbol{\phi }^{\prime}\boldsymbol{\phi}$ which determines the return on the portfolio.}

\singlespacing

table[table omitted — 1,310 chars of source]

\onehalfspacing

The strengths of the selected factors are also closely related to the average measures of fit often used in the literature. Here we consider both $Ave\bar{R}^{2}=n^{-1}\sum_{i=1}^{n}\bar{R}_{i}^{2}$, a simple average of the fit of the individual return regressions adjusted for degrees of freedom, $\bar{R}_{i}^{2}=1-(T-K-1)^{-1}\sum_{t=1}^{T}\hat{u}_{it}^{2}/T^{-1}\sum _{t=1}^{T}\left( r_{it}-\bar{r}_{i\circ}\right) ^{2}$, and the adjusted pooled $R^{2}$ defined by $\overline{PR}^{2}=1-\widehat{\bar{\sigma}}_{nT} ^{2}/s_{r,nT}^{2}$, where $\widehat{\bar{\sigma}}_{nT}^{2}$ is the bias-corrected estimator of $\bar{\sigma}_{n}^{2}$ defined by ((ref)) and $=$ $\left( nT\right) ^{-1}\sum_{i=1}^{n}\sum_{t=1}^{T}\left( r_{it}-\bar{r}_{i\circ}\right) ^{2}$. Both of these measures behave very similarly, but the pooled version is less sensitive to outliers. As shown in Appendix (ref), for sufficiently large $n$ and $T$, $\overline{PR}^{2}$ is dominated by the contribution of the most strong factor(s). Since the only strong factor selected is the market factor in Table (ref)\ we report the $Ave\bar{R}^{2}$ and $\overline {PR}^{2}$ in the case of return regressions which just include the market factor and those which include all other factors with strength in excess of $0.75$. First, we note that the $\overline{PR}^{2}$ values are generally lower than the $Ave\bar{R}^{2}$. Second, the additional factors do add to the fit, but their relative contributions vary considerably across sample sizes and periods. In general, the marginal contribution of non-market factors tend to be smaller when $T$ is larger, which is consistent with the theory for adjusted pooled $R^{2}$ set out in Section (ref) of the online supplement A.

\singlespacing

table[table omitted — 1,743 chars of source]

\onehalfspacing

Testing for non-zero $\boldsymbol{\phi}$

In principle, if $\boldsymbol{\phi}\neq0$ there are potentially exploitable excess returns. In practice, to construct an effective phi-portfolio a large number of securities is required and rebalancing such long-short portfolios for so many securities may not be feasible or may incur high transactions costs. In addition, model uncertainty, estimation uncertainty, time variation in both $\boldsymbol{\beta}_{i}$ and in conditional volatility pose additional difficulties in implementing a strategy to exploit the potential returns revealed by $\boldsymbol{\phi}$. We will abstract from such practical difficulties to provide some indication of the performance of phi-portfolios relative to alternatives which would face similar difficulties.

As our preferred asset pricing model, we consider the seven factors selected by pooled Lasso using the sample ending in $2015m12,$ which we label as PL7.\footnote{The selection of the PL7 model was reported in the earlier version of the paper submitted for publication and was not informed by the performance the phi-portfolio that we report in this version of the paper.} Recall that the PL7 includes the 3 Fama-French factors (Mkt., SMB, and HML) plus the four risk factors, Short Selling, CBOP, Beta Tail Risk, and Sin Stocks. But given uncertainties that surround the problem of model selection we also considered the two popular FF factor models, namely FF3 and FF5. The latter augments FF3 with RMW (robust minus weak operating profitability) and CMA (conservative minus aggressive investment portfolios). The three models are estimated using twenty-year rolling windows covering the $84$ months from $2015m12$ to $2022m11$, so that we can generate out of sample return forecasts for the months $2016m1-2022m12$.\footnote{Due to entry and exit of securities the number of securities included in our analysis varied across the rolling sample periods. We started with $n=1,090$ securities for the first rolling sample ending in $2015m12$, with the number of available securities with $240$ months of data falling to $953$ by $2017$, $838$ by $2019$, $767$ by $2021$, and $736$ by $2022$.} For all $3\times84$ model-sample interactions we computed the following Wald test statistics \[ W_{t\left\vert T\right. }^{2}=\boldsymbol{\tilde{\phi}}_{t\left\vert T\right. }^{\prime}\left[ \widehat{Var\left( \boldsymbol{\tilde{\phi} }_{t\left\vert T\right. }\right) }\right] ^{-1}\boldsymbol{\tilde{\phi} }_{t\left\vert T\right. }, \] for testing the null hypothesis $\boldsymbol{\phi}=0,$ where $\boldsymbol{\tilde{\phi}}_{t\left\vert T\right. }$ and $\widehat{Var\left( \boldsymbol{\tilde{\phi}}_{t\left\vert T\right. }\right) }$ denote the rolling versions of ((ref)) and ((ref)).\footnote{The formulae for the rolling estimates are provided in the sub-section (ref) of the online appendix.} For the PL7 the range of the test statistic was from 141.8 to 25.8, as compared to the 5 per cent $\chi^{2}(7)$ critical value of 14.07. Thus the hypothesis that $\boldsymbol{\phi}=0$ is strongly rejected in all the 84 rolling sample for PL7. This is also true for the FF5 and FF3 models where the test statistic ranged from 162.7 to 25.0, and from 67.9 to 25.8, compared to the 5 per cent critical values of 11.07 and 7.8, respectively.

The rolling values of the Wald statistics for testing $\boldsymbol{\phi}=0$ for the three models are shown in Figure 1. The horizontal line (pink) represents the critical value of the $\chi_{7}^{2}$ distribution at the 5 per cent level. The time profiles of these test statistics clearly show that $\boldsymbol{\phi}=0$ is rejected for all rolling samples and for all three models. But there is also a clear downward trend showing that the evidence against $\boldsymbol{\phi}=0$ has been getting weaker over time, irrespective of the choice of asset pricing model.

figure[figure omitted — 228 chars of source]

Comparative performance of phi and MV portfolios

Having established that most likely $\boldsymbol{\phi\neq0}$ for the asset pricing models we have considered, we now turn to the performance of phi-portfolios based on these models. Using the recursive version of the phi-portfolio given by ((ref)), we consider the following phi-portfolio returns

\[ \hat{\rho}_{t+1,\phi}=\boldsymbol{\tilde{\phi}}_{t\left\vert T\right. }\left[ \left( \mathbf{\hat{B}}_{t\left\vert T\right. }^{\prime} \mathbf{M}_{n}\mathbf{\hat{B}}_{t\left\vert T\right. }\right) ^{-1} \mathbf{\hat{B}}_{t\left\vert T\right. }\mathbf{M}_{n}\mathbf{r}_{\circ ,t+1}-\mathbf{f}_{t+1}\right] , \] for $t=2016m1,2016m2,...,2022m12$, using the rolling estimates $\mathbf{\hat {B}}_{t\left\vert T\right. }=\left( \mathbf{\hat{\beta}}_{1t\left\vert T\right. },\mathbf{\hat{\beta}}_{2t\left\vert T\right. }....,\mathbf{\hat {\beta}}_{nt\left\vert T\right. }\right) ^{\prime}$ with $T=240$, for each of the three factor models, FF3, FF5 and PL7. We compare the annualised Sharpe ratios of phi-portfolios with the ones based on associated MV\ portfolios, given by $\rho_{t+1,MV}=\boldsymbol{\mu}_{R}^{\prime}\mathbf{V}_{R} ^{-1}\mathbf{r}_{\circ,t+1}$.\footnote{Given our focus on the Sharpe ratios, we have set the scaling of the MV portfolio to unity.} Although in principle, MV portfolios can be constructed without a reference to a particular factor model, reliable estimation of $\boldsymbol{\mu}_{R}$ and $\mathbf{V}_{R}^{-1}$ are challenging when $n$ is relatively large. For example, the rolling sample covariance matrix estimator of $\mathbf{V}_{R}$, given by $\mathbf{\mathring {V}}_{R,t\left\vert T\right. }=T^{-1}\sum_{\tau=t-T+1}^{t}\left( \mathbf{r}_{\circ,\tau}-\mathbf{\bar{r}}_{\circ,t\left\vert T\right. }\right) \left( \mathbf{r}_{\circ,\tau}-\mathbf{\bar{r}}_{\circ,t\left\vert T\right. }\right) ^{\prime}$, with $\mathbf{\bar{r}}_{\circ,t\left\vert T\right. }=T^{-1}\sum_{\tau=t-T+1}^{t}\mathbf{r}_{\circ,\tau}$, will be singular when $n>T$, and can be very poorly estimated if $T$ is not sufficiently large relative to $n$. There is a vast literature on consistent estimation of high dimensional covariance matrices like $\mathbf{V}_{R}.$ fan2011high use observed factors while fan2013large use principal components to filter out the effects of strong factors, in both cases assuming $\mathbf{V}_{u}$ is sparse, and then using a threshold method to estimate it. Shrinkage estimators of $\mathbf{V}_{R}$ are also proposed in the literature with a recent survey provided by LedoitWolfsurvey2022. However, the shrinkage estimators require $n$ and $T$ to be of the same order of magnitude and do not work well when $n$ is much larger than $T$, as in the present application. We follow fan2011high and base our estimation of $\boldsymbol{\mu}_{R}$ and $\mathbf{V}_{R}$ on the same factor model used to construct the phi-portfolio. For a given factor model, characterized by $c\mathbf{,B}$, and $\mathbf{F}$, we compute the MV portfolio returns as

\[ \hat{\rho}_{t+1.MV}=\mathbf{\hat{\mu}}_{R,t\left\vert T\right. }^{\prime }\mathbf{\hat{V}}_{R,t\left\vert T\right. }^{-1}\mathbf{r}_{\circ,t+1}, \] where $\mathbf{\hat{\mu}}_{R,t\left\vert T\right. }=\hat{c}_{t\left\vert T\right. }\mathbf{\tau}_{n}+\mathbf{\hat{B}}_{t\left\vert T\right. }\boldsymbol{\tilde{\lambda}}_{t\left\vert T\right. }$, $\boldsymbol{\tilde {\lambda}}_{t\left\vert T\right. }=\boldsymbol{\tilde{\phi}}_{t\left\vert T\right. }+\boldsymbol{\hat{\mu}}_{t\left\vert T\right. },$ \[ \mathbf{\hat{V}}_{R,t\left\vert T\right. }=\mathbf{\hat{B}}_{t\left\vert T\right. }^{\prime}\left( \frac{\mathbf{F}_{t\left\vert T\right. }^{\prime }\mathbf{M}_{T}\mathbf{F}_{t\left\vert T\right. }}{T}\right) ^{-1} \mathbf{\hat{B}}_{t\left\vert T\right. }\mathbf{+\mathbf{\tilde{V}} }_{u,t\left\vert T\right. }\mathbf{,} \] $\boldsymbol{\hat{\mu}}_{t\left\vert T\right. }=\boldsymbol{\bar{f} }_{t\left\vert T\right. }=T^{-1}\sum_{\tau=t-T+1}^{t}\mathbf{f}_{\tau}$, and $\mathbf{\mathbf{\tilde{V}}}_{u,t\left\vert T\right. }$ is the rolling estimate of $\mathbf{V}_{u}$. The algorithms used to compute the recursive estimates for the MV portfolio can be found in the sub-section (ref) in the online supplement A.

Table (ref) presents annualised SR of the phi-portfolios, for the FF3, FF5, and PL7 models, and their corresponding MV portfolios for two samples, both beginning in 2016m1, one ending in 2019m12, pre Covid-19, and one ending in 2022m12. \ In five of the six SR ratios reported in this table, the phi-portfolio has a higher SR than the corresponding MV portfolio. The exception is the SR associated to the FF5 model for the pre Covid-19 sample. This illustrates that if $\boldsymbol{\phi}\neq\mathbf{0},$ it is possible to construct portfolios that outperform the mean-variance portfolio. Amongst the 3 models considered, the phi-portfolio based on the PL7 model performed best during the pre Covid-19 and the full sample, even beating the S&P 500. The Sharpe ratio of the phi-portfolio based on the Pl7 model was 1.95 compared to 0.94 for the S&P 500 during the pre Covid-19 period, and fell sharply to 0.65 for the full sample as compared to 0.58 for the S&P 500. The SR for the MV portfolio using the same model were 0.87 and 0.42 for the two samples, respectively.

The sharp decline in the SRs as we add the post Covid-19 years is in line with the strong downward trend in the Wald statistics for the test of $\boldsymbol{\phi}=\mathbf{0}$ shown in Figure 1. As is well known SRs have large standard errors, and in the case of our application that are around 1, so none of the SRs are significantly different from zero, with the possible except of the largest SR of 1.95 for phi-portfolio based on PL7 model for the pre Covid-19 period. We also note that the reported Sharpe ratios do not allow for transaction costs and the fact that shorting might not be feasible for all the securities.

table[table omitted — 1,147 chars of source]

Concluding remarks \onehalfspacing

In this paper we have highlighted the importance of decomposing the risk premia, $\boldsymbol{\lambda}$, into the the factor mean, $\boldsymbol{\mu}$, and $\boldsymbol{\phi}$, and writing the alpha of security $i$, $\mathit{\alpha}_{i}$, in terms of $\boldsymbol{\phi}$ and the idiosyncratic pricing errors. We have shown that when $\boldsymbol{\phi\neq0}$, it is possible to construct a portfolio, denoted as phi-portfolio, that dominates the associated mean-variance portfolio when the number securities, $n$, is sufficiently large and the risk factors are sufficiently strong. Given the pivotal role played by $\boldsymbol{\phi}$ for estimating the risk premia, for formation of large $n$ portfolios, and for tests of market efficiency, we have focussed on estimation of $\boldsymbol{\phi}$, and its asymptotic distribution under quite a general setting that allows for missing factors and idiosyncratic pricing errors. Since factor means, $\boldsymbol{\mu}$, can be estimated at the regular rate of $T^{-1/2}$ from time series data, it is relatively straightforward to develop a mixed strategy for estimation of $\boldsymbol{\lambda}$ by adding a time series estimate of $\boldsymbol{\mu}$ to the bias-corrected estimator of $\boldsymbol{\phi}$. If we use the same time series sample, such an estimator reduces to the Shanken bias-corrected estimator of $\boldsymbol{\lambda}$. But in practice, given the concern over the instability of the factor loadings $\beta_{ik}$ over time, one could use relatively long time series, say $T_{\mu}$, when estimating $\boldsymbol{\mu} $, and a shorter time series, say $T_{\phi}<T_{\mu}$, when estimating $\boldsymbol{\phi}$. The distributional and small sample properties of such a mixed estimator of risk premia is a topic for further research.

Our theoretical and Monte Carlo results further highlight the important role played by factor strengths in estimation and inference on $\boldsymbol{\phi}$, and hence on $\boldsymbol{\lambda}$. For a fixed $T_{\phi}$, factors with strength below $2/3$ lead to estimates of $\boldsymbol{\phi}$ \ with convergence rate of $n^{-1/3}$ or worse, and their use in asset pricing models can be justified only when $n$ is very large. We have also shown that weak factors, with strength below $1/2$ are best treated as missing and absorbed in the error term. We have shown that estimation of $\boldsymbol{\phi}$ for strong or semi-strong factors is robust to weak missing factors, and the explicit inclusion of weak factors in the empirical analysis is likely to have adverse spill over effects on the estimates of $\boldsymbol{\phi\ }$\ for strong and semi-strong factors. In view of these results we have proposed a factor selection procedure where only factors with strength above $1/2+1/3$ are included in the asset pricing model. Developing a formal statistical theory for the proposed selection is another topic for future research.

The paper also provides an empirical application to a large number of U.S. securities with risk factors selected from a large number of potential risk factors according to their strength, and use a pooled Lasso approach to select $7$ risk factors out of over $180$. We find strong statistical evidence against $\boldsymbol{\phi=0}$ for the selected model as well as for the two popular Fama-French models (FF3 and FF5). Using rolling estimates of $\boldsymbol{\phi}$ we also construct phi portfolios with better Sharpe ratios as compared to associated mean-variance portfolios. But we also warn that these portfolio comparisons are preliminary and need to be further investigated by allowing for transaction costs, and the feasibility of the long-short trading strategies that are involved.

\singlespacing