The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
162,164 characters
Identifying and exploiting alpha in linear asset pricing models with strong, semi-strong, and latent factors
\title{Identifying and exploiting alpha in linear asset pricing models with strong,
semi-strong, and latent factors\thanks{We are grateful to Juan Espinosa Torres
and Jing Kong for carrying out the computations and to Hayun Song for
excellent research assistance, all are PhD students at USC. We are also
grateful for helpful comments from the editor, Fabio Trojani, two anonymous
referees, Natalia Bailey, Alessio Sancetta, Takashi Yamagata and participants
at the 2023 Kansas Econometrics Institute, the 2023 NBER Summer Insitute, the
2023 IAAE Annual Conference in Oslo, University of California, Riverside,
April 2024, and ECMFE Workshop at University of Essex, October 2024.}}
\author{M. Hashem Pesaran\\University of Southern California,\ and Trinity College, Cambridge
\and Ron P. Smith\thanks{Corresponding author, email [email removed]. }\\Birkbeck, University of London}
\maketitle
\begin{abstract}
The risk premia of traded factors are the sum of factor means and a parameter
vector we denote by $\boldsymbol{\phi}$\textbf{ }$\ $which is identified from
the cross-section regression of $\alpha_{i}$ on the vector of factor loadings,
$\boldsymbol{\beta}_{i}$.\textbf{ }If $\boldsymbol{\phi}$ is non-zero, then
$\alpha_{i}$ are non-zero and one can construct "phi-portfolios" which exploit
the systematic components of non-zero alpha. We show that for known values of
$\boldsymbol{\beta}_{i}$ and when $\boldsymbol{\phi}\ \ $is non-zero there
exist phi-portfolios that dominate mean-variance portfolios. The paper then
proposes a two-step bias corrected estimator of $\boldsymbol{\phi}$ \ and
derives its asymptotic distribution allowing for idiosyncratic pricing errors,
weak missing factors, and weak error cross-sectional dependence. Small sample
results from extensive Monte Carlo experiments show that the proposed
estimator has the correct size with good power properties. The paper also
provides an empirical application to a large number of U.S. securities with
risk factors selected from a large number of potential risk factors according
to their strength and constructs phi-portfolios and compares their Sharpe
ratios to mean variance and S\&P portfolios.
\end{abstract}
\textbf{JEL Classifications: }{\small C38, G10}
\textbf{Key Words: }F{\small actor strength, pricing errors, risk premia,
missing factors, pooled Lasso, mean-variance and phi-portfolios.}
\thispagestyle{empty}
\newpage
\pagenumbering{arabic}
\onehalfspacing
\setlength{\abovedisplayskip}{0pt}
\setlength{\belowdisplayskip}{0pt}
\setlength{\abovedisplayshortskip}{0pt}
\setlength{\belowdisplayshortskip}{0pt}
\section{Introduction\label{Intro}}
There are two approaches to the estimation of risk premia and testing of
market efficiency, often referred as the beta and the SDF (stochastic discount
factor) methods, see \cite{Jagannathan2002}. This paper adopts the beta
method, and following the literature uses a linear factor pricing model (LFPM)
to explain the time series of excess returns on individual securities,
$r_{it}=R_{it}-r_{t}^{f}$, where $R_{it}$ is the return and $r_{t}^{f}$ the
risk free rate for $i=1,2,...,n;$ $t=1,2,....,T,$ by a set of observed
tradable risk factors. We use individual securities rather than portfolios
since, as we will show, if risk factors are not strong, large $n$ is required
for accurate estimation. \cite{Ang2020} discuss the general issues in the
choice between portfolios and individual stocks. \cite{Pesaran2023} discuss
both the use of portfolios and the relationship between the SDF and LFPM approaches.
The LFPM\ explains the excess return on each security, $r_{it},$ by an
intercept, $\mathit{\alpha}_{i}$, labelled alpha, and a $K\times1$ vector of
traded risk factors, $\mathbf{f}_{t}=(f_{1t},f_{2t},...,f_{Kt})^{\prime}$ with
loadings, $\boldsymbol{\beta}_{i}$:
\begin{equation}
r_{it}=\mathit{\alpha}_{i}+\boldsymbol{\beta}_{i}^{\prime}\mathbf{f}
_{t}+u_{it}. \label{LFPM}
\end{equation}
Under the Arbitrage Pricing Theory (APT) due to
\citet{ROSS1976341}
, we have
\begin{equation}
E\left( r_{it}\right) =c+\boldsymbol{\beta}_{i}^{\prime}\boldsymbol{\lambda
}+\eta_{i}, \label{APT}
\end{equation}
where $\eta_{i}$ is a firm-specific idiosyncratic pricing error. Ross allowed
$\eta_{i}$ to be non-zero for some $i$ but required that these errors are
bounded in the sense that $\sum_{i=1}^{n}\eta_{i}^{2}<C<\infty$. We relax this
condition and impose weaker restrictions on $\eta_{i}$'s as discussed in
sub-section \ref{pricing}. But the paper's main contribution lies in the fact
that even if there are no idiosyncratic pricing errors, it is still possible
to have non-zero $\mathit{\alpha}_{i}$. To see this, taking unconditional
expectations of (\ref{LFPM}) and using (\ref{APT}) we note that
\[
E(r_{it})=\mathit{\alpha}_{i}+\boldsymbol{\beta}_{i}^{\prime}\mathbf{\mu
}=c+\boldsymbol{\beta}_{i}^{\prime}\boldsymbol{\lambda}+\eta_{i}.
\]
where $E(\mathbf{f}_{t})=\boldsymbol{\mu}$, which in turn yields
\begin{equation}
\mathit{\alpha}_{i}=c+\boldsymbol{\beta}_{i}^{\prime}\left(
\boldsymbol{\lambda}-\boldsymbol{\mu}\right) +\eta_{i},\text{ for
}i=1,2,...,n \label{alphai}
\end{equation}
We denote $\boldsymbol{\lambda}-\boldsymbol{\mu}$ by $\boldsymbol{\phi}$, and
consider the implications of a non-zero $\boldsymbol{\phi}$ for portfolio
optimization. We show that when $\boldsymbol{\phi\neq0}$, one can construct
what we call "phi-portfolios" that exploit the systematic components of
$\mathit{\alpha}_{i}$, as captured by the non-zero elements of
$\boldsymbol{\phi}\ $\ in (\ref{alphai}).\footnote{We are grateful to one of
the anonymous reviewers for suggesting the idea of phi-portfolios, as a way of
motivating and interpreting the role of $\boldsymbol{\phi}$ in asset pricing
models.} In contrast, the idiosyncratic pricing errors, $\eta_{i}$, cannot be
exploited, as it is likely to involve insider trading. But our proposed
phi-portfolios can be implemented and applies irrespective of whether
$\eta_{i}=0$ or not.
The extent to which alpha can be exploited is discussed in sub-section
(\ref{pricing}) where it is shown that $\sum_{i=1}^{n}\left( \mathit{\alpha
}_{i}-\mathit{\bar{\alpha}}\right) ^{2}$ depends on the magnitude of
$\boldsymbol{\phi}^{\prime}\boldsymbol{\phi}$ and the strength of the risk
factors. Since this alpha can be exploited by the construction of the
phi-portfolios, this is not consistent with a no arbitrage condition. While it
is true that $E(r_{it})=\boldsymbol{\beta}_{i}^{\prime}\boldsymbol{\lambda}$,
estimating $\boldsymbol{\lambda}$\ does not tell us whether $\boldsymbol{\phi
}\neq0,$\ which can be exploited to obtain Sharpe ratios that are larger than
the Sharpe of the tangency portfolio. The claim that tangency portfolio has
the highest Sharpe ratio rests on the implicit assumption that
$\boldsymbol{\phi=0}$, even if there are no idiosyncratic pricing errors
($\eta_{i}=0$ for all $i$).
The focus of much of the literature has been on testing $\mathit{\alpha}
_{i}=0$, or on estimating the risk premium $\boldsymbol{\lambda}$. But given
that estimating a non-zero $\boldsymbol{\phi}$ \ provides a way to identify
and exploit the alpha in a linear factor pricing model for large $n,$
$\boldsymbol{\phi}$ is an interesting parameter in itself, which will be the
focus of this paper. Furthermore, the tangency condition used to construct
mean-variance (MV) portfolios has the implicit assumption that $\mathit{\alpha
}_{i}=0$, and the MV portfolio implied by the tangency condition is not
efficient unless $\boldsymbol{\phi=0}$. If $\boldsymbol{\phi\neq0}$ there is
no guarantee that the tangency portfolio will achieve a maximal Sharpe ratio,
and is likely to be dominated by the phi-portfolio. Specifically, abstracting
from model and parameter uncertainties, we show that when $\boldsymbol{\phi
\neq0}$ the Sharpe of the tangency portfolio is bounded in the number of
securities, whilst the Sharpe of the phi portfolio rises in $n$ and there
exists a sufficiently large $n$ that ensures that the Sharpe of the
phi-portfolio will exceed that of the tangency portfolio.
More specifically, first we introduce our proposed phi-portfolio in terms of
the factor loadings $\mathbf{B}_{n}=\left( \boldsymbol{\beta}_{1}
,\boldsymbol{\beta}_{2},...,\boldsymbol{\beta}_{n}\right) ^{\prime}$ and
$\boldsymbol{\phi}$, and compare its limiting properties (as $n\rightarrow
\infty$) with the standard MV portfolio. We show that for known factor
loadings\ and non-zero $\boldsymbol{\phi}$ there exist phi-portfolios with
strictly positive returns, given by $\boldsymbol{\phi}^{\prime}
\boldsymbol{\phi}>0$, that are fully diversified (their variance tends to zero
with $n$). The rate at which the variance of phi-portfolio returns tends to
zero will depend on the strength of the traded risk factors compared to the
strength of the missing (latent) factors, highlighting the importance of
factor strengths in portfolio analysis. Since the phi-portfolios have Sharpe
ratios that tend to infinity with $n$, they dominate MV portfolios whose
Sharpe ratio is bounded in $n$. Note that unlike MV\ portfolios, the
phi-portfolios do not require knowledge of the inverse of the covariance
matrix of returns, which is particularly difficult to estimate accurately when
$n$ is large. As is usual, these theoretical results apply to population
values where there are no restrictions on trading many individual securities.
Next we consider estimation of and inference about $\boldsymbol{\phi}$ using a
large number of individual securities, taking account of firm-specific pricing
errors, traded factors that are not strong, missing latent factors, and panel
data sets where the time dimension, $T$, is small relative to $n$. Factor
strength plays a central role in our analysis of $\boldsymbol{\phi}$. We use a
measure of factor strength, $\alpha_{k},$ developed in
\citet{bailey2016exponent, bailey2021measurement}
, which is defined in terms of the proportion of non-zero factor loadings,
$\beta_{ik}$.\footnote{Alpha is used both for the LFPM intercepts and the
measure of factor strength, because these are established usages, but it will
be clear from context and subscripts which is being referred to.} A factor is
strong if this proportion is very close to unity, it is semi-strong if
$1>\alpha_{k}>1/2$, and it is weak if $\alpha_{k}<1/2$. Use of this measure
allows us to be precise about the degree of pervasiveness and show how the
strengths of the observed factors, the missing factors and the pricing errors,
each influence estimation and inference about $\boldsymbol{\phi}$.
In practice, exploitation of a non-zero $\boldsymbol{\phi}$ requires $n$ to be
large and rebalancing such long-short portfolios for so many securities may
incur high transactions costs or not be feasible. In addition, model
uncertainty, estimation uncertainty, time variation in both $\boldsymbol{\beta
}_{i}$ and in conditional volatility pose additional difficulties in
implementing a strategy to exploit the potential returns revealed by
$\boldsymbol{\phi}$. In developing the theory we will abstract from such
practical difficulties, but in the empirical section we illustrate some of
these issues with a comparison of the performance of $\boldsymbol{\phi}$ based
portfolios relative to MV portfolios which would face similar difficulties.
We estimate $\boldsymbol{\phi}=(\phi_{1},\phi_{2},...,\phi_{K})^{\prime}$,
using a two-step estimator. In the first step we estimate the intercepts,
$\hat{\alpha}_{i}$, and the factor loadings, $\boldsymbol{\hat{\beta}}_{i}$,
from least squares regressions of excess returns on an intercept and risk
factors. In the second step, $\boldsymbol{\phi}$ is estimated from the cross
section regression of $\hat{\alpha}_{i}$ on $\boldsymbol{\hat{\beta}}_{i}$. As
with the two-step estimator of $\boldsymbol{\lambda}$, such a two-step
estimator of $\phi_{k}$ will also be biased, and requires bias-correction.
Following
\citet{shanken1992estimation}
, we consider a bias-corrected version of the two-step estimator of
$\boldsymbol{\phi}$, which we denote by $\boldsymbol{\tilde{\phi}}_{nT}$. We
develop the asymptotic distribution of $\boldsymbol{\tilde{\phi}}_{nT}$ under
quite general set of assumptions regarding the idiosyncratic pricing errors,
error cross-sectional dependence, and the presence of missing (latent)
factors. The paper also investigates the implications of factor strengths for
the precision with which $\boldsymbol{\phi}$ can be estimated. The LFPM,
following
\citet{cham1983arbi}
, assumes that all the observed factors are strong and the eigenvalues of the
covariance matrix of the errors are bounded.
In developing the arbitrage pricing theory, APT,
\citet{ROSS1976341}
, whose concerns were primarily theoretical, assumed the factors had mean
zero: $\boldsymbol{\mu}=0,$ so $\boldsymbol{\phi}=\boldsymbol{\lambda}$, is
the risk premium. For traded factors under market efficiency, where
$\boldsymbol{\phi}=\mathbf{0}$, the risk premium is the factor mean
$\boldsymbol{\mu}=\boldsymbol{\lambda}$. Were one interested in estimating
$\boldsymbol{\lambda}$ there may be statistical reasons to estimate $\phi_{k}$
and $\mu_{k}$ separately, then summing them to obtain an estimate of
$\lambda_{k}$. The factor mean, $\mu_{k}=E(f_{kt})$, can be estimated
consistently at the regular $\sqrt{T}$ rate directly using time series data on
the risk factors, $f_{kt}$, for $t=1,2,...,T,$ and does not require knowledge
of the factor loadings or $n$. In contrast, estimation of $\phi_{k}$ requires
panel data to estimate the factor loadings and hence both $n$ and $T$
dimensions are important. In some cases it may be beneficial to use different
time series dimensions, $T_{\mu}$ and $T_{\phi},$ to estimate $\mu_{k}$ and
$\phi_{k}$, respectively. For example, when factor loadings are subject to
breaks it is advisable to use a relatively short sample, and when some factors
are not sufficiently strong a large $n$ is required. Thus one could use large
$T$ to estimate $\mu_{k}$ and small $T$ to estimate $\phi_{k}$ and
$\lambda_{k}$ can be simply estimated by adding the estimates of $\mu_{k}$ and
$\phi_{k}.$ We do not pursue the use of different $T$, but if the same $T$ is
used to estimate the mean of factor $k,$ say $\hat{\mu}_{k,T},$ and the
bias-corrected $\tilde{\phi}_{k,nT}$ then their sum is the same as the
\citet{shanken1992estimation}
bias-corrected risk premium for factor $k,$ $\tilde{\lambda}_{k,nT}.$ This
decomposition can be used to obtain the rate at which $\tilde{\lambda}_{k,nT}$
converges to its true value, $\lambda_{0,k}=\phi_{0,k}+\mu_{0,k}$.
The main theoretical results of the paper are set out around five theorems
under a number of key assumptions, with proofs provided in the mathematical
appendix. The small sample properties of $\boldsymbol{\tilde{\phi}}_{nT}$ are
investigated using extensive Monte Carlo experiments, allowing for a mixture
of strong and semi-strong observed factors, latent factors, pricing errors,
GARCH effects and non-Gaussian errors. Small sample results are in line with
our theoretical findings, and confirm that the bias-corrected estimator,
$\boldsymbol{\tilde{\phi}}_{nT}$, has the correct size and good power
properties for samples with time series dimensions of $T=120$ and $T=240$.
They also show, in accord with the theory, that the precision with which
$\phi_{k}$ is estimated falls with $\alpha_{k}$, the strength of the $k^{th}$ factor.
Our theoretical derivations and Monte Carlo simulations assume that the list
of relevant observed factors is known. But in practice the relevant factors
must be selected. Extending the theory to the high dimensional case where
factors are selected, rather than given, is beyond the scope of the present
paper. Since the rate of convergence of $\boldsymbol{\tilde{\phi}}_{nT}$ to
its true value is given by $\sqrt{T}n^{(\alpha_{k}+\alpha_{\min}-1)/2}$, and
the Monte Carlo confirms the crucial role of factor strengths in estimation
and inference on $\boldsymbol{\phi,}$ it seems sensible to select factors on
the basis of their factor strength. Weak factors whose strength is around
$1/2$ can be ignored and absorbed into the error term.
The above selection procedure is applied in a high dimensional setting with
both a large number of securities ($n$ from $1,090$ to $1,175)$ and a large
number of potential risk factors ($m$ from $177$ to $189)$, taken from the
\citet{chen2022open}
, which can all be traded. We used monthly data over the period
$1996m1-2022m12$ and considered balanced panels obtained by including
\textit{all} existing stocks in a given month for which there are $T$
observations. We considered $T=120$ and $240$ months, and focus on the latter
which we found to be more reliable, given the large number of securities being
considered. Various procedures could be used to select risk factors for a
given security, $i$. We used Lasso for this purpose and then selected a subset
of these factors that were chosen by a sufficiently large number of securities
in the sample, and whose estimated strength were above the given threshold
value of $\alpha_{k}>0.75$. We refer to this selection procedure as pooled
Lasso. Using this procedure with $T=240$ we ended up with $7$ risk factors for
the sample ending in $2015$, declining to $4$ in $2021$. Interestingly, the
three Fama-French factors were always included in the set of factors selected
by pooled Lasso.
Accordingly, we considered three linear asset pricing models for our portfolio
analysis: the pooled Lasso selected at the start of our evaluation sample,
denoted by PL7, the Fama-French three factors model, FF3, and the Fama-French
five factors model, FF5, which includes two factors not in PL7. The three
models are estimated using rolling samples of size $T=240$, starting with a
sample ending in $2015m12$ and finishing with a sample ending in $2022m11$.
The hypothesis that $\boldsymbol{\phi}=\mathbf{0}$ was rejected for all $84$
rolling samples and all three models, albeit less strongly over the post
Covid-19 period. The test results suggested possible unexploited return
opportunities, and to investigate this possibility further, we constructed
phi-portfolios and compared their Sharpe ratios with the ones based on
standard MV portfolios over the full sample evaluation sample,
$2016m1-2022m12$, and sample ending $2019m12$, that excludes the Covid-19
period. We find that in five out of the six cases (3 models 2 samples) the
phi-portfolio has a higher SR than the corresponding MV portfolio. The
exception is the FF5 pre Covid-19. This illustrates that if $\boldsymbol{\phi
}\neq0,$ it is possible to construct a portfolio that outperforms the mean
variance portfolio. In both the pre Covid-19 sample and the full sample the
highest SR was obtained by the PL7 phi-portfolios, which also outperformed the
S\&P500. The SRs for the sample ending in $2022$ were substantially lower than
the sample ending in $2019$, consistent with a falling value of the
probability that $\boldsymbol{\phi}$ was non-zero.
\textbf{Related literature: }On estimation of risk premia, following Fama and
MacBeth and
\citet{shanken1992estimation}
, estimation of risk premia is further examined by
\citet{shanken2007estimating}
,
\citet{kan2013pricing}
, and
\citet{BAI201531}
. The survey paper by
\citet{jagannathan2010analysis}
provide further references.
Testing for market efficiency dates back to
\citet{jensen1968performance}
who proposes testing a$_{i}=0$ for each $i$ separately.
\citet{gibbons1989test}
provide a joint test for the case where the errors are Gaussian and $n<T$.
\citet{gagliardini2016time}
develop two-pass regressions of individual stock returns, allowing
time-varying risk premia, and propose a standardised Wald test.
\citet{raponi2019testing}
propose a test of pricing error in cross section regression for fixed number
of time series observations. They use a bias-corrected estimator of
\citet{shanken1992estimation}
to standardise their test statistic.
\citet{ma2020testing}
employ polynomial spline techniques to allow for time variations in factor
loadings when testing for alphas.
\citet{feng2022high}
propose a max-of-square type test of the intercepts instead of the average
used in the literature, and recommend using a combination of the two testing
procedures.
\citet{he2022testing}
propose two statistics, a Wald type statistic which require $n$ and $T$ to be
of the same order of magnitude and a standardised t-ratio.
\citet{kleibergen2009tests}
considers testing in the case where the loadings are small.
\citet{pesaran2012testing, pesaran2023testing}
consider testing that the intercepts in the LFPM are zero when $n$ is large
relative to $T$ and there may be non Gaussian errors and weakly
cross-correlated errors.
A large number of risk factors have been considered in the empirical
literature. We use the
\citet{fama1993common}
three factors in our Monte Carlo design. In our empirical application we use
the five factors proposed by
\citet{fama2015five}
and the large set of factors provided by
\citet{chen2022open}
.
\citet{harvey2019census}
document over 400 factors published in top finance journals.
\citet{dello2022missing}
argue that despite the hundreds of systematic risk factors considered in the
literature, there is still a sizable pricing error and that this can be
explained by asset specific risk that reflects market frictions and behavioral
biases. There is a large Bayesian literature, including
\citet{chib2020comparing}
, and
\citet{hwang2022bayesian}
on selecting factor models. The issue of factor selection is also addressed
by
\citet{fama2018choosing}
.
Strong and weak factors in asset returns are considered by
\citet{ANATOLYEV2022103}
,
\citet{connor2022semi}
, and \cite{giglio2023test}.
\citet{beaulieu2020arbitrage}
discuss the lack of identification of risk premia when many of the loadings
are zero. There has also been concern about the consequences of omitted
factors.
\citet{giglio2021asset}
discuss the problem and try to deal with it using a three-pass method which is
valid even when not all factors in the model are specified or observed using
principal components of the test assets.
\citet{onatski2012asymptotics}
and
\citet{lettau2020estimating, lettau2020factors}
provide extensive discussions of weak factor and latent factors, respectively.
More recent contributions include \cite{BaiNg2023Weak} and
\cite{uematsu2023inference}.
There is also a large literature on portfolio construction.
\cite{Herskovic2019} discuss low cost methods of hedging risk factors.
\cite{Preite2024} derive an SDF in which there is compensation for
unsystematic risk within the framework of the APT. \cite{Korsaye2021}
introduce model-free smart SDFs which give rise to non-parametric SDF bounds
for testing asset pricing models.\ \ \cite{Daniel2020} discuss the common
practice of creating characteristic portfolios by sorting on characteristics
associated with average returns and show that these portfolios capture not
only the priced risk associated with the characteristic but also unpriced
risk. \cite{Quaini2023} propose an estimator of tradable factor risk premia.
\textbf{Paper's outline}: The rest of the paper is organized as follows:
Section \ref{LFPMAPT} provides the framework for estimation of
$\boldsymbol{\phi}.$ Section \ref{Assumptions} sets out the assumptions and
states the main theorems. Theorem \ref{TFMbias} shows that the standard
Fama-MacBeth estimator is valid only when there are no pricing errors and
$n/T\rightarrow0$. Theorem \ref{Thzig} shows that the Shanken bias-corrected
estimator of $\lambda_{k}$ continues to be consistent for a fixed $T$ as
$n\rightarrow\infty$, even in presence of weak pricing errors and weak missing
common factors. Theorem \ref{Tfi} provides conditions under which the
bias-corrected estimator, $\boldsymbol{\tilde{\phi}}_{nT}$, is consistent for
$\boldsymbol{\phi}_{0}$, and derives the asymptotic distribution of
$\boldsymbol{\tilde{\phi}}_{nT},$ assuming the observed factors are strong
$(\alpha_{k}=1$ for all $k$). Theorem \ref{Tsemi} extends the results to cases
where one or more risk factors are semi-strong, and establishes the rate at
which $\boldsymbol{\tilde{\phi}}_{nT}$ converges to its true value, assuming
the idiosyncratic pricing errors are sufficiently weak, as discussed after the
theorem. For example, for a factor with strength $\alpha_{k}$ we show that
$\tilde{\phi}_{k,nT}-\phi_{0,k}=O_{p}\left( T^{-1/2}n^{-(\alpha_{k}
+\alpha_{\min}-1)/2}\right) $, and as a consequence
\[
\tilde{\lambda}_{k,nT}-\lambda_{0,k}=O_{p}\left( T^{-1/2}n^{-(\alpha
_{k}+\alpha_{\min}-1)/2}\right) +O_{p}\left( T^{-1/2}\right) ,
\]
where $\lambda_{0,k}$ is the true value of the risk premia associated to
factor $f_{kt}$, and $\alpha_{\min}$ is the strength of the least strong
factor included. This consistency condition is weaker than the one derived by
\cite{giglio2023test} who also assume $\eta_{i}=0$, for all $i$. Finally,
Theorem \ref{Var} gives conditions for consistent estimation of the asymptotic
variance of $\boldsymbol{\tilde{\phi}}_{nT},$ using a suitable threshold
estimator of the covariance matrix. Section \ref{Simulations} presents the
Monte Carlo (MC) design, its calibration and a summary of the main findings.
Section \ref{FacSel}\ discusses the problem of factor selection from a large
number of potential factors. Section \ref{Empirical} gives the empirical
application using monthly data on a large number of individual US stocks and
risk factors over the period 1996-2021. It selects factors, estimates
$\boldsymbol{\phi}\mathbf{,}$ and compares the performance of MV and
phi-portfolios. Section \ref{Conclusion} provides some concluding remarks.
Detailed mathematical proofs are provided in a mathematical appendix. Further
information on data sources, MC calibration plus some supplementary material
for the empirical application are provided in the online supplement A. To save
space all MC results are given in the online supplement B.
\section{Identification and estimation of $\boldsymbol{\phi}$\label{LFPMAPT}}
Let $R_{it}$ denote the holding period return on traded security $i$, which
can be bought long or short without transaction costs, $r_{t}^{f}$ \ is the
risk free rate, and $r_{it}=R_{it}-r_{t}^{f}$ is the excess return. We start
with the linear factor pricing model (LFPM)
\begin{equation}
r_{it}=\mathit{\alpha}_{i}+\boldsymbol{\beta}_{i}^{\prime}\mathbf{f}
_{t}+u_{it}, \label{StatM}
\end{equation}
for $i=1,2,...,n$ and $t=1,2,...,T$, where $r_{it}$ is explained in terms of
the $K\times1$ vector of factors $\mathbf{f}_{t}=(f_{1t},f_{2t},...,f_{Kt}
)^{\prime}$. The intercept $\mathit{\alpha}_{i}$ and the $K\times1$ vector of
factor loadings, $\boldsymbol{\beta}_{i}=(\beta_{i1},\beta_{i2},...,\beta
_{iK})^{\prime}$, are unknown. The idiosyncratic errors, $u_{it}$ have zero
means and are assumed to be serially uncorrelated. The factors, $\mathbf{f}
_{t}$, are assumed to be covariance stationary with the constant mean
$\boldsymbol{\mu}=E\left( \mathbf{f}_{t}\right) $, and $\boldsymbol{\Sigma
}_{f}=E\left[ \left( \mathbf{f}_{t}-\boldsymbol{\mu}\right) \left(
\mathbf{f}_{t}-\boldsymbol{\mu}\right) ^{\prime}\right] $.
Under the Arbitrage Pricing Theory (APT) due to Ross (1976) the pricing
errors, $\eta_{i}$ for $i=1,2,...,n$ defined by
\begin{equation}
\eta_{i}=E\left( r_{it}\right) -c-\boldsymbol{\beta}_{i}^{\prime
}\boldsymbol{\lambda}\mathbf{,} \label{APTmean}
\end{equation}
are bounded such that
\begin{equation}
\sum_{i=1}^{n}\eta_{i}^{2}<C<\infty, \label{APTRoss}
\end{equation}
where $c$ is zero-beta expected excess return, $\boldsymbol{\lambda}$ is the
$K\times1$ vector of risk premia. Under APT, we have
\begin{equation}
E\left( r_{it}\right) =c+\boldsymbol{\beta}_{i}^{\prime}\boldsymbol{\lambda
}_{0}+\eta_{i}, \label{Urit}
\end{equation}
and reduces to the standard beta representation of an unconditional asset
pricing model, if $c+\eta_{i}=0$. In this case, $\boldsymbol{\beta}
_{i}=-Cov(r_{it},m_{t})/var(m_{t})$, where $m_{t}$ is the stochastic discount
factor (SDF), which satisfies the moment condition $E_{t}(r_{i,t+1}m_{t+1}
)=0$. See \cite{Pesaran2023}\textbf{ }for a discussion of the link. For
empirical assessment, we consider the more general unconditional pricing model
(\ref{Urit}), and relate to the linear factor pricing model. To this end,
taking unconditional expectations of (\ref{StatM}) and using the
APT\ condition we have
\begin{equation}
E(r_{it})=\mathit{\alpha}_{i}+\boldsymbol{\beta}_{i}^{\prime}\boldsymbol{\mu
}=\boldsymbol{\beta}_{i}^{\prime}\boldsymbol{\lambda}+c+\eta_{i}.
\label{meuiR}
\end{equation}
which in turn yields
\begin{equation}
\mathit{\alpha}_{i}=c+\boldsymbol{\beta}_{i}^{\prime}\boldsymbol{\phi}
+\eta_{i}\text{, for }i=1,2,...,n. \label{ai}
\end{equation}
where
\begin{equation}
\boldsymbol{\phi}=\boldsymbol{\lambda}-\boldsymbol{\mu}, \label{phi0}
\end{equation}
Under the APT condition, (\ref{APTRoss}), $\boldsymbol{\phi}$ can be
identified from the cross section regression of $\alpha_{i}$ on
$\boldsymbol{\beta}_{i}$. We relax the APT\ condition and derive restrictions
on the degree to which idiosyncratic pricing errors, $\eta_{i}$, can be
pervasive to achieve identification.
The focus of the literature has been on testing for alpha, $\mathit{\alpha
}_{i}=0$, and the estimation of the risk premia, $\boldsymbol{\lambda}$, using
panel data on excess returns, $\left\{ r_{it},1,2,...,n\text{; }
t=1,2,...,T\right\} ,$ and $\mathbf{F}$, the $T\times K$ matrix of
observations on the factors. It is clear that $\boldsymbol{\phi}$ plays an
important role in tests for alpha in LFPM, and a non-zero $\boldsymbol{\phi}$
implies non-zero alphas, which in turn implies exploitable excess profitable
opportunities. More specifically, we show that for known values of
$\boldsymbol{\beta}_{i},$ $i=1,2,...,n\,$\ and when $\boldsymbol{\phi}$ is
non-zero, there exists phi-based portfolios with non-zero means that are fully
diversified (their variance tends to zero with $n$), namely have Sharpe ratios
that tend to infinity with $n$, and hence dominate the MV portfolios. For the
MV portfolios to be efficient it is necessary that $\boldsymbol{\phi=0}$.
\subsection{Why $\boldsymbol{\phi}$ matters: introduction of the
phi-portfolios}
Substitute the expression for $\mathit{\alpha}_{i}$ given by (\ref{ai}) in
(\ref{StatM}) to obtain
\begin{equation}
r_{it}=c+\boldsymbol{\beta}_{i}^{\prime}\boldsymbol{\phi}+\eta_{i}
+\boldsymbol{\beta}_{i}^{\prime}\mathbf{f}_{t}+u_{it},\text{ for }i=1,2,...,n,
\label{rit}
\end{equation}
and write the $n$ return equations more compactly as
\begin{equation}
\mathbf{r}_{\circ t}=c\mathbf{\tau}_{n}+\mathbf{B}_{n}\boldsymbol{\phi
}+\mathbf{B}_{n}\mathbf{f}_{t}+\boldsymbol{\eta}_{n}+\mathbf{u}_{\circ t},
\label{ritV}
\end{equation}
where $\mathbf{r}_{\circ t}=(r_{1t},r_{2t},....,r_{nt})^{\prime},$
$\boldsymbol{\tau}_{n}$ is an $n$-dimensional vector of ones, $\mathbf{B}
_{n}=(\boldsymbol{\beta}_{\circ1},\boldsymbol{\beta}_{\circ2}
,...,\boldsymbol{\beta}_{\circ K}),$ $\boldsymbol{\beta}_{\circ k}=(\beta
_{1k},\beta_{2k},...,\beta_{nk})^{\prime},$ $\mathbf{u}_{\circ t}
=(u_{1t},u_{2t},....,u_{nt})^{\prime}$, $\boldsymbol{\eta}_{n}=\left(
\eta_{1},\eta_{2},...,\eta_{n}\right) ^{\prime}$, and $\mathbf{V}
_{u}=E\left( \mathbf{u}_{\circ t}\mathbf{u}_{\circ t}^{\prime}\right) $.
Suppose that the factors, $\mathbf{f}_{t},$ are traded, and $\boldsymbol{\phi
}^{\prime}\boldsymbol{\phi}>0$. Consider the $n\times1$ vector of
phi-portfolio weights,$\mathbf{w}_{\phi}$, given by
\begin{equation}
\mathbf{w}_{\phi}=\mathbf{M}_{n}\mathbf{B}_{n}\left( \mathbf{B}_{n}^{\prime
}\mathbf{M}_{n}\mathbf{B}_{n}\right) ^{-1}\boldsymbol{\phi}, \label{wphi}
\end{equation}
where $\mathbf{M}_{n}=\mathbf{I}_{n}-n^{-1}\boldsymbol{\tau}_{n}
\boldsymbol{\tau}_{n}^{\prime}$. Finally, consider the long-short hedged
portfolio return
\begin{equation}
\rho_{t,\phi}=\mathbf{w}_{\phi}^{\prime}\mathbf{r}_{\circ t}-\boldsymbol{\phi
}^{\prime}\mathbf{f}_{t}=\boldsymbol{\phi}^{\prime}\left[ \left(
\mathbf{B}_{n}^{\prime}\mathbf{M}_{n}\mathbf{B}_{n}\right) ^{-1}
\mathbf{B}_{n}^{\prime}\mathbf{M}_{n}\mathbf{r}_{\circ t}-\mathbf{f}
_{t}\right] , \label{phiport}
\end{equation}
and using (\ref{ritV}) note that
\begin{equation}
\rho_{t,\phi}=\boldsymbol{\phi}^{\prime}\boldsymbol{\phi}+\mathbf{w}_{\phi
}^{\prime}\boldsymbol{\eta}_{n}+\mathbf{w}_{\phi}^{\prime}\mathbf{u}_{\circ
t}\mathbf{.} \label{phiport1}
\end{equation}
For a given vector of pricing errors, $\boldsymbol{\eta}_{n}$, $E\left(
\rho_{t,\phi}\right) =\boldsymbol{\phi}^{\prime}\boldsymbol{\phi}
+\mathbf{w}_{\phi}^{\prime}\boldsymbol{\eta}_{n}$, and $Var\left(
\rho_{t,\phi}\right) =\mathbf{w}_{\phi}^{\prime}\mathbf{V}_{u}\mathbf{w}
_{\phi}\mathbf{,}$ where $\mathbf{V}_{u}=E(\mathbf{u}_{\circ t}\mathbf{u}
_{\circ t}^{\prime})$. Using (\ref{wphi}) it follows that
\[
\left\Vert \mathbf{w}_{\phi}^{\prime}\boldsymbol{\eta}_{n}\right\Vert ^{2}
\leq\left\Vert \boldsymbol{\eta}_{n}\right\Vert ^{2}\left\Vert \mathbf{w}
_{\phi}\right\Vert ^{2}\leq\left\Vert \boldsymbol{\eta}_{n}\right\Vert
^{2}\left\Vert \boldsymbol{\phi}\right\Vert ^{2}\lambda_{max}\left[ \left(
\mathbf{B}_{n}^{\prime}\mathbf{M}_{n}\mathbf{B}_{n}\right) ^{-1}\right]
=\frac{\left\Vert \boldsymbol{\eta}_{n}\right\Vert ^{2}\left\Vert
\boldsymbol{\phi}\right\Vert ^{2}}{\lambda_{min}\left( \mathbf{B}_{n}
^{\prime}\mathbf{M}_{n}\mathbf{B}_{n}\right) }.
\]
Similarly,
\begin{equation}
Var\left( \rho_{t,\phi}\right) \leq\left\Vert \mathbf{w}_{\phi}\right\Vert
^{2}\lambda_{max}\left( \mathbf{V}_{u}\right) =\frac{\lambda_{max}\left(
\mathbf{V}_{u}\right) \left\Vert \boldsymbol{\phi}\right\Vert ^{2}}
{\lambda_{min}\left( \mathbf{B}_{n}^{\prime}\mathbf{M}_{n}\mathbf{B}
_{n}\right) }. \label{varphi}
\end{equation}
Hence, $E\left( \boldsymbol{\rho}_{t,\phi}\right) \rightarrow
\boldsymbol{\phi}^{\prime}\boldsymbol{\phi}$, and $Var\left( \rho_{t,\phi
}\right) \rightarrow0$, as $n\rightarrow\infty$, if
\begin{equation}
\frac{\left\Vert \boldsymbol{\eta}_{n}\right\Vert ^{2}}{\lambda_{min}\left(
\mathbf{B}_{n}^{\prime}\mathbf{M}_{n}\mathbf{B}_{n}\right) }\rightarrow
0\text{, and }\frac{\lambda_{max}\left( \mathbf{V}_{u}\right) }
{\lambda_{min}\left( \mathbf{B}_{n}^{\prime}\mathbf{M}_{n}\mathbf{B}
_{n}\right) }\rightarrow0. \label{fulldiv}
\end{equation}
These conditions are met if $\lambda_{min}\left( \mathbf{B}_{n}^{\prime
}\mathbf{M}_{n}\mathbf{B}_{n}\right) \rightarrow\infty$, $\lambda
_{max}\left( \mathbf{V}_{u}\right) <C$, and the APT condition (\ref{APTRoss}
) holds. The first two conditions follow if the LFPM given by (\ref{StatM}) is
an approximate factor model as assumed by
\citet{cham1983arbi}
, namely when the factors are strong and the errors, $u_{it}$, are weakly
cross correlated.\footnote{The diversification conditions in (\ref{fulldiv})
are met more generally, and accommodate the presence of semi-strong traded
factors in $f_{t}$, and less restrictive conditions on $\mathbf{V}_{u}$.and
the pricing errors $\eta_{i}.$ See Remark \ref{Remfulldiv} below.}
Therefore, knowledge of $\boldsymbol{\phi}^{\prime}\boldsymbol{\phi}$ and its
statistical significance can play an important role in portfolio analysis. To
illustrate this point suppose that $c=0$ and $\boldsymbol{\eta}_{n}
=\mathbf{0}$, but $\boldsymbol{\phi}^{\prime}\boldsymbol{\phi}>0$, and
consider the hedged return $\rho_{t,\phi}=\boldsymbol{\phi}^{\prime}\left[
\left( \mathbf{B}_{n}^{\prime}\mathbf{V}_{u}^{-1}\mathbf{B}_{n}\right)
^{-1}\mathbf{B}_{n}\mathbf{V}_{u}^{-1}\mathbf{r}_{\circ t}-\mathbf{f}
_{t}\right] $. Then under LFPM we have\footnote{When $c\neq0$, it is not
possible to eliminate $c$ (which is unpriced), and at the same time exploit
the error covariance matrix $\mathbf{V}_{u}$.}
\begin{align*}
\rho_{t,\phi} & =\boldsymbol{\phi}^{\prime}\left[ \left( \mathbf{B}
_{n}^{\prime}\mathbf{V}_{u}^{-1}\mathbf{B}_{n}\right) ^{-1}\mathbf{B}
_{n}^{\prime}\mathbf{V}_{u}^{-1}\left( \mathbf{B}_{n}\boldsymbol{\phi}
_{0}+\mathbf{B}_{n}\mathbf{f}_{t}+\mathbf{u}_{\circ t}\right) -\mathbf{f}
_{t}\right] \\
& =\boldsymbol{\phi}^{\prime}\boldsymbol{\phi}+\boldsymbol{\phi}^{\prime
}\left( \mathbf{B}_{n}^{\prime}\mathbf{V}_{u}^{-1}\mathbf{B}_{n}\right)
^{-1}\mathbf{B}_{n}^{\prime}\mathbf{V}_{u}^{-1}\mathbf{u}_{\circ t},
\end{align*}
and its squared Sharpe ratio is given by
\begin{equation}
SR_{\phi}^{2}=\frac{\left( \boldsymbol{\phi}^{\prime}\boldsymbol{\phi
}\right) ^{2}}{\boldsymbol{\phi}^{\prime}\left( \mathbf{B}_{n}^{\prime
}\mathbf{V}_{u}^{-1}\mathbf{B}_{n}\right) ^{-1}\boldsymbol{\phi}
}.\label{SRsquared1}
\end{equation}
Since,
\[
\boldsymbol{\phi}^{\prime}\left( \mathbf{B}_{n}^{\prime}\mathbf{V}_{u}
^{-1}\mathbf{B}_{n}\right) ^{-1}\boldsymbol{\phi}\leq\left( \boldsymbol{\phi
}^{\prime}\boldsymbol{\phi}\right) \lambda_{\max}\left[ \left(
\mathbf{B}_{n}^{\prime}\mathbf{V}_{u}^{-1}\mathbf{B}_{n}\right) ^{-1}\right]
=\frac{\boldsymbol{\phi}^{\prime}\boldsymbol{\phi}}{\lambda_{min}\left(
\mathbf{B}_{n}^{\prime}\mathbf{V}_{u}^{-1}\mathbf{B}_{n}\right) },
\]
then
\begin{equation}
SR_{\phi}^{2}\geq\left( \boldsymbol{\phi}^{\prime}\boldsymbol{\phi}\right)
\text{ }\lambda_{min}\left( \mathbf{B}_{n}^{\prime}\mathbf{V}_{u}
^{-1}\mathbf{B}_{n}\right) .\label{SRsquared2}
\end{equation}
As a result, if $\lambda_{min}\left( \mathbf{B}_{n}^{\prime}\mathbf{V}
_{u}^{-1}\mathbf{B}_{n}\right) \rightarrow\infty$ as $n\rightarrow\infty$,
$SR_{\phi}^{2}$ increases in $n$ without bounds if $\boldsymbol{\phi}^{\prime
}\boldsymbol{\phi}>0$. In contrast, it is easily seen that the Sharpe ratio of
the standard mean-variance (MV) portfolio, defined by $SR_{MV}^{2}
=\boldsymbol{\mu}_{R}^{\prime}\mathbf{V}_{R}^{-1}\boldsymbol{\mu}_{R}$ is
bounded in $n$, and in consequence $SR_{\phi}^{2}$ will eventually dominate
$SR_{MV}^{2}$ if $\boldsymbol{\phi}^{\prime}\boldsymbol{\phi}>0$. To see this,
note that under the LFPM given by (\ref{rit}) with $c=0$ and $\boldsymbol{\eta
}_{n}=\mathbf{0}$, the MV portfolio is given by $\rho_{MV,t}=\boldsymbol{\mu
}_{R}^{\prime}\mathbf{V}_{R}^{-1}\mathbf{r}_{\circ t}$, where $\boldsymbol{\mu
}_{R}=\mathbf{B}_{n}\boldsymbol{\lambda}$ and $\mathbf{V}_{R}=\mathbf{B}
_{n}\mathbf{\Sigma}_{f}\mathbf{B}_{n}^{\prime}+\mathbf{V}_{u}$. Also, since
$\mathbf{B}_{n}\mathbf{\Sigma}_{f}\mathbf{B}_{n}^{\prime}$ is rank deficient
then
\[
\mathbf{V}_{R}^{-1}=\mathbf{V}_{u}^{-1}\mathbf{-V}_{u}^{-1}\mathbf{B}\left(
\mathbf{\Sigma}_{f}^{-1}+\mathbf{B}_{n}^{\prime}\mathbf{V}_{u}^{-1}
\mathbf{B}_{n}\right) ^{-1}\mathbf{B}_{n}^{\prime}\mathbf{V}_{u}^{-1},
\]
and it follows that
\begin{align}
SR_{MV}^{2} & =\boldsymbol{\mu}_{R}^{\prime}\mathbf{V}_{R}^{-1}
\boldsymbol{\mu}_{R}-\boldsymbol{\lambda}^{\prime}\mathbf{\Sigma}_{f}
^{-1}\left( \mathbf{\Sigma}_{f}^{-1}+\mathbf{B}_{n}^{\prime}\mathbf{V}
_{u}^{-1}\mathbf{B}_{n}\right) ^{-1}\mathbf{\Sigma}_{f}^{-1}
\boldsymbol{\lambda}\label{SRsqMV}\\
& \leq\boldsymbol{\lambda}^{\prime}\mathbf{\Sigma}_{f}^{-1}
\boldsymbol{\lambda}=\left( \boldsymbol{\mu+\phi}\right) ^{\prime
}\mathbf{\Sigma}_{f}^{-1}\left( \boldsymbol{\mu+\phi}\right) <C\text{.}
\nonumber
\end{align}
Hence, the Sharpe ratio of the MV portfolio continues to be bounded in $n$
even if $\boldsymbol{\phi\neq0}$. Whether $\boldsymbol{\phi=0}$ or not affects
the magnitude of $SR_{MV}^{2}$ but does not alter the fact that $SR_{MV}^{2}$
will be bounded if $\lambda_{min}\left( \mathbf{B}_{n}^{\prime}\mathbf{V}
_{u}^{-1}\mathbf{B}_{n}\right) \rightarrow\infty$ as $n\rightarrow\infty$.
Furthermore, it is not possible to improve over the MV portfolio by using only
the factor loadings. The optimal portfolio formed using the $K$ beta-based
portfolios, {\small $\boldsymbol{\rho}_{B,t}=$ }$\left( \mathbf{B}
_{n}^{\prime}\mathbf{V}_{u}^{-1}\mathbf{B}_{n}\right) ^{-1}\mathbf{B}
_{n}^{\prime}\mathbf{V}_{u}^{-1}${\small $\mathbf{r}_{\circ t}$, }has the same
Sharpe ratio as the mean-variance portfolio. See Lemma \ref{Sharpe} for a proof.
In short, to make sure that the mean-variance portfolio is efficient we must
have $\boldsymbol{\phi}^{\prime}\boldsymbol{\phi}=0$, and it is of special
interest to estimate $\boldsymbol{\phi}$ and develop reliable procedures for
testing its statistical significance, using a large number of securities. In
practice, the number of tradeable securities, $n$, might not be sufficiently
large, and there are important specification and estimation uncertainties, and
the Sharpe ratio of the phi-based portfolio, $\boldsymbol{\rho}_{t,\phi}$, is
likely to be bounded even if $\boldsymbol{\phi}^{\prime}\boldsymbol{\phi}>0$.
We turn to these issues in the empirical application provided in\ Section
\ref{Empirical}.
\subsection{Fama-MacBeth and Shanken estimators of risk premia\label{FM}}
It will prove convenient to write (\ref{StatM}) in matrix notation by stacking
the excess returns by $t=1,2,...,T$, for each security $i$
\begin{equation}
\mathbf{r}_{i\circ}=\mathit{\alpha}_{i}\boldsymbol{\tau}_{T}+\mathbf{F}
\boldsymbol{\beta}_{i}+\mathbf{u}_{i\circ},\text{ for }i=1,2,...,n,
\label{mfi}
\end{equation}
where $\mathbf{r}_{i\circ}=(r_{i1},r_{i2},...,r_{iT})^{\prime}$,
$\mathbf{F=(f}_{1},\mathbf{f}_{2},...,\mathbf{f}_{T})^{\prime}$,
$\mathbf{u}_{i\circ}=\left( u_{i1},u_{i2},...,u_{iT}\right) ^{\prime}$, and
$\boldsymbol{\tau}_{T}$ is a $T\times1$ vector of ones. Similarly, stacking
the excess returns by $i$ for each $t$ we have (\ref{ritV}) which we rewrite
as
\begin{equation}
\mathbf{r}_{\circ t}=\boldsymbol{\alpha}_{n}+\boldsymbol{B}_{n}\mathbf{f}
_{t}+\mathbf{u}_{\circ t},\text{ for }t=1,2,...,T, \label{mfT}
\end{equation}
where $\boldsymbol{\alpha}_{n}=(\alpha_{1},\alpha_{2},...,\alpha_{n})^{\prime
}=c\mathbf{\tau}_{n}+\mathbf{B}_{n}\boldsymbol{\phi}+\mathbf{\eta}_{n}$.
The risk premia are usually estimated using a two-pass procedure suggested by
\citet{fama1973risk}
. The first-pass runs time series regressions of excess returns, $r_{it}$, on
the $K$ observed factors to give estimates of the factor loadings,
$\boldsymbol{\beta}_{i}:$
\begin{equation}
\text{ }\boldsymbol{\hat{\beta}}_{iT}=\left( \mathbf{F}^{\prime}
\mathbf{M}_{T}\mathbf{F}\right) ^{-1}\mathbf{F}^{\prime}\mathbf{M}
_{T}\mathbf{r}_{i\circ}. \label{betaihat}
\end{equation}
The second-pass runs a cross section regression of\ average returns, $\bar
{r}_{i\circ}=T^{-1}\sum_{t=1}^{T}r_{it}$ on the estimated factor loadings, to
obtain the FM estimator of $\boldsymbol{\lambda}$, namely
\begin{equation}
\boldsymbol{\hat{\lambda}}_{nT}=\left( \mathbf{\hat{B}}_{nT}^{\prime
}\mathbf{M}_{n}\mathbf{\hat{B}}_{nT}\right) ^{-1}\mathbf{\hat{B}}
_{nT}^{\prime}\mathbf{M}_{n}\mathbf{\bar{r}}_{n\circ}, \label{lambdahat}
\end{equation}
where $\mathbf{\hat{B}}_{nT}=(\boldsymbol{\hat{\beta}}_{1T},\boldsymbol{\hat
{\beta}}_{2T},...,\boldsymbol{\hat{\beta}}_{nT})^{\prime},$ $\mathbf{\bar{r}
}_{n\circ}=(\bar{r}_{1T},\bar{r}_{2T},...,\bar{r}_{nT})^{\prime},$
$\mathbf{M}_{T}=\mathbf{I}_{T}-T^{-1}\boldsymbol{\tau}_{T}\boldsymbol{\tau
}_{T}^{\prime}$, $\boldsymbol{\tau}_{T}$ is a $T$-dimensional vector of ones,
$\mathbf{M}_{n}=\mathbf{I}_{n}-n^{-1}\boldsymbol{\tau}_{n}\boldsymbol{\tau
}_{n}^{\prime}$, and $\boldsymbol{\tau}_{n}$ is an n-dimensional vector of ones.
As is well known, when $T$ is finite FM's two-pass estimator is biased due the
errors in estimation of factor loadings that do not vanish. The small $T$ bias
of the two-pass estimator of $\boldsymbol{\lambda}$ has been a source of
concern in the empirical literature. Under standard regularity conditions and
as $n\rightarrow\infty$, we have
\begin{equation}
\boldsymbol{\hat{\lambda}}_{nT}-\boldsymbol{\lambda}_{0}\rightarrow_{p}\left[
\mathbf{\Sigma}_{\beta\beta}+\frac{\overline{\sigma}^{2}}{T}\left(
\frac{\mathbf{F}^{\prime}\mathbf{M}_{T}\mathbf{F}}{T}\right) ^{-1}\right]
^{-1}\left( \mathbf{\Sigma}_{\beta\beta}\mathbf{d}_{fT}-\frac{\overline
{\sigma}^{2}}{T}\left( \frac{\mathbf{F}^{\prime}\mathbf{M}_{T}\mathbf{F}}
{T}\right) ^{-1}\boldsymbol{\lambda}_{0}\right) , \label{Bias}
\end{equation}
where $\mathbf{d}_{fT}=\mathbf{\hat{\mu}}_{T}-\mathbf{\mu}_{0},$
$\boldsymbol{\lambda}_{0}$ and $\mathbf{\mu}_{0}$ are the true values of
$\boldsymbol{\lambda}$ and $\mathbf{\mu}$, respectively, $\mathbf{\Sigma
}_{\beta\beta}=\lim_{n\rightarrow\infty}\left( n^{-1}\mathbf{B}_{n}^{\prime
}\mathbf{M}_{n}\mathbf{B}_{n}\right) $, and $\overline{\sigma}^{2}
=\lim_{n\rightarrow\infty}n^{-1}\sum_{i=1}^{n}\sigma_{i}^{2}>0$. Following
\citet{shanken1992estimation}
, $\overline{\sigma}_{n}^{2}$ can be consistently estimated (for a fixed $T$)
by
\begin{equation}
\widehat{\bar{\sigma}}_{nT}^{2}=\frac{\sum_{t=1}^{T}\sum_{i=1}^{n}\hat{u}
_{it}^{2}}{n(T-K-1)}, \label{zigbar}
\end{equation}
where
\begin{equation}
\hat{u}_{it}=r_{it}-\hat{\alpha}_{iT}-\boldsymbol{\hat{\beta}}_{iT}^{\prime
}\mathbf{f}_{t}, \label{res}
\end{equation}
and as before $\hat{\alpha}_{iT}$ and $\boldsymbol{\hat{\beta}}_{iT}$ are the
OLS estimators of $\mathit{\alpha}_{i}$ and $\boldsymbol{\beta}_{i}$. Using
these results the bias-corrected version of the two-pass estimator is given
by\footnote{See also
\citet{shanken2007estimating}
,
\citet{kan2013pricing}
, and
\citet{BAI201531}
, and the survey paper by
\citet{jagannathan2010analysis}
for further references.}
\begin{equation}
\boldsymbol{\tilde{\lambda}}_{nT}=\mathbf{H}_{nT\ }^{-1}\left( \frac
{\mathbf{\hat{B}}_{nT}^{\prime}\mathbf{M}_{n}\mathbf{\bar{r}}_{n\circ}}
{n}\right) , \label{lambda_BC}
\end{equation}
where
\begin{equation}
\mathbf{H}_{nT\ }=\frac{\mathbf{\hat{B}}_{nT}^{\prime}\mathbf{M}
_{n}\mathbf{\hat{B}}_{nT}}{n}-T^{-1}\widehat{\bar{\sigma}}_{nT}^{2}\left(
\frac{\mathbf{F}^{\prime}\mathbf{M}_{T}\mathbf{F}}{T}\right) ^{-1}.
\label{HnT1}
\end{equation}
When all the risk factors are strong, under certain regularity conditions,
there exists a fixed $T_{0}$ such that for all $T>T_{0}$, then
\begin{equation}
p\lim_{n\rightarrow\infty}\left( \boldsymbol{\tilde{\lambda}}_{nT}\right)
=\boldsymbol{\lambda}_{T}^{\ast}=\boldsymbol{\lambda}_{0}+\left(
\boldsymbol{\hat{\mu}}_{T}-\boldsymbol{\mu}_{0}\right) , \label{lambda*}
\end{equation}
where $\boldsymbol{\mu}_{0}$ indicates the true value of the factor mean.
Shanken refers to $\boldsymbol{\lambda}_{T}^{\ast}$ as "ex-post" risk premia
to be distinguished from $\boldsymbol{\lambda}_{0}$, referred to as "ex ante"
risk premia. See also Section 3.7 of
\citet{jagannathan2010analysis}
.
In this paper we exploit Shanken's bias correction procedure by applying it to
$\boldsymbol{\phi}=$ $\boldsymbol{\lambda}-\boldsymbol{\mu}$ which we identify
directly using (\ref{ai}) from the regression of $\mathit{\alpha}_{i}$ on
$\boldsymbol{\beta}_{i}$ for $i=1,2,...,n$, assuming the idiosyncratic pricing
errors, $\eta_{i}$, are sufficiently weak relative to the strengths of the
risk factors in a sense which will be made precise below.
\subsection{Estimation of $\boldsymbol{\phi}$\label{ESTphi}}
In view of (\ref{ai}), the estimation of $\boldsymbol{\phi}$ can be carried
out following a two-step procedure whereby in the first step $\mathit{\alpha
}_{i}$ and $\boldsymbol{\beta}_{i}$ are estimated from the least squares
regressions of $r_{it}$ on an intercept and $\mathbf{f}_{t}$, and these are
then used in a second step regression to estimate $\boldsymbol{\phi}$,
namely,
\begin{equation}
\boldsymbol{\hat{\phi}}_{nT}\text{ }=\left( \mathbf{\hat{B}}_{nT}^{\prime
}\mathbf{M}_{n}\mathbf{\hat{B}}_{nT}\right) ^{-1}\mathbf{\hat{B}}
_{nT}^{\prime}\mathbf{M}_{n}\boldsymbol{\hat{\alpha}}_{nT}, \label{phihat1}
\end{equation}
where $\boldsymbol{\hat{\alpha}}_{nT}=(\hat{\alpha}_{1T},\hat{\alpha}
_{1T},...,\hat{\alpha}_{nT})^{\prime}=\mathbf{\bar{r}}_{nT}-\mathbf{\hat{B}
}_{nT}\boldsymbol{\hat{\mu}}_{T},$ and as before $\mathbf{\hat{B}}
_{nT}=(\boldsymbol{\hat{\beta}}_{1T},\boldsymbol{\hat{\beta}}_{2T}
,...,\boldsymbol{\hat{\beta}}_{nT})^{\prime}$. This estimator is consistent
for $\boldsymbol{\phi}_{0}$ so long as $n$ and $T\rightarrow\infty$, and
bias-corrections are necessary to ensure the large $n$ consistency of the
estimator when $T$ is fixed. A Shanken type bias-corrected estimator of
$\boldsymbol{\phi}_{0}$ is given by
\begin{equation}
\boldsymbol{\tilde{\phi}}_{nT}=\mathbf{H}_{nT\ }^{-1}\left[ \frac
{\mathbf{\hat{B}}_{nT}^{\prime}\mathbf{M}_{n}\hat{\alpha}_{nT}}{n}
+T^{-1}\widehat{\bar{\sigma}}_{nT}^{2}\left( \frac{\mathbf{F}^{\prime
}\mathbf{M}_{T}\mathbf{F}}{T}\right) ^{-1}\boldsymbol{\hat{\mu}}_{T}\right]
, \label{phitilda}
\end{equation}
where $\mathbf{H}_{nT\ }$ and $\widehat{\bar{\sigma}}_{nT}^{2}$ are given by
(\ref{HnT1}) and (\ref{zigbar}), respectively. It is also easily established
that
\begin{equation}
\boldsymbol{\tilde{\phi}}_{nT}=\boldsymbol{\tilde{\lambda}}_{nT}
-\boldsymbol{\hat{\mu}}_{T}, \label{phitilda2}
\end{equation}
and for a fixed $T$ and as $n\rightarrow\infty$, we have
\[
p\lim_{n\rightarrow\infty}\boldsymbol{\tilde{\phi}}_{nT}=p\lim_{n\rightarrow
\infty}\boldsymbol{\tilde{\lambda}}_{nT}-\boldsymbol{\hat{\mu}}_{T}.
\]
Hence, upon using (\ref{lambda*})
\begin{equation}
p\lim_{n\rightarrow\infty}\boldsymbol{\tilde{\phi}}_{nT}=\boldsymbol{\lambda
}_{0}+\left( \boldsymbol{\hat{\mu}}_{T}-\boldsymbol{\mu}_{0}\right)
-\boldsymbol{\hat{\mu}}_{T}=\boldsymbol{\lambda}_{0}-\boldsymbol{\mu}
_{0}=\boldsymbol{\phi}_{0}\text{, } \label{Conphitilt}
\end{equation}
and there exists a fixed $T_{0}$ such that for all $T>T_{0},$
$\boldsymbol{\tilde{\phi}}_{nT}$ converges to $\boldsymbol{\phi}_{0}$ as
$n\rightarrow\infty$. \ Also using (\ref{lambda*}) and (\ref{phitilda2}), and
noting that $\boldsymbol{\lambda}_{0}-\boldsymbol{\mu}_{0}=\boldsymbol{\phi
}_{0}$, interestingly we have
\[
\boldsymbol{\tilde{\lambda}}_{nT}-\boldsymbol{\lambda}_{T}^{\ast
}=\boldsymbol{\tilde{\phi}}_{nT}+\boldsymbol{\hat{\mu}}_{T}
-\boldsymbol{\lambda}_{T}^{\ast}=\boldsymbol{\tilde{\phi}}_{nT}
-\boldsymbol{\phi}_{0}.
\]
So inference using the Shanken bias-corrected estimator of
$\boldsymbol{\lambda}$ around $\boldsymbol{\lambda}_{T}^{\ast}$, is the same
as making inference using $\boldsymbol{\tilde{\phi}}_{nT}$ around
$\boldsymbol{\phi}_{0}$.
The asymptotic distribution of $\boldsymbol{\tilde{\phi}}_{nT}$ depends on
both $n$ and $T$. Assuming the observed factors are strong and under certain
regularity conditions, to be introduced below, we have
\begin{equation}
\sqrt{nT}\left( \boldsymbol{\tilde{\phi}}_{nT}-\boldsymbol{\phi}_{0}\right)
\rightarrow_{d}N\left( \mathbf{0,\Sigma}_{\beta\beta}^{-1}\mathbf{V}_{\xi
}\mathbf{\Sigma}_{\beta\beta}^{-1}\right) , \label{Dphitilt}
\end{equation}
where
\[
\mathbf{V}_{\xi}=\left( 1+\boldsymbol{\lambda}_{0}^{\prime}\mathbf{\Sigma
}_{f}^{-1}\boldsymbol{\lambda}_{0}\right) p\lim_{n\rightarrow\infty}\left[
n^{-1}\mathbf{B}_{n}^{\prime}\mathbf{M}_{n}\mathbf{V}_{u}\mathbf{M}
_{n}\mathbf{B}_{n\ }\right] \emph{.}
\]
The variance of $\boldsymbol{\tilde{\phi}}_{nT}$ is consistently estimated by
\begin{equation}
\widehat{Var\left( \boldsymbol{\tilde{\phi}}_{nT}\right) }=T^{-1}
n^{-1}\mathbf{H}_{nT\ }^{-1}\mathbf{\hat{V}}_{\xi,nT}\mathbf{H}_{nT\ }^{-1},
\label{Varphitilt}
\end{equation}
where $\mathbf{H}_{nT\ }$ is given by (\ref{HnT1}),
\begin{equation}
\mathbf{\hat{V}}_{\xi,nT}=\left( 1+\hat{s}_{nT}\right) \left(
n^{-1}\mathbf{\hat{B}}_{n}^{\prime}\mathbf{M}_{n}\mathbf{\tilde{V}}
_{u}\mathbf{M}_{n}\mathbf{\hat{B}}_{n}\right) , \label{GnT}
\end{equation}
and
\begin{equation}
\hat{s}_{nT}=\boldsymbol{\tilde{\lambda}}_{nT}^{^{\prime}}\left(
\frac{\mathbf{F}^{\prime}\mathbf{M}_{T}\mathbf{F}}{T}\right) ^{-1}
\boldsymbol{\tilde{\lambda}}_{nT}, \label{shatnT}
\end{equation}
$\boldsymbol{\tilde{\lambda}}_{nT}$ is defined by (\ref{lambda_BC}), and
$\mathbf{\tilde{V}}_{u}\mathbf{\ }$is a suitable estimator of $\mathbf{V}
_{u}=E\left( \mathbf{u}_{\circ t}\mathbf{u}_{\circ t}^{\prime}\right) $. How
to estimate $\mathbf{V}_{u}$ and the conditions under which $Var\left(
\boldsymbol{\tilde{\phi}}_{nT}\right) $ is consistently estimated by
$\widehat{Var\left( \boldsymbol{\tilde{\phi}}_{nT}\right) }$ is discussed in
sub-section \ref{Vphihat}.
Note that even when $\mathbf{V}_{u}=\sigma^{2}\mathbf{I}_{T}$ the variance of
$\boldsymbol{\tilde{\phi}}_{nT}$ does not reduce to $\sigma^{2}
\boldsymbol{\Sigma}_{\beta\beta}^{-1}$, the standard least squares formula
used for the case of known factor loadings. When the loadings are estimated
the scaling term $\left( 1+\boldsymbol{\lambda}_{0}^{\prime}
\boldsymbol{\Sigma}_{f}^{-1}\boldsymbol{\lambda}_{0}\right) $ is required and
its neglect can lead to serious over-rejection even if $n/T\rightarrow0$ as
$n$ and $T\rightarrow\infty$.
\subsection{Factor strength \label{facstrength}}
In this paper we deviate from the standard literature and allow the observed
and latent factors to have different degrees of strength, depending on how
pervasively they impact the security returns.
\citet{bailey2021measurement}
define the strength of factor, $f_{kt}$, in terms of the number of its
non-zero factor loadings. For a factor to be strong almost all of its $n$
loadings must differ from zero. Given our focus on estimation of risk premia,
we adopt the following definition which directly relates to the covariance of
$\boldsymbol{\beta}_{i}.$ See also
\citet{chudik2011weak}
.
\begin{definition}
\label{strengths}(Factor strengths) The strength of factor $f_{kt}$ is
measured by its degree of pervasiveness as defined by the exponent
$\alpha_{_{k}}$ in
\begin{equation}
\sum_{i=1}^{n}\left( \beta_{ik}-\bar{\beta}_{k}\right) ^{2}=\ominus
(n^{\alpha_{k}}), \label{factor strength}
\end{equation}
and $0<\alpha_{k}\leq1$. We refer to $\left\{ \alpha_{k},\text{
}k=1,2,...,K\right\} $ as factor strengths. Factor $f_{kt}$ is said to be
strong if $\alpha_{k}=1$, semi-strong if $1>\alpha_{k}>1/2$, and weak if
$0\leq\alpha_{k}\leq1/2$. Condition (\ref{factor strength}) applies
irrespective of whether the loadings, $\beta_{ik}$, are viewed as
deterministic or stochastic.\
\end{definition}
In the above definition $\ominus_{p}\left( n^{\alpha_{k}}\right) $ denotes
the rate at which additional securities add to the factor's strength and
$\alpha_{k}$ can be viewed as a logarithmic expansion rate in terms of $n$ and
relates to the proportion of non-zero factor loadings. In the literature it is
commonly assumed that the covariance matrix of factor loadings defined by
\begin{equation}
\boldsymbol{\Sigma}_{\beta\beta}=p\lim_{n\rightarrow\infty}\left[ n^{-1}
\sum_{i=1}^{n}\left( \boldsymbol{\beta}_{i}-\boldsymbol{\bar{\beta}}
_{n}\right) \left( \boldsymbol{\beta}_{i}-\boldsymbol{\bar{\beta}}
_{n}\right) ^{\prime}\right] , \label{Zigbeta}
\end{equation}
is positive definite, where $\boldsymbol{\bar{\beta}}_{n}=n^{-1}\sum_{i=1}
^{n}\boldsymbol{\beta}_{i}=(\bar{\beta}_{1},\bar{\beta}_{2},...,\bar{\beta
}_{k})^{\prime}$. For $\mathbf{\Sigma}_{\beta\beta}$ to be positive definite
matrix it is \textit{necessary} that all the $K$ risk factors under
consideration are strong in the sense that
\begin{equation}
p\lim_{n\rightarrow\infty}\left[ n^{-1}\sum_{i=1}^{n}\left( \beta_{ik}
-\bar{\beta}_{k}\right) ^{2}\right] >0\text{, for }k=1,2,...,K.
\label{Sfact}
\end{equation}
In terms of our definition of factor strength, $\mathbf{\Sigma}_{\beta\beta}
$\ will be positive definite if all the observed factors are strong, namely if
$\alpha_{k}=1$ for $k=1,2,...,K$. However, such an assumption is quite
restrictive and is unlikely to be satisfied for many risk factors being
considered in the literature.
\citet{bailey2021measurement}
show that, apart from the market factor, only a handful of 144 factors in the
literature considered by
\citet{feng2020taming}
come close to being strong. \cite{giglio2023test} consider the estimation of
PCA-based risk premia in presence of weak factors. However, their definition
of factor strength involves both $n~$and $T$, and is best viewed as a
consistency condition rather than factor strength as such. See the discussion
following Theorem \ref{Tsemi}. Our notion of factor strength, $\alpha_{k}$, is
in line with the recent literature. See, for example, \cite{BaiNg2023Weak} and
\cite{uematsu2023inference}.
\subsection{Missing factor}
We now turn to the structure of the errors, $u_{it}$, in the returns
equations, and consider two possible sources of error cross-sectional
dependence: a missing or latent factor and production networks. The issue of
missing factors has been investigated in the recent literature by
\citet{giglio2021asset}
and
\citet{ANATOLYEV2022103}
. The issue of production networks has been investigated in the recent
literature by
\citet{herskovic2018networks}
, who derives two risk factors based on the changes in network concentration
and network sparsity, and
\citet{gofman2020production}
, who focus on the vertical dimension of production by modelling a supply
chain, in terms of supplier-customer links. They find that the further away a
firm is from final consumers the higher its return. They use this to create a
factor TMB (top minus bottom). Both sources of cross-sectional error
dependence could be important, since network dependence cannot be represented
using latent factor models. See Section 3 of \cite{chudik2011weak}.
To allow for both forms of error cross-sectional dependence we consider the
following decomposition of $u_{it}$
\begin{equation}
u_{it}=\gamma_{i}g_{t}+v_{it}, \label{uit}
\end{equation}
where $g_{t}$ is the missing (latent) factor and $v_{it}$ is weakly
cross-correlated in the sense of approximate factor models due to
\cite{Chamberlain1983} and \cite{cham1983arbi}. Here we allow for a single
missing factor to simplify the exposition, but note that increasing the number
of missing factors has little impact on our analysis, so long as the number of
missing factors is fixed. Using the normalization $E(g_{t}^{2})=1$, and
assuming that $\gamma_{i}g_{t}$ and $v_{it}$ are independently distributed
then $E(u_{it}u_{jt})=\sigma_{ij}=\gamma_{i}\gamma_{j}+\sigma_{v,ij}$, and as
shown in Lemma \ref{L1}, $n^{-1}\sum_{i=1}^{n}\sum_{j=1}^{n}\left\vert
\sigma_{ij}\right\vert =O(1)$ so long as the strength of $g_{t},$
$\alpha_{\gamma}<1/2$ and $\lambda_{\max}(\mathbf{V}_{v})<\infty.$ This is
despite the fact that $\lambda_{\max}(\mathbf{V}_{u})=O(n^{\alpha_{\gamma}})$,
where $\mathbf{V}_{v}=E\left( \mathbf{v}_{i}\mathbf{v}_{i}^{\prime}\right) $
and $\mathbf{V}_{u}=E\left( \mathbf{u}_{i}\mathbf{u}_{i}^{\prime}\right)
$.\footnote{Note that Chamberlain's approximate factor model specification
requires $\lambda_{\max}(\mathbf{V}_{u})=O(1)$ and is violated if
$\alpha_{\gamma}>0$.} In the Monte Carlo experiments, we consider the
possibility of missing factors, as well as weak spatial and network
cross-dependence that satisfy conditions of approximate factor models.
\subsection{Pricing errors and market efficiency\label{pricing}}
The APT condition (\ref{APTRoss}), given by (18)\ in Theorem II of
\citet{ROSS1976341}
, ensures that under APT the idiosyncratic pricing errors are sparse. In this
paper we relax the\ Ross's condition to
\begin{equation}
\sum_{i=1}^{n}\eta_{i}^{2}=O(n^{\alpha_{\eta}}), \label{APTg}
\end{equation}
where the exponent $\alpha_{\eta}$ measures the degrees of pervasiveness of
pricing errors. Deviations from APT are measured in terms of $\alpha_{\eta}$
($0\leq\alpha_{\eta}<1$). We investigate the robustness of our proposed
estimator of $\boldsymbol{\phi}$ to $\alpha_{\eta}$. This extension is
important for tests of market efficiency where the null of interest is
$H_{0}:$ $\mathit{\alpha}_{i}=c$ for all $i$ in (\ref{ai}). We note that under
the alternative hypothesis $H_{1}:$ $\mathit{\alpha}_{i}=c+\boldsymbol{\beta
}_{i}^{\prime}\boldsymbol{\phi}+\eta_{i}$, therefore it is desirable to
develop a test of $\boldsymbol{\phi}=\boldsymbol{0}$ which is robust to a
wider class of pricing errors than those entertained originally by Ross, where
$\alpha_{\eta}=0$.
Under the alternative hypothesis the power of testing $H_{0}$ will depend on
the rate at which $\sum_{i=1}^{n}\left( \mathit{\alpha}_{i}-\mathit{\bar
{\alpha}}\right) ^{2}$ rises with $n$. This in turn depends on the degree of
pervasiveness of the idiosyncratic pricing errors, $\alpha_{\eta}$, the
strength of the factors, $\alpha_{j}$, for $j=1,2,...,K$ and the magnitude of
$\boldsymbol{\phi}^{\prime}\boldsymbol{\phi}\,$. Using $\mathit{\alpha}
_{i}=c+\boldsymbol{\beta}_{i}^{\prime}\boldsymbol{\phi}+\eta_{i}$
\[
\sum_{i=1}^{n}\left( \mathit{\alpha}_{i}-\mathit{\bar{\alpha}}\right)
^{2}=\boldsymbol{\phi}^{\prime}\left( \sum_{i=1}^{n}\left( \boldsymbol{\beta
}_{i}-\boldsymbol{\bar{\beta}}_{n}\right) \left( \boldsymbol{\beta}
_{i}-\boldsymbol{\bar{\beta}}_{n}\right) ^{\prime}\right) \boldsymbol{\phi
}+\sum_{i=1}^{n}\left( \eta_{i}-\bar{\eta}\right) ^{2}+2\boldsymbol{\phi
}^{\prime}\sum_{i=1}^{n}\left( \boldsymbol{\beta}_{i}-\boldsymbol{\bar{\beta
}}_{n}\right) \eta_{i},
\]
and in the absence of idiosyncratic pricing errors
\begin{equation}
\sum_{i=1}^{n}\left( \mathit{\alpha}_{i}-\mathit{\bar{\alpha}}\right)
^{2}\geq n\left( \boldsymbol{\phi}^{\prime}\boldsymbol{\phi}\right)
\lambda_{\min}\left( n^{-1}\sum_{i=1}^{n}\left( \boldsymbol{\beta}
_{i}-\boldsymbol{\bar{\beta}}_{n}\right) \left( \boldsymbol{\beta}
_{i}-\boldsymbol{\bar{\beta}}_{n}\right) ^{\prime}\right) . \label{alphaE}
\end{equation}
When the risk factors are all strong $\lambda_{\min}\left( n^{-1}\sum
_{i=1}^{n}\left( \boldsymbol{\beta}_{i}-\boldsymbol{\bar{\beta}}_{n}\right)
\left( \boldsymbol{\beta}_{i}-\boldsymbol{\bar{\beta}}_{n}\right) ^{\prime
}\right) >0$, then $\sum_{i=1}^{n}\left( \mathit{\alpha}_{i}-\mathit{\bar
{\alpha}}\right) ^{2}=\ominus(n)$ if and only if $\boldsymbol{\phi}
\neq\boldsymbol{0}$. Namely, the extent to which alpha can be exploited will
depend on the magnitudes of $\phi_{j}$ and the strength of the risk factors.
\begin{remark}
\label{Remfulldiv}Having formalized the concepts of factor strength, missing
factors, and the less restrictive APT condition given by (\ref{APTg}), it is
now of interest to revisit the conditions under which the phi-portfolio fully
diversifies. Consider (\ref{fulldiv}) and note that
\[
\frac{\left\Vert \boldsymbol{\eta}_{n}\right\Vert ^{2}}{\lambda_{min}\left(
\mathbf{B}_{n}^{\prime}\mathbf{M}_{n}\mathbf{B}_{n}\right) }=O\left[
n^{-\left( \alpha_{min}-\alpha_{\eta}\right) }\right] \text{, and }
\frac{\lambda_{max}\left( \mathbf{V}_{u}\right) }{\lambda_{min}\left(
\mathbf{B}_{n}^{\prime}\mathbf{M}_{n}\mathbf{B}_{n}\right) }=O\left[
n^{-\left( \alpha_{min}-\alpha_{\gamma}\right) }\right] .
\]
Using these results in (\ref{phiport1})\ and (\ref{varphi}) and we have
\[
E\left( \rho_{t,\phi}\right) =\boldsymbol{\phi}^{\prime}\boldsymbol{\phi
}+O\left[ n^{-\left( \alpha_{min}-\alpha_{\eta}\right) }\right] \text{,
and }Var\left( \rho_{t,\phi}\right) =O\left[ n^{-\left( \alpha
_{min}-\alpha_{\gamma}\right) }\right]
\]
Therefore, for phi-portfolio to dominate the MV portfolio in addition to
$\boldsymbol{\phi}^{\prime}\boldsymbol{\phi}>0$, it is also required that the
strengths of the traded factors $\alpha_{k}$, $k=1,2,...,K$ are strictly
larger than the strength of the missing factor, $\alpha_{\gamma},$ as well as
the strength of the idiosyncratic pricing errors, $\alpha_{\eta}$.
\end{remark}
\begin{remark}
Similarly, If we allow for idiosyncratic pricing errors, missing factors and
non-strong risk factors the limiting expression for $\sum_{i=1}^{n}\left(
\mathit{\alpha}_{i}-\mathit{\bar{\alpha}}\right) ^{2}$ becomes
\begin{equation}
\sum_{i=1}^{n}\left( \mathit{\alpha}_{i}-\mathit{\bar{\alpha}}\right)
^{2}=\ominus(\sum_{j=1}^{K}n^{\alpha_{j}}\phi_{j}^{2})+O\left( n^{\alpha
_{\eta}}\right) +O\left( n^{\alpha_{\gamma}}\right) . \label{alphaG}
\end{equation}
In this more general case for alpha to be exploitable it is necessary that
$\phi_{j}$ associated with the strongest factor is non-zero, and $\alpha
_{\max}=\max_{j}(\alpha_{j})$ is larger than $\alpha_{\eta}$ and
$\alpha_{\gamma}$.
\end{remark}
\section{Assumptions and theorems\label{Assumptions}}
We make the following standard assumptions about $\mathbf{f}_{t},g_{t}$,
$v_{it}$, $\boldsymbol{\beta}_{i}$, $\eta_{i}$, and $\gamma_{i}$ (the drivers
of asset returns):
\begin{assumption}
\label{factors} (Observed common factors) (a) The $K\times1$ vector of
observed risk factors, $\mathbf{f}_{t}$, follows the general linear process
\begin{equation}
\mathbf{f}_{t}=\boldsymbol{\mu}+\sum_{\ell=0}^{\infty}\mathbf{\Psi}_{\ell
}\mathbf{\zeta}_{t-\ell}, \label{ft}
\end{equation}
where $\left\Vert \boldsymbol{\mu}\right\Vert <C$, $\mathbf{\zeta}
_{t}\thicksim IID(\mathbf{0},\mathbf{I}_{K})$, and $\mathbf{\Psi}_{\ell}$ are
$K\times K$ exponentially decaying matrices such that $\left\Vert
\mathbf{\Psi}_{\ell}\right\Vert <C\rho^{\ell}$ for some $C>0$ and $0<\rho<1$.
(b) The $T\times K$ data matrix $\mathbf{F=(f}_{1},\mathbf{f}_{2}
,...,\mathbf{f}_{T})^{\prime}$ is full column rank and there exists $T_{0}$
such that for all $T>T_{0}$, $\boldsymbol{\hat{\Sigma}}_{f}=T^{-1}
\mathbf{F}^{\prime}\mathbf{M}_{T}\mathbf{F}$ is a positive definite matrix,
$\lambda_{\max}\left[ (T^{-1}\mathbf{F}^{\prime}\mathbf{M}_{T}\mathbf{F)}
^{-1}\right] <C$, $\boldsymbol{\hat{\Sigma}}_{f}\rightarrow_{p}
\boldsymbol{\Sigma}_{f}=E\left( \mathbf{f}_{t}-\boldsymbol{\mu}_{0}\right)
\left( \mathbf{f}_{t}-\boldsymbol{\mu}_{0}\right) ^{\prime}>\mathbf{0},$
where $\boldsymbol{\mu}_{0}$ is the true value of $\boldsymbol{\mu}$.
\end{assumption}
\begin{assumption}
\label{loadings}(Observed factor loadings) (a) The factor loadings $\beta
_{ik}$ for $i=1,2,...,n$ and $k=1,2,...,K$ are stochastically bounded such
that $sup_{ik}E\left( \beta_{ik}^{2}\right) <C$,
\begin{equation}
\sum_{i=1}^{n}\left( \beta_{ik}-\bar{\beta}_{k}\right) ^{2}=\ominus
_{p}(n^{\alpha_{k}})\text{, for }k=1,2,...,K. \label{As}
\end{equation}
(b)The $n\times K$ matrix of factor loadings, $\mathbf{B}_{n}
=(\boldsymbol{\beta}_{\circ1},\boldsymbol{\beta}_{\circ2}
,...,\boldsymbol{\beta}_{\circ K}),$ where $\boldsymbol{\beta}_{\circ
k}=(\beta_{1k},\beta_{2k},...,\beta_{nk})^{\prime}$ satisfy
\begin{equation}
0<c<\lambda_{min}\left( \mathbf{D}_{\alpha}^{-1}\mathbf{B}_{n}^{\prime
}\mathbf{M}_{n}\mathbf{B}_{n}\mathbf{D}_{\alpha}^{-1}\right) <\lambda
_{max}\left( \mathbf{D}_{\alpha}^{-1}\mathbf{B}_{n}^{\prime}\mathbf{M}
_{n}\mathbf{B}_{n}\mathbf{D}_{\alpha}^{-1}\right) <C<\infty, \label{lmax}
\end{equation}
for some small and large positive constants, $c$ and $C$, where $\mathbf{M}
_{n}=\mathbf{I}_{n}-n^{-1}\boldsymbol{\tau}_{n}\boldsymbol{\tau}_{n}^{\prime}
$, $\boldsymbol{\tau}_{n}=(1,1,...,1)^{\prime},$ and $\mathbf{D}_{\alpha}
\,$\ is the $n\times n$ diagonal matrix
\begin{equation}
\mathbf{D}_{\alpha}=Diag(n^{\alpha_{1}/2},n^{\alpha_{2}/2},....,n^{\alpha
_{K}/2}). \label{Dn}
\end{equation}
\end{assumption}
\begin{assumption}
\label{Latent factor} (latent factor) (a) The latent factor, $g_{t}$, in
(\ref{uit}) is distributed independently of $\mathbf{f}_{t^{\prime}},$ for all
$t$ and $t^{\prime}$, $g_{t}$ is serially independent with mean zero,
$E(g_{t})=0,$ $E(g_{t}^{2})=1$, and a finite fourth order moment,
$sup_{t}E(g_{t}^{4})<C$. (b) The loadings $\gamma_{i}$ are such that
$sup_{i}\left\vert \gamma_{i}\right\vert <C$ and
\begin{equation}
\sum_{i=1}^{n}\left\vert \gamma_{i}\right\vert =O(n^{\alpha_{\gamma}}).
\label{normg}
\end{equation}
\end{assumption}
\begin{assumption}
\label{Errors} (idiosyncratic errors) (a) The errors $\left\{ v_{it}\text{,
}i=1,2,...,n;\text{ }t=1,2,...,T\right\} $ are distributed independently of
the factors $f_{k,t^{\prime}}$, and $g_{t}$, for all $i,t,t^{\prime}$ and
$k=1,2,...,K$, and their associated loadings $\beta_{ik},$ and $\gamma_{i}$.
They are serially independent with $E(v_{it})=0$ and finite fourth order
moments $E(v_{it}^{4})<\infty$, and covariances $E(v_{it}v_{jt})=\sigma
_{v,ij}$, such that
\begin{equation}
\sup_{i}\sum_{j=1}^{n}\left\vert \sigma_{v,ij}\right\vert <\infty,\text{ and
}\sup_{i}\sum_{j=1}^{n}Cov(v_{it}^{2},v_{jt}^{2})<\infty,\text{ } \label{Covv}
\end{equation}
with $\lambda_{\min}\left( \mathbf{V}_{v}\right) >0$, where $\mathbf{V}
_{v}=(\sigma_{v,ij})$. (b) The degree of cross-sectional dependence of
$v_{it}$ is sufficiently weak so that
\begin{equation}
T^{-1/2}n^{-1/2}\sum_{t=1}^{T}\sum_{i=1}^{n}(\beta_{ik}-\bar{\beta}_{k}
)v_{it}\rightarrow_{d}N(0,\omega_{k}^{2})\text{, for }k=1,2,...,K,
\label{Disv}
\end{equation}
where
\begin{equation}
\omega_{k}^{2}=p\lim_{n\rightarrow\infty}n^{-\alpha_{k}}\sum_{i=1}^{n}
\sum_{j=1}^{n}(\beta_{ik}-\bar{\beta}_{k})(\beta_{jk}-\bar{\beta}_{k}
)\sigma_{v,ij}. \label{Varv}
\end{equation}
\end{assumption}
\begin{assumption}
\label{PriceError} (Pricing errors) The pricing errors, $\eta_{i}$, for
$i=1,2,...,n$ are individually bounded, $sup_{j}\left\vert \eta_{j}\right\vert
<C$\thinspace, and are distributed independently of the factor loadings,
$\beta_{jk}$, and $\gamma_{j}$ for all $i,j$ and $k=1,2,...,K$, as well as
satisfying the condition
\begin{equation}
\sum_{i=1}^{n}\left\vert \eta_{i}\right\vert =O\left( n^{\alpha_{\eta}
}\right) , \label{NormetaD}
\end{equation}
with $\alpha_{\eta}<1/2$.
\end{assumption}
\begin{remark}
Under Assumption \ref{factors} $E\left( \mathbf{f}_{t}\right) =\mathbf{\mu
},$ and $Var(\mathbf{f}_{t})=\boldsymbol{\Sigma}_{f}=\sum_{\ell=0}^{\infty
}\mathbf{\Psi}_{\ell}\mathbf{\Psi}_{\ell}^{\prime}$. Also since $\left\Vert
\boldsymbol{\Sigma}_{f}\right\Vert \leq\sum_{\ell=0}^{\infty}\left\Vert
\mathbf{\Psi}_{\ell}\right\Vert ^{2}$ it then follows from part (a) of
Assumption \ref{factors} that $\left\Vert \boldsymbol{\Sigma}_{f}\right\Vert
<C$.
\end{remark}
\begin{remark}
\label{remarkstrength} Under Assumption \ref{loadings}
\begin{equation}
\mathbf{D}_{\alpha}^{-1}\mathbf{B}_{n}^{\prime}\mathbf{M}_{n}\mathbf{B}
_{n}\mathbf{D}_{\alpha}^{-1}\rightarrow_{p}\mathbf{\Sigma}_{\beta\beta
}(\boldsymbol{\alpha}\mathbf{)}>0, \label{zigbb}
\end{equation}
where $\mathbf{\Sigma}_{\beta\beta}(\boldsymbol{\alpha}\mathbf{)}$ is a
$k\times k$ symmetric positive definite matrix which is a function of
$\boldsymbol{\alpha}=\mathbf{(}\alpha_{1},\alpha_{2},...,\alpha_{K})^{\prime}
$. This follows from (\ref{lmax}) since for any non-zero $n\times1$ vector
$\boldsymbol{\kappa}$,
\[
\boldsymbol{\kappa}^{\prime}\mathbf{D}_{\alpha}^{-1}\mathbf{B}_{n}^{\prime
}\mathbf{M}_{n}\mathbf{B}_{n}\mathbf{D}_{\alpha}^{-1}\boldsymbol{\kappa}
\geq\left( \mathbf{\kappa}^{\prime}\mathbf{\kappa}\right) \lambda
_{min}\left( \mathbf{D}_{\alpha}^{-1}\mathbf{B}_{n}^{\prime}\mathbf{M}
_{n}\mathbf{B}_{n}\mathbf{D}_{\alpha}^{-1}\right) >0.
\]
In the standard case where the factors are all strong ($\alpha_{k}=1$ for all
$k$), the above limit reduces to $n^{-1}\mathbf{B}_{n}^{\prime}\mathbf{M}
_{n}\mathbf{B}_{n}\rightarrow_{p}\mathbf{\Sigma}_{\beta\beta}(\boldsymbol{\tau
}_{K}\mathbf{)=\Sigma}_{\beta\beta}>0$.
\end{remark}
\begin{remark}
The high level condition (\ref{Disv}) in Assumption \ref{Errors} is required
for establishing the asymptotic normality of the estimator of
$\boldsymbol{\phi}_{0}$, and is clearly met when $v_{it}$ and/or $\beta_{ik}$
are independently distributed. It is also possible to establish (\ref{Disv})
under weaker conditions assuming that $v_{it}$ and/or $\beta_{ik}$ satisfy
some time-series type mixing conditions applied to cross section.
\end{remark}
\begin{remark}
The exponent parameter, $\alpha_{\eta}$, of the pricing condition in
(\ref{NormetaD}), can be viewed as the degree to which pricing errors are
pervasive in large economies (as $n\rightarrow\infty$). Letting
$\boldsymbol{\eta}_{n}=(\eta_{1},\eta_{2},...,\eta_{n})^{\prime}\,$\ we have
\begin{equation}
\sum_{i=1}^{n}\eta_{i}^{2}=\left\Vert \boldsymbol{\eta}_{n}\right\Vert
^{2}\leq\left\Vert \boldsymbol{\eta}_{n}\right\Vert _{\infty}\left\Vert
\boldsymbol{\eta}_{n}\right\Vert _{1}=sup_{j}\left\vert \eta_{j}\right\vert
\left( \sum_{i=1}^{n}\left\vert \eta_{i}\right\vert \right) , \label{sqeta}
\end{equation}
and under Assumption (\ref{PriceError}) it also follows that
\begin{equation}
\sum_{i=1}^{n}\eta_{i}^{2}=O\left( n^{\alpha_{\eta}}\right) . \label{Snorm}
\end{equation}
Similarly
\begin{equation}
\sum_{i=1}^{n}\gamma_{i}^{2}=O(n^{\alpha_{\gamma}}). \label{Snormg}
\end{equation}
\end{remark}
\begin{remark}
Whilst (\ref{NormetaD}) implies (\ref{Snorm}), the reverse does not follow. By
allowing for $\alpha_{\eta}>0$ we are relaxing the Ross's boundedness
condition that requires setting $\alpha_{\eta}=0$.
\end{remark}
\begin{remark}
The assumption that the observed and missing factors, $\mathbf{f}_{t}$ and
$g_{t^{\prime}},$ are distributed independently is not restrictive and can be
relaxed. For example, suppose that
\[
g_{t}=\mu_{g}+\boldsymbol{\theta}^{\prime}\mathbf{f}_{t}+v_{gt},
\]
where $\boldsymbol{f}_{t}$ and $v_{gt}$ are independently distributed. Then
using (\ref{uit}) we have
\[
u_{it}=\gamma_{i}\mu_{g}+\gamma_{i}\left( \boldsymbol{\theta}^{\prime
}\mathbf{f}_{t}\right) +\gamma_{i}v_{gt}+v_{it},
\]
and the return equation (\ref{StatM}) can be written as
\[
r_{it}=\left( \mathit{\alpha}_{i}+\gamma_{i}\mu_{g}\right) +\left(
\boldsymbol{\beta}_{i}+\gamma_{i}\boldsymbol{\theta}\right) ^{\prime
}\mathbf{f}_{t}+\gamma_{i}v_{gt}+v_{it},
\]
with $v_{gt}$ now acting as the missing common factor, which, by construction,
is distributed independently of $\mathbf{f}_{t}$.
\end{remark}
\begin{remark}
Assumptions \ref{Latent factor} and \ref{Errors} allow $u_{it}$ to be
cross-sectionally weakly correlated, but require the errors to be serially
uncorrelated. This requirement is not strong for asset pricing models, since
realized returns are only mildly serially correlated and most likely such
serial dependence will be captured by the serial correlation in the observed factors.
\end{remark}
As we shall see, to estimate and conduct inference on the risk premia
associated with the observed factors, $f_{kt}$, we require $\alpha_{k}
>\alpha_{\gamma}<1/2$, where $\alpha_{\gamma}$ denotes the strength of the
latent factor, $g_{t},$ and similarly defined by $\sum_{i=1}^{n}\gamma_{i}
^{2}=\ominus(n^{\alpha_{\gamma}})$. Namely, the latent factor must be
sufficiently weak so that ignoring it will be inconsequential, and observed
factors sufficiently strong so that they can be distinguished from the weak
latent factor.
The main theoretical results of the paper are set out around five theorems.
Theorem \ref{TFMbias} considers the Fama-MacBeth two-step estimator and
derives its limiting property as $n$ and $T\rightarrow\infty$. To eliminate
the bias of Fama-MacBeth estimator we require $n/T\rightarrow0$, and to
eliminate the effects of pricing errors we need $Tn^{\alpha_{\eta}
}/n\rightarrow0$, which results in a contradiction. Thus the Fama-MacBeth
estimator is valid only when there are no pricing errors ($\eta_{i}=0$ for all
$i$) and when $n/T\rightarrow0$. Theorem \ref{Thzig} provides a proof that the
estimator of $\bar{\sigma}_{n}^{2}$ (denoted by $\widehat{\bar{\sigma}}
_{nT}^{2}$) proposed by
\citet{shanken1992estimation}
continues to be unbiased for a fixed $T$ as $n\rightarrow\infty$, even under
the general setting of the current paper that allows for missing factors as
well as pricing errors. Theorem \ref{Thzig} also establishes that
$\widehat{\bar{\sigma}}_{nT}^{2}-\bar{\sigma}_{n}^{2}\rightarrow
O_{p}(n^{-1/2}T^{-1/2})$, which is essential for establishing the results for
the bias-corrected estimator of $\boldsymbol{\phi}_{0}$, namely
$\boldsymbol{\tilde{\phi}}_{nT}$ given by (\ref{phitilda}), summarized in
Theorem \ref{Tfi}. This theorem provides conditions under which
$\boldsymbol{\tilde{\phi}}_{nT}$ is a consistent estimator of
$\boldsymbol{\phi}_{0}$, and derives its asymptotic distribution assuming the
observed factors are strong, again allowing for pricing errors, a missing
factor, and other forms of weak error cross-sectional dependence. Theorem
\ref{Tsemi} extends the results of Theorem \ref{Tfi} to the case where one or
more of the observed risk factors are semi-strong and shows how factor
strength impacts the precision with which the elements of $\boldsymbol{\phi
}_{0}$ are estimated. Finally, Theorem \ref{Var} presents the conditions under
which the asymptotic variance of $\boldsymbol{\tilde{\phi}}_{nT}$ can be
consistently estimated.
\begin{theorem}
\label{TFMbias}(Small $T$ bias of Fama-MacBeth estimator of
$\boldsymbol{\lambda}$) Consider the multi-factor linear return model
(\ref{mfT}) with the missing factor $g_{t}$ in $u_{it}$ as defined by
(\ref{uit}) and the associated risk premia, $\boldsymbol{\lambda}$, defined by
(\ref{APTmean}). Suppose that Assumptions \ref{factors}, \ref{loadings},
\ref{Errors}, \ref{Latent factor} and \ref{PriceError} hold and all observed
factors are strong. Suppose further that the true value of the risk premia,
$\boldsymbol{\lambda}_{0}$, is estimated by Fama-MacBeth two-pass estimator,
$\boldsymbol{\hat{\lambda}}_{nT}$, defined by (\ref{lambdahat}). Then for any
fixed $T>T_{0}$ such that $\lambda_{\min}\left( T^{-1}\mathbf{F}^{\prime
}\mathbf{M}_{T}\mathbf{F}\right) >0,$ we have (as $n\rightarrow\infty$)
\begin{equation}
\boldsymbol{\hat{\lambda}}_{nT}-\boldsymbol{\lambda}_{0}=\left(
\boldsymbol{\hat{\mu}}_{T}-\boldsymbol{\mu}_{0}\right) -\frac{\bar{\sigma
}^{2}}{T}\left[ \boldsymbol{\Sigma}_{\beta\beta}+\bar{\sigma}^{2}\frac{1}
{T}\left( \frac{\mathbf{F}^{\prime}\mathbf{M}_{T}\mathbf{F}}{T}\right)
^{-1}\right] ^{-1}\left( \frac{\mathbf{F}^{\prime}\mathbf{M}_{T}\mathbf{F}
}{T}\right) ^{-1}\boldsymbol{\lambda}_{T}^{\ast}+o_{p}(1), \label{bias}
\end{equation}
where $\boldsymbol{\hat{\mu}}_{T}=T^{-1}\sum_{t=1}^{T}\mathbf{f}_{t}$,
\[
\mathbf{\Sigma}_{\beta\beta}=\lim_{n\rightarrow\infty}\left( \frac
{\mathbf{B}_{n}^{\prime}\mathbf{M}_{n}\mathbf{B}_{n}}{n}\right) \text{, and
}\overline{\sigma}^{2}=\lim_{n\rightarrow\infty}n^{-1}\sum_{i=1}^{n}\sigma
_{i}^{2}>0.
\]
\end{theorem}
The proof is provided in Section \ref{PFMbias} of the mathematical appendix.
To derive the asymptotic distribution of $\boldsymbol{\hat{\lambda}}
_{nT}-\boldsymbol{\lambda}_{0}$ it is required that both $n$ and
$T\rightarrow\infty$, jointly. Also, noting that
\[
\boldsymbol{\hat{\lambda}}_{nT}-\boldsymbol{\lambda}_{0}=\left(
\boldsymbol{\hat{\mu}}_{T}-\boldsymbol{\mu}_{0}\right) +\left(
\boldsymbol{\hat{\phi}}_{nT}-\boldsymbol{\phi}_{0}\right) ,
\]
it is clear that increasing $n$ is not relevant for the distribution of
$\boldsymbol{\hat{\mu}}_{T}-\boldsymbol{\mu}_{0}$, but joint $n$ and $T$
asymptotics are required when investigating the distribution of
$\boldsymbol{\hat{\phi}}_{nT}-\boldsymbol{\phi}_{0}$. Focussing on the latter,
and using result (\ref{phidis1})\ in the Appendix, we have
\begin{align*}
& \left( n^{-1}\mathbf{\hat{B}}_{nT}^{\prime}\mathbf{M}_{n}\mathbf{\hat{B}
}_{nT}\right) \sqrt{nT}\left( \boldsymbol{\hat{\phi}}_{nT}-\boldsymbol{\phi
}_{0}\right) =n^{-1/2}T^{1/2}\mathbf{B}_{n}^{\prime}\mathbf{M}_{n}
\boldsymbol{\eta}_{n}+n^{-1/2}T^{1/2}\mathbf{G}_{T}^{\prime}\mathbf{U}
_{nT}^{\prime}\mathbf{M}_{n}\boldsymbol{\eta}_{n}\\
& +n^{-1/2}T^{1/2}\mathbf{G}_{T}^{\prime}\mathbf{U}_{nT}^{\prime}
\mathbf{M}_{n}\mathbf{\bar{u}}_{n\circ}-n^{-1/2}T^{1/2}\mathbf{G}_{T}^{\prime
}\mathbf{U}_{nT}^{\prime}\mathbf{U}_{nT}\mathbf{G}_{T}\boldsymbol{\lambda}
_{T}^{\ast}.
\end{align*}
Where $\boldsymbol{U}_{nT}=\left( \mathbf{u}_{1\circ},\mathbf{u}_{2\circ
},...,\mathbf{u}_{n\circ}\right) ^{\prime},$ $\mathbf{u}_{i\circ}=\left(
u_{i1},u_{i2},...,u_{iT}\right) ^{\prime},$ $\mathbf{G}_{T}=$ $\mathbf{M}
_{T}\mathbf{F}\left( \mathbf{F}^{\prime}\mathbf{M}_{T}\mathbf{F}\right)
^{-1}$, $\overline{\mathbf{u}}_{n\circ}=(\overline{u}_{1\circ},\overline
{u}_{2\circ},...,\overline{u}_{n\circ})^{\prime},$ and $\overline{u}_{i\circ
}=T^{-1}\sum_{t=1}^{T}u_{it}.$ Consider first the terms that include the
pricing errors, $\boldsymbol{\eta}_{n}$, and using the results in Lemma
\ref{L2} note that
\[
n^{-1/2}T^{1/2}\mathbf{B}_{n}^{\prime}\mathbf{M}_{n}\boldsymbol{\eta}
_{n}=O_{p}\left( T^{1/2}n^{-1/2+\alpha_{\eta}}\right) ,\text{ }
n^{-1/2}T^{1/2}\mathbf{G}_{T}^{\prime}\mathbf{U}_{nT}^{\prime}\mathbf{M}
_{n}\boldsymbol{\eta}_{n}=O_{p}\left( n^{-1/2+\frac{\alpha_{\eta}
+\alpha_{\gamma}}{2}}\right) .
\]
It is clear that the effects of pricing errors on the distribution of
$\boldsymbol{\hat{\phi}}_{nT}$ vanish only if $T^{1/2}n^{-1/2+\alpha_{\eta}
}\rightarrow0$, and $\alpha_{\eta}+\alpha_{\gamma}<1$. Also
\begin{align*}
n^{-1/2}T^{1/2}\mathbf{G}_{T}^{\prime}\mathbf{U}_{nT}^{\prime}\mathbf{M}
_{n}\mathbf{\bar{u}}_{n\circ} & =O_{p}\left( T^{-1/2}\right) \text{, }\\
n^{-1/2}T^{1/2}\mathbf{G}_{T}^{\prime}\mathbf{U}_{nT}^{\prime}\mathbf{U}
_{nT}\mathbf{G}_{T}\boldsymbol{\lambda}_{T}^{\ast} & =\sqrt{\frac{n}{T}}
\bar{\sigma}_{n}^{2}\left( \frac{\mathbf{F}^{\prime}\mathbf{M}_{T}\mathbf{F}
}{T}\right) ^{-1}\boldsymbol{\lambda}_{T}^{\ast}+O_{p}\left( T^{-1/2}
\right) .
\end{align*}
Finally, for the first two terms involving $\mathbf{B}_{n}$ and $\mathbf{U}
_{nT}$ we have
\begin{equation}
n^{-1/2}T^{1/2}\mathbf{B}_{n}^{\prime}\mathbf{M}_{n}\left( \mathbf{\bar{u}
}_{n\circ}-\mathbf{U}_{nT}\mathbf{G}_{T}\boldsymbol{\lambda}_{T}^{\ast
}\right) =O_{p}\left( 1\right) . \label{Dis}
\end{equation}
It is clear that the small $T$ bias of the asymptotic distribution of the
two-step estimator, given by $\sqrt{\frac{n}{T}}\bar{\sigma}_{n}^{2}\left(
\frac{\mathbf{F}^{\prime}\mathbf{M}_{T}\mathbf{F}}{T}\right) ^{-1}
\boldsymbol{\lambda}_{T}^{\ast}$, does not vanish unless, $n/T\rightarrow0$.
At the same time for the pricing errors to have no impact on the distribution
of the two-step estimator we must have $Tn^{a_{\eta}}/n\rightarrow0$. Both
conditions cannot be met simultaneously. It is possible to derive the
asymptotic distribution of $\boldsymbol{\hat{\phi}}_{nT}$, and hence that of
$\boldsymbol{\hat{\lambda}}_{nT}$, when $n/T\rightarrow0$ and $\mathbf{\eta
=0}$, but these are quite restrictive conditions, and to avoid them we follow
\citet{shanken1992estimation} and instead consider a bias-corrected version of
$\boldsymbol{\hat{\phi}}_{nT}$, namely $\boldsymbol{\tilde{\phi}}_{nT}$ given
by (\ref{phitilda}). As noted earlier $\boldsymbol{\tilde{\phi}}
_{nT}=\boldsymbol{\tilde{\lambda}}_{nT}-\boldsymbol{\hat{\mu}}_{T}$, where
$\boldsymbol{\tilde{\lambda}}_{nT}$ is the bias-corrected version of
$\boldsymbol{\hat{\lambda}}_{nT}$ originally proposed by Shanken.
To investigate the asymptotic properties of $\boldsymbol{\tilde{\phi}}_{nT}$
we first need to establish conditions under which $\widehat{\bar{\sigma}}
_{nT}^{2}$, defined by (\ref{zigbar}), is a consistent estimator of
$\bar{\sigma}_{n}^{2}=n^{-1}\sum_{i=1}^{n}\sigma_{i}^{2}$, which enters the
bias-corrected estimator. The proof of consistency in the literature does not
allow for missing factors or pricing errors and only considers the case where
$T$ is fixed as $n\rightarrow\infty$. For derivation of asymptotic
distribution of $\boldsymbol{\tilde{\phi}}_{nT}$ we also need to consider the
limiting properties of $\widehat{\bar{\sigma}}_{nT}^{2}$ under joint $n$ and
$T$ asymptotics. The following theorem provides the required results for
$\widehat{\bar{\sigma}}_{nT}^{2}$ as an estimator of $\bar{\sigma}_{n}^{2}$.
\begin{theorem}
\label{Thzig} Consider $\widehat{\bar{\sigma}}_{nT}^{2}$, the estimator of
$\bar{\sigma}_{n}^{2}$ given by (see (\ref{zigbar})),
\begin{equation}
\widehat{\bar{\sigma}}_{nT}^{2}=\frac{\sum_{t=1}^{T}\sum_{i=1}^{n}\hat{u}
_{it}^{2}}{n(T-K-1)}, \label{AdjZig}
\end{equation}
and suppose that Assumptions \ref{factors}, \ref{Latent factor}, and
\ref{Errors}, are satisfied. Then for a fixed $T$
\begin{equation}
\lim_{n\rightarrow\infty}E\left( \widehat{\bar{\sigma}}_{nT}^{2}\right)
=\bar{\sigma}^{2}, \label{unbiasZig}
\end{equation}
where $\bar{\sigma}^{2}=\lim_{n\rightarrow\infty}\bar{\sigma}_{n}^{2}$, and
$\bar{\sigma}_{n}^{2}=n^{-1}\sum_{i=1}^{n}\sigma_{i}^{2}$. Furthermore
\begin{equation}
\widehat{\bar{\sigma}}_{nT}^{2}-\bar{\sigma}_{n}^{2}=O_{p}\left(
T^{-1/2}n^{-1/2}\right) . \label{orderzig}
\end{equation}
\end{theorem}
For a proof see sub-section \ref{ProofThzig} in the Appendix.
Result (\ref{orderzig}) shows that $\widehat{\bar{\sigma}}_{nT}^{2}$ continues
to be a consistent estimator of $\bar{\sigma}^{2}=\lim_{n\rightarrow\infty
}\bar{\sigma}_{n}^{2}$ for a fixed $T$ as $n\rightarrow\infty$, even in the
presence of pricing errors and a missing common factor. This result also holds
when one or more of the factors are semi-strong.
Equipped with the above result we are now in a position to present the theorem
that sets out the asymptotic distribution of $\boldsymbol{\tilde{\phi}}_{nT}$.
\begin{theorem}
\label{Tfi}Consider, $\boldsymbol{\tilde{\phi}}_{nT}$, the bias-corrected
estimators of $\boldsymbol{\phi}_{0}$ given by (\ref{phitilda}). Suppose
Assumptions \ref{factors}, \ref{loadings}, \ref{Errors}, \ref{Latent factor}
and \ref{PriceError} hold, all the observed factors are strong, ($\alpha
_{k}=1$, for $k=1,2,...,K$), and the strength of the missing factor,
$\alpha_{\gamma}$ defined by (\ref{normg}), satisfies $\alpha_{\gamma}$
$<1/2$.
\begin{align}
\boldsymbol{\tilde{\phi}}_{nT}-\boldsymbol{\phi}_{0} & =O_{p}\left(
T^{-1/2}n^{^{-1/2}}\right) +O_{p}\left( T^{-1/2}n^{-1+\frac{\alpha_{\eta
}+\alpha_{\gamma}}{2}}\right) \label{phigap}\\
& +O_{p}\left( n^{-1+\alpha_{\eta}}\right) +O_{p}\left( T^{-1}
n^{-1/2}\right) ,\nonumber
\end{align}
where $\alpha_{\eta}\,$denotes the degree of pervasiveness of the pricing
errors defined by (\ref{APTg}). (a) When $T$ is fixed, $\alpha_{\gamma}<1/2$
and $\alpha_{\eta}<1$, then there exists $T_{0}$ such that for all $T>T_{0}$
\begin{equation}
p\lim_{n\rightarrow\infty}\left( \boldsymbol{\tilde{\phi}}_{nT}\right)
=\boldsymbol{\phi}_{0}\text{.} \label{phicon}
\end{equation}
Also
\begin{equation}
\sqrt{nT}\left( \boldsymbol{\tilde{\phi}}_{nT}-\boldsymbol{\phi}_{0}\right)
=\mathbf{\Sigma}_{\beta\beta}^{-1}\boldsymbol{\xi}_{nT}+O_{p}\left(
n^{-\frac{1}{2}+\frac{\alpha_{\eta+\alpha_{\gamma}}}{2}}\right) +O_{p}\left(
T^{1/2}n^{-1/2+\alpha_{\eta}}\right) +O_{p}\left( T^{-1/2}\right) ,
\label{phiDis}
\end{equation}
where $\mathbf{\Sigma}_{\beta\beta}=p\lim_{n\rightarrow\infty}\left(
n^{-1}\mathbf{B}_{n}^{\prime}\mathbf{M}_{n}\mathbf{B}_{n\ }\right) ,$
\begin{equation}
\boldsymbol{\xi}_{nT}=n^{-1/2}T^{-1/2}\mathbf{B}_{n}^{\prime}\mathbf{M}
_{n}\mathbf{U}_{nT}\mathbf{a}_{T}, \label{Tegzi}
\end{equation}
and $\mathbf{a}_{T}=\mathbf{\tau}_{T}-\mathbf{M}_{T}\mathbf{F}(T^{-1}
\mathbf{F}^{\prime}\mathbf{M}_{T}\mathbf{F)}^{-1}\boldsymbol{\lambda}
_{T}^{\ast}$. (b) If $\alpha_{\gamma}<1/2$, $\alpha_{\eta}<1/2,$ and
$\sqrt{\frac{T}{n}}n^{\alpha_{\eta}}\rightarrow0$, as $n$ and $T\rightarrow
\infty$ jointly, then
\begin{equation}
\sqrt{nT}\left( \boldsymbol{\tilde{\phi}}_{nT}-\boldsymbol{\phi}_{0}\right)
\rightarrow_{d}N\left( \mathbf{0,\Sigma}_{\beta\beta}^{-1}\mathbf{V}_{\xi
}\mathbf{\Sigma}_{\beta\beta}^{-1}\right) , \label{Dphi}
\end{equation}
where
\begin{equation}
\mathbf{V}_{\xi}=\left( 1+\boldsymbol{\lambda}_{0}^{\prime}\mathbf{\Sigma
}_{f}^{-1}\boldsymbol{\lambda}_{0}\right) \text{ }p\lim_{n\rightarrow\infty
}\left( n^{-1}\mathbf{B}_{n}^{\prime}\mathbf{M}_{n}\mathbf{V}_{u}
\mathbf{M}_{n}\mathbf{B}_{n\ }\right) . \label{Vegzi}
\end{equation}
\end{theorem}
For a proof see sub-section \ref{ProofTfi} in the Appendix.
Result (\ref{phigap}) establishes the finite $T$ consistency of
$\boldsymbol{\tilde{\phi}}_{nT}$ for $\boldsymbol{\phi}_{0}$ so long as
$\alpha_{\gamma}<1/2$ and $\alpha_{\eta}<1$, thus extending the Shanken result
to a much more general setting. To the best of our knowledge the asymptotic
distribution in (\ref{Dphi}) is new and shows that the asymptotic covariance
matrix of $\boldsymbol{\tilde{\phi}}_{nT}$ includes the term
$\boldsymbol{\lambda}_{T}^{\ast\prime}(T^{-1}\mathbf{F}^{\prime}\mathbf{M}
_{T}\mathbf{F)}^{-1}\boldsymbol{\lambda}_{T}^{\ast}$, that arises from the
first stage estimation of the factor loadings, and must be included in the
analysis for valid inference. It is also clear that this additional term does
not vanish with $T\rightarrow\infty$, and tends to $\boldsymbol{\lambda}
_{0}^{\prime}\mathbf{\Sigma}_{f}^{-1}\boldsymbol{\lambda}_{0}\geq\left(
\boldsymbol{\lambda}_{0}^{\prime}\boldsymbol{\lambda}_{0}\right)
\lambda_{\max}\left( \mathbf{\Sigma}_{f}^{-1}\right) =\left(
\boldsymbol{\lambda}_{0}^{\prime}\boldsymbol{\lambda}_{0}\right)
\lambda_{\min}\left( \mathbf{\Sigma}_{f}\right) >0$, which is strictly
non-zero unless $\boldsymbol{\lambda}_{0}=\mathbf{0}$. Shanken type bias
correction addresses the mean of the asymptotic distribution of
$\boldsymbol{\tilde{\phi}}_{nT}$, but not its covariance.
The $O_{p}\left( T^{-1/2}\right) $ term in (\ref{phiDis}) arises from the
sampling errors involved in the estimation of the factor loadings and
$\bar{\sigma}_{n}^{2}$, and tends to zero at the regular $\sqrt{T}$ rate. But
$n$ has to be sufficiently large to eliminate the effects of pricing errors on
identification of $\boldsymbol{\phi}_{0}$, as dictated by condition
$\sqrt{\frac{T}{n}}n^{\alpha_{\eta}}\rightarrow0,$ as $n$ and $T\rightarrow
\infty$.\footnote{The condition $\sqrt{\frac{T}{n}}n^{\alpha_{\eta}
}\rightarrow0$ can be weakened somewhat to $\sqrt{\frac{T}{n}}n^{\alpha_{\eta
}/2}\rightarrow0$ if we also assume that $\beta_{ik}-\bar{\beta}_{k}$ are
independently distributed over $i$, but will still require $n$ to be larger
than $T$.} The requirement that $T$ need not be too large relative to $n$ for
estimation of $\boldsymbol{\phi}_{0}$ is consistent with separating the
estimation of $\boldsymbol{\phi}_{0}$ from that of $\boldsymbol{\mu}_{0}$,
allowing the use a relatively small $T$ and a large $n$ to estimate
$\boldsymbol{\phi}_{0}$ and a relatively large $T$ when estimating
$\boldsymbol{\mu}_{0}.$
\subsection{What if one or more of the risk factors are semi-strong?}
We now turn to an intermediate case where one or more of the observed factors
are semi-strong, in the sense that their factor strength, $\alpha_{k}$ lies
between $1/2$ and $1$. The case of weak risk factors is already covered in the
proceeding analysis, and such factors can be included in the error term,
$u_{it}$, with little consequence for the estimation of risk premia of the
remaining factors that are strong or semi-strong. Weak factors do not have any
explanatory power and can be dropped from the analysis.
When one or more of the observed factors is semi-strong $\mathbf{\Sigma
}_{\beta\beta}$ is no longer positive definite and Theorem \ref{Tfi} does not
apply, but it is possible to adapt the proofs to establish the limiting
properties of $\tilde{\phi}_{k,nT}\boldsymbol{\ }$(the $k^{th}$ element of
$\boldsymbol{\tilde{\phi}}_{nT}$)\ for different values of \thinspace
$\alpha_{k}\,$.
To this end, analogously to $\boldsymbol{\tilde{\phi}}_{nT}$, we introduce the
following estimator of $\boldsymbol{\phi}_{0}$
\begin{equation}
\boldsymbol{\tilde{\phi}}_{nT}\left( \boldsymbol{\alpha}\right)
=\mathbf{H}_{nT}^{-1}\left( \boldsymbol{\alpha}\right) \left[
\mathbf{D}_{\alpha}^{-1}\mathbf{\hat{B}}_{nT}^{\prime}\mathbf{M}
_{n}\mathbf{\hat{\alpha}}_{nT}+\frac{n}{T}\widehat{\bar{\sigma}}_{nT}
^{2}\mathbf{D}_{\alpha}^{-1}\left( \frac{\mathbf{F}^{\prime}\mathbf{M}
_{T}\mathbf{F}}{T}\right) ^{-1}\boldsymbol{\hat{\mu}}\right] ,
\label{phialpha0}
\end{equation}
where
\begin{equation}
\mathbf{H}_{nT}\left( \boldsymbol{\alpha}\right) =\mathbf{D}_{\alpha}
^{-1}\mathbf{\hat{B}}_{nT}^{\prime}\mathbf{M}_{n}\mathbf{\hat{B}}
_{nT}\mathbf{D}_{\alpha}^{-1}-\frac{n}{T}\widehat{\bar{\sigma}}_{nT}
^{2}\mathbf{D}_{\alpha}^{-1}\left( \frac{\mathbf{F}^{\prime}\mathbf{M}
_{T}\mathbf{F}}{T}\right) ^{-1}\mathbf{D}_{\alpha}^{-1}. \label{Halpha}
\end{equation}
It is now easily seen that
\begin{equation}
\mathbf{D}_{\alpha}\left( \boldsymbol{\tilde{\phi}}_{nT}\left(
\boldsymbol{\alpha}\right) -\boldsymbol{\phi}_{0}\right) =\mathbf{H}
_{nT}^{-1}\left( \boldsymbol{\alpha}\right) \mathbf{q}_{nT}\left(
\boldsymbol{\alpha}\right) , \label{phialpha}
\end{equation}
where
\begin{equation}
\mathbf{q}_{nT}\left( \boldsymbol{\alpha}\right) =\mathbf{D}_{\alpha}
^{-1}\mathbf{\hat{B}}_{nT}^{\prime}\mathbf{M}_{n}\mathbf{\hat{\alpha}}
_{nT}+\frac{n}{T}\widehat{\bar{\sigma}}_{nT}^{2}\mathbf{D}_{\alpha}
^{-1}\left( \frac{\mathbf{F}^{\prime}\mathbf{M}_{T}\mathbf{F}}{T}\right)
^{-1}\boldsymbol{\hat{\mu}}_{T}-\mathbf{H}_{nT}(\boldsymbol{\alpha}
)\mathbf{D}_{\alpha}\boldsymbol{\phi}_{0}. \label{qnT}
\end{equation}
$\mathbf{D}_{\alpha}$ is defined by (\ref{Dn}), and $\boldsymbol{\alpha
}=\mathbf{(}\alpha_{1},\alpha_{2},...,\alpha_{K})^{\prime}$. It is easily
established that numerically $\boldsymbol{\tilde{\phi}}_{nT}\left(
\boldsymbol{\alpha}\right) $ is identical to $\boldsymbol{\tilde{\phi}}_{nT}
$, and its introduction is primarily for the purpose of establishing the
limiting properties of $\tilde{\phi}_{k,nT}-\phi_{0,k}$ that do depend on
$\alpha_{k}$. Note that
\[
\mathbf{H}_{nT}\left( \boldsymbol{\alpha}\right) =n\mathbf{D}_{\alpha}
^{-1}\mathbf{H}_{nT}\mathbf{D}_{\alpha}^{-1},\text{ and }\mathbf{q}
_{nT}\left( \boldsymbol{\alpha}\right) =n\mathbf{D}_{\alpha}^{-1}
\mathbf{s}_{nT}
\]
where $\mathbf{s}_{nT}$ and $\mathbf{H}_{nT}$ are already defined by
(\ref{snT0}) and (\ref{A-HnT}). Using these in (\ref{phialpha}) we have
\[
\mathbf{D}_{\alpha}\left( \boldsymbol{\tilde{\phi}}_{nT}\left(
\boldsymbol{\alpha}\right) -\boldsymbol{\phi}_{0}\right) =\left(
n\mathbf{D}_{\alpha}^{-1}\mathbf{H}_{nT}\mathbf{D}_{\alpha}^{-1}\right)
^{-1}n\mathbf{D}_{\alpha}^{-1}\mathbf{s}_{nT}=\mathbf{D}_{\alpha}
\mathbf{H}_{nT}^{-1}\mathbf{s}_{nT},
\]
and it follows that $\boldsymbol{\tilde{\phi}}_{nT}\left( \boldsymbol{\alpha
}\right) -\boldsymbol{\phi}_{0}=\mathbf{H}_{nT}^{-1}\mathbf{s}_{nT}
=\boldsymbol{\tilde{\phi}}_{nT}\left( \boldsymbol{\tau}_{K}\right)
=\boldsymbol{\tilde{\phi}}_{nT}$. See (\ref{A-phitilda}).
The convergence results for $\boldsymbol{\tilde{\phi}}_{nT}\left(
\boldsymbol{\alpha}\right) $ are set out in the following theorem.
\begin{theorem}
\label{Tsemi}Consider, $\boldsymbol{\tilde{\phi}}_{nT}\left(
\boldsymbol{\alpha}\right) $, the bias-corrected estimators of
$\boldsymbol{\phi}_{0}$ given by (\ref{phialpha0}), and suppose Assumptions
\ref{factors}, \ref{loadings}, \ref{Errors}, \ref{Latent factor} and
\ref{PriceError} hold, the strength of observed factors, $\boldsymbol{f}
_{t}=(f_{1t},f_{2t},...,f_{Kt})^{\prime},$ is given by $\boldsymbol{\alpha
}\mathbf{=(}\alpha_{1},\alpha_{2},...,\alpha_{K})^{\prime}$, and the strength
of the missing factor, $g_{t}$, defined by (\ref{normg}) is $\alpha_{\gamma}$.
Let $\alpha_{\min}=\min_{k}(\alpha_{k})$ and suppose that $\alpha_{\gamma
}<1/2$. Then
\[
\mathbf{H}_{nT}\left( \boldsymbol{\alpha}\right) =\mathbf{D}_{\alpha}
^{-1}\mathbf{B}_{n}^{\prime}\mathbf{M}_{n}\mathbf{B}_{n}\mathbf{D}_{\alpha
}^{-1}+O_{p}\left( T^{-1}n^{-\alpha_{\min}+1/2}\right) ,
\]
where $\mathbf{H}_{nT}\left( \boldsymbol{\alpha}\right) $ is given by
(\ref{Halpha}), and by part (b) of Assumption \ref{loadings}, $\mathbf{H}
_{nT}\left( \boldsymbol{\alpha}\right) \rightarrow_{p}\mathbf{\Sigma}
_{\beta\beta}(\boldsymbol{\alpha}\mathbf{)}>0$, for any fixed $T>T_{0}$ such
that $\lambda_{\max}\left( \frac{\mathbf{F}^{\prime}\mathbf{M}_{T}\mathbf{F}
}{T}\right) ^{-1}<C$ and $\alpha_{\min}>1/2>\alpha_{\gamma}$. Also
\begin{align}
\tilde{\phi}_{k,nT}\left( \boldsymbol{\alpha}\right) -\phi_{0,k} &
=O_{p}\left( n^{-(\alpha_{k}+\alpha_{\min})/2+1/2}T^{-1/2}\right)
+O_{p}\left( n^{\frac{-\left( \alpha_{k}+\alpha_{\min}\right) +\left(
\alpha_{\eta+\alpha_{\gamma}}\right) }{2}}T^{-1/2}\right) \label{Psemi}\\
& +O_{p}\left( n^{-\left( \alpha_{k}+\alpha_{\min}\right) /2+\alpha_{\eta
}}\right) +O_{p}\left( n^{-(\alpha_{k}+\alpha_{\min})/2+1/2}T^{-1}\right)
.\nonumber
\end{align}
\end{theorem}
See sub-section \ref{ProofThsemi} of the Appendix for a proof.
The result in (\ref{Psemi}) establishes the consistency of $\tilde{\phi
}_{k,nT}\left( \boldsymbol{\alpha}\right) =\tilde{\phi}_{k,nT}$ even if
$f_{kt}$ is semi-strong so long as $n\rightarrow\infty$, and $\alpha_{\min
}>1/2,$ $\alpha_{k}+\alpha_{\min}>\alpha_{\eta}+\alpha_{\gamma}$, and
$\alpha_{k}+\alpha_{\min}>2\alpha_{\eta}$. Clearly, these results reduce to
the case of strong factors where $\alpha_{\min}=\alpha_{k}=1$. Turning to the
asymptotic distribution of $\tilde{\phi}_{k,nT}\left( \boldsymbol{\alpha
}\right) $, again only convergence rates are affected, and instead of the
regular rate of $\sqrt{nT}$, we have $\ \sqrt{T}n^{(\alpha_{k}+\alpha_{\min
}-1)/2}$, and using (\ref{Psemi}) we have
\begin{align}
& \sqrt{T}n^{(\alpha_{k}+\alpha_{\min}-1)/2}\left( \tilde{\phi}
_{k,nT}\left( \boldsymbol{\alpha}\right) -\phi_{0,k}\right) \label{phikrate}
\\
& =O_{p}(1)+O_{p}\left( n^{-1/2+\frac{\left( \alpha_{\eta}+\alpha_{\gamma
}\right) }{2}}\right) +O_{p}\left( \sqrt{T}n^{-1/2+\alpha_{\eta}}\right)
+O_{p}\left( T^{-1/2}\right) .\nonumber
\end{align}
The conditions needed for eliminating the effects of the pricing errors are
the same as before and are given by $\alpha_{\eta}+\alpha_{\gamma}<1$ and
$\sqrt{T}n^{-1/2+\alpha_{\eta}}\rightarrow0$. The asymptotic distribution is
unaffected except for the slower rate of convergence alluded to above. It is
also of interest to note that adding semi-strong factors can adversely affect
the convergence rate of the strong factor with $\alpha_{k}=1$. As an example
suppose the asset pricing model contains two factors, one strong, $\alpha
_{1}=1$ and one semi strong with $\alpha_{2}<1$.$\,\ $Then the convergence
rate of $\tilde{\phi}_{1,nT}\left( \boldsymbol{\alpha}\right) -\phi_{0,1}$
is given by$\sqrt{T}$ $n^{(1+\alpha_{2}-1)/2}$ which is slower than the rate
we would have obtained for $\tilde{\phi}_{1,nT}\left( \boldsymbol{\alpha
}\right) -\phi_{0,1}$ if both factors were strong ($\alpha_{min}=\alpha
_{k}=1)$, namely the regular rate of $\sqrt{nT}$.
Furthermore, when conditions $\alpha_{\eta}+\alpha_{\gamma}<1$ and $\sqrt
{T}n^{-1/2+\alpha_{\eta}}\rightarrow0$ are met we have
\begin{equation}
\tilde{\phi}_{k,nT}\left( \boldsymbol{\alpha}\right) -\phi_{0,k}
=O_{p}\left( T^{-1/2}n^{-(\alpha_{k}+\alpha_{\min}-1)/2}\right) ,
\label{phiorder}
\end{equation}
and $\phi_{0,k}$ is consistently estimated if $T^{-1/2}n^{-(\alpha_{k}
+\alpha_{\min}-1)/2}\rightarrow0$. Also, using (\ref{phitilda2}), an estimator
of the risk premia, $\lambda_{k}$, is given by $\tilde{\lambda}_{k,nT}\left(
\boldsymbol{\alpha}\right) =\tilde{\phi}_{k,nT}\left( \boldsymbol{\alpha
}\right) +\hat{\mu}_{k,T}$, where $\hat{\mu}_{k,T}=T^{-1}\sum_{t=1}^{T}
f_{kt}$. Hence
\[
\tilde{\lambda}_{k,nT}\left( \boldsymbol{\alpha}\right) -\lambda
_{0,k}=\left[ \tilde{\phi}_{k,nT}\left( \boldsymbol{\alpha}\right)
-\phi_{0,k}\right] +\left( \hat{\mu}_{k,T}-\mu_{0,k}\right) .
\]
Under Assumption \ref{factors} $\hat{\mu}_{k,T}-\mu_{0,k}=O_{p}\left(
T^{-1/2}\right) $, and using (\ref{phiorder}) it then follows that
\begin{equation}
\tilde{\lambda}_{k,nT}\left( \boldsymbol{\alpha}\right) -\lambda_{0,k}
=O_{p}\left( T^{-1/2}n^{-(\alpha_{k}+\alpha_{\min}-1)/2}\right)
+O_{p}\left( T^{-1/2}\right) , \label{ratelambda}
\end{equation}
and $\tilde{\lambda}_{k,nT}\left( \boldsymbol{\alpha}\right) $ is a
consistent estimator of $\lambda_{0,k}$ if $T\rightarrow\infty$ as well as
$T^{-1/2}n^{-(\alpha_{k}+\alpha_{\min}-1)/2}\rightarrow0$. More specifically,
suppose $T=\ominus\left( n^{d}\right) $ for some $d>0$, where $\ominus
\left( \cdot\right) $ denotes $T$ and $n^{d}$ are of the same order of
magnitude. Then for any $d>0$, the condition for consistency of $\tilde
{\lambda}_{k,nT}\left( \boldsymbol{\alpha}\right) $ is given by $\alpha
_{k}+\alpha_{min}+d>1$ and $d>0$. In the case where all risk factors have the
same strength, $\alpha$, the consistency condition reduces to $\alpha>\left(
1-d\right) /2$, which is weaker than the one derived by \cite{giglio2023test}
, namely $n/(\left\Vert \beta\right\Vert ^{2}T)\rightarrow0$, where, in terms
of our notation, $\left\Vert \beta\right\Vert ^{2}=\ominus\left( n^{\alpha
}\right) $. This latter condition will be met if $\alpha>1-d$. In practice
where $T$ is small relative to $n$, the accuracy of $\tilde{\lambda}
_{k,nT}\left( \boldsymbol{\alpha}\right) $ as an estimator $\lambda_{0,k}$
does depend on $\alpha$, and our weaker condition on $\alpha>\left(
1-d\right) /2$ is advantageous.
\subsection{Consistent estimation of the variance of $\boldsymbol{\tilde{\phi
}}_{nT}$\label{Vphihat}}
To carry out inference on $\boldsymbol{\phi}_{0}$, or any of its elements
individually, we require a consistent estimator of $Var\left(
\boldsymbol{\tilde{\phi}}_{nT}\right) $. Using (\ref{Dphi}) and (\ref{Vegzi})
we first note that $\mathbf{\Sigma}_{\beta\beta}$ is consistently estimated by
$\mathbf{H}_{nT\ }$ given by (\ref{HnT1}). Therefore, it is sufficient to find
a suitable estimator of $\mathbf{V}_{u}=(\sigma_{ij})$ such that
$\mathbf{V}_{\xi}$ given by (\ref{Vegzi}) is consistently estimated. Under
suitable sparsity restrictions $\mathbf{V}_{u}$ can be consistently estimated
using the various thresholding procedures advanced in the statistical
literature by
\citet{bickel2008covariance, bickel2008regularized}
,
\citet{cai2011adaptive}
, and
\citet[BPS]{BPS2019multiple}
.
\citet{fan2011high, fan2013large}
also show that the adaptive threshold technique of Cai and Liu applies equally
to the residuals from an approximate factor model. Here we consider the
threshold estimator proposed by BPS which does not require cross-validation
and is shown to have desirable small sample properties. It is given by
$\mathbf{\tilde{V}}_{u}=\left( \tilde{\sigma}_{ij}\right) $
\begin{align}
\tilde{\sigma}_{ii} & =\hat{\sigma}_{ii}\nonumber\\
\tilde{\sigma}_{ij} & =\hat{\sigma}_{ij}\mathbf{1}\left[ \left\vert
\hat{\rho}_{ij}\right\vert >T^{-1/2}c_{\alpha}(n,\delta)\right] ,\text{
}i=1,2,\ldots,n-1,\text{ }j=i+1,\ldots,n, \label{threshold}
\end{align}
where
\begin{equation}
\hat{\sigma}_{ij}=\frac{1}{T}\sum_{t=1}^{T}\hat{u}_{it}\hat{u}_{jt},\text{
}\hat{\rho}_{ij}=\frac{\hat{\sigma}_{ij}}{\sqrt{\hat{\sigma}_{ii}\hat{\sigma
}_{jj}}},\text{ \ }\hat{u}_{it}=r_{it}-\hat{\alpha}_{i,T}-\boldsymbol{\hat
{\beta}}_{i,T}^{\prime}\mathbf{f}_{t}, \label{sij}
\end{equation}
and $c_{p}(n,d)=\Phi^{-1}\left( 1-\frac{p}{2n^{d}}\right) ,$ is a normal
critical value function, $p$ is the the nominal size of testing of
$\sigma_{ij}=0$, ($i\neq j$) and $d$ is chosen to take account of the
$n(n-1)/2$ multiple tests being carried out. Monte Carlo experiments carried
out by BPS suggest setting $d=2$. The variance estimator given by
(\ref{threshold}) does not require a knowledge of the factor strength and
applies to risk factors of differing degrees.
Under Assumptions \ref{factors}, \ref{Latent factor}, and \ref{Errors},
$\left\Vert \mathbf{V}_{u}\right\Vert =O\left( n^{\alpha_{\gamma}}\right) $,
and using results in
\citet{fan2011high, fan2013large}
we have
\begin{equation}
\left\Vert \mathbf{\tilde{V}}_{u}-\mathbf{V}_{u}\right\Vert =O_{p}\left(
n^{\alpha_{\gamma}}\sqrt{\frac{\ln(n)}{T}}\right) . \label{NormVu}
\end{equation}
Consider the following estimator of $\mathbf{V}_{\xi}$
\[
\mathbf{\hat{V}}_{\xi,nT}=\left( 1+\hat{s}_{nT}\right) \left(
n^{-1}\mathbf{\hat{B}}_{nT}^{\prime}\mathbf{M}_{n}\mathbf{\tilde{V}}
_{u}\mathbf{M}_{n}\mathbf{\hat{B}}_{nT}\right) \text{.}
\]
where $\hat{s}_{nT}=\boldsymbol{\tilde{\lambda}}_{nT}^{^{\prime}}\left(
T^{-1}\mathbf{F}^{\prime}\mathbf{M}_{T}\mathbf{F}\right) ^{-1}
\boldsymbol{\tilde{\lambda}}_{nT}$. Under Assumption \ref{factors},
$T^{-1}\mathbf{F}^{\prime}\mathbf{M}_{T}\mathbf{F\rightarrow}_{p}
\mathbf{\Sigma}_{f}$ $\ $and using the results above we have
$\boldsymbol{\tilde{\lambda}}_{nT}=\boldsymbol{\tilde{\phi}}_{nT}
+\boldsymbol{\hat{\mu}}_{T}\rightarrow_{p}\boldsymbol{\phi}_{0}
+\boldsymbol{\mu}_{0}=\boldsymbol{\lambda}_{0}$. Hence, $\hat{s}
_{nT}\rightarrow_{p}\boldsymbol{\lambda}_{0}^{\prime}\mathbf{\Sigma}_{f}
^{-1}\boldsymbol{\lambda}_{0}$ as $n,T\rightarrow\infty$, jointly, and it is
sufficient to show that
\begin{equation}
n^{-1}\mathbf{\hat{B}}_{nT}^{\prime}\mathbf{M}_{n}\mathbf{\tilde{V}}
_{u}\mathbf{M}_{n}\mathbf{\hat{B}}_{nT}-n^{-1}\mathbf{B}_{n}^{\prime
}\mathbf{M}_{n}\mathbf{V}_{u}\mathbf{M}_{n}\mathbf{B}_{n\ }\rightarrow
_{p}\mathbf{0}\text{.} \label{VarCon}
\end{equation}
The following theorem provides a formal statement of the conditions under
which $\mathbf{\hat{V}}_{\xi,nT}$ is a consistent estimator of $\mathbf{V}
_{\xi}$.
\begin{theorem}
\label{Var}Suppose Assumptions \ref{factors}, \ref{loadings}, \ref{Errors},
\ref{Latent factor} and \ref{PriceError} hold, and all the observed factors
are strong, ($\alpha_{k}=1$, for $k=1,2,...,K$), and the strength of the
missing factor, $g_{t}$, defined by (\ref{normg}), $\alpha_{\gamma}$ $<1/2$.
Then
\begin{equation}
\left\Vert \mathbf{\hat{V}}_{\xi,nT}-\mathbf{V}_{\xi}\right\Vert =O_{p}\left(
n^{\alpha_{\gamma}}\sqrt{\frac{\ln(n)}{T}}\right) , \label{NormVaregzi}
\end{equation}
where
\begin{equation}
\mathbf{\hat{V}}_{\xi,nT}=\left( 1+\hat{s}_{nT}\right) \left(
n^{-1}\mathbf{\hat{B}}_{nT}^{\prime}\mathbf{M}_{n}\mathbf{\tilde{V}}
_{u}\mathbf{M}_{n}\mathbf{\hat{B}}_{nT}\right) , \label{Varegzihat}
\end{equation}
$\mathbf{\tilde{V}}_{u}=\left( \tilde{\sigma}_{ij}\right) $, $\tilde{\sigma
}_{ij}$ is the threshold estimator of $\sigma_{ij}$ given by (\ref{threshold}
), and
\begin{equation}
\mathbf{V}_{\xi}=\left( 1+\boldsymbol{\lambda}_{0}^{\prime}\mathbf{\Sigma
}_{f}^{-1}\boldsymbol{\lambda}_{0}\right) \text{ }p\lim_{n\rightarrow\infty
}\left( n^{-1}\mathbf{B}_{n}^{\prime}\mathbf{M}_{n}\mathbf{V}_{u}
\mathbf{M}_{n}\mathbf{B}_{n\ }\right) , \label{Varegzi}
\end{equation}
\end{theorem}
For a proof see sub-section \ref{ProofTvar} in the Appendix.
This theorem shows that consistent estimation of $Var\left(
\boldsymbol{\tilde{\phi}}_{nT}\right) $ can be achieved by using a suitable
threshold estimator of $\mathbf{V}_{u}$, so long as the strength of the
missing factor, $\alpha_{\gamma}$, is sufficiently weak in the sense that
$n^{\alpha_{\gamma}}\sqrt{\ln(n)/T}\rightarrow0$ as $n,T\rightarrow\infty$.
\section{Small sample properties of the estimators and tests for $\boldsymbol{\phi}$\label{Simulations}}
\subsection{Monte Carlo Design}
This section presents Monte Carlo simulations to investigate the small sample
properties of estimators and tests for $\boldsymbol{\phi}_{0}$. In the
empirical application of the next section the factors are selected from a
large list. But here we assume $K=3$ and mimic the 3 Fama-French factors,
namely the market return minus the risk free rate, MKT, the value factor (high
minus low book to market portfolios, HML) and the size factor (small minus big
portfolios, SMB). These are denoted by $f_{kt},$ $k=M,H,S$.\footnote{Data on
factors and the risk free rate are downloaded from Kenneth French's data
library:
https://mba.tuck.dartmouth.edu/pages/faculty/ken.french/data\_library.html}
For further details see Section \ref{EMC} of the online supplement A.
\subsubsection{Loadings and factor strengths}
To calibrate the loadings, $\beta_{ik}$, we used excess returns on a large
number securities observed over the shorter sample covering the 20 years
$2002m1$ $-2021m12$ ($T=240$). Monthly returns for NYSE and NASDAQ stocks code
10 and 11 from CRSP were downloaded from Wharton Research Data Services and
converted to excess returns over the risk free rate, taken from Kenneth
French's webpages, in percent per month. Only stocks with available data for
the full sample were included, yielding a balanced panel, and to avoid
outliers influencing the results, stocks with a kurtosis greater than 16 were
excluded. There were $1289$ stocks before exclusion on the basis of kurtosis
and $1175$ after. The summary statistics giving mean, median, standard
deviation of the estimates of $\beta_{ik}$ and their histograms are provided
in the online supplement A.
For factor strength, we considered a range of DGPs. Given the evidence that
most factors, other than the market factor, are not strong, we focus on the
case where there is one strong factor, namely the market factor with
$\alpha_{M}=1,$ plus two semi-strong factors, with the value factor, $HML$,
being quite strong with $\alpha_{H}=0.85,$ and the size factor, $SML$, being
only moderately strong with $\alpha_{S}=0.65$. These estimates are also
informed by the results provided in
\citet{bailey2021measurement}
who propose methods for estimation of factor strength. For a given factor
strength, $\alpha_{k}$, the associated loadings, $\beta_{ik}$, are generated
as $\boldsymbol{\beta}_{k}=(\beta_{1k},\beta_{2k},,...,\beta_{\alpha_{k}
},0,0,...0)$ where $n_{\alpha_{k}}=\lfloor n^{\alpha_{k}}\rfloor$ the integer
part of $n^{\alpha_{k}}$, with non-zero and zero values of $\boldsymbol{\beta
}_{k}$ given by
\begin{align*}
\beta_{ik} & \sim IIDN(\mu_{\beta_{k}},\sigma_{\beta_{k}}^{2})\text{, for
}i=1,2,....,\lfloor n^{\alpha_{k}}\rfloor,\\
\beta_{ik} & =0\text{ for }i=\lfloor n^{\alpha_{k}}\rfloor+1,\lfloor
n^{\alpha_{k}}\rfloor+2,...,n,
\end{align*}
where $\lfloor n^{\alpha_{k}}\rfloor$ denotes the integer part of
$n^{\alpha_{k}}$. \ Since the security returns are randomly generated, it does
not matter how zero and non-zero values of $\beta_{ik}$ are distributed across
$i$. Also, the zero loadings can also be replaced by an exponentially decaying
sequence without any implications for the simulation results.\footnote{See
also footnote 5 of
\citet
*[p.942]{bailey2016exponent}.} We also set
\begin{align*}
\mu_{\beta_{M}} & =1\text{, }\sigma_{\beta_{M}}=0.4;\text{ }\mu_{\beta_{H}
}=0.2\text{, }\sigma_{\beta_{H}}=0.5\\
\mu_{\beta_{S}} & =0.6\text{, }\sigma_{\beta_{S}}=0.5,
\end{align*}
which match the mean and standard deviation of the estimates of $\beta_{ik}$.
See above.
\subsubsection{Generation of pricing errors}
The pricing errors in (\ref{APTRoss}) can be considered as firm-specific
characteristics and are set as $\boldsymbol{\eta}_{n}=\left( \eta_{1}
,\eta_{2},...,\eta_{n_{\eta}},0,0,...,0\right) ^{\prime}.$ The non-zero
loadings of $\boldsymbol{\eta}_{n}$ for $i\leq n_{\eta}=\lfloor n^{\alpha
_{\eta}}\rfloor$\ are drawn from $IIDU(0.7,0.9)$, and $\eta_{i}=0$ for
$i=n_{\eta}+1,n_{\eta}+2,....,n.$ We consider $\alpha_{\eta}=(0,0.3).$\textbf{
}When\textbf{ }$\alpha_{\eta}=0$ we have $\eta_{i}=0$ ,\textbf{\ }for all $i.$
As in the case of factor loadings the non-zero values of $\boldsymbol{\eta
}_{n}$ must be randomly allocated to different groups.
\subsubsection{Generation of return equation errors}
The return equation errors, $u_{it}$, are generated following (\ref{uit}) as a
combination of a missing factor, $g_{t}\sim IIDN\left( 0,1\right) $ plus an
idiosyncratic error, $v_{jt}.$ The loadings $\boldsymbol{\gamma}=\left(
\gamma_{1},\gamma_{2},...,\gamma_{n_{\gamma}},0,0,...,0\right) ^{\prime}$ of
the missing factor are set as
\begin{align*}
\gamma_{i} & \sim IIDU(0.7,0.9)\text{, for }i=1,2,....,\lfloor
n^{\alpha_{\gamma}}\rfloor,\\
\gamma_{i} & =0\text{, for }i=\lfloor n^{\alpha_{\gamma}}\rfloor+1,\lfloor
n^{\alpha_{\gamma}}\rfloor+2,...,n,
\end{align*}
where $\alpha_{\gamma}$ is the strength of the missing factor $g_{t}$. We
consider $\alpha_{\gamma}=1/4$ and $1/2$.
For the idiosyncratic errors, $v_{it}$, we consider spatial as well as a block
diagonal specification, with the spatial specification including a diagonal
specification as the special case. Under the spatial specification the
idiosyncratic errors are generated as the first order spatial autoregressive
model $v_{it}=\rho_{\varepsilon}\sum_{j=1}^{n}w_{ij}v_{jt}+\kappa
\varepsilon_{it},$ which can be written in matrix notation as $\mathbf{v}
_{t}=\rho_{\varepsilon}\mathbf{Wv}_{t}+\kappa\boldsymbol{\varepsilon}_{t},$
and solved for as $\mathbf{v}_{t}=\kappa\left( \mathbf{I}_{n}-\rho
_{\varepsilon}\mathbf{W}\right) ^{-1}\boldsymbol{\varepsilon}_{t}$. Adding
the missing factor now yields
\begin{equation}
\mathbf{u}_{t}=\boldsymbol{\gamma}\text{ }g_{t}+\kappa\left( \mathbf{I}
_{n}-\rho_{\varepsilon}\mathbf{W}\right) ^{-1}\boldsymbol{\varepsilon}_{t}.
\label{vSAR}
\end{equation}
The spatial coefficient $\rho_{\varepsilon}$ is such that $|\rho_{\varepsilon
}|<1$, $\mathbf{W=(}w_{ij})$ with $w_{ii}=0$, and $\sum_{j=1}^{n}w_{ij}=1$.
The diagonal case is obtained by setting $\rho_{\varepsilon}=0$, with
$\rho_{\varepsilon}=0.5$ characterizing the SAR specification. The weight
matrix $\mathbf{W}=(w_{ij})$ is set to follow the familiar rook pattern where
all its elements are set to zero except for $w_{i+1,i}=w_{j-1,j}=0.5$ for
$i=1,2,...,n-2$ and $j=3,4...,n$, with $w_{1,2}=w_{n,n-1}=1$.
Under the block error covariance specification, $\mathbf{v}_{t}$ is generated
as $\mathbf{v}_{t}=\kappa\mathbf{\hat{S}}\boldsymbol{\varepsilon}_{t},$ where
$\mathbf{\hat{S}}$ is a block diagonal matrix with its $b^{th}$ block given by
$\mathbf{\hat{S}}_{b}$ for $b=1,2,...,B$, and $\boldsymbol{\varepsilon}
_{t}=(\mathbf{\varepsilon}_{1t}^{\prime},\mathbf{\varepsilon}_{2t}^{\prime
},...,\mathbf{\varepsilon}_{Bt}^{\prime})^{\prime}$, and
$\boldsymbol{\varepsilon}_{bt}=(\varepsilon_{b,1t},\varepsilon_{b,2t}
,...,\varepsilon_{b,n_{b},t})^{\prime}$. $\mathbf{\hat{S}}$ is set as a
Cholesky factor of the correlation matrix of $\mathbf{u}_{t}$. Denoting this
correlation matrix by $\mathbf{\hat{R}}_{u},$
\[
\mathbf{\hat{R}}_{u}=\left[ Diag(\mathbf{\hat{V}}_{Bu})\right]
^{-1/2}\mathbf{\hat{V}}_{Bu}\left[ Diag(\mathbf{\hat{V}}_{Bu})\right]
^{-1/2}=Diag(\mathbf{\hat{R}}_{bu},\text{ }b=1,2,...,B),
\]
where $\mathbf{\hat{V}}_{Bu}$ is the threshold estimator of $\mathbf{V}_{u}$
subject to the additional restriction that $\mathbf{V}_{u}$ is block diagonal.
For each block $\mathbf{\hat{R}}_{bu}$ we set the number of distinct non-zero
elements of this block equal to the integer part of $[n_{b}(n_{b}-1)/2]\times
q_{b}$ where $q_{b}$ is the proportion of non-zero distinct elements in block
$b$ of our calibrated sample and computed by the calibration over the sample
$2001m10-2021m9$. The non-zero elements are drawn randomly from $IIDU(0,0.5)$.
Similarly, adding the missing factor, we have
\begin{equation}
\mathbf{u}_{t}=\boldsymbol{\gamma}g_{t}+\kappa\mathbf{\hat{S}}
\boldsymbol{\varepsilon}_{t}. \label{vblock}
\end{equation}
The block diagonal structure is intended to capture possible within industry
correlations not picked up by observed or weak missing factors, with each
block representing an industry or sector. To calibrate the block structure
estimates of the pair-wise correlations between the residuals of the return
regressions using the Fama-French three factors of the $T=240$ sample ending
in $2021$ were obtained. Then all the statistically insignificant correlations
were set to zero, allowing for the multiple testing nature of the tests. For
the majority of securities (668 out of the 1168), the pair-wise return
correlations were not statistically significant. The securities with a
relatively large number of non-zero correlations were either in the banking or
energy related industries. Considering stocks by 2-digit SIC classifications,
a division into $B=14$ contiguous groups ranging in size from $33$ to $145$
stocks, seemed sensible. More detail on the process is given in Section
\ref{BlockCov} of the online supplement A.
The primitive errors, $\varepsilon_{it}$ for $i=1,2,...,n$ in (\ref{vSAR}) and
(\ref{vblock})\ are generated as $\varepsilon_{it}=\sqrt{\sigma_{ii}}
\varpi_{it}$, where $\varpi_{it}\sim IIDN\left( 0,1\right) $, and
$\varepsilon_{it}=\sqrt{\sigma_{ii}}\left[ \sqrt{\frac{v-2}{v}}\varpi
_{it}\right] ,$ where $\varpi_{it}\sim IID$ $t(v)$, with $t(v)$ denotes a
standard $t$ distributed variate with $v=5$ degrees of freedom. Also
$\sigma_{ii}\thicksim IID$ \ $0.5(1+\chi_{1}^{2})$ $,\ for$\ $i=1,2,...,n_{b}$
and $b=1,2,...,B$. In this way, it is ensured that $Var(\varepsilon
_{b,it})=\sigma_{b,ii}$, and on average $E\left[ Var(\varepsilon
_{b,it})\right] =E(\sigma_{ii})=1$, under both Gaussian and t-distributed
errors. Note that $Var\left( \nu_{b,it}\right) =v/(v-2)$. All the
experiments are designed to give an $R^{2}$ of about $0.3,$ similar to that
obtained in the empirical applications. For further details see sub-section
\ref{Regfit} of the online supplement A.
\subsubsection{Experiments}
In total, we consider 12 experimental designs: six designs with Gaussian
errors and six with $t(5)$ distributed errors. We considered designs with
GARCH effects, with and without pricing errors, $\eta_{i}$, and with and
without the missing factor, $g_{t}$. We also considered designs with spatial
patterns in the idiosyncratic errors, $v_{it}$. All experiments are
implemented using $R=2,000$ replications. Details of of the 12 experiments are
summarized in Table S-1 of the online supplement B (MC results).
\subsubsection{Alternative estimators of $\mathbf{V}_{u}$}
Subsection \ref{Vphihat} considered consistent estimation of the variance of
$\tilde{\phi}_{nT}$ using $\mathbf{\tilde{V}}_{u}$ a threshold estimator for
$\mathbf{V}_{u},$ given by equation (\ref{threshold})$.$ For comparison
purposes we also considered two other estimators of $\mathbf{V}_{u}.$ These
were the sample covariance matrix $\mathbf{\hat{V}}_{u}=\sum_{t=1}
^{T}\mathbf{\hat{u}}_{t}\mathbf{\hat{u}}_{t}^{^{\prime}}/T$ and a diagonal
covariance matrix, where the off-diagonal elements of $\mathbf{\hat{V}}_{u},$
$\hat{\sigma}_{ij},$ are set to zero. Thus we have three designs for the
return error covariance matrix, $\mathbf{V}_{u},$ and three different
estimators of it. A comparison of the results for the different covariance
matrices is available on request. The diagonal estimator, as to be expected,
performed poorly when the true covariance matrix was not diagonal,
particularly for the spatial error covariance matrix and when the strength of
the missing factor was close to $1/2$. For these designs the sample and
threshold estimators of the covariance matrix generally performed similarly
and given that there is a theoretical justification for the threshold
estimator and there are structures of the error covariance matrix for which
the sample estimator is unlikely to perform well we report the results using
the threshold estimator in the simulations below.
\subsection{Monte Carlo results}
We focus on a comparison of two-step (defined by (\ref{phihat1}) and the
bias-corrected (BC) estimator (defined by (\ref{phitilda})), and report bias,
root mean square error (RMSE) and size for testing $H_{0j}\,:\phi_{0k}=0,$
$k=M,H,S$ at the five per cent nominal level, for all $n=100,500,1,000,3,000$
and $T=60,$ $120,$ $240$ combinations. The results for all 12 experiments are
summarized in Tables S-A-E1 to S-A-E12 in the online supplement B. In terms of
bias and RMSE the two-step estimator does much better than the bias-corrected
(BC) estimator when $T=60$ and $n=100$, but this gap closes quickly as $n$ is
increased. In fact for $T=60$ and $n=3,000,$ the bias and RMSE of the BC
estimator (at $0.0010$ and $0.0607)$ are much less than those of the two-step
estimator (at -$0.0080$ and $0.1489)$. This pattern continues to hold when
$T=120$ and $240$. Bias correction can cause the RMSE to "blow up" for small
samples, such as $n=100$ and $T=60$, but for $n=500$ and above the
bias-corrected estimator always has a smaller RMSE than the two-step
estimator. As discussed in the theoretical section, having a large $n$ is
important for the properties of the estimators.
But most importantly, the two-step estimator is subject to substantial size
distortions, particularly when $T$ is small relative to $n$. As predicted by
the theory, the degree of over-rejection of the tests based on FM estimator
falls with $T,$ but increases with $n$. For example, the two-step test sizes
rise from $11.1\%$ when $T=60$ and $n=100$ to $60.9\%$ when $T=60$ and
$n=3,000$. Increasing $T$ reduces the size distortion of the two-step
estimator but test sizes are still substantially above the $5\%$ nominal value
when $n$ is large. The strong tendency of the tests based on the
two-step\ estimator to over-reject could be an important contributory factor
leading to false discovery of a large number of apparently significant factors
in the literature. In contrast, sizes of the tests based on the BC estimator,
using the variance estimator given by (\ref{Varphitilt}), are all close to its
nominal value, irrespective of the factor strength or sample size
combinations. We only note some elevated test sizes in the case of the
experimental design 12, and when we consider the semi-strong factors. The
highest test size of $7.85$ per cent is obtained for the least strong factor,
$f_{st}$, when $n=3000$ and $T=60$. See Table S-A-E12 of the online supplement B.
We also experimented with raising $\alpha_{\eta}$ from $0.3$ to $0.5,$ making
the pricing errors much more pervasive and $\rho_{\varepsilon}$ from $0.5$ to
$0.85,$ introducing more spatial correlation$.$ This increased the rejection
rate in experiments 9 and 10.
\subsubsection{Empirical power functions}
Plots of the empirical power functions for testing the null hypothesis
$H_{0j}\,:\phi_{0k}=0,$ $k=M,H,S$, are also provided in the online supplement
B for the 12 experiments (Figures S-A-E1 to S-A-E12). All power functions have
the familiar bell curve shape and tend to unity as $n$ and/or $T$ are
increased, showing the test has satisfactory power, particularly for $n$
sufficiently large even when $T=60$. Again the power functions are quite
similar across the 12 different experiments and show similar patterns for
strong and semi-strong factors. However, this similarity hides the fact that
the test of $\phi_{k}=0$ for the strong factor is much more powerful than
corresponding tests for the semi-strong factors, with the test power declining
as factor strength is reduced.
\subsubsection{Differences in performance of strong and semi-strong factors}
These differences in the effects of factor strength on the power of the test
of $\phi_{k}=0$ are in line with our theoretical results, and are also
reflected in the rate at which the RMSE of the estimators of $\phi_{M}$,
$\phi_{H}$ and $\phi_{S}$ fall with $n$. For example, using results in Tables
S-B-E10-12 in the online supplement B for the bias-corrected estimator in the
case of design 12 with $T=240,$ we note that the ratio of RMSE of $n=3,000$ to
$n=100$ is $17\%$ for the strong factor $\left( \alpha_{M}=1\right) $,
$23\%$ for the first semi-strong factor with $\alpha_{H}=0.85$, and $36\%$ for
the second semi-strong factor with $\alpha_{S}=0.65.$ As strength falls one
needs larger cross section samples of securities to attain the same level of
precision. In the case of the two-step estimator there was the same pattern,
but the fall in the RMSE with $n$ was much slower. The ratio for $T=240$ of
RMSE of $n=3,000$ to $n=100$ is $25\%$ for $\phi_{M}$, rather than $17\%$ (for
the bias-corrected estimator); $48\%$ rather than $23\%$ for $\phi_{H}$.
\subsubsection{Misspecification: Semi-strong versus weak factors}
So far, we have assumed that the DGP is correctly specified with two
semi-strong and no observed weak factors. Here we consider the implications of
incorrectly excluding semi-strong factors or correctly including weak factors
on the small sample properties of the bias-corrected\ estimator of $\phi_{M},$
the coefficient of the strong factor. Using the same DGP (which includes one
strong factor and two semi-strong factors), we carried out additional MC
experiments (designs 1-12) where we also estimated $\phi_{M}$ without the
semi-strong factors being included in the regressions. Comparative results,
with and without the semi-strong factors, are summarized in Tables S-C-E1-3 to
S-C-E10-12 of the online supplement B. We find that incorrectly excluding
semi-strong factors can be quite costly, both in terms of bias and RMSE as
well as size distortions. In terms of RMSE it was almost always better to
estimate the model with the semi strong factors included. The exception was
for the case of $T=60,$ $n=100,$ where including the semi-strong factors
caused the RMSE to blow up. Size distortions resulting from the exclusion of
the semi-strong factors tended to be more pronounced for large $n$ and $T$
samples. These conclusions were not sensitive to the choice of the
experimental design. For these experiments the lesson seems to be that it is
important to have $n$ large and include relevant semi-strong factors provided
that they are sufficiently strong.
When the DGP includes one strong factor ($\alpha_{M}=1$) and two weak factors
($\alpha_{H}=$ $\alpha_{S}=0.5)$, in terms of the bias and RMSE for $\phi_{M}$
it is unambiguously better to exclude the weak factors from the regression,
even though they are in the DGP. Weak factors are best treated as missing and
absorbed in the error term.
\subsection{Main conclusions from MC experiments}
The conclusions from the Monte Carlo simulations are that the bias corrected
estimator of $\boldsymbol{\phi}_{0}\boldsymbol{,}$ generally works well.
Although, it can generate a large RMSE for small $n$ and $T$, this can be
solved by increasing $n.$ This performance is robust to non-Gaussian errors,
GARCH effects, missing weak factors, pricing errors and weak cross-sectional
dependence. Test sizes are generally correct and the power good. The rate at
which RMSEs decline with $n$ depends on the strength of the underlying
factors. Semi-strong factors need much larger values of $n$ for precise
estimation. Tests of the joint significance of $\phi_{M}=\phi_{H}=\phi_{S}=0,$
not reported here, also performed well, as might be expected given the good
power performance of the separate induced tests which are reported. Including
weak factors could be harmful, but there are potential advantages of adding
semi-strong factors, although the issue of how best to select such factors is
an open question to which we now turn.
\section{Factor selection when the number of securities and the number of factors are both large\label{FacSel}}
Our theoretical derivations and Monte Carlo simulations both assume that the
number of risk factors included in the return regressions is fixed and the
factors are known. For the empirical application we face the additional
challenge of selecting a small number of relevant risk factors from a possibly
large number of potential factors, $m$. This problem has been the subject of a
number of recent studies.
\citet{harvey2016and}
propose a multiple testing approach aimed at controlling the false discovery
rate in the process of factor selection, and in a more recent paper
\citet{harvey2020false}
suggest using a double-bootstrap method to calibrate the t-statistic used for
controlling the desired level of the false discovery rate.
\citet{giglio2021asset}
suggest applying the double-selection Lasso procedure by
\citet{belloni2014inference}
to second pass regressions. None of these methods distinguish between strong,
semi-strong or weak factors in their selection process, whereas the theory and
simulations presented above indicate the importance of factor strength for
estimation and inference.
Here we propose an alternative selection procedure where we first estimate the
strength of all the $m$ factors under consideration, and then select factors
with strength above a given threshold, the value of which is informed by the
convergence results of Theorem \ref{Tsemi}. This theorem showed that if factor
$f_{kt}$ has strength $\alpha_{k},$ then for a given $T$ the BC estimator of
$\phi_{k}$ converges to its true value, $\phi_{0k}$, at the rate of
$n^{(\alpha_{k}+\alpha_{\min}-1)/2}$, where $\alpha_{\min}=\min_{i}(\alpha
_{i})$. As is recognized in non-parametric estimation literature, if the rate
of convergence is less than $1/3,$ the gain in precision with $n$ is so slow
that the estimator may not be that useful.\footnote{The
\citet{manski1985semiparametric}
maximum score estimator for a binary response model has $n^{1/3}$ convergence
and this is regarded as very slow and there are suggested modifications such
as
\citet{horowitz1992smoothed}
to increase the rate of convergence to $n^{2/5}$.} To achieve rate of
$n^{1/3}$ we need to set the threshold value of $\alpha_{k}$, denoted by
$\underline{\alpha}$, such that $\underline{\alpha}+\alpha_{\min}>1+2/3$. The
smallest value of such a threshold is obtained when $\alpha_{\min
}=\underline{\alpha}$ or if $\underline{\alpha}>1/2+1/3.$ Given the threshold
the main issue is how to estimate factor strength. This problem is already
addressed in \citet*[BKP]{bailey2021measurement} when $m$ is fixed. In this
setting they base their estimation on the statistical significance of $f_{kt}$
in the first stage time-series regressions of excess returns on all the
factors under consideration, whilst allowing for the $n$ multiple testing
problem which their approach entails. When $m\,$(the number of factors) is
also large the first stage regressions will also be subject to the multiple
testing problem and penalized regression techniques such as Lasso or the one
covariate at a time (OCMT) selection procedure technique proposed by
\citet{chudik2018one}
could be used. Irrespective of selection technique used at the level of
individual security returns, we end up with $n$ different subsets of the $m$
factors under consideration. Factor strengths can then be estimated similarly
to BKP\ from their selection frequencies across the $n$ securities.
To be more specific, denote the set of $m$ factors under consideration by
$\mathcal{S}$ and denote the set of selected factors for security $i$ by
$\hat{S}_{i}$ and their numbers by $\hat{m}_{i}=|\hat{S}_{i}|$. Clearly
$\hat{S}_{i}\subseteq\mathcal{S}$, and $\hat{m}_{i}\leq m$, for $i=1,2,...,n$.
Then compute the proportion of stocks in which the $k^{th}$ factor is
selected, $\hat{\pi}_{k},$ for $k\in\{1,2,...,m\}$ based on $\hat{S}_{i},$
$i=1,2,...,n$ by $\hat{\pi}_{k}=\frac{1}{n}\sum_{i=1}^{n}\mathcal{I}\{k\in
\hat{S}_{i}\}$. Then the strength of the $k^{th}$ factor is measured by
\begin{equation}
\hat{\alpha}_{k}=
\begin{cases}
1+\frac{\ln\hat{\pi}_{k}}{\ln n},\text{ if }\hat{\pi}_{k}>0,\\
0,\quad\quad\quad\text{ if }\hat{\pi}_{k}=0.
\end{cases}
\label{alphaj}
\end{equation}
The transformation from $\hat{\pi}_{k}$ to $\hat{\alpha}_{k}$ is explained and
justified in BKP, where it is shown that considering strength aids
interpretation because it is not dependent on $n$. It is beyond the scope of
the present paper to provide theoretical justification for the proposed factor
selection procedure, but using extensive Monte Carlo experiments
\citet{yoo2022factor}
has shown that the proposed method has desirable small sample properties
whether Lasso or OCMT is used for factor selection at the level of individual
security returns.
\section{An empirical application using a large number of U.S. securities and a large number of risk factors\label{Empirical}}
This section uses the results above in the explanation of monthly returns for
a large number, $n,$ of U.S. securities, by a large active set of $m$
potential risk factors. We first briefly describe the sources and
characteristics of the data for the stock returns and factors, which cover
different sub-samples over the period $1996m1-2022m12.$ We then consider the
selection of a subset of $K$ factors from the active set. Finally we test
$\boldsymbol{\phi}_{0}=0$, and construct and evaluate phi-portfolios and
corresponding mean-variance\textbf{ } portfolios for alternative models.
Monthly returns (inclusive of dividends) for NYSE and NASDAQ stocks from CRSP
with codes $10$ and $11$ were downloaded from Wharton Research Data Services.
They were converted to excess returns by subtracting the risk free rate, which
was taken from Kenneth French's data base. To obtain balanced panels of stock
returns and factors, only variables for which there was data for the full
sample under consideration were used. Excess returns are measured in percent
per month. To avoid outliers influencing the results, stocks with a kurtosis
greater than 16 were excluded. To examine factor selection, four samples were
considered, each had $20$ years of data, $T=240$, ending in $2015m12$,
$2017m12$, $2019m12$, $2021m12$. Filtering out the stocks with kurtosis larger
than $16$ removed about $100$ of the roughly $1200$ stocks. The number of
stocks ($n$) considered for each of the four $T=240$ samples are given in
panel A of Table \ref{FactorSelPL240}. Further detail is given in Section
\ref{DEMP} of the online supplement A.\footnote{Summary statistics for the
excess returns across the different samples are given in Table
\ref{tab:ret_sumtat} of the online supplement A.} For analysis of the
phi-portfolios, the factors selected in the sample ending in $2015m12$ were
used to construct portfolios up to $2022m12.$
\subsection{Factor selection}
For factor selection we used a sample of $T=240$ observations\footnote{Some
results for $T=120$ are included in the online supplement.}. The set of
factors considered combine the $5$ Fama-French factors with the $207$ factors
from the
\citet{chen2022open}
, Open Source Asset Pricing webpages, both downloaded July 6 2022. Only
factors with data for the full sample were considered so the return
regressions constitute a balanced panel. The number of factors in each of the
four $20$ year samples ending in the years $2015$, $2017$, $2019$, $2021$ is
also given in panel A of Table \ref{FactorSelPL240}, and range between $187$
to $199$. Summary statistics for the factors in the active set are given in
\ref{SumFac} of the online supplement A.
To implement the factor selection procedure set out in Section \ref{FacSel},
Lasso is used to carry out selection in the return regressions for individual
securities and we refer to the factor selection procedure as pooled Lasso
(PL). As is well known, Lasso does not work well with too many highly
correlated regressors, therefore, factors with an absolute correlation with
the market factor greater than $0.70$ were dropped. This still left between
$177$ and $190$ risk factors in the active set $\mathcal{S}$, depending on the
sample period (see panel A of Table \ref{FactorSelPL240}).
Specifically, Lasso was applied $n$ times to the regressions of excess
returns, $r_{it}=R_{it}-r_{t}^{f}$, for $i=1,2,...,n$, on the $177$ to $190$
factors in the active set, $\mathcal{S}$, to select the sub-set $\hat{S}_{i}$
for each $i$ over the four 20-year samples, separately. Following the
literature, the tuning parameters in the Lasso algorithm were set by ten-fold
cross-validation.\footnote{The post-Lasso and one covariate multiple testing
(OCMT) approach of
\citet
*[CKP]{chudik2018one}. were also investigated, but Lasso seemed to work
reasonably well. The details of the Lasso procedure used are given in Section
2.2 of the online supplement of CKP paper.} Interestingly, the market factor
was selected by Lasso for almost all the securities, thus confirming the
pervasive nature of the market factor. No other factor came close to being
selected for all the securities. Lasso tended to choose a lot of non-market
factors and every factor got chosen in at least one return regression. The
mean number of non-market factors chosen by Lasso fell from $11.6$ in the
$2015$ sample to $9.9$ in the $2021$ sample. The median was lower, falling
from $10$ to $8$. There was a long right tail because Lasso tended to choose a
very large number of non-market factors for some securities, ranging from a
maximum of $48$ in the 2021 sample to $54$ in the $2017$ sample.
Apart from the market factor, there are no systematic patterns for the rest of
selected factors in the return equations for the individual securities.
Following the theory set out in Section \ref{FacSel}, the $K$ factors used to
estimate $\boldsymbol{\phi}$ are chosen on the basis of their factor
strength.\footnote{The idea of using factor strength could also be viewed as a
kind of averaging of the factors selected in individual return regressions.} A
minimum threshold of $0.7$ was used. This is below the threshold value of
$\underline{\alpha}=1/2+1/3$ required to achieve the convergence rate of
$n^{1/3}$, and is intended to capture borderline semi-strong factors. We also
consider the values of $0.75,$and $0.80$ that are quite close to
$\underline{\alpha}$. The number of selected factors for different choices of
factor strength threshold is given in panel B of Table \ref{FactorSelPL240},
for the four different samples.
Using the threshold of $0.70$, $17$ factors (inclusive of the market factor)
were selected for the sample ending in $2015$, with the number of selected
factors declining to $15,13$ and $11,$ for the samples ending $2017,$ $2019$
and $2021$, respectively. At the other extreme, setting the threshold at
$0.80$, the number of selected factors dropped to $4$ for the samples ending
in $2015$, $2017$ and $2021$, and $2$ for the sample ending in $2021$. Since
$17$ factors seemed too large and $2$ factors too small, the threshold value
was set at the intermediate value of $0.75$. We considered always conditioning
on the market factor, but since Lasso almost always selected it, this was unnecessary.
\floatstyle{plaintop}
\restylefloat{table}
\begin{table}[H]
\begin{center}
\begin{tabular}
[c]{ccccc}\hline\hline
$T=240$ with \ end dates & 2021 & 2019 & 2017 & 2015\\\hline
\multicolumn{5}{c}{}\\
\multicolumn{5}{c}{Panel A: Number of stocks and factors under consideration}
\\
Number of stocks & 1289 & 1276 & 1243 & 1181\\
Number of stocks with kurtosis $<16$ & 1175 & 1143 & 1132 & 1090\\
Number of non-market factors & 187 & 198 & 199 & 197\\
Number of non-market factors with $r<0.70$ & 177 & 189 & 190 & 189\\
\multicolumn{5}{c}{}\\
\multicolumn{5}{c}{Panel B: Number of selected factors by strength
threshold}\\
Number with strength $>0.80$ & 2 & 4 & 4 & 4\\
Number with strength $>0.75$ & 4 & 6 & 7 & 7\\
Number with strength $>0.70$ & 11 & 13 & 15 & 17\\\hline\hline
\end{tabular}
\end{center}
{\footnotesize \textit{Note:} Panel A shows the number of stocks and risk
factors used before and after filtering by the specified criterion. Panel B
shows the number of risk factors selected with strength greater than the
specified threshold level using Lasso to select factors at the level of
individual securities. }
\caption
{Summary statistics for the number of stocks and number of selected factors using the factor strength threshold of 0.75, for four twenty years ($T=240$) samples ending in 2021, 2019, 2017 and 2015}
\label{FactorSelPL240}
\end{table}
\onehalfspacing
The list of factors selected by pooled Lasso for the four samples are given in
Table \ref{TabRiskF}. The three Fama-French factors, Market, HML and SMB, are
all selected in all four periods.\footnote{We also considered selecting the
risk factors using the generalized one covariate at a time (OCMT) method
proposed by
\citet{sharifvaghefi2022variable}
. Using GOCMT the Fama-French three factors were again amongst the five
strongest factors selected. The use of GOCMT\ for factor selection is also
investigated by
\citet{yoo2022factor}
, using Monte Carlo and empirical applications.} Of the Fama-French three,
only the market factor is strong, with estimated strength in excess of $0.98$
across the four periods.\footnote{These results also support the choice of 3
Fama-French factors and their strength used in our Monte Carlo simulations.}
The other factor which is selected across all the four periods is "short
selling". This is proposed by
\citet{dechow2001short}
who argue that short-sellers target firms that are priced high relative to
fundamentals. It measures the extent to which investors are shorting the
market as reflected in Compustat data. Two additional factors are selected in
periods ending in $2019$ and earlier. One is "Beta Tail Risk" proposed by
\citet{kelly2014tail}
which estimates a time-varying tail exponent from the cross section of
returns. The other is "Cash Based Operating Profitability" (CBOP) suggested
by
\citet{ball2016accruals}
. This is operating profit less accruals, with working capital and R\&D
adjustments. For periods ending in 2017 and 2015 the "Sin Stock" indicator
proposed by
\citet{hong2009price}
is also selected. It takes the value of unity if the stock in question is
involved in producing alcohol, tobacco, and gaming. They find that such stocks
are held less by norm-constrained institutions such as pension plans.
The factor strengths are relatively stable across the periods, with many of
the estimates close to the threshold value of $0.75$. Apart from the market
factor only SMB, Short Selling and Beta Tail Risk factors have strengths in
excess of $0.85$ when averaged across the four periods. From the large number
of factors in the active set we have ended up with relatively few factors that
are reasonably strong and for which $\boldsymbol{\phi}_{0}$ can be estimated
reasonably accurately.\footnote{We do not report estimates of individual
$\phi_{k}.$ Because of correlations between the loadings, the sign, size and
significance of the coefficients are difficult to interpret and for
phi-portfolio construction, discussed below, what matters is $\boldsymbol{\phi
}^{\prime}\boldsymbol{\phi}$ which determines the return on the portfolio.}
\singlespacing
\begin{table}[htbp]
\begin{center}
\begin{tabular}
[c]{ccccccccc}\hline\hline
End date & 2021 \ & \multicolumn{2}{c}{2019} & \multicolumn{2}{c}{2017} &
\multicolumn{3}{c}{2015}\\
Selected Factors & \multicolumn{8}{c}{Estimated strength ($\alpha$)}\\\hline
Mkt. & 0.99 & & 0.98 & & 0.98 & & 0.98 & \\
SMB & 0.90 & & 0.84 & & 0.86 & & 0.86 & \\
Short Selling & 0.77 & & 0.85 & & 0.85 & & 0.83 & \\
HML & 0.76 & & 0.77 & & 0.76 & & 0.75 & \\
BetaTailRisk & & & 0.87 & & 0.86 & & 0.86 & \\
CBOP & & & 0.76 & & 0.77 & & 0.77 & \\
Sin Stock & & & . & & 0.76 & & 0.76 & \\\hline\hline
\end{tabular}
\end{center}
{\footnotesize \textit{Note:} The risk factors listed are the market factor
(Mkt.), size (SMB), Short Selling that measures the extent of short sales in
the market, the value factor (HML), the cash-based operating profitability
factor (CBOP), Beta Tail Risk, and Sin Stock which is a binary indicator
taking the value of unity if the stock in question is involved in so called
"Sin" industries producing alcohol, tobacco, and gaming. Further details on
these risk factors are provided by
\citet{chen2022open}
.}
\caption{Selected factors with estimated strength in excess of the threshold 0.75 for the samples of size
T=240 ending in 2021, 2019, 2017 and 2015 }\label{TabRiskF}
\end{table}
\onehalfspacing
The strengths of the selected factors are also closely related to the average
measures of fit often used in the literature. Here we consider both
$Ave\bar{R}^{2}=n^{-1}\sum_{i=1}^{n}\bar{R}_{i}^{2}$, a simple average of the
fit of the individual return regressions adjusted for degrees of freedom,
$\bar{R}_{i}^{2}=1-(T-K-1)^{-1}\sum_{t=1}^{T}\hat{u}_{it}^{2}/T^{-1}\sum
_{t=1}^{T}\left( r_{it}-\bar{r}_{i\circ}\right) ^{2}$, and the adjusted
pooled $R^{2}$ defined by $\overline{PR}^{2}=1-\widehat{\bar{\sigma}}_{nT}
^{2}/s_{r,nT}^{2}$, where $\widehat{\bar{\sigma}}_{nT}^{2}$ is the
bias-corrected estimator of $\bar{\sigma}_{n}^{2}$ defined by (\ref{AdjZig})
and $=$ $\left( nT\right) ^{-1}\sum_{i=1}^{n}\sum_{t=1}^{T}\left(
r_{it}-\bar{r}_{i\circ}\right) ^{2}$. Both of these measures behave very
similarly, but the pooled version is less sensitive to outliers. As shown in
Appendix \ref{appPRsquared}, for sufficiently large $n$ and $T$,
$\overline{PR}^{2}$ is dominated by the contribution of the most strong
factor(s). Since the only strong factor selected is the market factor in Table
\ref{tab:RsquaredFAandPL}\ we report the $Ave\bar{R}^{2}$ and $\overline
{PR}^{2}$ in the case of return regressions which just include the market
factor and those which include all other factors with strength in excess of
$0.75$. First, we note that the $\overline{PR}^{2}$ values are generally lower
than the $Ave\bar{R}^{2}$. Second, the additional factors do add to the fit,
but their relative contributions vary considerably across sample sizes and
periods. In general, the marginal contribution of non-market factors tend to
be smaller when $T$ is larger, which is consistent with the theory for
adjusted pooled $R^{2}$ set out in Section \ref{appPRsquared} of the online
supplement A.
\singlespacing
\begin{table}[htbp]
\begin{center}
\begin{tabular}
[c]{cccccccccccc}\hline\hline
\multicolumn{3}{c}{End Year} & & \multicolumn{2}{c}{2021} &
\multicolumn{2}{c}{2019} & \multicolumn{2}{c}{2017} & \multicolumn{2}{c}{2015}
\\\cline{5-12}
\multicolumn{3}{c}{No. of stocks, $n$} & & \multicolumn{2}{c}{1175} &
\multicolumn{2}{c}{1143} & \multicolumn{2}{c}{1132} & \multicolumn{2}{c}{1090}
\\
\multicolumn{3}{c}{No. of selected factors} & & \multicolumn{2}{c}{4} &
\multicolumn{2}{c}{6} & \multicolumn{2}{c}{7} & \multicolumn{2}{c}{7}\\\hline
& & & & Mkt. & Selected & Mkt. & Selected & Mkt. & Selected & Mkt. &
Selected\\
& & $T$ & & & & & & & & & \\
$Ave\bar{R}^{2}$ & & 240 & & 0.23 & 0.29 & 0.19 & 0.28 & 0.17 & 0.27 &
0.17 & 0.27\\
& & 120 & & 0.24 & 0.33 & 0.25 & 0.35 & 0.26 & 0.38 & 0.26 & 0.37\\
& & & & & & & & & & & \\
$\overline{PR}^{2}$ & & 240 & & 0.20 & 0.26 & 0.17 & 0.26 & 0.16 & 0.26 &
0.16 & 0.26\\
& & 120 & & 0.19 & 0.27 & 0.20 & 0.29 & 0.23 & 0.34 & 0.23 &
0.34\\\hline\hline
\end{tabular}
\end{center}
{\footnotesize \textit{Note:} This table shows for each of the four end years
and the two sample sizes the adjusted average and pooled }${\footnotesize R}
^{2}${\footnotesize (}${\footnotesize Ave\bar{R}}^{2}${\footnotesize and
}$\overline{{\footnotesize PR}}^{{\footnotesize 2}}${\footnotesize ) for the
return regressions when using market factor alone or factors selected with
strength higher than }${\footnotesize 0.75}${\footnotesize . The list of
selected factors are given in Table \ref{TabRiskF}. }
\caption{Average and pooled R squared for the return regressions when using market factor alone or factors chosen by Pooled Lasso plus 0.75 threshold }\label{tab:RsquaredFAandPL}
\end{table}
\onehalfspacing
\subsection{Testing for non-zero $\boldsymbol{\phi}$}
In principle, if $\boldsymbol{\phi}\neq0$ there are potentially exploitable
excess returns. In practice, to construct an effective phi-portfolio a large
number of securities is required and rebalancing such long-short portfolios
for so many securities may not be feasible or may incur high transactions
costs. In addition, model uncertainty, estimation uncertainty, time variation
in both $\boldsymbol{\beta}_{i}$ and in conditional volatility pose additional
difficulties in implementing a strategy to exploit the potential returns
revealed by $\boldsymbol{\phi}$. We will abstract from such practical
difficulties to provide some indication of the performance of phi-portfolios
relative to alternatives which would face similar difficulties.
As our preferred asset pricing model, we consider the seven factors selected
by pooled Lasso using the sample ending in $2015m12,$ which we label as
PL7.\footnote{The selection of the PL7 model was reported in the earlier
version of the paper submitted for publication and was not informed by the
performance the phi-portfolio that we report in this version of the paper.}
Recall that the PL7 includes the 3 Fama-French factors (Mkt., SMB, and HML)
plus the four risk factors, Short Selling, CBOP, Beta Tail Risk, and Sin
Stocks. But given uncertainties that surround the problem of model selection
we also considered the two popular FF factor models, namely FF3 and FF5. The
latter augments FF3 with RMW (robust minus weak operating profitability) and
CMA (conservative minus aggressive investment portfolios). The three models
are estimated using twenty-year rolling windows covering the $84$ months from
$2015m12$ to $2022m11$, so that we can generate out of sample return forecasts
for the months $2016m1-2022m12$.\footnote{Due to entry and exit of securities
the number of securities included in our analysis varied across the rolling
sample periods. We started with $n=1,090$ securities for the first rolling
sample ending in $2015m12$, with the number of available securities with $240$
months of data falling to $953$ by $2017$, $838$ by $2019$, $767$ by $2021$,
and $736$ by $2022$.} For all $3\times84$ model-sample interactions we
computed the following Wald test statistics
\[
W_{t\left\vert T\right. }^{2}=\boldsymbol{\tilde{\phi}}_{t\left\vert
T\right. }^{\prime}\left[ \widehat{Var\left( \boldsymbol{\tilde{\phi}
}_{t\left\vert T\right. }\right) }\right] ^{-1}\boldsymbol{\tilde{\phi}
}_{t\left\vert T\right. },
\]
for testing the null hypothesis $\boldsymbol{\phi}=0,$ where
$\boldsymbol{\tilde{\phi}}_{t\left\vert T\right. }$ and $\widehat{Var\left(
\boldsymbol{\tilde{\phi}}_{t\left\vert T\right. }\right) }$ denote the
rolling versions of (\ref{phitilda}) and (\ref{Varphitilt}).\footnote{The
formulae for the rolling estimates are provided in the sub-section \ref{Roll}
of the online appendix.} For the PL7 the range of the test statistic was from
141.8 to 25.8, as compared to the 5 per cent $\chi^{2}(7)$ critical value of
14.07. Thus the hypothesis that $\boldsymbol{\phi}=0$ is strongly rejected in
all the 84 rolling sample for PL7. This is also true for the FF5 and FF3
models where the test statistic ranged from 162.7 to 25.0, and from 67.9 to
25.8, compared to the 5 per cent critical values of 11.07 and 7.8, respectively.
The rolling values of the Wald statistics for testing $\boldsymbol{\phi}=0$
for the three models are shown in Figure 1. The horizontal line (pink)
represents the critical value of the $\chi_{7}^{2}$ distribution at the 5 per
cent level. The time profiles of these test statistics clearly show that
$\boldsymbol{\phi}=0$ is rejected for all rolling samples and for all three
models. But there is also a clear downward trend showing that the evidence
against $\boldsymbol{\phi}=0$ has been getting weaker over time, irrespective
of the choice of asset pricing model.
\begin{figure}[ptb]
\centering
\caption{Rolling chi-squared statistics for testing $\boldsymbol{\phi}=0$
using a window of size $240$ for FF3, FF5, and PL7 models. }
\includegraphics[
height=2.8357in,
width=4.5861in
]
{CHISQ.png}
\end{figure}
\bigskip
\subsection{Comparative performance of phi and MV portfolios}
Having established that most likely $\boldsymbol{\phi\neq0}$ for the asset
pricing models we have considered, we now turn to the performance of
phi-portfolios based on these models. Using the recursive version of the
phi-portfolio given by (\ref{phiport}), we consider the following
phi-portfolio returns
\[
\hat{\rho}_{t+1,\phi}=\boldsymbol{\tilde{\phi}}_{t\left\vert T\right.
}\left[ \left( \mathbf{\hat{B}}_{t\left\vert T\right. }^{\prime}
\mathbf{M}_{n}\mathbf{\hat{B}}_{t\left\vert T\right. }\right) ^{-1}
\mathbf{\hat{B}}_{t\left\vert T\right. }\mathbf{M}_{n}\mathbf{r}_{\circ
,t+1}-\mathbf{f}_{t+1}\right] ,
\]
for $t=2016m1,2016m2,...,2022m12$, using the rolling estimates $\mathbf{\hat
{B}}_{t\left\vert T\right. }=\left( \mathbf{\hat{\beta}}_{1t\left\vert
T\right. },\mathbf{\hat{\beta}}_{2t\left\vert T\right. }....,\mathbf{\hat
{\beta}}_{nt\left\vert T\right. }\right) ^{\prime}$ with $T=240$, for each
of the three factor models, FF3, FF5 and PL7. We compare the annualised Sharpe
ratios of phi-portfolios with the ones based on associated MV\ portfolios,
given by $\rho_{t+1,MV}=\boldsymbol{\mu}_{R}^{\prime}\mathbf{V}_{R}
^{-1}\mathbf{r}_{\circ,t+1}$.\footnote{Given our focus on the Sharpe ratios,
we have set the scaling of the MV portfolio to unity.} Although in principle,
MV portfolios can be constructed without a reference to a particular factor
model, reliable estimation of $\boldsymbol{\mu}_{R}$ and $\mathbf{V}_{R}^{-1}$
are challenging when $n$ is relatively large. For example, the rolling sample
covariance matrix estimator of $\mathbf{V}_{R}$, given by $\mathbf{\mathring
{V}}_{R,t\left\vert T\right. }=T^{-1}\sum_{\tau=t-T+1}^{t}\left(
\mathbf{r}_{\circ,\tau}-\mathbf{\bar{r}}_{\circ,t\left\vert T\right.
}\right) \left( \mathbf{r}_{\circ,\tau}-\mathbf{\bar{r}}_{\circ,t\left\vert
T\right. }\right) ^{\prime}$, with $\mathbf{\bar{r}}_{\circ,t\left\vert
T\right. }=T^{-1}\sum_{\tau=t-T+1}^{t}\mathbf{r}_{\circ,\tau}$, will be
singular when $n>T$, and can be very poorly estimated if $T$ is not
sufficiently large relative to $n$. There is a vast literature on consistent
estimation of high dimensional covariance matrices like $\mathbf{V}_{R}.$
\cite{fan2011high} use observed factors while \cite{fan2013large} use
principal components to filter out the effects of strong factors, in both
cases assuming $\mathbf{V}_{u}$ is sparse, and then using a threshold method
to estimate it. Shrinkage estimators of $\mathbf{V}_{R}$ are also proposed in
the literature with a recent survey provided by \cite{LedoitWolfsurvey2022}.
However, the shrinkage estimators require $n$ and $T$ to be of the same order
of magnitude and do not work well when $n$ is much larger than $T$, as in the
present application. We follow \cite{fan2011high} and base our estimation of
$\boldsymbol{\mu}_{R}$ and $\mathbf{V}_{R}$ on the same factor model used to
construct the phi-portfolio. For a given factor model, characterized by
$c\mathbf{,B}$, and $\mathbf{F}$, we compute the MV portfolio returns as
\[
\hat{\rho}_{t+1.MV}=\mathbf{\hat{\mu}}_{R,t\left\vert T\right. }^{\prime
}\mathbf{\hat{V}}_{R,t\left\vert T\right. }^{-1}\mathbf{r}_{\circ,t+1},
\]
where $\mathbf{\hat{\mu}}_{R,t\left\vert T\right. }=\hat{c}_{t\left\vert
T\right. }\mathbf{\tau}_{n}+\mathbf{\hat{B}}_{t\left\vert T\right.
}\boldsymbol{\tilde{\lambda}}_{t\left\vert T\right. }$, $\boldsymbol{\tilde
{\lambda}}_{t\left\vert T\right. }=\boldsymbol{\tilde{\phi}}_{t\left\vert
T\right. }+\boldsymbol{\hat{\mu}}_{t\left\vert T\right. },$
\[
\mathbf{\hat{V}}_{R,t\left\vert T\right. }=\mathbf{\hat{B}}_{t\left\vert
T\right. }^{\prime}\left( \frac{\mathbf{F}_{t\left\vert T\right. }^{\prime
}\mathbf{M}_{T}\mathbf{F}_{t\left\vert T\right. }}{T}\right) ^{-1}
\mathbf{\hat{B}}_{t\left\vert T\right. }\mathbf{+\mathbf{\tilde{V}}
}_{u,t\left\vert T\right. }\mathbf{,}
\]
$\boldsymbol{\hat{\mu}}_{t\left\vert T\right. }=\boldsymbol{\bar{f}
}_{t\left\vert T\right. }=T^{-1}\sum_{\tau=t-T+1}^{t}\mathbf{f}_{\tau}$, and
$\mathbf{\mathbf{\tilde{V}}}_{u,t\left\vert T\right. }$ is the rolling
estimate of $\mathbf{V}_{u}$. The algorithms used to compute the recursive
estimates for the MV portfolio can be found in the sub-section \ref{Roll} in
the online supplement A.
Table \ref{tab:portfolios} presents annualised SR of the phi-portfolios, for
the FF3, FF5, and PL7 models, and their corresponding MV portfolios for two
samples, both beginning in 2016m1, one ending in 2019m12, pre Covid-19, and
one ending in 2022m12. \ In five of the six SR ratios reported in this table,
the phi-portfolio has a higher SR than the corresponding MV portfolio. The
exception is the SR associated to the FF5 model for the pre Covid-19 sample.
This illustrates that if $\boldsymbol{\phi}\neq\mathbf{0},$ it is possible to
construct portfolios that outperform the mean-variance\textbf{ } portfolio.
Amongst the 3 models considered, the phi-portfolio based on the PL7 model
performed best during the pre Covid-19 and the full sample, even beating the
S\&P 500. The Sharpe ratio of the phi-portfolio based on the Pl7 model was
1.95 compared to 0.94 for the S\&P 500 during the pre Covid-19 period, and
fell sharply to 0.65 for the full sample as compared to 0.58 for the S\&P 500.
The SR for the MV portfolio using the same model were 0.87 and 0.42 for the
two samples, respectively.
The sharp decline in the SRs as we add the post Covid-19 years is in line with
the strong downward trend in the Wald statistics for the test of
$\boldsymbol{\phi}=\mathbf{0}$ shown in Figure 1. As is well known SRs have
large standard errors, and in the case of our application that are around 1,
so none of the SRs are significantly different from zero, with the possible
except of the largest SR of 1.95 for phi-portfolio based on PL7 model for the
pre Covid-19 period. We also note that the reported Sharpe ratios do not allow
for transaction costs and the fact that shorting might not be feasible for all
the securities.
\begin{table}[H]
\caption
{Annualised Sharpe ratios of realized monthly returns for alternative portfolios based on 240 month rolling window estimates.}
\begin{centering}
\begin{tabular}
[c]{cccccc}\hline\hline
& \multicolumn{2}{c}{Mean-variance portfolios} & &
\multicolumn{2}{c}{Phi-portfolios}\\\cline{2-3}\cline{5-6}
& 2016$m$1 - 2019$m$12 & 2016$m$1 - 2022$m$12 & & 2016$m$1-2019$m$12 &
2016$m$1-2022$m$12\\
Models & Pre Covid-19 & Full sample & & Pre Covid-19 & Full sample\\\hline
& & & & & \\
FF3 & 0.57 & 0.35 & & 0.86 & 0.51\\
FF5 & 0.49 & 0.43 & & 0.35 & 0.57\\
Lasso 7 & 0.87 & 0.42 & & 1.95 & 0.65\\\hline\hline
\end{tabular}
\end{centering}
\label{tab:portfolios}
\vspace{3mm}
\emph{{\footnotesize {}\textit{Note}}}{\footnotesize {}: The annualised Sharpe
ratio (SR) is computed as }$\sqrt{12}\bar{\rho}/s_{\rho}${\footnotesize ,
where }$\bar{\rho}$ {\footnotesize is the mean of monthly returns, and
}$s_{\rho}$ {\footnotesize is the standard deviation of monthly returns. For
comparison the SR of the monthly returns on S\&P 500 were 0.94 and 0.58 over
the periods 2016m1-2019m12, and 2016m1-2022m12, respectively.}
\end{table}
\section{Concluding remarks\label{Conclusion}
\onehalfspacing
}
In this paper we have highlighted the importance of decomposing the risk
premia, $\boldsymbol{\lambda}$, into the the factor mean, $\boldsymbol{\mu}$,
and $\boldsymbol{\phi}$, and writing the alpha of security $i$,
$\mathit{\alpha}_{i}$, in terms of $\boldsymbol{\phi}$ and the idiosyncratic
pricing errors. We have shown that when $\boldsymbol{\phi\neq0}$, it is
possible to construct a portfolio, denoted as phi-portfolio, that dominates
the associated mean-variance portfolio when the number securities, $n$, is
sufficiently large and the risk factors are sufficiently strong. Given the
pivotal role played by $\boldsymbol{\phi}$ for estimating the risk premia, for
formation of large $n$ portfolios, and for tests of market efficiency, we have
focussed on estimation of $\boldsymbol{\phi}$, and its asymptotic distribution
under quite a general setting that allows for missing factors and
idiosyncratic pricing errors. Since factor means, $\boldsymbol{\mu}$, can be
estimated at the regular rate of $T^{-1/2}$ from time series data, it is
relatively straightforward to develop a mixed strategy for estimation of
$\boldsymbol{\lambda}$ by adding a time series estimate of $\boldsymbol{\mu}$
to the bias-corrected estimator of $\boldsymbol{\phi}$. If we use the same
time series sample, such an estimator reduces to the Shanken bias-corrected
estimator of $\boldsymbol{\lambda}$. But in practice, given the concern over
the instability of the factor loadings $\beta_{ik}$ over time, one could use
relatively long time series, say $T_{\mu}$, when estimating $\boldsymbol{\mu}
$, and a shorter time series, say $T_{\phi}<T_{\mu}$, when estimating
$\boldsymbol{\phi}$. The distributional and small sample properties of such a
mixed estimator of risk premia is a topic for further research.
Our theoretical and Monte Carlo results further highlight the important role
played by factor strengths in estimation and inference on $\boldsymbol{\phi}$,
and hence on $\boldsymbol{\lambda}$. For a fixed $T_{\phi}$, factors with
strength below $2/3$ lead to estimates of $\boldsymbol{\phi}$ \ with
convergence rate of $n^{-1/3}$ or worse, and their use in asset pricing models
can be justified only when $n$ is very large. We have also shown that weak
factors, with strength below $1/2$ are best treated as missing and absorbed in
the error term. We have shown that estimation of $\boldsymbol{\phi}$ for
strong or semi-strong factors is robust to weak missing factors, and the
explicit inclusion of weak factors in the empirical analysis is likely to have
adverse spill over effects on the estimates of $\boldsymbol{\phi\ }$\ for
strong and semi-strong factors. In view of these results we have proposed a
factor selection procedure where only factors with strength above $1/2+1/3$
are included in the asset pricing model. Developing a formal statistical
theory for the proposed selection is another topic for future research.
The paper also provides an empirical application to a large number of U.S.
securities with risk factors selected from a large number of potential risk
factors according to their strength, and use a pooled Lasso approach to select
$7$ risk factors out of over $180$. We find strong statistical evidence
against $\boldsymbol{\phi=0}$ for the selected model as well as for the two
popular Fama-French models (FF3 and FF5). Using rolling estimates of
$\boldsymbol{\phi}$ we also construct phi portfolios with better Sharpe ratios
as compared to associated mean-variance\textbf{ } portfolios. But we also warn
that these portfolio comparisons are preliminary and need to be further
investigated by allowing for transaction costs, and the feasibility of the
long-short trading strategies that are involved.
\pagebreak
\singlespacing
\bibliographystyle{econometrica}
\bibliography{pslapmainV2}
\newpage