EconBase
← Back to paper

Optimal Portfolio Using Factor Graphical Lasso

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

92,460 characters · 17 sections · 92 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Optimal Portfolio Using Factor Graphical Lasso footnotefootnote mpfootnotempfootnote The authors would like to thank the editor Fabio Trojani and three anonymous referees for helpful and constructive comments on the paper.

\bgroup \renewcommand\fnsymbol{footnote}{\fnsymbol{footnote}} \renewcommand\fnsymbol{mpfootnote}{\fnsymbol{mpfootnote}} \footnotetext[0]{The authors would like to thank the editor Fabio Trojani and three anonymous referees for helpful and constructive comments on the paper. } \egroup }

\setcounter{page}{1}

abstract\begin{spacing}{2} Graphical models are a powerful tool to estimate a high-dimensional inverse covariance (precision) matrix, which has been applied for a portfolio allocation problem. The assumption made by these models is a sparsity of the precision matrix. However, when stock returns are driven by common factors, such assumption does not hold. We address this limitation and develop a framework, Factor Graphical Lasso (FGL), which integrates graphical models with the factor structure in the context of portfolio allocation by decomposing a precision matrix into low-rank and sparse components. Our theoretical results and simulations show that FGL consistently estimates the portfolio weights and risk exposure and also that FGL is robust to heavy-tailed distributions which makes our method suitable for financial applications. FGL-based portfolios are shown to exhibit superior performance over several prominent competitors including equal-weighted and Index portfolios in the empirical application for the S&P500 constituents. \end{spacing} \vskip 2mm Keywords: High-dimensionality, Portfolio optimization, Graphical Lasso, Approximate Factor Model, Sharpe Ratio, Elliptical distributions \vskip 2mm JEL Classifications: C13, C55, C58, G11, G17

{22pt} \setstretch{2}

Introduction

Estimating the inverse covariance matrix, or precision matrix, of excess stock returns is crucial for constructing weights of financial assets in a portfolio and estimating the out-of-sample Sharpe Ratio. In high-dimensional setting, when the number of assets, $p$, is greater than or equal to the sample size, $T$, using an estimator of covariance matrix for obtaining portfolio weights leads to unstable investment allocations. This is known as the Markowitz’ curse: a higher number of assets increases correlation between the investments, which calls for a more diversified portfolio, and yet unstable corner solutions for weights become more likely. The reason behind this curse is the need to invert a high-dimensional covariance matrix to obtain the optimal weights from the quadratic optimization problem: when $p\geq T$, the condition number of the covariance matrix (i.e., the absolute value of the ratio between maximal and minimal eigenvalues of the covariance matrix) is high. Hence, the inverted covariance matrix yields an unstable estimator of the precision matrix. To circumvent this issue one can estimate precision matrix directly, rather than inverting an estimated covariance matrix.

Graphical models were shown to provide consistent estimates of the precision matrix (GLASSO,meinshausen2006,cai2011constrained). goto2015 estimated a sparse precision matrix for portfolio hedging using graphical models. They found out that their portfolio achieves significant out-of-sample risk reduction and higher return, as compared to the portfolios based on equal weights, shrunk covariance matrix, industry factor models, and no-short-sale constraints. AwoyePhD used Graphical Lasso (GLASSO) to estimate a sparse covariance matrix for the Markowitz mean-variance portfolio problem and reduce the realized portfolio risk. Millington conducted an empirical study that applies Graphical Lasso for the estimation of covariance for the portfolio allocation. Their empirical findings suggest that portfolios using Graphical Lasso enjoy lower risk and higher returns compared to those using empirical covariance matrix. Millington also construct a financial network using the estimated precision matrix to explore the relationship between the companies and show how the constructed network helps to make investment decisions. Caner2019 use the nodewise-regression method of meinshausen2006 to establish consistency of the estimated covariance matrix, weights and risk of high-dimensional financial portfolio. Their empirical application demonstrates that the precision matrix estimator based on the nodewise-regression outperforms the principal orthogonal complement thresholding estimator (POET) (fan2013POET) and linear shrinkage (Ledoit2004). cai2020high use constrained $\ell_1$-minimization for inverse matrix estimation (Clime) of the precision matrix (cai2011constrained) to develop a consistent estimator of the minimum variance for high-dimensional global minimum-variance portfolio. It is important to note that all the aforementioned methods impose some sparsity assumption on the precision matrix of excess returns.

An alternative strategy to handle high-dimensional setting uses factor models to acknowledge common variation in the stock prices, which was documented in many empirical studies (see campbell1997 among many others). A common approach decomposes covariance matrix of excess returns into low-rank and sparse parts, the latter is further regularized since, after the common factors are accounted for, the remaining covariance matrix of the idiosyncratic components is still high-dimensional (fan2013POET,Fan2011,fan2018elliptical). This stream of literature, however, focuses on the estimation of a covariance matrix. The accuracy of precision matrices obtained from inverting the factor-based covariance matrix was investigated by ait2017using, but they did not study a high-dimensional case. Factor models are generally treated as competitors to graphical models: as an example, Caner2019 find evidence of superior performance of nodewise-regression estimator of precision matrix over a factor-based estimator POET (fan2013POET) in terms of the out-of-sample Sharpe Ratio and risk of financial portfolio. The root cause why factor models and graphical models are treated separately is the sparsity assumption on the precision matrix made in the latter. Specifically, as pointed out in koike2019biased, when asset returns have common factors, the precision matrix cannot be sparse because all pairs of assets are partially correlated conditional on other assets through the common factors. One attempt to integrate factor modeling and high-dimensional precision estimation was made by fan2018elliptical (Section 5.2): the authors referred to such class of models as “conditional graphical models". However, this was not the main focus of their paper which concentrated on covariance estimation through elliptical factor models. As fan2018elliptical pointed out, “though substantial amount of efforts have been made to understand the graphical model, little has been done for estimating conditional graphical model, which is more general and realistic". Concretely, to the best of our knowledge there are no studies that examine theoretical and empirical performance of graphical models integrated with the factor structure in the context of portfolio allocation.

In this paper we fill this gap and develop a new conditional precision matrix estimator for the excess returns under the approximate factor model that combines the benefits of graphical models and factor structure. We call our algorithm the Factor Graphical Lasso (FGL). We use a factor model to remove the co-movements induced by the factors, and then we apply the Weighted Graphical Lasso for the estimation of the precision matrix of the idiosyncratic terms. We prove consistency of FGL in the spectral and $\ell_{1}$ matrix norms. In addition, we prove consistency of the estimated portfolio weights and risk exposure for three formulations of the optimal portfolio allocation.

Our empirical application uses daily and monthly data for the constituents of the S&P500: we demonstrate that FGL outperforms equal-weighted portfolio, index portfolio, portfolios based on other estimators of precision matrix (Clime, cai2011constrained) and covariance matrix, including POET (fan2013POET) and the shrinkage estimators adjusted to allow for the factor structure (Ledoit2004, ledoit2017nonlinear), in terms of the out-of-sample Sharpe Ratio. Furthermore, we find strong empirical evidence that relaxing the constraint that portfolio weights sum up to one leads to a large increase in the out-of-sample Sharpe Ratio, which, to the best of our knowledge, has not been previously well-studied in the empirical finance literature.

From the theoretical perspective, our paper makes several important contributions to the existing literature on graphical models and factor models. First, to the best of out knowledge, there are no equivalent theoretical results that establish consistency of the portfolio weights and risk exposure in a high-dimensional setting without assuming sparsity on the covariance or precision matrix of stock returns. Second, we extend the theoretical results of POET (fan2013POET) to allow the number of factors to grow with the number of assets. Concretely, we establish uniform consistency for the factors and factor loadings estimated using PCA. Third, we are not aware of any other papers that provide convergence results for estimating a high-dimensional precision matrix using the Weighted Graphical Lasso under the approximate factor model with unobserved factors. Furthermore, all theoretical results established in this paper hold for a wide range of distributions: Sub-Gaussian family (including Gaussian) and elliptical family. Our simulations demonstrate that FGL is robust to very heavy-tailed distributions, which makes our method suitable for the financial applications. Finally, we demonstrate that in contrast to POET, the success of the proposed method does not heavily depend on the factor pervasiveness assumption: FGL is robust to the scenarios when the gap between the diverging and bounded eigenvalues decreases.

This paper is organized as follows: Section 2 reviews the basics of the Markowitz mean-variance portfolio theory. Section 3 provides a brief summary of the graphical models and introduces the Factor Graphical Lasso. Section 4 contains theoretical results and Section 5 validates these results using simulations. Section 6 provides empirical application. Section 7 concludes.

\phantomsection

Notation

\addcontentsline{toc}{section}{Notation} For the convenience of the reader, we summarize the notation to be used throughout the paper. Let $\mathcal{S}_p$ denote the set of all $p \times p$ symmetric matrices, and $\mathcal{S}_{p}^{++}$ denotes the set of all $p \times p$ positive definite matrices. For any matrix ${\mathbf C}$, its $(i,j)$-th element is denoted as $c_{ij}$. Given a vector ${\mathbf u}\in \mathbb{R}^d$ and parameter $a\in \lbrack1,\infty)$, let $\@ifstar{\oldnorm}{\oldnorm*}{{\mathbf u}}_a$ denote $\ell_a$-norm. Given a matrix ${\mathbf U} \in\mathcal{S}_p$, let $\Lambda_{\text{max}}({\mathbf U}) \equiv \Lambda_1({\mathbf U}) \geq \Lambda_2({\mathbf U})\geq \ldots \geq \Lambda_{\text{min}}({\mathbf U}) \equiv \Lambda_p({\mathbf U})$ be the eigenvalues of ${\mathbf U}$, and $\text{eig}_K({\mathbf U}) \in \mathbb{R}^{K\times p}$ denote the first $K\leq p$ normalized eigenvectors corresponding to $\Lambda_1({\mathbf U}), \ldots ,\Lambda_K({\mathbf U})$. Given parameters $a,b\in \lbrack1,\infty)$, let ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert {\mathbf U} \right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{a,b}\equiv \max_{\@ifstar{\oldnorm}{\oldnorm*}{{\mathbf y}}_a=1}\@ifstar{\oldnorm}{\oldnorm*}{{\mathbf U}{\mathbf y}}_{b}$ denote the induced matrix-operator norm. The special cases are ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert {\mathbf U} \right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_1\equiv \max_{1\leq j\leq N}\sum_{i=1}^{N}\@ifstar{\oldabs}{\oldabs*}{u_{i,j}}$ for the $\ell_1/\ell_1$-operator norm; the operator norm ($\ell_2$-matrix norm) ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert {\mathbf U} \right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{2}^{2}\equiv\Lambda_{\text{max}}({\mathbf U}{\mathbf U}')$ is equal to the maximal singular value of ${\mathbf U}$; ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert {\mathbf U} \right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{\infty}\equiv \max_{1\leq j\leq N}\sum_{i=1}^{N}\@ifstar{\oldabs}{\oldabs*}{u_{j,i}}$ for the $\ell_{\infty}/\ell_{\infty}$-operator norm. Finally, $\@ifstar{\oldnorm}{\oldnorm*}{{\mathbf U}}_{\text{max}}\equiv \max_{i,j}\@ifstar{\oldabs}{\oldabs*}{u_{i,j}}$ denotes the element-wise maximum, and ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert {\mathbf U} \right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{F}^{2}\equiv\sum_{i,j}u_{i,j}^{2}$ denotes the Frobenius matrix norm.

Optimal Portfolio Allocation

Suppose we observe $p$ assets (indexed by $i$) over $T$ period of time (indexed by $t$). Let $\widetilde{{\mathbf r}}_t=(\widetilde{r}_{1t}, \widetilde{r}_{2t},\ldots,\widetilde{r}_{pt})' \sim \mathcal{D} ({\mathbf m}, {\bm \Sigma})$ be a $p \times 1$ vector of excess returns drawn from a distribution $\mathcal{D}$, where ${\mathbf m}$ and ${\bm \Sigma}$ are the unconditional mean and covariance matrix of the returns. The goal of the Markowitz theory is to choose asset weights in a portfolio optimally. We will study two optimization problems: the well-known Markowitz weight-constrained (MWC) optimization problem, and the Markowitz risk-constrained (MRC) optimization that relaxes the constraint on portfolio weights.

The first optimization problem searches for asset weights such that the portfolio achieves a desired expected rate of return with minimum risk, under the restriction that all weights sum up to one. This can be formulated as the following quadratic optimization problem:

equation[equation omitted — 189 chars of source]

where ${\mathbf w}$ is a $p \times 1$ vector of asset weights in the portfolio, ${\bm \iota}_p$ is a $p \times 1$ vector of ones, and $\mu$ is a desired expected rate of portfolio return. Let ${\bm \Theta}\equiv{\bm \Sigma}^{-1}$ be the precision matrix.

If ${\mathbf m}'{\mathbf w}>\mu$, then the solution to (ref) yields the global minimum-variance (GMV) portfolio weights ${\mathbf w}_{GMV}$:

equation[equation omitted — 120 chars of source]

If ${\mathbf m}'{\mathbf w}=\mu$, the solution to (ref) is a well-known two-fund separation theorem introduced by Tobin:

align[align omitted — 90 chars of source]

where ${\mathbf w}_{MWC}$ denotes the portfolio allocation with the constraint that the weights need to sum up to one, ${\mathbf w}_{M}=({\bm \iota}_p'{\bm \Theta}{\mathbf m})^{-1}{\bm \Theta}{\mathbf m}$, and $a_1=[\mu({\mathbf m}'{\bm \Theta}{\bm \iota}_p)({\bm \iota}_p'{\bm \Theta}{\bm \iota}_p)-({\mathbf m}'{\bm \Theta}{\bm \iota}_p)^2]/[({\mathbf m}'{\bm \Theta}{\mathbf m})({\bm \iota}_p'{\bm \Theta}{\bm \iota}_p)-({\mathbf m}'{\bm \Theta}{\bm \iota}_p)^2]$.

The MRC problem maximizes Sharpe Ratio (SR) subject to either target risk or target return constraints, but portfolio weights are not required to sum up to one:

equation[equation omitted — 260 chars of source]

When $\mu=\sigma\sqrt{{\mathbf m}'{\bm \Theta}{\mathbf m}}$, the solution to either of the constraints is given by

equation[equation omitted — 130 chars of source]

Equation (ref) tells us that once an investor specifies the desired return, $\mu$, and maximum risk-tolerance level, $\sigma$, the MRC weight maximizes the Sharpe Ratio of the portfolio.

Therefore, we have three alternative portfolio allocations commonly used in the existing literature: GMV in (ref), MWC in (ref) and MRC in (ref). It is clear that all formulations require an estimate of the precision matrix ${\bm \Theta}$.

Factor Graphical Lasso

In this section we introduce a framework for estimating precision matrix for the aforementioned financial portfolios which accounts for the fact that the returns follow approximate factor structure. We examine how to solve the Markowitz mean-variance portfolio allocation problems using factor structure in the returns. We also develop Factor Graphical Lasso Algorithm that uses the estimated common factors to obtain a sparse precision matrix of the idiosyncratic component. The resulting estimator is used to obtain the precision of the asset returns necessary to form portfolio weights.

The arbitrage pricing theory (APT), developed by APTRoss, postulates that the expected returns on securities should be related to their covariance with the common components or factors. The goal of the APT is to model the tendency of asset returns to move together via factor decomposition. Assume that the return generating process ($\widetilde{{\mathbf r}}_t$) follows a $K$-factor model:

align[align omitted — 187 chars of source]

where ${\mathbf f}_t=(f_{1t},\ldots, f_{Kt})'$ are the factors, ${\mathbf B}$ is a $p \times K$ matrix of factor loadings, and ${\bm \varepsilon}_t$ is the idiosyncratic component that cannot be explained by the common factors. Without loss of generality, we assume throughout the paper that unconditional means of factors and idiosyncratic component are zero. Factors in (ref) can be either observable, such as in Fama3Factor,Fama5Factor, or can be estimated using statistical factor models. Unobservable factors and loadings are usually estimated by the principal component analysis (PCA), as studied in Bai2003, Bai2002, Connor1988, and Stock2002.

In this paper our main interest lies in establishing asymptotic properties of the estimators of precision matrix, portfolio weights and risk-exposure for the high-dimensional case. We assume that the number of common factors, $K=K_{p,T}\rightarrow\infty$ as $p\rightarrow\infty$, or $T\rightarrow\infty$, or both $p,T \rightarrow \infty$, but we require that $\max\{K/p,K/T\}\rightarrow0$ as $p,T \rightarrow \infty$.

Our setup is similar to the one studied in fan2013POET: we consider a spiked covariance model when the first $K$ principal eigenvalues of ${\bm \Sigma}$ are growing with $p$, while the remaining $p-K$ eigenvalues are bounded.

Rewrite equation (ref) in matrix form:

equation[equation omitted — 173 chars of source]

where ${\bm \iota}_T$ is a $T\times 1$ vector of ones. We further demean the returns using the sample mean, $\widehat{{\mathbf m}}$, to obtain ${\mathbf R} \equiv \widetilde{{\mathbf R}} - \widehat{{\mathbf m}}{\bm \iota}'_T$. We assume that $\@ifstar{\oldnorm}{\oldnorm*}{\widehat{{\mathbf m}}-{\mathbf m}}_{\text{max}}=\mathcal{O}_P(\sqrt{\log p/T})$, which was proven to hold in CHANG2018 (see their Lemma 1).

Let ${\bm \Sigma}_{\varepsilon}=T^{-1}{\mathbf E}{\mathbf E}'$ and ${\bm \Sigma}_{f}=T^{-1}{\mathbf F}{\mathbf F}'$ be covariance matrices of the idiosyncratic components and factors, and let ${\bm \Theta}_{\varepsilon}={\bm \Sigma}_{\varepsilon}^{-1}$ and ${\bm \Theta}_{f}={\bm \Sigma}_{f}^{-1}$ be their inverses. The factors and loadings in (ref) are estimated by solving the following minimization problem: $(\widehat{{\mathbf B}},\widehat{{\mathbf F}})=\arg\!\min_{{\mathbf B},{\mathbf F}}\@ifstar{\oldnorm}{\oldnorm*}{{\mathbf R}-{\mathbf B}{\mathbf F}}^{2}_{F}$ s.t. $\frac{1}{T}{\mathbf F}{\mathbf F}'={\mathbf I}_K, \ {\mathbf B}'{\mathbf B}\ \text{is diagonal}$. The constraints are needed to identify the factors (fan2018elliptical). It was shown (Stock2002) that $\widehat{{\mathbf F}}=\sqrt{T}\text{eig}_K({\mathbf R}'{\mathbf R})$ and $\widehat{{\mathbf B}}=T^{-1}{\mathbf R}\widehat{{\mathbf F}}'$. Given $\widehat{{\mathbf F}},\widehat{{\mathbf B}}$, define $\widehat{{\mathbf E}}={\mathbf R}-\widehat{{\mathbf B}}\widehat{{\mathbf F}}$. Given a sample of the estimated residuals $\{\widehat{{\bm \varepsilon}}_t={\mathbf r}_t-\widehat{{\mathbf B}}\widehat{{\mathbf f}_t}\}_{t=1}^{T}$ and the estimated factors $\{\widehat{{\mathbf f}}_t\}_{t=1}^{T}$, let $\widehat{{\bm \Sigma}}_{\varepsilon} = (1/T)\sum_{t=1}^{T}\widehat{{\bm \varepsilon}}_t\widehat{{\bm \varepsilon}}_t'$ and $\widehat{{\bm \Sigma}}_{f}=(1/T)\sum_{t=1}^{T}\widehat{{\mathbf f}}_t\widehat{{\mathbf f}}_t'$ be the sample counterparts of the covariance matrices. Since our interest is in constructing portfolio weights, our goal is to estimate a precision matrix of the excess returns ${\bm \Theta}$.

We impose a sparsity assumption on the precision matrix of the idiosyncratic errors, ${\bm \Theta}_{\varepsilon}$, which is obtained using the estimated residuals after removing the co-movements induced by the factors (see Brownlees2018EJS,Brownlees2018JAE,koike2019biased).

Let us elaborate on three reasons justifying the assumption of sparsity on the precision matrix of residuals. First, from the technical viewpoint, this assumption is widely used in high-dimensional settings when $p>T$. Second, a more intuitive rationale for the sparsity assumption on ${\bm \Theta}_{\varepsilon}$ stems from its implication for the structure of corresponding optimal portfolios. Let $r_{t}^{\text{portf}}\equiv \widetilde{{\mathbf r}}_{t}'{\mathbf w}_t$ be the optimal portfolio. Plugging in the definition of $\widetilde{{\mathbf r}}_{t}$ from (ref), we get $r_{t}^{\text{portf}} = ({\mathbf m}+{\bm \varepsilon}_{t})'{\mathbf w}_t + {\mathbf f}_{t}'{\mathbf B}{\mathbf w}_t$. Hence, after hedging factor risk, we can isolate the excess return component only loading on non-factor risk. In this context, since ${\mathbf w}_t$ is a function of ${\bm \Theta}_{\varepsilon}$, imposing sparsity on ${\bm \Theta}_{\varepsilon}$ translates into reducing the contribution of more volatile non-factor risk on the optimal portfolio and thus leading to less sensitive (more robust) investment strategies.

Third, another rationale comes from relatively high “concentration" of S&P 500 Composite Index: as evidenced from \href{https://www.spglobal.com/spdji/en/governance/methodologies/#methodology-information}{SP Global Index methodology} and \href{https://www.slickcharts.com/sp500}{financial data on S&P 500 constituents by weight}, 15 large companies (top 3%) comprise 30% of the total index weights (starting from Apple that has the highest weight of nearly 7%). As the number of firms, $p$, increases, one reasonable assumption is that the number of large firms increases at a rate slower than $p$ (Chudik,gabaix2011granular). This suggests that one could divide the firms into dominant ones and followers. After the effect of common factors is accounted for, dominant firms still have significant idiosyncratic movements that influence other firms and must be taken into account when constructing a portfolio. When it comes to fringe firms (or market followers), idiosyncratic movements are smaller in magnitude and might be less relevant for portfolio allocation purposes. Hence, the network of the idiosyncratic returns is sparse and the sparsity increases with $p$. By imposing sparsity, we only keep relatively large partial correlations among idiosyncratic components: as illustrated in Supplemental Appendix (ref), in our empirical application the estimated number of zeroes in off-diagonal elements of ${\bm \Theta}_{\varepsilon}$ varies over time from 74.5%-98.8%.

Henceforth, having established the need for a sparse precision of errors, we search for a tool that would help us recover its entries. This brings us to consider a family of graphical models, which have evolved from the connection between partial correlations and the entries of an adjacency matrix. The adjacency matrix has zero or one in its entries, with a zero entry indicating that two variables are independent conditional on the rest. The adjacency matrix is sometimes referred to as a “graph". Graphical Lasso procedure (GLASSO) described in Supplemental Appendix (ref) is a representative member of graphical models family: its theoretical and empirical properties have been thoroughly examined in a standard sparse setting (GLASSO, DPGLASSO, Sara2018). One of the goals of our paper is to augment graphical models to non-sparse settings through integrating them with factor modeling. By doing so, graphical models would become adequate for applications in economics and finance.

A common way to induce sparsity is by utilizing Lasso-type penalty. This strategy is used in the Graphical Lasso (GL) together with the objective function based on the Bregman divergence for estimating inverse covariance. The discussion of GL is presented in Supplemental Appendix (ref). We now elaborate on the Bregman divergence class which unifies many commonly used loss functions, including the quasi-likelihood function. Let ${\mathbf W}_{\varepsilon}$ be an estimate of ${\bm \Sigma}_{\varepsilon}$. ravikumar2011 showed that Bregman divergence of the form $\text{trace}({\mathbf W}_{\varepsilon}{\bm \Theta}_{\varepsilon})-\log\det({\bm \Theta}_{\varepsilon})$, known as the log-determinant Bregman function, is suitable to be used as a measure of the quality of constructed sparse approximations of signals such as precision matrices. As pointed out by ravikumar2011, in principle one could use other Bregman divergences including the von Neumann Entropy or the Frobenius divergence which would lead to alternative forms of divergence minimizations for estimating precision matrix. We proceed with the log-determinant Bregman function since (i) it ensures positive definite estimator of precision matrix; (ii) the population optimization problem involves only the population covariance and not its inverse; (iii) the log-determinant divergence gives rise to the likelihood function in the multivariate Gaussian case. At the same time, despite its resemblance with the Gaussian log-likelihood, Bregman divergence was shown to be applicable for non-Gaussian distributions (ravikumar2011). Let $\widehat{{\mathbf D}}_{\varepsilon}^{2}\equiv \textup{diag}({\mathbf W}_{\varepsilon})$. To sparsify entries of precision matrix of the idiosyncratic errors ${\bm \Theta}_{\varepsilon}$, we use the following penalized Bregman divergence with the Weighted Graphical Lasso penalty:

align[align omitted — 357 chars of source]

The subscript $\lambda$ in $\widehat{{\bm \Theta}}_{\varepsilon,\lambda}$ means that the solution of the optimization problem in (ref) will depend upon the choice of the tuning parameter which is discussed below. Section 4 establishes sparsity requirements that guarantee convergence of (ref). In order to simplify notation, we will omit the subscript $\lambda$.

The objective function in (ref) extends the family of linear shrinkage estimators of the first moment to linear shrinkage estimators of the inverse of the second moments. Instead of restricting the number of regressors for estimating conditional mean, equation (ref) restricts the number of edges in a graph by shrinking some off-diagonal entries of precision matrix to zero. Note that shrinkage occurs adaptively with respect to partial covariances.

Let us discuss the choice of the tuning parameter $\lambda$ in (ref). Let $\widehat{{\bm \Theta}}_{\varepsilon,\lambda}$ be the solution to (ref) for a fixed $\lambda$. Following koike2019biased, we minimize the following Bayesian Information Criterion (BIC) using grid search:

equation[equation omitted — 324 chars of source]

The grid $\mathcal{G}\equiv \{\lambda_1,\ldots,\lambda_{M}\}$ is constructed as follows: the maximum value in the grid, $\lambda_{M}$, is set to be the smallest value for which all the off-diagonal entries of $\widehat{{\bm \Theta}}_{\varepsilon,\lambda_{M}}$ are zero. The smallest value of the grid, $\lambda_{1}\in \mathcal{G}$, is determined as $\lambda_{1}\equiv\vartheta\lambda_{M}$ for a constant $0<\vartheta<1$. The remaining grid values $\lambda_1,\ldots,\lambda_{M}$ are constructed in the ascending order from $\lambda_{1}$ to $\lambda_{M}$ on the log scale:

equation*[equation* omitted — 132 chars of source]

We use $\vartheta=\omega_{3T}$ which is defined in Theorem 2 of the next section and $M=10$ in the simulations and the empirical exercise.

Having estimated factors, factor loadings and precision matrix of the idiosyncratic components, we combine them using the Sherman-Morrison-Woodbury formula to estimate the precision matrix of excess returns:

equation[equation omitted — 333 chars of source]

To solve (ref) we use the procedure based on the GL. However, the original algorithm developed by GLASSO is not suitable under the factor structure. Our procedure called Factor Graphical Lasso (FGL), which is summarized in Procedure (ref), augments the standard GL: it starts with estimating factors, loadings (low-rank part) and error terms (sparse part), then it proceeds by recovering sparse precision matrix of the errors using GL, and, finally, low-rank and sparse components are combined through Shermann-Morrison-Woodbury formula in (ref).

spacing{1.6} \begin{algorithm}[H] \floatname{algorithm}{Procedure} \caption{Factor Graphical Lasso} \begin{algorithmic}[1] \STATE (Factor Model) Estimate $\widehat{{\mathbf f}}_t$ and $\widehat{{\mathbf b}}_i$ (Theorem (ref)). Get $\widehat{{\bm \varepsilon}}_t={\mathbf r}_t-\widehat{{\mathbf B}}\widehat{{\mathbf f}_t}$, $\widehat{{\bm \Sigma}}_{\varepsilon}$, $\widehat{{\bm \Sigma}}_f$ and $\widehat{{\bm \Theta}}_f=\widehat{{\bm \Sigma}}_{f}^{-1}$. \STATE (GL) Use GL from GLASSO (see Supplemental Appendix (ref) for more details) to get $\widehat{{\bm \Theta}}_{\varepsilon}$. (Theorem (ref)) \STATE (FGL) Use $\widehat{{\bm \Theta}}_{\varepsilon}$, $\widehat{{\bm \Theta}}_f$ and $\widehat{{\mathbf b}}_i$ from Steps 1-2 to get $\widehat{{\bm \Theta}}$ in Equation (ref). (Theorem (ref)) \STATE Use $\widehat{{\bm \Theta}}$ to get $\widehat{{\mathbf w}}_{\xi}$, $\xi\in\{\text{GMV, MWC, MRC}\}$. (Theorem (ref)) \STATE Use $\widehat{{\bm \Sigma}}=\widehat{{\bm \Theta}}^{-1}$ and $\widehat{{\mathbf w}}_{\xi}$ to get portfolio exposure $\widehat{{\mathbf w}}_{\xi}^{'}\widehat{{\bm \Sigma}}\widehat{{\mathbf w}}_{\xi}$. (Theorem (ref)) \end{algorithmic} \end{algorithm}

The estimator produced by GL in general and FGL in particular is guaranteed to be positive definite. We have verified it in the simulations (Section 5) and the empirical application (Section 6). In Section 4, consistency properties of estimators are established for the factors and loadings (Theorem (ref)), the precision matrix of ${\bm \varepsilon}$ (Theorem (ref)), the precision matrix ${\bm \Theta}$ (Theorem (ref)), portfolio weights (Theorem (ref)), and the portfolio risk exposure (Theorem (ref)). We can use $\widehat{{\bm \Theta}}$ obtained from (ref) using Step 4 of Procedure (ref) to estimate portfolio weights in (ref), (ref) and (ref):

Asymptotic Properties

In this section we first provide a brief review of the terminology used in the literature on graphical models and the approaches to estimate a precision matrix. After that we establish consistency of the Factor Graphical Lasso in Procedure (ref). We also study consistency of the estimators of weights in (ref), (ref) and (ref) and the implications on the out-of sample Sharpe Ratio. Throughout the main text we assume that errors and factors have exponential-type tails ((ref)\red{(c)}). Supplemental Appendix (ref) proves that the conclusions of all theorems studied in Section 4 continue to hold when this assumption is relaxed.

The review of the Gaussian graphical models is based on ESLII and Bishop2006. A graph consists of a set of vertices (nodes) and a set of edges (arcs) that join some pairs of the vertices. In graphical models, each vertex represents a random variable, and the graph visualizes the joint distribution of the entire set of random variables. The edges in a graph are parameterized by potentials (values) that encode the strength of the conditional dependence between the random variables at the corresponding vertices. Sparse graphs have a relatively small number of edges. Among the main challenges in working with the graphical models are choosing the structure of the graph (model selection) and estimation of the edge parameters from the data.

Let $A\in \mathcal{S}_p$. Define the following set for $j=1,\ldots,p$:

align[align omitted — 150 chars of source]

where $d_j(A)$ is the number of edges adjacent to the vertex $j$ (i.e., the degree of vertex $j$), and $d(A)$ measures the maximum vertex degree. Define $S(A)\equiv \bigcup_{j=1}^{p}D_j(A)$ to be the overall off-diagonal sparsity pattern, and $s(A)\equiv \sum_{j=1}^{p}d_j(A)$ is the overall number of edges contained in the graph. Note that $\text{card}(S(A)) \leq s(A)$: when $s(A)=p(p-1)/2$ this would give a fully connected graph.

Assumptions

We now list the assumptions on the model (ref):

enumerate[({A}.1)] • (Spiked covariance model) As $p \rightarrow \infty$, $\Lambda_1({\bm \Sigma})>\Lambda_2({\bm \Sigma})>\ldots>\Lambda_K({\bm \Sigma})\gg \Lambda_{K+1}({\bm \Sigma})\geq \ldots \geq \Lambda_p({\bm \Sigma}) \geq 0$, where $\Lambda_j({\bm \Sigma})=\mathcal{O}(p)$ for $j \leq K$, while the non-spiked eigenvalues are bounded, that is, $c_0 \leq \Lambda_j({\bm \Sigma}) \leq C_0$, $j > K$ for constants $c_0, C_0 > 0$.
enumerate[({A}.2)] • (Pervasive factors) There exists a positive definite $K \times K$ matrix $\breve{{\mathbf B}}$ such that ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert p^{-1}{\mathbf B}'{\mathbf B}-\breve{{\mathbf B}} \right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{2}\rightarrow 0$ and $\Lambda_{\text{min}}(\breve{{\mathbf B}})^{-1}=\mathcal{O}(1)$ as $p \rightarrow \infty$.
enumerate[({A}.3)] • \begin{enumerate}[label=(\alph*)] • $\{{\bm \varepsilon}_t,{\mathbf f}_{t}\}_{t\geq 1}$ is strictly stationary. Also, $\operatorname*{\mathbb{E}}{\varepsilon_{it}}=\operatorname*{\mathbb{E}}{\varepsilon_{it}f_{it}}=0$ $\forall i\leq p$, $j\leq K$ and $t\leq T$. • There are constants $c_1, c_2 >0$ such that $\Lambda_{\text{min}}({\bm \Sigma}_{\varepsilon})>c_1$, ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert {\bm \Sigma}_{\varepsilon} \right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_1<c_2$ and $\text{min}_{i\leq p, j\leq p} \text{var}(\varepsilon_{it}\varepsilon_{jt})>c_1$. • There are $r_1,r_2>0$ and $b_1,b_2>0$ such that for any $s>0$, $i\leq p$, $j\leq K$, \begin{align*} \Pr{(\@ifstar{\oldabs}{\oldabs*}{\varepsilon_{it}}>s)\leq \exp\{-(s/b_1)^{r_1} \}}, \ \Pr{(\@ifstar{\oldabs}{\oldabs*}{f_{jt}}>s)\leq \exp\{-(s/b_2)^{r_2} \}}. \end{align*} \end{enumerate}

We also impose the strong mixing condition. Let $\mathcal{F}_{-\infty}^{0}$ and $\mathcal{F}_{T}^{\infty}$ denote the $\sigma$-algebras that are generated by $\{({\mathbf f}_t,{\bm \varepsilon}_{t}):t\leq 0\}$ and $\{({\mathbf f}_t,{\bm \varepsilon}_{t}):t\geq T\}$ respectively. Define the mixing coefficient

equation[equation omitted — 145 chars of source]
enumerate[({A}.4)] • (Strong mixing) There exists $r_3>0$ such that $3r_{1}^{-1}+1.5r_{2}^{-1}+3r_{3}^{-1}>1$, and $C>0$ satisfying, for all $T\in \mathbb{Z}^{+}$, $\alpha(T)\leq \exp (-CT^{r_3})$.
enumerate[({A}.5)] • (Regularity conditions) There exists $M>0$ such that, for all $i\leq p$, $t\leq T$ and $s\leq T$, such that: \begin{enumerate}[label=(\alph*)] • $\@ifstar{\oldnorm}{\oldnorm*}{{\mathbf b}_i}_{\text{max}}<M$$\operatorname*{\mathbb{E}}{p^{-1/2}\{{\bm \varepsilon}'_{s}{\bm \varepsilon}_t-\operatorname*{\mathbb{E}}{{\bm \varepsilon}'_{s}{\bm \varepsilon}_t}\}}^4<M$ and • $\operatorname*{\mathbb{E}}{\@ifstar{\oldnorm}{\oldnorm*}{p^{-1/2}\sum_{i=1}^{p}{\mathbf b}_i\varepsilon_{it}}^4}<K^2M$. \end{enumerate}

Some comments regarding the aforementioned assumptions are in order. Assumptions (ref)-(ref) are the same as in fan2013POET, and assumption (ref) is modified to account for the increasing number of factors. Assumption (ref) divides the eigenvalues into the diverging and bounded ones. Without loss of generality, we assume that $K$ largest eigenvalues have multiplicity of 1. The assumption of a spiked covariance model is common in the literature on approximate factor models. However, we note that the model studied in this paper can be characterized as a \enquote{very spiked model}. In other words, the gap between the first $K$ eigenvalues and the rest is increasing with $p$. As pointed out by fan2018elliptical, (ref) is typically satisfied by the factor model with pervasive factors, which brings us to Assumption (ref): the factors impact a non-vanishing proportion of individual time-series. Supplemental Appendix (ref) explores the sensitivity of portfolios constructed using FGL when the pervasiveness assumption is relaxed, that is, when the gap between the diverging and bounded eigenvalues decreases. Assumption (ref)\red{(a)} is slightly stronger than in Bai2003, since it requires strict stationarity and non-correlation between $\{{\bm \varepsilon}_{t}\}$ and $\{{\mathbf f}_t\}$ to simplify technical calculations. In (ref)\red{(b)} we require ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert {\bm \Sigma}_{\varepsilon} \right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_1<c_2$ instead of $\lambda_{\textup{max}}({\bm \Sigma}_{\varepsilon})=\mathcal{O}(1)$ to estimate $K$ consistently. When $K$ is known, as in koike2019biased,Fan2011, this condition can be relaxed. (ref)\red{(c)} requires exponential-type tails to apply the large deviation theory to $(1/T)\sum_{t=1}^{T}\varepsilon_{it}\varepsilon_{jt}-\sigma_{\varepsilon,ij}$ and $(1/T)\sum_{t=1}^{T}f_{jt}\varepsilon_{it}$. However, in Supplemental Appendix (ref) we discuss the extension of our results to the setting with elliptical distribution family which is more appropriate for financial applications. Specifically, we discuss the appropriate modifications to the initial estimator of the covariance matrix of returns such that the bounds derived in this paper continue to hold. (ref)-(ref) are technical conditions which are needed to consistently estimate the common factors and loadings. The conditions (ref)\red{(a-b)} are weaker than those in Bai2003 since our goal is to estimate a precision matrix, and (ref)\red{(c)} differs from Bai2003 and Bai2006 in that the number of factors is assumed to slowly grow with $p$.

In addition, the following structural assumption on the population quantities is imposed:

enumerate[({B}.1)] • $\@ifstar{\oldnorm}{\oldnorm*}{{\bm \Sigma}}_{\text{max}}=\mathcal{O}(1)$, $\@ifstar{\oldnorm}{\oldnorm*}{{\mathbf B}}_{\text{max}}=\mathcal{O}(1)$, and $\@ifstar{\oldnorm}{\oldnorm*}{{\mathbf m}}_{\infty}=\mathcal{O}(1)$.

The sparsity of ${\bm \Theta}_{\varepsilon}$ is controlled by the deterministic sequences $s_T$ and $d_T$: $s({\bm \Theta}_{\varepsilon})=\mathcal{O}_{P}(s_T)$ for some sequence $s_T\in (0,\infty), \ T=1,2,\ldots$, and $d({\bm \Theta}_{\varepsilon})=\mathcal{O}_{P}(d_T)$ for some sequence $d_T\in (0,\infty), \ T=1,2,\ldots$. We will impose restrictions on the growth rates of $s_T$ and $d_T$. Note that assumptions on $d_T$ are weaker since they are always satisfied when $s_T=d_T$. However, $d_T$ can generally be smaller than $s_T$. In contrast to fan2013POET we do not impose sparsity on the covariance matrix of the idiosyncratic component. Instead, it is more realistic and relevant for error quantification in portfolio analysis to impose conditional sparsity on the precision matrix after the common factors are accounted for.

The FGL Procedure

Recall the definition of the Weighted Graphical Lasso estimator in (ref) for the precision matrix of the idiosyncratic components. Also, recall that to estimate ${\bm \Theta}$ we used equation (ref). Therefore, in order to obtain the FGL estimator $\widehat{{\bm \Theta}}$ we take the following steps: (1): estimate unknown factors and factor loadings to get an estimator of ${\bm \Sigma}_{\varepsilon}$. (2): use $\widehat{{\bm \Sigma}}_{\varepsilon}$ to get an estimator of ${\bm \Theta}_{\varepsilon}$ in (ref). (3): use $\widehat{{\bm \Theta}}_{\varepsilon}$ together with the estimators of factors and factor loadings from Step 1 to obtain the final precision matrix estimator $\widehat{{\bm \Theta}}$, portfolio weight estimator $\widehat{{\mathbf w}}_{\xi}$, and risk exposure estimator $\widehat{\Phi}_{\xi} = \widehat{{\mathbf w}}'_{\xi}\widehat{{\bm \Theta}}^{-1}\widehat{{\mathbf w}}_{\xi}$ where $\xi\in\{\text{GMV, MWC, MRC}\}$.

Subsection 4.3 examines the theoretical foundations of the first step, and Subsections 4.4-4.5 are devoted to Steps 2 and 3.

Convergence in Estimation of Factors and Loadings

As pointed out in Bai2003 and fan2013POET, $K\times 1$-dimensional factor loadings $\{{\mathbf b}_i\}_{i=1}^{p}$, which are the rows of the factor loadings matrix ${\mathbf B}$, and $K\times 1$-dimensional common factors $\{{\mathbf f}_{t}\}_{t=1}^{T}$, which are the columns of ${\mathbf F}$, are not separately identifiable. Concretely, for any $K \times K$ matrix ${\mathbf H}$ such that ${\mathbf H}'{\mathbf H}={\mathbf I}_K$, ${\mathbf B}{\mathbf f}_t={\mathbf B}{\mathbf H}'{\mathbf H}{\mathbf f}_t$, therefore, we cannot identify the tuple $({\mathbf B}, {\mathbf f}_t)$ from $({\mathbf B}{\mathbf H}',{\mathbf H}{\mathbf f}_{t})$. Let $\widehat{K}\in \{1,\ldots,K_{\text{max}}\}$ denote the estimated number of factors, where $K_{\text{max}}$ is allowed to increase at a slower speed than $\min\{p,T\}$ such that $K_{\text{max}}=o(\min\{p^{1/3},T\})$ (see Li2017_Increasing_Factors for the discussion about the rate).

Define ${\mathbf V}$ to be a $\widehat{K}\times \widehat{K}$ diagonal matrix of the first $\widehat{K}$ largest eigenvalues of the sample covariance matrix in decreasing order. Further, define a $\widehat{K}\times \widehat{K}$ matrix ${\mathbf H}=(1/T){\mathbf V}^{-1}\widehat{{\mathbf F}}'{\mathbf F}{\mathbf B}'{\mathbf B}$. For $t\leq T$, ${\mathbf H}{\mathbf f}_t=T^{-1}{\mathbf V}^{-1}\widehat{{\mathbf F}}'({\mathbf B}{\mathbf f}_{1},\ldots,{\mathbf B}{\mathbf f}_{T})'{\mathbf B}{\mathbf f}_{t}$, which depends only on the data ${\mathbf V}^{-1}\widehat{{\mathbf F}}'$ and an identifiable part of parameters $\{{\mathbf B}{\mathbf f}_{t}\}_{t=1}^{T}$. Hence, ${\mathbf H}{\mathbf f}_{t}$ does not have an identifiability problem regardless of the imposed identifiability condition.

Let $\gamma^{-1}=3r_{1}^{-1}+1.5r_{2}^{-1}+r_{3}^{-1}+1$. The following theorem is an extension of the results in fan2013POET for the case when the number of factors is unknown and is allowed to grow. Proofs of all the theorems are in Supplemental Appendix (ref).

thmSuppose that $K_{\textup{max}}=o(\min\{p^{1/3},T\})$, $K^3\log p=o(T^{\gamma/6})$, $KT=o(p^2)$ and Assumptions (ref)-(ref) and (ref) hold. Let $\omega_{1T}\equiv K^{3/2}\sqrt{\log p/T} +K/\sqrt{p}$ and $\omega_{2T}\equiv K/\sqrt{T}+KT^{1/4}/\sqrt{p}$. Then $\max_{i\leq p} \@ifstar{\oldnorm}{\oldnorm*}{\widehat{{\mathbf b}}_i-{\mathbf H}{\mathbf b}_i}=\mathcal{O}_P(\omega_{1T})$ and $\max_{t\leq T} \@ifstar{\oldnorm}{\oldnorm*}{\widehat{{\mathbf f}}_t-{\mathbf H}{\mathbf f}_t}=\mathcal{O}_P(\omega_{2T})$.

The conditions $K^3\log p=o(T^{\gamma/6})$, $KT=o(p^2)$ are similar to fan2013POET, the difference arises due to the fact that we do not fix $K$, hence, in addition to the factor loadings, there are $KT$ factors to estimate. Therefore, the number of parameters introduced by the unknown growing factors should not be \enquote{too large}, such that we can consistently estimate them uniformly. The growth rate of the number of factors is controlled by $K_{\text{max}}=o(\min\{p^{1/3},T\})$.

The bounds derived in Theorem (ref) help us establish the convergence properties of the estimated idiosyncratic covariance, $\widehat{{\bm \Sigma}}_{\varepsilon}$, and precision matrix $\widehat{{\bm \Theta}}_{\varepsilon}$ which are presented in the next theorem:

thmLet $\omega_{3T}\equiv K^{2}\sqrt{\log p/T} +K^3/\sqrt{p}$. Under the assumptions of Theorem (ref) and with $\lambda \asymp \omega_{3T}$ (where $\lambda$ is the tuning parameter in (ref)), the estimator $\widehat{{\bm \Sigma}}_{\varepsilon}$ obtained by estimating factor model in (ref) satisfies $\@ifstar{\oldnorm}{\oldnorm*}{\widehat{{\bm \Sigma}}_{\varepsilon}-{\bm \Sigma}_{\varepsilon}}_{\textup{max}}=\mathcal{O}_P(\omega_{3T}).$ Let $\varrho_{T}$ be a sequence of positive-valued random variables such that $\varrho_{T}^{-1}\omega_{3T}\xrightarrow{p}0$. If $s_T\varrho_{T}\xrightarrow{\text{p}}0$, then ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \widehat{{\bm \Theta}}_{\varepsilon}-{\bm \Theta}_{\varepsilon} \right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{l}=\mathcal{O}_P(\varrho_{T}s_T)$ as $T \rightarrow \infty$ for any $l \in [1,\infty]$.

Note that the term containing $K^3/\sqrt{p}$ arises due to the need to estimate unknown factors. Fan2011 obtained a similar rate but for the case when factors are observable (in their work, $\omega_{3T}= K^{1/2}\sqrt{\log p/T}$). The second part of Theorem (ref) is based on the relationship between the convergence rates of the estimated covariance and precision matrices established in Sara2018 (Theorem 14.1.3). koike2019biased obtained the convergence rate when factors are observable: the rate obtained in our paper is slower due to the fact that factors need to be estimated (concretely, the rate under observable factors would satisfy $\varrho_{T}^{-1}\sqrt{K\log p /T}\xrightarrow{p}0$ ). We now comment on the optimality of the rate in Theorem (ref): as pointed out in koike2019biased, in the standard Gaussian setting without factor structure, the minimax optimal rate is $d({\bm \Theta}_{\varepsilon})\sqrt{\log p/T}$, which can be faster than the rate obtained in Theorem (ref) if $d({\bm \Theta}_{\varepsilon})<s_T$. Using penalized nodewise regression could help achieve this faster rate. However, our empirical application to the monthly stock returns demonstrated superior performance of the Weighted Graphical Lasso compared to the nodewise regression in terms of the out-of-sample Sharpe Ratio and portfolio risk. Hence, in order not to divert the focus of this paper, we leave the theoretical properties of the nodewise regression for future research.

Convergence in Estimation of Precision Matrix and Portfolio Weights

Having established the convergence properties of $\widehat{{\bm \Sigma}}_{\varepsilon}$ and $\widehat{{\bm \Theta}}_{\varepsilon}$, we now move to the estimation of the precision matrix of the factor-adjusted returns in equation (ref).

thmUnder the assumptions of Theorem (ref), if $d_Ts_T\varrho_{T}\xrightarrow{\text{p}}0$, then ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \widehat{{\bm \Theta}}-{\bm \Theta} \right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{2}=\mathcal{O}_P(\varrho_{T}s_T)$ and ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \widehat{{\bm \Theta}}-{\bm \Theta} \right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{1}=\mathcal{O}_P(\varrho_{T}d_TK^{3/2}s_T )$.

Note that since, by construction, the precision matrix obtained using the Factor Graphical Lasso is symmetric, ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert \widehat{{\bm \Theta}}-{\bm \Theta} \right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{\infty}$ can be trivially obtained from the above theorem.

Using Theorem (ref), we can then establish the consistency of the estimated weights of portfolios based on the Factor Graphical Lasso.

thmUnder the assumptions of Theorem (ref), we additionally assume ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert {\bm \Theta} \right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_2=\mathcal{O}(1)$ (this additional requirement essentially imposes $\Lambda_p({\bm \Sigma})>0$ in (ref)), and $\varrho_{T}d_{T}^{2}s_T=o(1)$. Procedure (ref) consistently estimates portfolio weights in (ref), (ref) and (ref):\\ $\@ifstar{\oldnorm}{\oldnorm*}{\widehat{{\mathbf w}}_{\text{GMV}}-{\mathbf w}_{\text{GMV}}}_1=\mathcal{O}_P\Big(\varrho_{T}d_{T}^2K^{3}s_T\Big)=o_P(1)$, $\@ifstar{\oldnorm}{\oldnorm*}{\widehat{{\mathbf w}}_{\text{MWC}}-{\mathbf w}_{\text{MWC}}}_1=\mathcal{O}_P(\varrho_{T}d_{T}^2K^{3}s_T)=o_P(1)$, and $\@ifstar{\oldnorm}{\oldnorm*}{\widehat{{\mathbf w}}_{\text{MRC}}-{\mathbf w}_{\text{MRC}}}_1= \mathcal{O}_P \Big( d_{T}^{3/2}K^{3}\cdot\lbrack\varrho_{T}s_T\rbrack^{1/2} \Big)=o_P(1)$.

We now comment on the rates in Theorem (ref): first, the rates obtained by Caner2019 for GMV and MWC formulations, when no factor structure of stock returns is assumed, require $s({\bm \Theta})^{3/2}\sqrt{\log p/T}=o_P(1)$, where the authors imposed sparsity on the precision matrix of stock returns, ${\bm \Theta}$. Therefore, if the precision matrix of stock returns is not sparse, portfolio weights can be consistently estimated only if $p$ is less than $T^{1/3}$ (since $(p-1)^{3/2}\sqrt{\log p/T}=o(1)$ is required to ensure consistent estimation of portfolio weights). Our result in Theorem (ref) improves this rate and shows that as long as $d_{T}^{2}s_TK^{3}\sqrt{\log p/T}=o_P(1)$ we can consistently estimate weights of the financial portfolio. Specifically, when the precision of the factor-adjusted returns is sparse, we can consistently estimate portfolio weights when $p>T$ without assuming sparsity on ${\bm \Sigma}$ or ${\bm \Theta}$. Second, note that GMV and MWC weights converge slightly slower than MRC weight. This result is further supported by our simulations presented in the next section.

Implications on Portfolio Risk Exposure

Having examined the properties of portfolio weights, it is natural to comment on the portfolio variance estimation error. It is determined by the errors in two components: the estimated covariance matrix and the estimated portfolio weights. Define $a = {\bm \iota}'_{p}{\bm \Theta}{\bm \iota}_p/p$, $b = {\bm \iota}'_{p}{\bm \Theta}{\mathbf m}/p$, $d = {\mathbf m}'{\bm \Theta}{\mathbf m}/p$, $g = \sqrt{{\mathbf m}'{\bm \Theta}{\mathbf m}}/p$ and $\widehat{a} = {\bm \iota}'_{p}\widehat{{\bm \Theta}}{\bm \iota}_p/p$, $\widehat{b} = {\bm \iota}'_{p}\widehat{{\bm \Theta}}\widehat{{\mathbf m}}/p$, $\widehat{d}=\widehat{{\mathbf m}}'\widehat{{\bm \Theta}}\widehat{{\mathbf m}}/p$, $\widehat{g}=\sqrt{\widehat{{\mathbf m}}'\widehat{{\bm \Theta}}\widehat{{\mathbf m}}}/p$. Define $\Phi_{\text{GMV}} = {\mathbf w}_{GMV}'{\bm \Sigma}{\mathbf w}_{GMV}=(pa)^{-1}$ to be the global minimum variance, $\Phi_{\text{MWC}} = {\mathbf w}_{MWC}'{\bm \Sigma}{\mathbf w}_{MWC}=p^{-1}\Big[\frac{a\mu^2-2b\mu+d}{ad-b^2}\Big]$ is the MWC portfolio variance, and $\Phi_{\text{MRC}} = {\mathbf w}_{MRC}'{\bm \Sigma}{\mathbf w}_{MRC}=\sigma^2(pg)$ is the MRC portfolio variance. We use the terms variance and risk exposure interchangeably. Let $\widehat{\Phi}_{\text{GMV}}$, $\widehat{\Phi}_{\text{MWC}}$, and $\widehat{\Phi}_{\text{MRC}}$ be the sample counterparts of the respective portfolio variances. The expressions for $\Phi_{\text{GMV}}$ and $\Phi_{\text{MWC}}$ were derived in FANLV2008 and Caner2019. Theorem (ref) establishes the consistency of a large portfolio’s variance estimator.

thmUnder the assumptions of Theorem (ref), FGL consistently estimates GMV, MWC, and MRC portfolio variance:\\ $\@ifstar{\oldabs}{\oldabs*}{\widehat{\Phi}_{\text{GMV}}/\Phi_{\text{GMV}} -1 }=\mathcal{O}_P(\varrho_{T}d_Ts_TK^{3/2} )=o_P(1)$,\\ $\@ifstar{\oldabs}{\oldabs*}{\widehat{\Phi}_{\text{MWC}}/\Phi_{\text{MWC}} -1 }=\mathcal{O}_P(\varrho_{T}d_Ts_TK^{3/2} )=o_P(1)$,\\ $\@ifstar{\oldabs}{\oldabs*}{\widehat{\Phi}_{\text{MRC}}/\Phi_{\text{MRC}} -1 }=\mathcal{O}_P\Big(\lbrack\varrho_{T}d_Ts_TK^{3/2}\rbrack^{1/2}\Big)=o_P(1)$.

Caner2019 derived a similar result for $\Phi_{\text{GMV}}$ and $\Phi_{\text{MWC}}$ under the assumption that precision matrix of stock returns is sparse. Also, DING2020 derived the bounds for $\Phi_{\text{GMV}}$ under the factor structure assuming sparse covariance matrix of idiosyncratic components and gross exposure constraint on portfolio weights which limits negative positions.

The empirical application in Section 6 reveals that the portfolios constructed using MRC formulation have higher risk compared with GMV and MWC alternatives: using monthly and daily returns of the components of S&P500 index, MRC portfolios exhibit higher out-of-sample risk and return compared to the alternative formulations. Furthermore, the empirical exercise demonstrates that the higher return of MRC portfolios outweighs higher risk for the monthly data which is evidenced by the increased out-of-sample Sharpe Ratio.

Monte Carlo

In order to validate our theoretical results, we perform several simulation studies which are divided into four parts. The first set of results computes the empirical convergence rates and compares them with the theoretical expressions derived in Theorems (ref)-(ref). The second set of results compares the performance of the FGL with several alternative models for estimating covariance and precision matrix. To highlight the benefit of using the information about factor structure as opposed to standard graphical models, we include Graphical Lasso by GLASSO (GL) that does not account for the factor structure. To explore the benefits of using FGL for error quantification in (ref), we consider several alternative estimators of covariance/precision matrix of the idiosyncratic component in (ref): (1) linear shrinkage estimator of covariance developed by Ledoit2004 further referred to as Factor LW or FLW; (2) nonlinear shrinkage estimator of covariance by ledoit2017nonlinear (Factor NLW or FNLW); (3) POET (fan2013POET); (4) constrained $\ell_{1}$-minimization for inverse matrix estimator, Clime (cai2011constrained) (Factor Clime or FClime). Furthermore, we discovered that in certain setups the estimator of covariance produced by POET is not positive definite. In such cases we use the matrix symmetrization procedure as in fan2018elliptical and then use eigenvalue cleaning as in Callot2017 and Hautsch2012. This estimator is referred to as Projected POET; it coincides with POET when the covariance estimator produced by the latter is positive definite. The third set of results examines the performance of FGL and Robust FGL (described in Supplemental Appendix (ref)) when the dependent variable follows elliptical distribution. The fourth set of results explores the sensitivity of portfolios constructed using different covariance and precision estimators of interest when the pervasiveness assumption (ref) is relaxed, that is, when the gap between the diverging and bounded eigenvalues decreases. All exercises in this section use 100 Monte Carlo simulations.

We consider the following setup: let $p = T^{\delta}$, $\delta = 0.85$, $K = 2(\log T)^{0.5}$ and $T = \lbrack 2^h \rbrack, \ \text{for} \ h=7,7.5,8,\ldots,9.5$. A sparse precision matrix of the idiosyncratic components is constructed as follows: we first generate the adjacency matrix using a random graph structure. Define a $p \times p$ adjacency matrix ${\mathbf A}_{\varepsilon}$ which is used to represent the structure of the graph:

align[align omitted — 140 chars of source]

Let $a_{\varepsilon,ij}$ denote the $i,j$-th element of the adjacency matrix ${\mathbf A}_{\varepsilon}$. We set $a_{\varepsilon,ij} = a_{\varepsilon,ji}=1, \ \text{for} \ i\neq j$ with probability $q$, and $0$ otherwise. Such structure results in $s_T = p(p-1)q/2$ edges in the graph. To control sparsity, we set $q = 1/(pT^{0.8})$, which makes $s_T = \mathcal{O}(T^{0.05})$. The adjacency matrix has all diagonal elements equal to zero. Hence, to obtain a positive definite precision matrix we apply the procedure described in HUGE: using their notation, ${\bm \Theta}_{\varepsilon}={\mathbf A}_{\varepsilon}\cdot v+{\mathbf I}(\@ifstar{\oldabs}{\oldabs*}{\tau}+0.1+u)$, where $u>0$ is a positive number added to the diagonal of the precision matrix to control the magnitude of partial correlations, $v$ controls the magnitude of partial correlations with $u$, and $\tau$ is the smallest eigenvalue of ${\mathbf A}_{\varepsilon}\cdot v$. In our simulations we use $u=0.1$ and $v=0.3$.

Factors are assumed to have the following structure:

align[align omitted — 241 chars of source]

where $m_i \sim \mathcal{N}(1,1)$ independently for each $i=1,\ldots,p$, ${\bm \varepsilon}_{t}$ is a $p \times 1$ random vector of idiosyncratic errors following $\mathcal{N}(\bm{0},{\bm \Sigma}_{\varepsilon})$, with sparse ${\bm \Theta}_{\varepsilon}$ that has a random graph structure described above, ${\mathbf f}_{t}$ is a $K \times 1$ vector of factors, $\phi_f$ is an autoregressive parameter in the factors which is a scalar for simplicity, ${\mathbf B}$ is a $p\times K$ matrix of factor loadings, ${\bm \zeta}_t$ is a $K \times 1$ random vector with each component independently following $\mathcal{N}(0,\sigma^{2}_{\zeta})$. To create ${\mathbf B}$ in (ref) we take the first $K$ rows of an upper triangular matrix from a Cholesky decomposition of the $p \times p$ Toeplitz matrix parameterized by $\rho$. For the first set of results we set $\rho = 0.2$, $\phi_f = 0.2$ and $\sigma^{2}_{\zeta} = 1$. The specification in (ref) leads to the low-rank plus sparse decomposition of the covariance matrix of stock returns ${\mathbf r}_t$.

As a first exercise, we compare the empirical and theoretical convergence rates of the precision matrix, portfolio weights and exposure. A detailed description of the procedure and the simulation results is provided in Supplemental Appendix (ref). We confirm that the empirical rates and theoretical rates from Theorems (ref)-(ref) are matched.

As a second exercise, we compare the performance of FGL with the alternative models listed at the beginning of this section. We consider two cases: Case 1 is the same as for the first set of simulations ($p<T$): $p = T^{\delta}$, $\delta = 0.85$, $K = 2(\log T)^{0.5}$, $s_T = \mathcal{O}(T^{0.05})$. Case 2 captures the cases when $p>T$ with $p = 3\cdot T^{\delta}$, $\delta = 0.85$, all else equal. The results for Case 2 are reported in (ref)-(ref), and Case 1 is located in Supplemental Appendix (ref). FGL demonstrates superior performance for estimating precision matrix and portfolio weights in both cases, exhibiting consistency for both Case 1 and Case 2 settings. Also, FGL outperforms GL for estimating portfolio exposure and consistently estimates the latter, however, depending on the case under consideration some alternative models produce lower averaged error.

As a third exercise, we examine the performance of FGL and Robust FGL (described in Supplemental Appendix (ref)) when the dependent variable follows elliptical distributions. A detailed description of the data generating process (DGP) and simulation results are provided in Supplemental Appendix (ref). We find that the performance of FGL for estimating the precision matrix is comparable with that of Robust FGL: this suggests that our FGL algorithm is robust to heavy-tailed distributions even without additional modifications.

As a final exercise, we explore the sensitivity of portfolios constructed using different covariance and precision estimators of interest when the pervasiveness assumption (ref) is relaxed. A detailed description of the data generating process (DGP) and simulation results are provided in Supplemental Appendix (ref). We verify that FGL exhibits robust performance when the gap between the diverging and bounded eigenvalues decreases. In contrast, POET and Projected POET are most sensitive to relaxing pervasiveness assumption which is consistent with our empirical findings and also with the simulation results by onatski2013POET.

Empirical Application

In this section we examine the performance of the Factor Graphical Lasso for constructing a financial portfolio using daily data. The description and empirical results for monthly data can be found in Supplemental Appendix (ref). We first describe the data and the estimation methodology, then we list four metrics commonly reported in the finance literature, and, finally, we present the results.

Data

We use daily returns of the components of the S&P500 index. The data on historical S&P500 constituents and stock returns is fetched from CRSP and Compustat using SAS interface. For the daily data the full sample size has 5040 observations on 420 stocks from January 20, 2000 - January 31, 2020. We use January 20, 2000 - January 24, 2002 (504 obs) as the first training (estimation) period and January 25, 2002 - January 31, 2020 (4536 obs) as the out-of-sample (OOS) test period. Supplemental Appendix (ref) examines the performance of different competing methods for longer training periods. We roll the estimation window (training periods) over the test sample to rebalance the portfolios monthly. At the end of each month, prior to portfolio construction, we remove stocks with less than 2 years of historical stock return data. The performance of the competing models is compared with the Index -- the composite S&P500 index listed as \textsuperscript{$\wedge$}GSPC. We take the risk-free rate and Fama/French factors from \href{https://mba.tuck.dartmouth.edu/pages/faculty/ken.french/data_library.html}{Kenneth R. French's data library.}

Performance Measures

Similarly to Caner2019, we consider four metrics commonly reported in the finance literature: the Sharpe Ratio, the portfolio turnover, the average return and the risk of a portfolio (which is defined as the square root of the out-of-sample variance of the portfolio). We consider two scenarios: with and without transaction costs. Let $T$ denote the total number of observations, the training sample consists of $m=504$ observations, and the test sample is $n=T-m$.

When transaction costs are not taken into account, the out-of-sample average portfolio return, variance and SR are

align[align omitted — 328 chars of source]

When transaction costs are considered, we follow ban2018machine, Caner2019, Demiguel2009optimal, and Li2015sparse to account for the transaction costs, further denoted as $\textup{tc}$. In line with the aforementioned papers, we set $\textup{tc}=10 \text{bps}$. Define the excess portfolio at time $t+1$ with transaction costs (tc) as

align[align omitted — 221 chars of source]

where

align[align omitted — 114 chars of source]

$r_{t+1,j}+r^{f}_{t+1}$ is sum of the excess return of the $j$-th asset and risk-free rate, and $r_{t+1,\text{portfolio}}+r^{f}_{t+1}$ is the sum of the excess return of the portfolio and risk-free rate. The out-of-sample average portfolio return, variance, Sharpe Ratio and turnover are defined accordingly:

align[align omitted — 417 chars of source]

Description of Empirical Design

In the empirical application for constructing financial portfolio we consider two scenarios, when the factors are unknown and estimated using the standard PCA (statistical factors), and when the factors are known. The number of statistical factors, $\hat{K}$, is estimated in accordance with Remark (ref) in Supplemental Appendix (ref). For the scenario with known factors we include up to 5 Fama-French factors: FF1 includes the excess return on the market, FF3 includes FF1 plus size factor (Small Minus Big, SMB) and value factor (High Minus Low, HML), and FF5 includes FF3 plus profitability factor (Robust Minus Weak, RMW) and risk factor (Conservative Minus Agressive, CMA).\\ We examine the performance of Factor Graphical Lasso for three alternative portfolio allocations (ref), (ref) and (ref) and compare it with the equal-weighted portfolio (EW), index portfolio (Index), FClime, FLW, FNLW (as in the simulations, we use alternative covariance and precision estimators that incorporate the factor structure through Sherman-Morrison inversion formula), POET, Projected POET, and factor models without sparsity restriction on the residual risk (FF1, FF3, and FF5).\\ In (ref) and Supplemental Appendix (ref), we report the daily and monthly portfolio performance for three alternative portfolio allocations in (ref), (ref) and (ref). We consider a relatively risk-averse investor in a sense that they are willing to tolerate no more risk than that incurred by holding the S&P500 Index: the target level of risk for the weight-constrained and risk-constrained Markowitz portfolio (MWC and MRC) is set at $\sigma=0.013$ which is the standard deviation of the daily excess returns of the S&P500 index in the first training set. A return target $\mu=0.0378\%$ which is equivalent to $10\%$ yearly return when compounded. Transaction costs for each individual stock are set to be a constant $0.1\%$. Supplemental Appendix (ref) provides the results for less risk-averse investors that have higher target levels of risk and return for both monthly and daily data.

To compare the relative performance of investment strategies induced by different precision matrix estimators, we use a stepwise multiple testing procedure developed in RomanoWolf2005 and further covered in RomanoWolf2016. Let $\text{SR}^{P} = \mu_{\text{test}}/\sigma_{\text{test}}$ be the population counterpart of the sample Sharpe Ratio defined in (ref). We compare each strategy $s$, $1\leq s \leq S$, with the benchmark (Index) strategy, indexed as $S+1$. Define $\chi_{s} \equiv \text{SR}_{s}^{P} - \text{SR}_{S+1}^{P}$. The test statistic is $\hat{\chi}_{s} \equiv \text{SR}_{s} - \text{SR}_{S+1}$. For a given strategy $s$, we consider the individual testing problem $\mathbb{H}_0: \chi_{s} \leq 0 \quad \text{vs.} \quad \mathbb{H}_A: \chi_{s} > 0$. Using the stepwise multiple testing procedure we aim at identifying as many strategies as possible for which $\chi_{s} > 0$: we relabel the strategies according to the size of the individual test statistics, from largest to smallest, and make the individual decisions in a stepdown manner starting with the null hypothesis that corresponds to the largest test statistic. P-values for competing methods are reported in the tables with empirical results. We note that by construction of the stepwise multiple testing procedure, the resulting p-values are relatively conservative, consistent with Remark 3.1 of RomanoWolf2005.

Empirical Results

This section explores the performance of the Factor Graphical Lasso for the financial portfolio using daily data.

Let us summarize the results for daily data in (ref): (1) MRC portfolios produce higher return and higher risk, compared to MWC and GMV. However, the out-of-sample Sharpe Ratio for MRC is lower than that of MWC and GMV, which implies that the higher risk of MRC portfolios is not fully compensated by the higher return. (2) FGL outperforms all the competitors, including EW and Index. Specifically, our method has the lowest risk and turnover (compared to FClime, FLW, FNLW and POET), and the highest out-of-sample Sharpe Ratio compared with all alternative methods. (3) The implementation of POET for MRC resulted in the erratic behavior of this method for estimating portfolio weights; many entries in the weight matrix had \enquote{NaN} entries. We elaborate on the reasons behind such performance below. (4) Using the observable Fama-French factors in the FGL, in general, produces portfolios with higher return and higher out-of-sample Sharpe Ratio compared to the portfolios based on statistical factors. Interestingly, this increase in return is not followed by higher risk. (5) FGL strongly dominates all factor models that do not impose sparsity on the precision of the idiosyncratic component. The results for monthly data are provided in Supplemental Appendix (ref): all the conclusions are similar to the ones for daily data.

We now examine possible reasons behind the observed puzzling behavior of POET and Projected POET. The erratic behavior of the former is caused by the fact that POET estimator of covariance matrix was not positive-definite which produced poor estimates of GMV and MWC weights and made it infeasible to compute MRC weights (recall, by construction MRC weight in (ref) requires taking a square root). To explore deteriorated behavior of Projected POET, let us highlight two findings outlined by the existing closely related literature. First, bailey2020factorstrength examined “pervasiveness" degree, or strength, of 146 factors commonly used in the empirical finance literature, and found that only the market factor was strong, while all other factors were semi-strong. This indicates that the factor pervasiveness assumption (ref) might be unrealistic in practice. Second, as pointed out by onatski2013POET, “the quality of POET dramatically deteriorates as the systematic-idiosyncratic eigenvalue gap becomes small". Therefore, being guided by the two aforementioned findings, we attribute deteriorated performance of POET and Projected POET to the decreased gap between the diverging and bounded eigenvalues documented in the past studies on financial returns. High sensitivity of these two covariance estimators in such settings was further supported by our additional simulation study (Supplemental Appendix (ref)) examining the robustness of portfolios constructed using different covariance and precision estimators.

(ref) compares the performance of MRC portfolios for the daily data for different time periods of interesting episodes in terms of the cumulative excess return (CER), risk, and SR. To demonstrate the performance of all methods during the periods of recession and expansion, we chose four periods and recorded CER for the whole year in each period of interest. Two years, 2002 and 2008 correspond to the recession periods, which is why we we refer to them as \enquote{Downturns}. We note that the references to Argentine Great Depression and The Financial Crisis do not intend to limit these economic downturns to only one year. They merely provide the context for the recessions. The other two years, 2017 and 2019, correspond to the years which were relatively favorable to the stock market (\enquote{Booms}). Overall, it is easier to beat the Index in Downturns than in Booms. In most cases FGL shows superior performance in terms of CER and SR for Downturn \#1, Boom \#1 and Boom \#2. For Downturn \#2, even though FGL has the highest CER, its SR is smaller than SR of some other competing methods. One explanation would be the following: as evidenced by high risk of the competing methods during Boom \#2, there were high positive and negative returns during the period, with high returns driving up the average used in computing the SR. However, if one were to use the alternative strategies ignoring CER statistics, then the return on the money deposited at the beginning of 2008 would either be negative (e.g. FClime, Projected POET) or smaller than the CER of FGL-based strategies. This exercise demonstrates that SR statistics alone, especially during recession periods characterized by higher volatility, could be misleading. Another interesting finding from such exercise is that FGL exhibits smaller risk compared to most competing methods even during the periods of recession, which holds for all portfolio formulations. This allows FGL to minimize cumulative losses during economic downturns. Subperiod analyses for MWC and GMV portfolio formulations is presented in Supplemental Appendix (ref).

Conclusion

In this paper, we propose a new conditional precision matrix estimator for the excess returns under the approximate factor model with unobserved factors that combines the benefits of graphical models and factor structure. We established consistency of FGL in the spectral and $\ell_{1}$ matrix norms. In addition, we proved consistency of the portfolio weights and risk exposure for three formulations of the optimal portfolio allocation without assuming sparsity on the covariance or precision matrix of stock returns. All theoretical results established in this paper hold for a wide range of distributions: sub-Gaussian family (including Gaussian) and elliptical family. Our simulations demonstrate that FGL is robust to very heavy-tailed distributions, which makes our method suitable for the financial applications. Furthermore, we demonstrate that in contrast to POET and Projected POET, the success of the proposed method does not heavily depend on the factor pervasiveness assumption: FGL is robust to the scenarios when the gap between the diverging and bounded eigenvalues decreases.

The empirical exercise uses the constituents of the S&P500 index and demonstrates superior performance of FGL compared to several alternative models for estimating precision (FClime) and covariance (FLW, FNLW, POET) matrices, Equal-Weighted (EW) portfolio and Index portfolio in terms of the OOS SR and risk. This result is robust to monthly and daily data. We examine three portfolio formulations and discover that the only portfolios that produce positive CER during recessions are the ones that relax the constraint requiring portfolio weights sum up to one.

\cleardoublepage

\phantomsection

\addcontentsline{toc}{section}{References} {14pt} \cleardoublepage

figure[figure omitted — 344 chars of source]
figure[figure omitted — 344 chars of source]
figure[figure omitted — 331 chars of source]
landscape\begin{table}[] \addcontentsline{toc}{section}{Tables} \caption{ {Daily portfolio returns, risk, SR and turnover. In the upper part corresponding to the results w/o transactions costs, p-values are in parentheses. In the lower part corresponding to the results with transaction costs, $^{***}$ indicates p-value $<$ 0.01, $^{**}$ indicates p-value $<$ 0.05, and $^{*}$ indicates p-value $<$ 0.10. In-sample: January 20, 2000 - January 24, 2002 (504 obs), Out-of-sample: January 17, 2002 - January 31, 2020 (4536 obs).}} \resizebox{0.9\textwidth}{!}{ \begin{tabular}{ccccccccccccc} \toprule & \multicolumn{4}{c}{Markowitz Risk-Constrained} & \multicolumn{4}{c}{Markowitz Weight-Constrained} & \multicolumn{4}{c}{Global Minimum-Variance} \\ \midrule & Return & Risk & SR & Turnover & Return & Risk & \textbf{SR} & \textbf{Turnover} & \textbf{Return} & \textbf{Risk} & \textbf{SR} & \textbf{Turnover} \\ \midrule \textbf{Without TC} & & & & & & & & & & & & \\ EW & 2.33E-04 & 1.90E-02 & 0.0123 & - & 2.33E-04 & 1.90E-02 & 0.0123 & - & 2.33E-04 & 1.90E-02 & 0.0123 & - \\ Index & 1.86E-04 & 1.17E-02 & 0.0159 & - & 1.86E-04 & 1.17E-02 & 0.0159 & - & 1.86E-04 & 1.17E-02 & 0.0159 & - \\ FGL & 8.12E-04 & 2.66E-02 & \begin{tabular}[c]{@c@}0.0305\\(0.0579)\end{tabular} & - & 2.95E-04 & 8.21E-03 & \begin{tabular}[c]{@c@}0.0360\\(0.024)\end{tabular} & - & 2.94E-04 & 7.51E-03 & \begin{tabular}[c]{@c@}0.0392\\(0.0279)\end{tabular} & - \\ FClime & 2.15E-03 & 8.46E-02 & \begin{tabular}[c]{@c@}0.0254\\(0.0758)\end{tabular} & - & 2.02E-04 & 9.85E-03 & \begin{tabular}[c]{@c@}0.0205\\(0.0299)\end{tabular} & - & 2.73E-04 & 1.07E-02 & \begin{tabular}[c]{@c@}0.0255\\(0.0419)\end{tabular} & - \\ FLW & 4.34E-04 & 2.65E-02 & \begin{tabular}[c]{@c@}0.0164\\(0.1782)\end{tabular} & - & 3.12E-04 & 9.96E-03 & \begin{tabular}[c]{@c@}0.0313\\(0.024)\end{tabular} & - & 3.10E-04 & 9.38E-03 & \begin{tabular}[c]{@c@}0.0330\\(0.0279)\end{tabular} & - \\ FNLW & 4.91E-04 & 6.66E-02 & \begin{tabular}[c]{@c@}0.0074\\(0.5515 )\end{tabular} & - & 2.98E-04 & 1.24E-02 & \begin{tabular}[c]{@c@}0.0241\\(0.0419)\end{tabular} & - & 3.06E-04 & 1.32E-02 & \begin{tabular}[c]{@c@}0.0231\\(0.0419)\end{tabular} & - \\ POET & NaN & NaN & NaN & - & -7.06E-04 & 2.74E-01 & \begin{tabular}[c]{@c@}-0.0026\\(0.9137)\end{tabular} & - & 1.07E-03 & 2.71E-01 & \begin{tabular}[c]{@c@}0.0039\\(0.7912)\end{tabular} & - \\ Projected POET & 1.20E-03 & 1.71E-01 & \begin{tabular}[c]{@c@}0.0070\\(0.5515)\end{tabular} & - & -8.06E-05 & 1.61E-02 & \begin{tabular}[c]{@c@}-0.0050\\(0.9337)\end{tabular} & - & -7.57E-05 & 1.93E-02 & \begin{tabular}[c]{@c@}-0.0039\\(0.9482)\end{tabular} & - \\ FGL (FF1) & 7.96E-04 & 2.80E-02 & \begin{tabular}[c]{@c@}0.0285\\(0.0758)\end{tabular} & - & 3.73E-04 & 8.73E-03 & \begin{tabular}[c]{@c@}0.0427\\(0.024)\end{tabular} & - & 3.52E-04 & 8.62E-03 & \begin{tabular}[c]{@c@}0.0408\\(0.0259)\end{tabular} & - \\ FGL (FF3) & 6.51E-04 & 2.74E-02 & \begin{tabular}[c]{@c@}0.0238\\(0.0758)\end{tabular} & - & 3.52E-04 & 8.96E-03 & \begin{tabular}[c]{@c@}0.0393\\(0.024)\end{tabular} & - & 3.39E-04 & 8.94E-03 & \begin{tabular}[c]{@c@}0.0379\\(0.022)\end{tabular} & - \\ FGL (FF5) & 5.87E-04 & 2.70E-02 & \begin{tabular}[c]{@c@}0.0217\\(0.0758)\end{tabular} & - & 3.47E-04 & 9.38E-03 & \begin{tabular}[c]{@c@}0.0370\\(0.024)\end{tabular} & - & 3.36E-04 & 9.29E-03 & \begin{tabular}[c]{@c@}0.0362\\(0.022)\end{tabular} & - \\ FF1 & 7.38E-04 & 1.11E-01 & \begin{tabular}[c]{@c@}0.0067\\(0.5821)\end{tabular} & - & 3.30E-05 & 1.62E-02 & \begin{tabular}[c]{@c@}0.0020\\(0.7139)\end{tabular} & - & 2.49E-05 & 1.61E-02 & \begin{tabular}[c]{@c@}0.0015\\(0.8430)\end{tabular} & - \\ FF3 & 7.52E-04 & 1.11E-01 & \begin{tabular}[c]{@c@}0.0068\\(0.5821)\end{tabular} & - & 2.68E-05 & 1.62E-02 & \begin{tabular}[c]{@c@}0.0017\\(0.7139)\end{tabular} & - & 2.06E-05 & 1.61E-02 & \begin{tabular}[c]{@c@}0.0013\\(0.8430)\end{tabular} & - \\ FF5 & 7.59E-04 & 1.11E-01 & \begin{tabular}[c]{@c@}0.0069\\(0.5821)\end{tabular} & - & 2.01E-05 & 1.62E-02 & \begin{tabular}[c]{@c@}0.0012\\(0.7139)\end{tabular} & - & 1.38E-05 & 1.61E-02 & \begin{tabular}[c]{@c@}0.0009\\(0.8430)\end{tabular} & - \\ \midrule \textbf{With TC} & & & & & & & & & & & & \\ EW & 2.01E-04 & 1.90E-02 & 0.0106 & 0.0292 & 2.01E-04 & 1.90E-02 & 0.0106 & 0.0292 & 2.01E-04 & 1.90E-02 & 0.0106 & 0.0292 \\ FGL & 4.47E-04 & 2.66E-02 & 0.0168 & 0.3655 & 2.30E-04 & 8.22E-03 & 0.0280* & 0.0666 & 2.32E-04 & 7.52E-03 & 0.0309* & 0.0633 \\ FClime & 1.18E-03 & 8.48E-02 & 0.0139 & 1.0005 & 1.67E-04 & 9.86E-03 & 0.0170 & 0.0369 & 2.46E-04 & 1.07E-02 & 0.0230* & 0.0290 \\ FLW & -5.54E-05 & 2.65E-02 & -0.0021 & 0.4874 & 1.92E-04 & 9.98E-03 & 0.0193 & 0.1207 & 1.92E-04 & 9.39E-03 & 0.0204* & 0.1194 \\ FNLW & -2.39E-03 & 7.03E-02 & -0.0340 & 3.6370 & 5.50E-05 & 1.25E-02 & 0.0044 & 0.2441 & 6.08E-05 & 1.33E-02 & 0.0046 & 0.2457 \\ POET & NaN & NaN & NaN & NaN & -2.28E-02 & 5.55E-01 & -0.0411 & 113.3848 & -2.81E-02 & 4.21E-01 & -0.0666 & 132.8215 \\ Projected POET & -1.59E-02 & 3.64E-01 & -0.0437 & 35.9692 & -1.03E-03 & 1.68E-02 & -0.0616 & 0.9544 & -1.37E-03 & 2.06E-02 & -0.0666 & 1.2946 \\ FGL (FF1) & 3.86E-04 & 2.80E-02 & 0.0138 & 0.4068 & 2.82E-04 & 8.74E-03 & 0.0323** & 0.0903 & 2.63E-04 & 8.63E-03 & 0.0305* & 0.0887 \\ FGL (FF3) & 2.47E-04 & 2.74E-02 & 0.0090 & 0.4043 & 2.60E-04 & 8.98E-03 & 0.0290** & 0.0928 & 2.49E-04 & 8.96E-03 & 0.0278* & 0.0911 \\ FGL (FF5) & 1.83E-04 & 2.71E-02 & 0.0068 & 0.4032 & 2.53E-04 & 9.40E-03 & 0.0269* & 0.0952 & 2.43E-04 & 9.30E-03 & 0.0262* & 0.0937 \\ FF1 & -6.69E-03 & 1.28E-01 & -0.0639 & 8.5721 & -5.27E-04 & 1.65E-02 & -0.0319 & 0.5704 & -5.30E-04 & 1.64E-02 & -0.0323 & 0.5641 \\ FF3 & -6.65E-03 & 1.28E-01 & -0.0635 & 8.5411 & -5.33E-04 & 1.65E-02 & -0.0323 & 0.5701 & -5.34E-04 & 1.64E-02 & -0.0326 & 0.5638 \\ FF5 & -6.63E-03 & 1.28E-01 & -0.0634 & 8.5262 & -5.40E-04 & 1.65E-02 & -0.0327 & 0.5703 & -5.41E-04 & 1.64E-02 & -0.0330 & 0.5646 \\ \bottomrule \end{tabular} } \end{table}
sidewaystable[ph!] \caption{Cumulative excess return (CER) and risk of MRC portfolios using daily data. Targeted risk is set at $\sigma=0.013$, daily targeted return is $0.0378\%$. P-values are in parentheses. In-sample: January 20, 2000 - January 24, 2002 (504 obs), Out-of-sample: January 17, 2002 - January 31, 2020 (4536 obs).} \resizebox{\textwidth}{!}{ \begin{tabular}{clllcccccccccc} \toprule & EW & Index & FGL & FClime & FLW & FNLW & ProjPOET & FGL(FF1) & FGL(FF3) & FGL(FF5) & FF1 & FF3 & FF5 \\ \midrule \multicolumn{7}{l}{Downturn \#1: Argentine Great Depression (2002)} & \multicolumn{1}{l} & \multicolumn{1}{l} & \multicolumn{1}{l} & \multicolumn{1}{l} & \multicolumn{1}{l} & \multicolumn{1}{l} & \multicolumn{1}{l} \\ CER & -0.1633 & -0.2418 & 0.2909 & -0.0079 & 0.0308 & 0.0728 & -0.6178 & 0.3375 & 0.3423 & 0.3401 & -0.0860 & -0.0860 & -0.0860 \\ Risk & 0.0160 & 0.0168 & 0.0206 & 0.0348 & 0.0231 & 0.0213 & 0.0545 & 0.0211 & 0.0211 & 0.0212 & 0.0495 & 0.0495 & 0.0495 \\ SR & -0.0393 & -0.0615 & \begin{tabular}[c]{@l@}0.0629\\(0.0619)\end{tabular} & \begin{tabular}[c]{@c@}0.0164\\(0.0759)\end{tabular} & \begin{tabular}[c]{@c@}0.0171\\(0.0759)\end{tabular} & \begin{tabular}[c]{@c@}0.0246\\(0.0759)\end{tabular} & \begin{tabular}[c]{@c@}-0.0467\\(0.4852)\end{tabular} & \begin{tabular}[c]{@c@}0.0689\\(0.0619)\end{tabular} & \begin{tabular}[c]{@c@}0.0696\\(0.0619)\end{tabular} & \begin{tabular}[c]{@c@}0.0692\\(0.0619)\end{tabular} & \begin{tabular}[c]{@c@}0.0169\\(0.0759)\end{tabular} & \begin{tabular}[c]{@c@}0.0169\\(0.0759)\end{tabular} & \begin{tabular}[c]{@c@}0.0169\\(0.0759)\end{tabular} \\ \midrule \multicolumn{7}{l}{Downturn \#2: Financial Crisis (2008)} & \multicolumn{1}{l} & \multicolumn{1}{l} & \multicolumn{1}{l} & \multicolumn{1}{l} & \multicolumn{1}{l} & \multicolumn{1}{l} & \multicolumn{1}{l} \\ CER & -0.5622 & -0.4746 & 0.2938 & -0.8912 & 0.2885 & 0.2075 & -0.9999 & 0.2665 & 0.2650 & 0.2560 & 0.0404 & 0.0404 & 0.0404 \\ Risk & 0.0310 & 0.0258 & 0.0282 & 0.1484 & 0.0315 & 0.0392 & 0.1963 & 0.0320 & 0.0319 & 0.0319 & 0.0986 & 0.0986 & 0.0986 \\ SR & -0.0857 & -0.0857 & \begin{tabular}[c]{@l@}0.0315\\(0.0889)\end{tabular} & \begin{tabular}[c]{@c@}0.1045\\(0.1079)\end{tabular} & \begin{tabular}[c]{@c@}0.0282\\(0.1079)\end{tabular} & \begin{tabular}[c]{@c@}0.0392\\(0.1079)\end{tabular} & \begin{tabular}[c]{@c@}0.1963\\(0.1079)\end{tabular} & \begin{tabular}[c]{@c@}0.0320\\(0.0889)\end{tabular} & \begin{tabular}[c]{@c@}0.0319\\(0.0889)\end{tabular} & \begin{tabular}[c]{@c@}0.0319\\(0.0889)\end{tabular} & \begin{tabular}[c]{@c@}0.0006\\(0.1079)\end{tabular} & \begin{tabular}[c]{@c@}0.0006\\(0.1079)\end{tabular} & \begin{tabular}[c]{@c@}0.0006\\(0.1079)\end{tabular} \\ \midrule \multicolumn{7}{l}{Boom \#1 (2017)} & \multicolumn{1}{l} & \multicolumn{1}{l} & \multicolumn{1}{l} & \multicolumn{1}{l} & \multicolumn{1}{l} & \multicolumn{1}{l} & \multicolumn{1}{l} \\ CER & 0.0627 & 0.1752 & 0.7267 & 0.5331 & 0.3164 & 0.5796 & -0.7599 & 0.6568 & 0.6607 & 0.6486 & -0.5070 & -0.5294 & -0.4755 \\ Risk & 0.0218 & 0.0042 & 0.0142 & 0.0383 & 0.0118 & 0.0497 & 0.1197 & 0.0135 & 0.0134 & 0.0132 & 0.0720 & 0.0721 & 0.0710 \\ SR & 0.0220 & 0.1536 & \begin{tabular}[c]{@l@}0.1606\\(0.5544)\end{tabular} & \begin{tabular}[c]{@c@}0.1231\\(0.6264)\end{tabular} & \begin{tabular}[c]{@c@}0.0987\\(0.6264)\end{tabular} & \begin{tabular}[c]{@c@}0.1008\\(0.5465)\end{tabular} & \begin{tabular}[c]{@c@}0.0151\\(0.9815)\end{tabular} & \begin{tabular}[c]{@c@}0.1563\\(0.5455)\end{tabular} & \begin{tabular}[c]{@c@}0.1581\\(0.5455)\end{tabular} & \begin{tabular}[c]{@c@}0.1575\\(0.5455)\end{tabular} & \begin{tabular}[c]{@c@}-0.0022\\(0.9985)\end{tabular} & \begin{tabular}[c]{@c@}-0.0046\\(0.9985)\end{tabular} & \begin{tabular}[c]{@c@}0.0002\\(0.9815)\end{tabular} \\ \midrule \multicolumn{7}{l}{Boom \#2 (2019)} & \multicolumn{1}{l} & \multicolumn{1}{l} & \multicolumn{1}{l} & \multicolumn{1}{l} & \multicolumn{1}{l} & \multicolumn{1}{l} & \multicolumn{1}{l} \\ CER & 0.1642 & 0.2934 & 0.6872 & 0.2346 & 0.5520 & 0.9315 & 1.8592 & 0.5166 & 0.5168 & 0.5037 & 0.2690 & 0.2682 & 0.2730 \\ Risk & 0.0185 & 0.0086 & 0.0263 & 0.0557 & 0.0287 & 0.0355 & 0.1177 & 0.0247 & 0.0248 & 0.0248 & 0.1094 & 0.1094 & 0.1094 \\ SR & 0.0418 & 0.1228 & \begin{tabular}[c]{@l@}0.0919\\(0.1738)\end{tabular} & \begin{tabular}[c]{@c@}0.0436\\(0.2298)\end{tabular} & \begin{tabular}[c]{@c@}0.0753\\(0.2298)\end{tabular} & \begin{tabular}[c]{@c@}0.0905\\(0.2298)\end{tabular} & \begin{tabular}[c]{@c@}0.0898\\(0.2298)\end{tabular} & \begin{tabular}[c]{@c@}0.0793\\(0.1728)\end{tabular} & \begin{tabular}[c]{@c@}0.0792\\(0.1728)\end{tabular} & \begin{tabular}[c]{@c@}0.0779\\(0.1728)\end{tabular} & \begin{tabular}[c]{@c@}0.0798\\(0.2298)\end{tabular} & \begin{tabular}[c]{@c@}0.0798\\(0.2298)\end{tabular} & \begin{tabular}[c]{@c@}0.0799\\(0.2298)\end{tabular} \\ \bottomrule \end{tabular} }