Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
117,094 characters · 21 sections · 87 citation commands
A Basket Half Full: Sparse Portfolios
\thispagestyle{empty}
{20pt} \setcounter{page}{1}
The search for the optimal portfolio weights reduces to the questions (i) which stocks to buy and (ii) how much to invest in these stocks. Depending on the strategy used to address the first question, the existing allocation approaches can be further broken down into the ones that invest in all available stocks, and the ones that select a subset out of the stock universe. The latter is referred to as a sparse portfolio, since some assets will be excluded and get a zero weight leading to sparse wealth allocations. Any portfolio optimization problem requires the inverse covariance matrix, or precision matrix, of excess stock returns as an input. In the era of big data, a search for the optimal portfolio becomes a high-dimensional problem: the number of assets, $p$, is comparable to or greater than the sample size, $T$. Constructing non-sparse portfolios in high dimensions has been the main focus of the existing research on asset management for a long time. In particular, many papers focus on developing an improved covariance or precision estimator to achieve desirable statistical properties of portfolio weights. In contrast, the literature on constructing sparse portfolio is scarce: it is limited to a low-dimensional framework and lacks theoretical analysis of the resulting sparse allocations. In this paper we fill this gap and propose a novel approach to construct sparse portfolios in high dimensions. We obtain the oracle bounds of sparse weight estimators and provide guidance regarding their distribution. From the empirical perspective, we examine the merit of sparse portfolios during the periods of economic growth, moderate market decline and severe economic downturns. We find that in contrast to non-sparse counterparts, our strategy is robust to recessions and can be used as a hedging vehicle during such times.
As pointed out above, estimating high-dimensional covariance or precision matrix to improve portfolio performance of non-sparse strategies has received a lot of attention in the existing literature. Ledoit2004,Ledoit2017 developed linear and non-linear shrinkage estimators of covariance matrix, fan2013POET,fan2018elliptical introduced a covariance matrix estimator when stock returns are driven by common factors under the assumption of a spiked covariance model. Once the covariance estimator is obtained, it is then inverted to get a precision matrix, the main input to any portfolio optimization problem. A parallel stream of literature has focused on estimating precision matrix directly, that is, avoiding the inversion step that leads to additional estimation errors, especially in high dimensions. GLASSO developed an iterative algorithm that estimates the entries of precision matrix column-wise using penalized Gaussian log-likelihood (Graphical Lasso); meinshausen2006 used the relationship between regression coefficients and the entries of precision matrix to estimate the elements of the latter column by column (nodewise regression). cai2011constrained use constrained $\ell_1$-minimization for inverse matrix estimation (CLIME). Caner2019 examined the performance of high-dimensional portfolios constructed using covariance and precision estimators and found that precision-based models outperform covariance-based counterparts in terms of the out-of-sample (OOS) Sharpe Ratio and portfolio return.
From a practical perspective, apart from enjoying favorable statistical properties a successful wealth allocation strategy should be easy to maintain and monitor and it should be robust to economic downturns such that investors could use it as a hedging vehicle. Having this motivation in mind, we chose several popular covariance and precision-based estimators to construct non-sparse portfolios and explore their performance during the recent COVID-19 outbreak. Using daily returns of 495 constituents of the S&P500 from May 25, 2018 -- September 24, 2020 (588 obs.), Table (ref) reports the performance of the selected strategies: we included equal-weighted (EW) and Index portfolios, as well as precision-based nodewise regression estimator by meinshausen2006 (motivated by the recent application of this statistical technique to portfolio studied in Caner2019), linear shrinkage covariance estimator by Ledoit2004 and CLIME by cai2011constrained. We use May 25, 2018 -- October 23, 2018 (105 obs.) as a training period and October 24, 2018 -- September 24, 2020 (483 obs.) as the out-of-sample test period. We roll the estimation window over the test sample to rebalance the portfolios monthly. The left panel of Table (ref) shows return, risk and Sharpe Ratio of portfolios over the training period, and the right panel reports cumulative excess return (CER) and risk over two sub-periods of interest: before the pandemic (January 2, 2019 -- December 31, 2019) and during the first wave of COVID-19 outbreak in the US (January 2, 2020 -- June 30, 2020). As evidenced by Table (ref), none of the portfolios was robust to the downturn brought by pandemic and yielded negative CER. We noticed that similar pattern pertained in several other historic episodes of mild and severe downturns, such as the Global Financial Crisis (GFC) of 2007-09.\footnote{Please see the Empirical Application section for more details.}
Studies that examine the relationship between portfolio performance and the number of stock holdings are scarce. Tidmore2019 used active US equity funds' quarterly data from January 2000 to December 2017 from Morningstar, Inc. to study the impact of concentration (measured by the number of holdings) on fund excess returns: they found that the effect was significant and fluctuated considerably over time. Notably, the relationship became negative in the period preceding and including the GFC. This indicates that holding sparse portfolios might be the key to hedging during downturns. To support this hypothesis, we further compare the performance of sparse vs non-sparse strategies in terms of utility gain to investors. Suppose we observe $i=1,\ldots,p$ excess returns over $t=1,\ldots,T$ period of time: ${\mathbf r}_t=(r_{1t},\ldots,r_{pt})' \ {\sim}\ \mathcal{D} ({\mathbf m}, {\bm \Sigma})$. Consider the following mean-variance utility problem: $\text{min}_{{\mathbf w}} \ -U \equiv \frac{\gamma}{2}{\mathbf w}{\bm \Sigma}{\mathbf w} - {\mathbf w}'{\mathbf m}, \ \text{s.t.} \ {\mathbf w}'{\bm \iota} = 1, \ \@ifstar{\oldabs}{\oldabs*}{\text{supp}({\mathbf w})}\leq \bar{p}, \ \bar{p} \leq p$, where ${\mathbf w}$ is a $p\times 1$ vector of portfolio weights, $\text{supp}({\mathbf w})=\{i:w_i>0\}$ is the cardinality constraint that controls sparsity, and $\gamma$ determines the risk of an investor under the assumption of a normal distribution. When $\bar{p}=p$ the portfolio is non-sparse and the respective utility is denoted as $U^{\text{Non-Sparse}}$, while when $\bar{p}<p$ the utility of such sparse portfolio is denoted as $U^{\text{Sparse}}$. Figure (ref) reports the ratio of utilities using monthly data from 2003:04 to 2009:12 on the constituents of the S&P100 as a function of $\bar{p}$: we set $\gamma=3$ and vary $\bar{p}=\{5,10,15,20,30,\ldots,90\}$\footnote{Since the optimization problem with a cardinality constraint is not convex, we find a solution using Lagrangian relaxation procedure of Shaw2008}. Our test sample includes two periods of particular interest: before the GFC (2004:01-2006:12) and during the GFC (2007:01-2009:12) As evidenced from Figure (ref): (1) for both time periods there exists a lower-dimensional subset of stocks which brings greater utility compared to non-sparse portfolios; (2) the number of stocks minimizing the ratio of utilities is smaller during the GFC compared to the period preceding it. Both findings are consistent with the empirical result of Tidmore2019 that including more stocks does not guarantee better performance and suggesting that holding a “basket half full" instead can help achieve superior performance even in stressed market scenarios.
In order to create a sparse portfolio, that is, a portfolio with many zero entries in the weight vector, we can use an $\ell_1$-penalty (Lasso) on the portfolio weights which shrinks some of them to zero (see Fan2019Sparse, Ao2019, Li2015sparse, Brodie2009sparse among others). caccioli2016liquidity proved the mathematical equivalence of adding an $\ell_1$-penalty and controlling transaction costs associated with the bid-ask spread impact of single and sequential trades executed in a very short time. This indicates another advantage of sparse portfolios: market liquidity dries up during economic downturns which increases bid-ask spreads, a measure of liquidity costs. Henceforth, regularizing portfolio positions accounts for the increased liquidity risk associated with acquiring and liquidating positions. The existing literature on sparse wealth allocations is scarce and has several drawbacks: (1) it is limited to low-dimensional setup when $p<T$, whereas sparsity becomes especially important in high-dimensional scenarios; (2) it lacks theoretical analysis of sparse wealth allocations and their impact on portfolio exposure; (3) the use of an $\ell_1$-penalty produces biased estimates (see Zhang2014,Javanmard2014confidence,Javanmard2014hypothesis,Buhlmann2014,Belloni2015uniform,Javanmard2018debiasing among others), however, this issue has been overlooked in the context of portfolio allocation. This paper addresses the aforementioned drawbacks and develops an approach to construct sparse portfolios in high dimensions. Our contribution is twofold: from the theoretical perspective, we establish the oracle bounds of sparse weight estimators and provide guidance regarding their distribution. From the empirical perspective, we examine the merit of sparse portfolios during different market scenarios. We find that in contrast to non-sparse counterparts, our strategy is robust to recessions and can be used as a hedging vehicle during such times. To illustrate, the last two rows of Table (ref) show the performance of two sparse strategies proposed in this paper: both approaches outperform non-sparse counterparts in terms of total OOS Sharpe Ratio, and they produce positive CER during the pandemic, as well as in the period preceding it. Figure (ref) shows the stocks selected by post-Lasso in August, 2019 and in May, 2020: the colors serve as a visual guide to identify groups of closely-related stocks (stocks of the same color do not necessarily correspond to the same sector). Our framework makes use of the tool from the network theory called nodewise regression which not only satisfies desirable statistical properties, but also allows us to study whether certain industries could serve as safe havens during recessions. We find that such non-cyclical industries as consumer staples, healthcare, retail and food were driving the returns of the sparse portfolios during both GFC and COVID-19 outbreak, whereas insurance sector was the least attractive investment in both periods.
This paper is organized as follows: Section 2 introduces sparse de-biased portfolio and sparse portfolio using post-Lasso. Section 3 develops a new high-dimensional precision estimator called Factor Nodewise regression. Section 4 develops a framework for factor investing. Section 5 contains theoretical results and Section 6 validates these results using simulations. Section 7 provides empirical application. Section 8 concludes. \phantomsection
\addcontentsline{toc}{section}{Notation} For the convenience of the reader, we summarize the notation to be used throughout the paper. Let $\mathcal{S}_p$ denote the set of all $p \times p$ symmetric matrices. For any matrix ${\mathbf C}$, its $(i,j)$-th element is denoted as $c_{ij}$. Given a vector ${\mathbf u}\in \mathbb{R}^d$ and parameter $a\in \lbrack1,\infty)$, let $\@ifstar{\oldnorm}{\oldnorm*}{{\mathbf u}}_a$ denote $\ell_a$-norm. Given a matrix ${\mathbf U} \in\mathcal{S}_p$, let $\Lambda_{\text{max}}({\mathbf U}) \equiv \Lambda_1({\mathbf U}) \geq \Lambda_2({\mathbf U})\geq \ldots \Lambda_{\text{min}}({\mathbf U}) \equiv \Lambda_p({\mathbf U})$ be the eigenvalues of ${\mathbf U}$, and $\text{eig}_K({\mathbf U}) \in \mathbb{R}^{K\times p}$ denote the first $K\leq p$ normalized eigenvectors corresponding to $\Lambda_1({\mathbf U}), \ldots \Lambda_K({\mathbf U})$. Given parameters $a,b\in \lbrack1,\infty)$, let ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert {\mathbf U} \right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{a,b}=\max_{\@ifstar{\oldnorm}{\oldnorm*}{{\mathbf y}}_a=1}\@ifstar{\oldnorm}{\oldnorm*}{{\mathbf U}{\mathbf y}}_{b}$ denote the induced matrix-operator norm. The special cases are ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert {\mathbf U} \right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_1\equiv \max_{1\leq j\leq p}\sum_{i=1}^{p}\@ifstar{\oldabs}{\oldabs*}{u_{i,j}}$ for the $\ell_1/\ell_1$-operator norm; the operator norm ($\ell_2$-matrix norm) ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert {\mathbf U} \right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{2}^{2}\equiv\Lambda_{\text{max}}({\mathbf U}{\mathbf U}')$ is equal to the maximal singular value of ${\mathbf U}$; ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert {\mathbf U} \right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{\infty}\equiv \max_{1\leq j\leq p}\sum_{i=1}^{p}\@ifstar{\oldabs}{\oldabs*}{u_{j,i}}$ for the $\ell_{\infty}/\ell_{\infty}$-operator norm. Finally, $\@ifstar{\oldnorm}{\oldnorm*}{{\mathbf U}}_{\text{max}}=\max_{i,j}\@ifstar{\oldabs}{\oldabs*}{u_{i,j}}$ denotes the element-wise maximum, and ${\left\vert\kern-0.25ex\left\vert\kern-0.25ex\left\vert {\mathbf U} \right\vert\kern-0.25ex\right\vert\kern-0.25ex\right\vert}_{F}^{2}=\sum_{i,j}u_{i,j}^{2}$ denotes the Frobenius matrix norm. We also use the following notations: $a\vee b=\max\{a,b\}$, and $a\wedge b=\min\{a,b\}$. For an event $A$, we say that $A \ \text{wp} \rightarrow 1$ when $A$ occurs with probability approaching $1$ as $T$ increases.
There exist several widely used portfolio weight formulations depending on the type of optimization problem solved by an investor. Suppose we observe $p$ assets (indexed by $i$) over $T$ period of time (indexed by $t$). Let ${\mathbf r}_t=(r_{1t}, r_{2t},\ldots,r_{pt})' \sim \mathcal{D} ({\mathbf m}, {\bm \Sigma})$ be a $p \times 1$ vector of excess returns drawn from a distribution $\mathcal{D}$, where ${\mathbf m}$ and ${\bm \Sigma}$ are unconditional mean and covariance of excess returns, and $\mathcal{D}$ belongs to either sub-Gaussian or elliptical families. When $\mathcal{D} = \mathcal{N}$, the precision matrix ${\bm \Sigma}^{-1}\equiv {\bm \Theta}$ contains information about conditional dependence between the variables. For instance, if $\theta_{ij}$, which is the $ij$-th element of the precision matrix, is zero, then the variables $i$ and $j$ are conditionally independent, given the other variables. The goal of the Markowitz theory is to choose assets weights in a portfolio optimally. We will study two criteria of optimality: the first is a well-known Markowitz weight-constrained optimization problem, and the second formulation relaxes constraints on portfolio weights.
The first optimization problem, which will be referred to as Markowitz weight-constrained problem (MWC), searches for assets weights such that the portfolio achieves a desired expected rate of return with minimum risk, under the restriction that all weights sum up to one. The aforementioned goal can be formulated as the following quadratic optimization problem:
where ${\mathbf w}$ is a $p \times 1$ vector of assets weights in the portfolio, ${\bm \iota}$ is a $p \times 1$ vector of ones, and $\mu$ is a desired expected rate of portfolio return. The constraint in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{sys1} \endgroup requires portfolio weights to sum up to one - this assumption can be easily relaxed and we will demonstrate the implications of this constraint on portfolio weights.
If ${\mathbf m}'{\mathbf w}>\mu$, then the solution to \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{sys1} \endgroup yields the global minimum-variance (GMV) portfolio weights ${\mathbf w}_{GMV}$:
If ${\mathbf m}'{\mathbf w}=\mu$, the solution to \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{sys1} \endgroup is
where ${\mathbf w}_{MWC}$ denotes the portfolio allocation with the constraint that the weights need to sum up to one and ${\mathbf w}_{M}$ captures all mean-related market information.
The second optimization problem, which will be referred to as Markowitz risk-constrained (MRC) problem, has the same objective as in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{sys1} \endgroup , but portfolio weights are not required to sum up to one:
It can be easily shown that the solution to \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{sys2} \endgroup is:
Alternatively, instead of searching for a portfolio with a specified desired expected rate of return and minimum risk, one can maximize expected portfolio return given a maximum risk-tolerance level:
In this case, the solution to \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{sys3} \endgroup yields:
To get the second equality in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{w2} \endgroup we used the definition of $\mu$ from \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{sys1} \endgroup and \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{sys2} \endgroup . It follows that if $\mu=\sigma\sqrt{\theta}$, where $\theta\equiv {\mathbf m}'{\bm \Theta}{\mathbf m}$ is the squared Sharpe Ratio, then the solution to \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{sys2} \endgroup and \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{sys3} \endgroup admits the following expression:
where ${\bm \alpha}\equiv {\bm \Theta}{\mathbf m}$. Equation \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{ee20} \endgroup tells us that once an investor specifies the desired return, $\mu$, and maximum risk-tolerance level, $\sigma$, this pins down the Sharpe Ratio of the portfolio which makes the optimization problems of minimizing risk and maximizing expected return of the portfolio in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{sys2} \endgroup and \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{sys3} \endgroup identical.
This brings us to three alternative portfolio allocations commonly used in the existing literature: Global Minimum-Variance Portfolio in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{eq2} \endgroup , weight-constrained Markowitz Mean-Variance in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{eq3} \endgroup and maximum-risk-constrained Markowitz Mean-Variance in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{ee20} \endgroup . Below we summarize the aforementioned portfolio weight expressions:
So far we have considered allocation strategies that put non-zero weights to all assets in the financial portfolio. As an implication, an investor needs to buy a certain amount of each security even if there are a lot of small weights. However, oftentimes investors are interested in managing a few assets which significantly reduces monitoring and transaction costs and was shown to outperform equal weighted and index portfolios in terms of the Sharpe Ratio and cumulative return (see Fan2019Sparse, Ao2019, Li2015sparse, Brodie2009sparse among others). This strategy is based on holding a sparse portfolio, that is, a portfolio with many zero entries in the weight vector.
Let us first introduce some notations. The sample mean and sample covariance matrix have standard formulas: $ \widehat{{\mathbf m}}=\dfrac{1}{T}\sum_{t=1}^{T}{\mathbf r}_{t}$ and $\widehat{{\bm \Sigma}}=\dfrac{1}{T}\sum_{t=1}^{T}({\mathbf r}_t-\widehat{{\mathbf m}})({\mathbf r}_t-\widehat{{\mathbf m}})^{'}$. Our empirical application shows that risk-constrained Markowitz allocation in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{e158} \endgroup outperforms GMV and MWC portfolios in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{e156} \endgroup - \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{e157} \endgroup . Therefore, we first study sparse MRC portfolios. Our goal is to construct a sparse vector of portfolio weights given by \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{e158} \endgroup . To achieve this we use the following equivalent and unconstrained regression representation of the mean-variance optimization in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{sys2} \endgroup and \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{sys3} \endgroup :
The sample counterpart of \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{ee65} \endgroup is written as:
Ao2019 prove that the weight allocation from \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{ee65} \endgroup is equivalent to \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{e158} \endgroup . The sparsity is introduced through Lasso which yields the following constrained optimization problem:
Now we propose two extensions to the setup \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{ee67} \endgroup . First, the estimator ${\mathbf w}_{\text{MRC, SPARSE}}$ is infeasible since $\theta$ used for constructing $y$ is unknown. Ao2019 construct an estimator of $\theta$ under normally distributed excess returns, assuming $p/T \rightarrow \rho \in (0,1)$ and the sample size $T$ is required to be larger than the number of assets $p$. Their paper uses an unbiased estimator proposed in Kan2007optimal: $\hat{\theta}=((T-p-2)\widehat{{\mathbf m}}'\widehat{{\bm \Sigma}}^{-1}\widehat{{\mathbf m}}-p)/T$, where $\widehat{{\mathbf m}}$ and $\widehat{{\bm \Sigma}}^{-1}$ are sample mean and inverse of the sample covariance matrix respectively. One of the limitations of the model studied by Ao2019 is that it cannot handle high dimensions. In both simulations and empirical application the maximum number of stocks used by the authors is limited to 100. Another limitation of Ao2019 approach is that they do not correct the bias introduced by imposing $\ell_1$-constraint in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{ee67} \endgroup . However, it is well-known that the estimator in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{ee67} \endgroup is biased and the existing literature proposes several de-biasing techniques (see Zhang2014,Javanmard2014confidence,Javanmard2014hypothesis,Buhlmann2014,Belloni2015uniform,Javanmard2018debiasing among others).
To address the first aforementioned limitation, we propose to use an estimator of a high-dimensional precision matrix discussed in the next section. The suggested estimator is appropriate for high-dimensional settings, it can handle cases when the sample size is less than the number of assets, and it is always non-negative by construction\footnote{Our empirical results suggest that the unbiased estimator $\hat{\theta}=((T-p-2)\widehat{{\mathbf m}}'\widehat{{\bm \Sigma}}^{-1}\widehat{{\mathbf m}}-p)/T$ is oftentimes negative even after using the adjusted estimator defined in Kan2007optimal (p. 2906).}. Consequently, the estimator of $y$ is
To approach the second limitation, motivated by Buhlmann2014, we propose the de-biasing technique that uses the nodewise regression estimator of the precision matrix. First, let ${\mathbf R}$ be a $T \times p$ matrix of excess returns stacked over time and $ \widehat{{\mathbf y}}$ be a $T \times 1$ constant vector. Consider a high-dimensional linear model
We study high-dimensional framework $p\geq T$ and in the asymptotic results we require $\log p/T=o(1)$. Let us rewrite \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{ee67} \endgroup :
The estimator in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{ee67} \endgroup satisfies the following KKT conditions:
where $\widehat{{\mathbf g}}$ is a $p \times 1$ vector arising from the subgradient of $\@ifstar{\oldnorm}{\oldnorm*}{{\mathbf w}}_{1}$. Let $\widehat{{\bm \Sigma}}={\mathbf R}'{\mathbf R}/T$, then we can rewrite the KKT conditions:
Multiply both sides of \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{ee72} \endgroup by $\widehat{{\bm \Theta}}$ obtained from Algorithm (ref), add and subtract $(\widehat{{\mathbf w}}-{\mathbf w})$, and rearrange the terms:
In the section with the theoretical results we show that $\Delta$ is asymptotically negligible under certain sparsity assumptions\footnote{Note that we cannot directly apply Theorem 2.2 of Buhlmann2014 since ${\mathbf r}_c$ needs to be estimated and we first need to show consistency of the respective estimator.}. Combining \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{ee70} \endgroup and \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{ee73} \endgroup brings us to the de-biased estimator of portfolio weights:
The properties of the proposed de-biased estimator are examined in Section 5.
One of the drawbacks of the de-biased portfolio weights in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{ee74} \endgroup is that the weight formula is tailored to a specific portfolio choice that maximizes an unconstrained Sharpe Ratio (i.e. MRC in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{e158} \endgroup ). However, it is desirable to accommodate preferences of different types of investors who might be interested in weight allocations corresponding to GMV \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{e156} \endgroup or MWC \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{e157} \endgroup portfolios. At the same time, we are willing to stay within the framework of sparse allocations. One of the difficulties that precludes us from pursuing a similar technique as in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{ee67} \endgroup is the fact that once the weight constraint is added, the optimization problem in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{ee67} \endgroup has two solutions depending on whether ${\bm \iota}'{\bm \Theta}{\mathbf m}$ is positive or negative. As shown in maller2003new, when ${\bm \iota}'{\bm \Theta}{\mathbf m} <0$, the minimum value cannot be achieved exactly for a specified portfolio allocation that satisfies the full investment constraint. Hence, one can design an approximate solution to approach the supremum as closely as desired.
To overcome this difficulty, we propose to use Lasso regression in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{ee69} \endgroup for selecting a subset of stocks, and then constructing a financial portfolio using any of the weight formulations in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{e156} \endgroup - \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{e158} \endgroup . The procedure to estimate sparse portfolio using post-Lasso is described in Algorithm (ref).
In this section we first review a nodewise regression (meinshausen2006), a popular approach to estimate a precision matrix. After that we propose a novel estimator which accounts for the common factors in the excess returns.
In the high-dimensional settings it is necessary to regularize the precision matrix, which means that some of the entries $\theta_{ij}$ will be zero. In other words, to achieve consistent estimation of the inverse covariance, the estimated precision matrix should be sparse.
One of the approaches to induce sparsity in the estimation of precision matrix is to solve for $\widehat{{\bm \Theta}}$ one column at a time via linear regressions, replacing population moments by their sample counterparts. When we repeat this procedure for each variable $j=1,\ldots,p$, we will estimate the elements of $\widehat{{\bm \Theta}}$ column by column using $\{{\mathbf r}_t\}_{t=1}^{T}$ via $p$ linear regressions. meinshausen2006 use this approach to incorporate sparsity into the estimation of the precision matrix. They fit $p$ separate Lasso regressions using each variable (node) as the response and the others as predictors to estimate $\widehat{{\bm \Theta}}$. This method is known as the \enquote{nodewise} regression and it is reviewed below based on Buhlmann2014 and Caner2019.
Let ${\mathbf r}_j$ be a $T \times 1$ vector of observations for the $j$-th regressor, the remaining covariates are collected in a $T \times (p-1)$ matrix ${\mathbf R}_{-j}$. For each $j=1,\ldots,p$ we run the following Lasso regressions:
where $\widehat{{\bm \gamma}}_j=\{\widehat{\gamma}_{j,k}; j=1,\ldots,p, k\neq j\}$ is a $(p-1)\times 1$ vector of the estimated regression coefficients that will be used to construct the estimate of the precision matrix, $\widehat{{\bm \Theta}}$. Define
For $j=1,\ldots,p$, define
and write
The approximate inverse is defined as
The procedure to estimate the precision matrix using nodewise regression is summarized in Algorithm (ref).
One of the caveats to keep in mind when using the nodewise regression method is that the estimator in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{e25} \endgroup is not self-adjoint. Caner2019 show (see their Lemma A.1) that $\widehat{{\bm \Theta}}$ in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{e25} \endgroup is positive definite with high probability, however, it could still occur that $\widehat{{\bm \Theta}}$ is not positive definite in finite samples. To resolve this issue we use the matrix symmetrization procedure as in fan2018elliptical and then use eigenvalue cleaning as in Callot2017 and Hautsch2012. First, the symmetric matrix is constructed as
where $\widehat{\theta}_{ij}$ is the $(i,j)$-th element of the estimated precision matrix from \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{e25} \endgroup . Second, we use eigenvalue cleaning to make $\widehat{{\bm \Theta}}^{s}$ positive definite: write the spectral decomposition $\widehat{{\bm \Theta}}^{s}=\widehat{{\mathbf V}}'\widehat{{\bm \Lambda}}\widehat{{\mathbf V}}$, where $\widehat{{\mathbf V}}$ is a matrix of eigenvectors and $\widehat{{\bm \Lambda}}$ is a diagonal matrix with $p$ eigenvalues $\widehat{{\bm \Lambda}}_{i}$ on its diagonal. Let ${\bm \Lambda}_{m}\equiv\min\{\widehat{{\bm \Lambda}}_{i}|\widehat{{\bm \Lambda}}_{i}>0\}$. We replace all $\widehat{{\bm \Lambda}}_{i}<{\bm \Lambda}_{m}$ with ${\bm \Lambda}_{m}$ and define the diagonal matrix with cleaned eigenvalues as $\widetilde{{\bm \Lambda}}$. We use $\widetilde{{\bm \Theta}}=\widehat{{\mathbf V}}'\widetilde{{\bm \Lambda}}\widehat{{\mathbf V}}$ which is symmetric and positive definite.
The arbitrage pricing theory (APT), developed by APTRoss, postulates that expected returns on securities should be related to their covariance with the common components or factors only. The goal of the APT is to model the tendency of asset returns to move together via factor decomposition. Assume that the return generating process (${\mathbf r}_t$) follows a $K$-factor model:
where ${\mathbf f}_t=(f_{1t},\ldots, f_{Kt})'$ are the factors, ${\mathbf B}$ is a $p \times K$ matrix of factor loadings, and ${\bm \varepsilon}_t$ is the idiosyncratic component that cannot be explained by the common factors. Factors in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{equ1} \endgroup can be either observable, such as in Fama3Factor,Fama5Factor, or can be estimated using statistical factor models.
In this subsection we examine how to approach the portfolio allocation problems in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{e156} \endgroup - \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{e158} \endgroup using a factor structure in the returns. Our approach, called Factor Nodewise Regression, uses the estimated common factors to obtain sparse precision matrix of the idiosyncratic component. The resulting estimator is used to obtain the precision of the asset returns necessary to form portfolio weights.
As in fan2013POET, we consider a spiked covariance model when the first $K$ principal eigenvalues of ${\bm \Sigma}$ are growing with $p$, while the remaining $p-K$ eigenvalues are bounded and grow slower than $p$.
Rewrite equation \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{equ1} \endgroup in matrix form:
Let ${\bm \Sigma}=T^{-1}{\mathbf R}{\mathbf R}'$, ${\bm \Sigma}_{\varepsilon}=T^{-1}{\mathbf E}{\mathbf E}'$ and ${\bm \Sigma}_{f}=T^{-1}{\mathbf F}{\mathbf F}'$ be covariance matrices of stock returns, idiosyncratic components and factors, and let ${\bm \Theta}={\bm \Sigma}^{-1}$, ${\bm \Theta}_{\varepsilon}={\bm \Sigma}_{\varepsilon}^{-1}$ and ${\bm \Theta}_{f}={\bm \Sigma}_{f}^{-1}$ be their inverses. The factors and loadings in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{5.2n} \endgroup are estimated by solving $ (\widehat{{\mathbf B}},\widehat{{\mathbf F}})=\arg\!\min_{{\mathbf B},{\mathbf F}}\@ifstar{\oldnorm}{\oldnorm*}{{\mathbf R}-{\mathbf B}{\mathbf F}}^{2}_{F}$ s.t. $\frac{1}{T}{\mathbf F}{\mathbf F}'={\mathbf I}_K, \ {\mathbf B}'{\mathbf B}\ \text{is diagonal}$. The constraints are needed to identify the factors (fan2018elliptical). It was shown (Stock2002) that $\widehat{{\mathbf F}}=\sqrt{T}\text{eig}_K({\mathbf R}'{\mathbf R})$ and $\widehat{{\mathbf B}}=T^{-1}{\mathbf R}\widehat{{\mathbf F}}'$. Given $\widehat{{\mathbf F}},\widehat{{\mathbf B}}$, define $\widehat{{\mathbf E}}={\mathbf R}-\widehat{{\mathbf B}}\widehat{{\mathbf F}}$.
Since our interest is in constructing portfolio weights, our goal is to estimate a precision matrix of the excess returns. However, as pointed out by koike2019biased, when common factors are present across the excess returns, the precision matrix cannot be sparse because all pairs of the returns are partially correlated given other excess returns through the common factors. Therefore, we impose a sparsity assumption on the precision matrix of the idiosyncratic errors, ${\bm \Theta}_{\varepsilon}$, which is obtained using the estimated residuals after removing the co-movements induced by the factors (see Brownlees2018EJS,Brownlees2018JAE,koike2019biased).
We use the nodewise regression as a shrinkage technique to estimate the precision matrix of residuals. Once the precision ${\bm \Theta}_{f}$ of the low-rank component is also obtained, similarly to Fan2011, we use the Sherman-Morrison-Woodbury formula to estimate the precision of excess returns:
To obtain $\widehat{{\bm \Theta}}_{f}=\widehat{{\bm \Sigma}}_{f}^{-1}$, we use the inverse of the sample covariance of the estimated factors $\widehat{{\bm \Sigma}}_{f}=T^{-1}\widehat{{\mathbf F}}\widehat{{\mathbf F}}'$. To get $\widehat{{\bm \Theta}}_{\varepsilon}$, we apply Algorithm (ref) to the estimated idiosyncratic errors, $\widehat{{\bm \varepsilon}}_t$. Once we have estimated $\widehat{{\bm \Theta}}_{f}$ and $\widehat{{\bm \Theta}}_{\varepsilon}$, we can get $\widehat{{\bm \Theta}}$ using a sample analogue of \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{equa18} \endgroup . The proposed procedure is called Factor Nodewise Regression and is summarized in Algorithm (ref).
Algorithm (ref) involves a tuning parameter $\lambda_j$ in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{e21a} \endgroup : we choose shrinkage intensity by minimizing the generalized information criterion (GIC). Let $\@ifstar{\oldabs}{\oldabs*}{\widehat{S}_j(\lambda_j)}$ denote the estimated number of nonzero parameters in the vector $\widehat{{\bm \gamma}}_j$:
We can use $\widehat{{\bm \Theta}}$ obtained in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{3.11} \endgroup to estimate $y$ in equation \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{e2.17} \endgroup and obtain sparse portfolio weights in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{ee74} \endgroup and Algorithm (ref).
In this section we allow an investor to hold a portfolio of assets and factors, in other words, factors are assumed to be tradable. Note that in contrast with Ao2019, the distinction between tradable and non-tradable factors is not pinned down by the fact that the excess returns are driven by the common factors. That is, factor structure of returns is allowed independently of whether factors are tradable or not. We assume that only observable factors can be tradable. Denote a $K_1 \times 1$ vector of observable factors as $\widetilde{{\mathbf f}}_t$, and $K_2 \times 1$ vector of unobservable factors as ${\mathbf f}^{PCA}_t$, where $K1+K2=K$. The goal of factor investing is to decide how much weight is allocated to factors $\widetilde{{\mathbf f}}_t$ and stocks ${\mathbf r}_t$. Let $r_{t,all}$ be the return of portfolio at time $t$:
where ${\mathbf x}_t=(\widetilde{{\mathbf f}}'_t,{\mathbf r}'_t)'$ is a $(p+K_1)\times 1$ vector of excess returns of observable factors and stocks and ${\mathbf w}_{all,t}=({\mathbf w}'_{ft}, {\mathbf w}'_t)'$ is a vector of weights with ${\mathbf w}_{ft}$ invested in $\widetilde{{\mathbf f}}_t$ and ${\mathbf w}_t$ invested in stocks. We treat $\widetilde{{\mathbf f}}_t$ as additional $K_1$ investments vehicles which will contribute to the return of the total portfolio. Now consider $K_2$-factor model for ${\mathbf x}_t$:
Rewrite equation \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{e6.2} \endgroup in matrix form:
which can be estimated using the standard PCA techniques as in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{5.2n} \endgroup :\\ $\widehat{{\mathbf F}}^{PCA}=\sqrt{T}\text{eig}_{K_2}({\mathbf X}'{\mathbf X})$ and $\widehat{{\mathbf B}}=T^{-1}{\mathbf X}\widehat{{\mathbf F}}^{'PCA}$. Given $\widehat{{\mathbf F}}^{PCA},\widehat{{\mathbf B}}$, define $\widehat{{\mathbf E}}={\mathbf X}-\widehat{{\mathbf B}}\widehat{{\mathbf F}}^{PCA}$.
Similarly to Algorithm (ref), we use \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{equa18} \endgroup to estimate the precision of the augmented excess returns, ${\bm \Theta}_x$. To get $\widehat{{\bm \Theta}}_{f^{PCA}}=\widehat{{\bm \Sigma}}_{f^{PCA}}^{-1}$, we use the inverse of the sample covariance of the estimated factors $\widehat{{\bm \Sigma}}_{f^{PCA}}=T^{-1}\widehat{{\mathbf F}}^{PCA}\widehat{{\mathbf F}}^{'PCA}$. To get $\widehat{{\bm \Theta}}_{e}$, we first apply Algorithm (ref) to the estimated idiosyncratic errors, $\widehat{{\mathbf e}}_t$ in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{e6.2} \endgroup . Once we have estimated $\widehat{{\bm \Theta}}_{f^{PCA}}$ and $\widehat{{\bm \Theta}}_{e}$, we can get $\widehat{{\bm \Theta}}_x$ using a sample analogue of \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{equa18} \endgroup . This procedure is summarized in Algorithm (ref).
We can use $\widehat{{\bm \Theta}}_x$ obtained from Algorithm (ref) to estimate portfolio weights ${\mathbf w}_{all,t}$ using either a de-biased technique from section 2.1 ( \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{ee74} \endgroup ), or post-Lasso (Algorithm (ref)). Once we obtain $\widehat{{\mathbf w}}_{all,t}=(\widehat{{\mathbf w}}'_{ft}, \widehat{{\mathbf w}}'_t)'$, we can test whether factor investing significantly contributes to the portfolio return by testing whether ${\mathbf w}_{ft}=0$.
In this section we study asymptotic properties of the de-biased estimator of weights for sparse portfolio in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{ee74} \endgroup and post-Lasso estimator from Algorithm (ref).
Denote $S_0\equiv \{ j; {\mathbf w}_j\neq 0 \}$ to be the active set of variables, where ${\mathbf w}$ is a vector of true portfolio weights in equation \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{e4.9} \endgroup . Also, let $s_0\equiv \@ifstar{\oldabs}{\oldabs*}{S_0}$. Further, let $S_j\equiv \{ k; \gamma_{j,k}\neq 0 \}$ be the active set for row ${\bm \gamma}_{j}$ for the nodewise regression in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{e21} \endgroup , and let $s_j\equiv \@ifstar{\oldabs}{\oldabs*}{S_j}$. Define $\bar{s}\equiv \max_{1\leq j\leq p}s_j$.
Consider a factor model from equation \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{equ1} \endgroup :
We study the case when the factors are not known, i.e. the only observable variable in equation \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{e5.1} \endgroup is the excess returns ${\mathbf r}_t$. In this paper our main interest lies in establishing asymptotic properties of sparse portfolio weights and the out-of-sample Sharpe Ratio for the high-dimensional case. We assume that the number of common factors, $K$, is fixed.
We now list the assumptions on the model \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{e5.1} \endgroup :
Similarly to CHANG2018 and Caner2019, we also impose beta mixing condition.
Some comments regarding the aforementioned assumptions are in order. Assumptions (ref)-(ref) are the same as in fan2018elliptical, and assumption (ref) is required to consistently estimate precision matrix for de-biasing portfolio weights. Assumption (ref) divides the eigenvalues into the diverging and bounded ones. Without loss of generality, we assume that $K$ largest eigenvalues have multiplicity of 1. The assumption of a spiked covariance model is common in the literature on approximate factor models, however, we note that the model studied in this paper can be characterized as a \enquote{very spiked model}. In other words, the gap between the first $K$ eigenvalues and the rest is increasing with $p$. As pointed out by fan2018elliptical, (ref) is typically satisfied by the factor model with pervasive factors, which brings us to the assumption (ref): the factors impact a non-vanishing proportion of individual time-series. Assumption (ref) allows for weak dependence in the residuals of the factor model in (ref): causal ARMA processes, certain stationary Markov chains and stationary GARCH models with finite second moments satisfy this assumption. We note that our Assumption (ref) is much weaker than in Caner2019, the latter requires weak dependence of the returns series, whereas we only restrict dependence of the idiosyncratic components.
Let ${\bm \Sigma}={\bm \Gamma}{\bm \Lambda}{\bm \Gamma}^{'}$, where ${\bm \Sigma}$ is the covariance matrix of returns that follow factor structure described in equation \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{e5.1} \endgroup . Define $\widehat{{\bm \Sigma}}, \widehat{{\bm \Lambda}}_K,\widehat{{\bm \Gamma}}_K$ to be the estimators of ${\bm \Sigma},{\bm \Lambda},{\bm \Gamma}$. We further let $\widehat{{\bm \Lambda}}_K=\text{diag}(\hat{\lambda}_1,\ldots,\hat{\lambda}_K)$ and $\widehat{{\bm \Gamma}}_K=(\hat{v}_1,\ldots,\hat{v}_K)$ to be constructed by the first $K$ leading empirical eigenvalues and the corresponding eigenvectors of $\widehat{{\bm \Sigma}}$ and $\widehat{{\mathbf B}}\widehat{{\mathbf B}}'=\widehat{{\bm \Gamma}}_K\widehat{{\bm \Lambda}}_K\widehat{{\bm \Gamma}}_{K}^{'}$. Similarly to fan2018elliptical, we require the following bounds on the componentwise maximums of the estimators:
Let $\widehat{{\bm \Sigma}}^{SG}$ be the sample covariance matrix, with $\widehat{{\bm \Lambda}}_{K}^{SG}$ and $\widehat{{\bm \Gamma}}_{K}^{SG}$ constructed with the first $K$ leading empirical eigenvalues and eigenvectors of $\widehat{{\bm \Sigma}}^{SG}$ respectively. Also, let $\widehat{{\bm \Sigma}}^{EL1} = \widehat{{\mathbf D}}\widehat{{\mathbf R}}_1\widehat{{\mathbf D}}$, where $\widehat{{\mathbf R}}_1$ is obtained using the Kendall's tau correlation coefficients and $\widehat{{\mathbf D}}$ is a robust estimator of variances constructed using the Huber loss. Furthermore, let $\widehat{{\bm \Sigma}}^{EL2} = \widehat{{\mathbf D}}\widehat{{\mathbf R}}_2\widehat{{\mathbf D}}$, where $\widehat{{\mathbf R}}_2$ is obtained using the spatial Kendall's tau estimator. Define $\widehat{{\bm \Lambda}}_{K}^{EL}$ to be the matrix of the first $K$ leading empirical eigenvalues of $\widehat{{\bm \Sigma}}^{EL1}$, and $\widehat{{\bm \Gamma}}_{K}^{EL}$ is the matrix of the first $K$ leading empirical eigenvectors of $\widehat{{\bm \Sigma}}^{EL2}$. For more details regarding constructing $\widehat{{\bm \Sigma}}^{SG}$, $\widehat{{\bm \Sigma}}^{EL1}$ and $\widehat{{\bm \Sigma}}^{EL2}$ see fan2018elliptical, Sections 3 and 4.
Theorem (ref) is essentially a rephrasing of the results obtained in fan2018elliptical, Sections 3 and 4. Since there is no separate statement of these results in their paper (it is rather a summary of several theorems), we separated it as a Theorem for the convenience of the reader. As evidenced from the above Theorem, $\widehat{{\bm \Sigma}}^{EL2}$ is only used for estimating the eigenvectors. This is necessary due to the fact that, in contrast with $\widehat{{\bm \Sigma}}^{EL2}$, the theoretical properties of the eigenvectors of $\widehat{{\bm \Sigma}}^{EL}$ are mathematically involved because of the sin function.
In addition, the following structural assumption on the model is imposed:
which is a natural assumption on the population quantities.
In contrast to fan2018elliptical, instead of estimating and inverting covariance matrix, we focus on obtaining precision matrix directly since it is the ultimate input to any portfolio optimization problem.
Recall that we used equation \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{equa18} \endgroup to estimate ${\bm \Theta}$. Therefore, in order to establish consistency of the estimator in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{equa18} \endgroup , we first show consistency of $\widehat{{\bm \Theta}}_{\varepsilon}$. Proofs of all the theorems are in Appendix.
Some comments are in order. First, the sparsity assumption $\bar{s}^2\omega_{T}=o(1)$ is stronger than that required for convergence of $\widehat{{\bm \Theta}}_{\varepsilon}$: this is necessary to ensure consistency for $\widehat{{\bm \Theta}}$ established in Theorem (ref), so we impose a stronger assumption at the beginning. We also note that at the first glance, our sparsity assumption in Theorem (ref) is stronger than that required by Buhlmann2014 and Caner2019, however, recall that we impose sparsity on ${\bm \Theta}_{\varepsilon}$, not ${\bm \Theta}$ as opposed to the two aforementioned papers. Hence, this assumption can be easily satisfied once the common factors have been accounted for and the precision of the idiosyncratic components is expected to be sparse. The bounds derived in Theorem (ref) help us establish the convergence properties of the precision matrix of stock returns in equation \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{equa18} \endgroup .
Using Theorem (ref) we can then establish the consistency of the non-sparse counterpart of the estimated MRC portfolio weight in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{ee69} \endgroup .
Note that the rate in Theorem (ref) depends on the sparsity of ${\bm \Theta}_{\varepsilon}$. If, instead, sparsity on ${\bm \Theta}$ is imposed, the rate becomes similar to the one derived by Caner2019: $\bar{s}({\bm \Theta})^{3/2}\omega_{T}^{1/2}=o_P(1)$, where $\bar{s}({\bm \Theta})$ is the maximum vertex degree of ${\bm \Theta}$. In their case, if the precision matrix of stock returns is not sparse, consistent estimation of portfolio weights is possible if $(p-1)^{3/2}(\sqrt{\log p/T}+1/\sqrt{p})=o(1)$. However, this excludes high-dimensional cases since $p$ is required to be less than $T^{1/3}$.
We now proceed to examining the properties of sparse MRC portfolio weights for de-biased portfolio, as summarized by the following Theorem:
Some comments are in order. Our Theorem (ref) is an extension of Theorem 2.4 of Buhlmann2014 for non-iid case, where the latter is achieved with a help of CHANG2018. Furthermore, there are several fundamental differences between Theorem (ref) and Theorem 2.4 of Buhlmann2014: first, we apply nodewise regression to estimate sparse precision matrix of factor-adjusted returns, which explains the difference in convergence rates. Concretely, Buhlmann2014 have $\omega_{T}=\sqrt{\log p/T}$, whereas we have $\omega_{T}=\sqrt{\log p/T}+1/\sqrt{p}$, where $1/\sqrt{p}$ arises due to the fact that factors need to be estimated. However, we note that since we deal with high-dimensional regime $p\geq T$, this additional term is asymptotically negligible, we only keep it for identification purposes. Second, in contrast with Buhlmann2014, the dependent variable in the Lasso regression in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{ee69} \endgroup is unknown and needs to be estimated. Lemma (ref) shows that $\widehat{y}$ constructed using the precision matrix estimator from Theorem (ref) is consistent and shares the same rate as the $\ell_1$-bound in Theorem (ref). Third, interestingly, the sparsity assumption on the Lasso regression in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{ee69} \endgroup is the same as in Buhlmann2014: as shown in the Appendix, this condition is still sufficient to ensure that the bias term is asymptotically negligible even when the stock returns follow factor structure with unknown factors. Once we impose Gaussianity of ${\mathbf e}$ in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{e4.9} \endgroup , we can infer the distribution of portfolio weights. Note that in this case normally distributed errors do not imply that the stock returns are also Gaussian: we did not assume $\varepsilon_t \sim \mathcal{N}_p(\cdot)$ in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{e5.1} \endgroup . The unknown $\sigma_{e}^{2}$ can be replaced by a consistent estimator. Finally, even when Gaussianity of ${\mathbf e}$ is relaxed, we can use the central limit theorem argument to obtain approximate Gaussianity of components of $W|{\mathbf R}$ of fixed dimension, or moderately growing dimensions (see Buhlmann2014 for more details), however, in order not to divert the focus of this paper, we leave it for future research.
To establish the properties of the post-Lasso estimator in Algorithm (ref), we focus on MRC weight formulation, since it satisfies the standard post-Lasso assumptions. For GMV and MWC formulations, the procedure described in Algorithm (ref) is not \enquote{post-Lasso} in the usual sense. Concretely, the latter assumes that both steps in Algorithm (ref) have the same objective function, which is violated for GMV and MWC. Consequently, we leave rigorous theoretical derivations of these two portfolio formulation for future research. For MRC, we use the post-model selection results established in chernozhukov2013. Specifically, we have the following theorem:
The proof of Theorem (ref) easily follows from the proof of Corollary 2 of chernozhukov2013 and is omitted here. Let us comment on the upper bounds for post-Lasso estimator: first, the term $(s_0\omega_{T})\vee (\bar{s}^2\omega_{T})$ appears since one needs to estimate the dependent variable in equation \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{e4.9} \endgroup , which creates the difference between the bound in chernozhukov2013 and our Theorem (ref). Second, similarly to chernozhukov2013, the upper bound undergoes a transition from the oracle rate enjoyed by the standard Lasso to the faster rate that improves on the latter when (1) the precision matrix of the idiosyncratic components is sparse enough and (2) the oracle model has well-separated coefficients. Noticeably, the upper bounds in Theorem (ref) hold despite the fact that the first-stage Lasso regression in Algorithm (ref) may fail to correctly select the oracle model $\Xi$ as a subset, that is, $\Xi \notin \hat{\Xi}$.
Finally, let us compare the rates of non-sparse MRC portfolio weights in Theorem (ref), de-biased weights in Theorem (ref), and post-lasso weights in Theorem (ref): de-biased estimator exhibits fastest convergence, followed by post-lasso and non-sparse weights. This result is further supported by our simulations presented in the next section.
We study the consistency for estimating portfolio weights in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{ee20} \endgroup of (i) sparse portfolios that use the standard Lasso without de-biasing in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{ee69} \endgroup , (ii) Lasso with de-biasing in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{ee74} \endgroup , (iii) post-Lasso in Algorithm (ref), and (iv) non-sparse portfolios that use FMB from Algorithm (ref). Our simulation results are divided into two parts: the first part examines the performance of models (i)-(iv) under the Gaussian setting, and the second part examines the robustness of performance under the elliptical distributions (to be described later). Each part is further subdivided into two cases: with $p<T$ (Case 1) and with $p>T$ (Case 2), in both cases we allow the number of stocks to increase with the sample size, i.e. $p=p_T \rightarrow \infty$ as $T \rightarrow \infty$. In Case 1 we let $p = T^{\delta}$, $\delta = 0.85$ and $T = \lbrack 2^h \rbrack, \ \text{for} \ h=7,7.5,8,\ldots,9.5$, in Case 2 we let $p = 3\cdot T^{\delta}$, $\delta = 0.85$, all else equal.
First, consider the following data generating process for stock returns:
where ${\mathbf m}_i \sim \mathcal{N}(1,1)$ independently for each $i=1,\ldots,p$, ${\bm \varepsilon}_{t}$ is a $p \times 1$ random vector of idiosyncratic errors following $\mathcal{N}(\bm{0},{\bm \Sigma}_{\varepsilon})$, with a Toeplitz matrix ${\bm \Sigma}_{\varepsilon}$ parameterized by $\rho$: that is, ${\bm \Sigma}_\varepsilon = ({\bm \Sigma}_\varepsilon)_{ij}$, where $({\bm \Sigma}_\varepsilon)_{ij}=\rho^{\@ifstar{\oldabs}{\oldabs*}{i-j}}$, $i,j\in 1,\ldots,p$ which leads to sparse ${\bm \Theta}_{\varepsilon}$, ${\mathbf f}_{t}$ is a $K \times 1$ vector of factors drawn from $\mathcal{N}(\mathbf{0},{\bm \Sigma}_f = {\mathbf I}_{K}/10)$, ${\mathbf B}$ is a $p\times K$ matrix of factor loadings drawn from $\mathcal{N}(\mathbf{0},{\mathbf I}_{K}/100)$. We set $\rho = 0.5$ and fix the number of factors $K=3$.
Let ${\bm \Sigma} = {\mathbf B}{\bm \Sigma}_f{\mathbf B}'+{\bm \Sigma}_{\varepsilon}$. To create sparse MRC portfolio weights we use the following procedure: first, we threshold the vector ${\bm \Sigma}^{-1}{\mathbf m}$ to keep the top $p/2$ entries with largest absolute values. This yields sparse vector ${\bm \alpha}={\bm \Sigma}^{-1}{\mathbf m}$ defined in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{e158} \endgroup . We use ${\bm \Sigma}{\bm \alpha}$ and ${\bm \Sigma}$ as the values for the mean and covariance matrix parameters to generate multivariate Gaussian returns in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{e52} \endgroup . Note that the low rank plus sparse structure of the covariance matrix is preserved under this transformation.
Figure (ref) shows the averaged (over Monte Carlo simulations) errors of the estimators of the weight ${\mathbf w}_{\text{MRC}}$ versus the sample size $T$ in the logarithmic scale (base 2). As evidenced by Figure (ref), (1) sparse estimators outperform non-sparse counterparts; (2) using de-biasing or post-Lasso improves the performance compared to the standard Lasso estimator. As expected from Theorems (ref)-(ref), the Lasso, de-biased Lasso and post-Lasso exhibit similar rates, but the two latter estimators enjoy lower estimation error. The ranking remains similar for Case 2, however, as illustrated in Figure (ref), the performance of all estimators slightly deteriorates.
Gaussian-tail assumption is too restrictive for modeling the behavior of financial returns. Hence, as a second exercise we check the robustness of our sparse portfolio allocation estimators under the elliptical distributions, which we briefly review based on fan2018elliptical. Elliptical distribution family generalizes the multivariate normal distribution and multivariate t-distribution. Let ${\mathbf m} \in \mathbb{R}^p$ and ${\bm \Sigma} \in \mathbb{R}^{p\times p}$. A $p$-dimensional random vector ${\mathbf r}$ has an elliptical distribution, denoted by ${\mathbf r} \sim \text{ED}_p({\mathbf m},{\bm \Sigma},\zeta)$, if it has a stochastic representation
where ${\mathbf U}$ is a random vector uniformly distributed on the unit sphere $\mathcal{S}^{q-1}$ in $\mathbb{R}^q$, $\zeta \geq 0$ is a scalar random variable independent of ${\mathbf U}$, ${\mathbf A} \in \mathbb{R}^{p\times q}$ is a deterministic matrix satisfying ${\mathbf A}{\mathbf A}'={\bm \Sigma}$. As pointed out in fan2018elliptical, the representation in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{eq6.2} \endgroup is not identifiable, hence, we require $\operatorname*{\mathbb{E}}{\zeta^2}=q$, such that $\text{Cov}({\mathbf r}) = {\bm \Sigma}$. We only consider continuous elliptical distributions with $\Pr[\zeta=0]=0$. The advantage of the elliptical distribution for the financial returns is its ability to model heavy-tailed data and the tail dependence between variables.
Having reviewed the elliptical distribution, we proceed to the second part of simulation results. The data generating process is similar to fan2018elliptical: let $({\mathbf f}_t, {\bm \varepsilon}_{t})$ from \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{e52} \endgroup jointly follow the multivariate t-distribution with the degrees of freedom $\nu$. When $\nu = \infty$, this corresponds to the multivariate normal distribution, smaller values of $\nu$ are associated with thicker tails. We draw $T$ independent samples of $({\mathbf f}_t, {\bm \varepsilon}_{t})$ from the multivariate t-distribution with zero mean and covariance matrix ${\bm \Sigma} = \text{diag}({\bm \Sigma}_{f}, {\bm \Sigma}_{\varepsilon})$, where ${\bm \Sigma}_{f} = {\mathbf I}_{K}$. To construct ${\bm \Sigma}_{\varepsilon}$ we use a Toeplitz structure parameterized by $\rho = 0.5$, which leads to the sparse ${\bm \Theta}_{\varepsilon} = {\bm \Sigma}^{-1}_{\varepsilon}$. The rows of ${\mathbf B}$ are drawn from $\mathcal{N}(\bm{0}, {\mathbf I}_{K}/100)$. Figure 2 reports the results for $\nu=4.2$\footnote{The results for larger degrees of freedom do not provide any additional insight, hence we do not report them here. However, they are available upon request.}: the performance of the standard Lasso estimator significantly deteriorates, which is further amplified in the high-dimensional case where it exhibits the worst performance. Noticeably, post-Lasso still achieves the lowest estimation error, followed by de-biased estimator.
This section is divided into three main parts. First, we examine the performance of several non-sparse portfolios, including the equal-weighted and Index portfolios (reported as the composite S&P500 index listed as \textsuperscript{$\wedge$}GSPC). Second, we study the performance of sparse portfolios that are based on de-biasing and post-Lasso. Third, we consider several interesting periods that include different states of the economy: we examine the merit of sparse vs non-sparse portfolios during the periods of economic growth, moderate market decline and severe economic downturns.
We use monthly returns of the components of the S&P500 index\footnote{The conclusions from using daily data are the same as those for monthly returns, hence we do not report them in the main manuscript text. However, they are available upon request.}. The data on historical S&P500 constituents and stock returns is fetched from CRSP and Compustat using SAS interface. The full sample has 480 observations on 355 stocks from January 1, 1980 - December 1, 2019. We use January 1, 1980 - December 1, 1994 (180 obs) as a training period and January 1, 1995 - December 1, 2019 (300 obs) as the out-of-sample test period. We roll the estimation window over the test sample to rebalance the portfolios monthly. At the end of each month, prior to portfolio construction, we remove stocks with less than 15 years of historical stock return data. For sparse portfolio we employ the following strategy to choose the tuning parameter $\lambda$ in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{ee67} \endgroup : we use the first two thirds of the training data (which we call the training window) to estimate weights and tune the shrinkage intensity $\lambda$ in the remaining one third of the training sample to yield the highest Sharpe Ratio which serves as a validation window. We estimate factors and factor loadings in the training window and validation window combined. The risk-free rate and Fama-French factors are taken from \href{https://mba.tuck.dartmouth.edu/pages/faculty/ken.french/data_library.html}{Kenneth R. French's data library.}
Similarly to Caner2019, we consider four metrics commonly reported in finance literature: the Sharpe Ratio, the portfolio turnover, the average return and risk of a portfolio. We consider two scenarios: with and without transaction costs. Let $T$ denote the total number of observations, the training sample consists of $m$ observations, and the test sample is $n=T-m$. When transaction costs are not taken into account, the out-of-sample average portfolio return, risk and Sharpe Ratio are
We follow Demiguel2009optimal,Caner2019,Li2015sparse,ban2018machine to account for transaction costs (tc). In line with the aforementioned papers, we set $c=50 \text{bps}$. Define the excess portfolio at time $t+1$ with transaction costs as
where $r_{t+1,j}+r^{f}_{t+1}$ is sum of the excess return of the $j$-th asset and risk-free rate, and $r_{t+1,\text{portfolio}}+r^{f}_{t+1}$ is the sum of the excess return of the portfolio and risk-free rate. The out-of-sample average portfolio return, risk, Sharpe Ratio and turnover are defined accordingly:
The first set of results explores the performance of several non-sparse portfolios: equal-weighted portfolio (EW), Index portfolio (Index), MB from Algorithm (ref), FMB from Algorithm (ref), linear shrinkage estimator of covariance that incorporates factor structure through the Sherman-Morrison inversion formula (Ledoit2004, further referred to as LW), CLIME (cai2011constrained). We consider two scenarios, when the factors are unknown and estimated using the standard PCA (statistical factors), and when the factors are known. For the statistical factors, we determine the number of factors, $K$, in a standard data-driven way using the information criteria discussed in Bai2002 and kapetanios2010testing among others. For the scenario with known factors we include up to 5 Fama-French factors: FF1 includes the excess return on the market, FF3 includes FF1 plus size factor (Small Minus Big, SMB) and value factor (High Minus Low, HML), and FF5 includes FF3 plus profitability factor (Robust Minus Weak, RMW) and risk factor (Conservative Minus Agressive, CMA). In Table (ref), we report monthly portfolio performance for three alternative portfolio allocations in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{e156} \endgroup , \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{e157} \endgroup and \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{e156} \endgroup . We set a return target $\mu= 0.7974\%$ which is equivalent to $10\%$ yearly return when compounded. The target level of risk for the weight-constrained and risk-constrained Markowitz portfolio (MWC and MRC) is set at $\sigma=0.05$ which is the standard deviation of the monthly excess returns of the S&P500 index in the first training set.
We now comment on the results which are presented in Table (ref): (1) accounting for the factor structure in stock returns improves portfolio performance in terms of the OOS Sharpe Ratio. Specifically, EW, Index, MB and CLIME which ignore factor structure perform worse than FMB and LW. (2) The models that use an improved estimator of covariance or precision matrix outperform EW and Index on the test sample. As a downside, such models have higher Turnover. This implies that superior performance is achieved at the cost of larger variability of portfolio positions over time and, as a consequence, increased risk associated with it.
The second set of results studies the performance of sparse portfolios: we include our proposed methods based on de-biasing and post-Lasso, as well as the approach studied in Ao2019 (Lasso) without factor investing. For post-Lasso we first use Lasso-based weight estimator in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{ee69} \endgroup for selecting stocks with absolute value of weights above a small threshold $\epsilon$ (we use $\epsilon=0.0001$), then we form portfolio with the selected stocks using three alternative portfolio allocations in \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{e156} \endgroup - \begingroup \leavevmode \color{newred} \hypersetup{linkbordercolor=[named]{newred}} \@msm@th@eqref{e158} \endgroup .
Let us comment on the results presented in Table (ref): (1) column one demonstrates that de-biasing leads to significant performance improvement in terms of the return and the OOS Sharpe Ratio. Note that even though the risk of de-biased portfolio is also higher, it still satisfies the risk-constraint. This result emphasizes the importance of correcting for the bias introduced by the $\ell_1$-regularization. (2) Comparing two bias-correction methods, de-biasing and post-Lasso, we find that the latter is characterized by higher return and higher risk. However, increase in portfolio return brought by post-Lasso is, overall, not sufficient to outperform de-biasing approach in terms of the out-of-sample Sharpe Ratio. (3) Sparse portfolios have lower return, risk and turnover compared to non-sparse counterparts in Table (ref), however, the OOS Sharpe Ratio is comparable, i.e. we do not see uniform superiority of either method. Therefore, incorporating sparsity allows investors to reduce portfolio risk at the cost of lower return while maintaining the Sharpe Ratio comparable to holding a non-sparse portfolio.
Tables (ref)-(ref) compare the performance of non-sparse and sparse (de-biased. “DL", and post-Lasso, “PL") portfolios for different time periods in terms of the cumulative excess return (CER) over the period of interest and risk. The first period of interest (1997-98), which will be referred to as “Period I", corresponds to economic growth since Index exhibited positive CER during this time. The second period of interest, “Period II", corresponds to moderate market decline since EW and Index had relatively small negative CER. Finally, “Period III", corresponds to severe economic downturn and significant drop in the performance of EW and Index. We note that the references to the specific crises in Tables (ref)-(ref) do not intend to limit these economic periods to these time spans. They merely provide the context for the time intervals of interest. Since the performance of MWC portfolios is similar to GMV, we only report MRC and GMV for the ease of presentation.
Let us summarize the findings from Tables (ref)-(ref): (1) In Period I non-sparse portfolios that rely on the estimation of covariance or precision matrix outperformed EW and Index in terms of CER for both MRC and GMV. However, in Period II GMV portfolios exhibited slightly negative CER, whereas MRC portfolios had higher risk but positive CER (albeit being lower compared to Period I). Note that in Period III none of the non-sparse portfolios generated positive CER and portfolio risk increased rapidly. Examining the performance of sparse portfolios in Table (ref), we see that (2) our proposed sparse portfolios produce positive CER during all three periods of interest. Furthermore, the return generated by PL is higher than that by non-sparse portfolios even during Periods I and II. Interestingly, DL produces positive CER without having high risk exposure. This suggests that our de-biased estimator of portfolio weights exhibits minimax properties. We leave the formal theoretical treatment of the latter for the future research.
This paper develops an approach to construct sparse portfolios in high dimensions that addresses the shortcomings of the existing sparse portfolio allocation techniques. We establish the oracle bounds of sparse weight estimators and provide guidance regarding their distribution. From the empirical perspective, we examine the merit of sparse portfolios during different market scenarios. We find that in contrast to non-sparse counterparts, our strategy is robust to recessions and can be used as a hedging vehicle during such times. Our framework makes use of the tool from the network theory called nodewise regression which not only satisfies desirable statistical properties, but also allows us to study whether certain industries could serve as safe havens during recessions. We find that such non-cyclical industries as consumer staples, healthcare, retail and food were driving the returns of the sparse portfolios during both the global financial crisis of 2007-09 and COVID-19 outbreak, whereas insurance sector was the least attractive investment in both periods. Finally, we develop a simple framework that provides clear guidelines how to implement factor investing using the methodology developed in this paper.
I greatly appreciate thoughtful comments and immense support from Tae-Hwy Lee, Jean Helwege, Jang-Ting Guo, Aman Ullah, Matthew Lyle, Varlam Kutateladze and UC Riverside Finance faculty. I also thank seminar participants at the 14th International CFE Conference (virtual), 2021 Southwestern Finance Association Annual Meeting, and Vilnius University. \cleardoublepage
\phantomsection
\addcontentsline{toc}{section}{References} {14pt} \cleardoublepage
\phantomsection