EconBase
← Back to paper

Dynamic Portfolio Allocation in High Dimensions using Sparse Risk Factors

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

95,053 characters · 16 sections · 109 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Dynamic Portfolio Allocation in High Dimensions using Sparse Risk Factors

abstractWe propose a fast and flexible method to scale multivariate return volatility predictions up to high-dimensions using a dynamic risk factor model. Our approach increases parsimony via time-varying sparsity on factor loadings and is able to sequentially learn the use of constant or time-varying parameters and volatilities. We show in a dynamic portfolio allocation problem with 452 stocks from the $S\&P$ 500 index that our dynamic risk factor model is able to produce more stable and sparse predictions, achieving not just considerable portfolio performance improvements but also higher utility gains for the mean-variance investor compared to the traditional Wishart benchmark and the passive investment on the market index. Keywords: Dynamic Factor Model; Time-Varying Sparsity; Portfolio Allocation; High-Dimension. J.E.L. codes: C32, C52, C53, C58, G11

\onehalfspace

Introduction

Portfolio allocation is one of the most common problems in finance. Since the seminal work of markowitz1952portfolio, mean-variance optimization has been the most traditional way to select stocks in the financial industry and is commonly applied in the academic literature. The covariance matrix of returns is the key input to generate optimal portfolio weights, which makes its forecast accuracy crucial for out-of-sample portfolio performance. However, the universe of assets available for allocation is vast nowadays, increasing the dimension of such covariance matrices potentially to hundreds or even thousands of stocks. Due to the fact that the number of parameters in a covariance matrix grows quadratically with the number of assets, it inserts the traditional Markowitz's problem into the curse of dimensionality.

Since large covariance matrices are extremely susceptible to estimation errors and instabilities, producing poor out-of-sample predictions for portfolio construction, we propose what we call a Dynamic Risk Factor Dependency Model (DRFDM). By the use of economically motivated risk factors and inducing time-varying sparsity on factor loadings, our approach is able to achieve higher model parsimony, dramatically reducing the parameter space and improving predictions for final portfolio decisions. The DRFDM combines a factor structure with sparsity in a conjugate and sequential fashion without the use of MCMC schemes, making the estimation process much faster and allowing investors to backtest a universe of hundreds of assets in a matter of few minutes.

The use of factor models in the financial literature is not new. With different applications in the asset pricing literature, since the CAPM of sharpe1964capital, the APT of ross1976arbitrage and the seminal work of fama1992cross, these models consider that all systematic variation in returns are driven by a set of common factors that can be observable financial indices or unknown latent variables. This framework is commonly used for evaluating return anomalies and portfolio manager performance but also for portfolios construction. Our paper will focus on the use of observable risk factor models to build optimal portfolios.

To illustrate the idea behind factor models, consider an exact K-risk factor model with a N-dimensional vector of assets returns:

equation[equation omitted — 277 chars of source]

where $\boldsymbol{B}_{t}$ is a N $\times$ K matrix of factor exposures (loadings) to the K risk factors $\boldsymbol{f}_{t} $ and $\boldsymbol{\Omega}_{t} = \operatorname{diag} \left(\sigma_{1 t}^{2}, \ldots, \sigma_{N t}^{2}\right)$. If $Var(\boldsymbol{f}_{t} ) = \boldsymbol{\Sigma_{t}^{f}}$, the model in ((ref)) implies an unconditional variance for return as:

equation[equation omitted — 188 chars of source]

which is divided between a systematic component related to the factor exposures and factor covariances and an idiosyncratic component for each individual asset return. Note that using K $<<$ N implies a strong reduction on the total number of parameters to be estimated in $\boldsymbol{\Sigma}_{t}^{r}$.

When latent factor models are considered, $\boldsymbol{f}_{t} $ is treated as an unobserved variable estimated from the data. Some important references are aguilar2000bayesian, han2006asset, lopes2007factor, carvalho2011dynamic, zhou2014bayesian, kastner2017efficient and kastner2019sparse. Both carvalho2011dynamic and zhou2014bayesian explore the notion of time-varying sparsity, with factor loadings being equal to zero for different periods of time. kastner2019sparse induces static sparsity by the use of shrinkage priors on factor loadings, pulling coefficients toward zero. Although his approach induces greater parsimony, shrinkage priors do not impose coefficients to be exactly zero, remaining a portion of estimation uncertainty that is still carried to the covariance matrix.

The major drawback of latent factor models is that it requires MCMC schemes to simulate from joint posteriors, imposing great computational burden to scale up models to high-dimensions. The problem is itensified by the fact that strategies in quantitative finance require a sequential analysis for backtests, so the MCMC needs to be repeated for each period of time, which can take not just several days but weeks to be completed. Therefore, latent factor models are prohibitive for sequential analysis in high-dimensions (see gruber2016gpu, gruber2017bayesian and west2020bayesian for a deeper discussion on model scalability).

In order to seek for fast sequential analysis and higher model flexibility, instead of estimating latent factors, we include observable risk factor commonly used in the financial literature to represent the common movements of returns. We use the traditional 5 factors from fama2015five as the main representation, also showing results for different subsets of risk factors. There are different papers in the literature that have already addressed the estimation of the covariance matrix of returns using observable risk factors. Recent key references include wang2011dynamic, brito2018forecasting, puelz2020portfolio and de2018factor. The main advantage of DRFDM in relation to those papers is its ability to take into account model uncertainty in a dynamic fashion in different model settings for individual asset returns. As we explain in the next section, the DRFDM uses what started to be known in the recent econometric literature as a "Decouple/Recouple" concept. The basic idea is to decouple the multivariate dynamic model into several univariate customized dynamic linear models (DLM) that can be solved in parallel and then be recouple for forecasting and decisions. It is strictly related to the popular Cholesky-style Multivariate Stochastic Volatility of lopes2016parsimony, shirota2017cholesky and primiceri2005time and also applied in a similar fashion in zhao2016dynamic, fisher2020optimal, lavine2020adaptive and levy2021dynamic. This framework allow us to take the model uncertainty problem into the univariate context, making highly flexible dynamic model choices.

Since it is well known that the environment of the economy and the financial market is continuously changing, model flexibility becomes a extremely appealing feature to be incorporated nowadays. The common patterns in stock returns in the 90s are different from those during the Great Financial Crisis or during the recent Covid-19 pandemic. Factor exposures can be lower or higher, depending on calm or stressed periods. Additionaly, we can consider the fact that for some periods a subset of stock returns are not loading on specific risk factors. This motivates our work to impose time-varying sparsity, where inspired by the works of raftery2010online, dangl2012predictive and koop2013large we use a Dynamic Model Selection (DMS) approach to sequentially select risk factors via dynamic model probabilities, where a risk factor is included if it is empirically wanted. We also consider a model space that is determined not just by different risk factors, but also by different degrees of variation in factor loadings, return volatilities and factor volatilities. Therefore, the DRFDM can adapt to environments of higher, lower or no variation in coefficients, for each specific entry of matrices in Equation ((ref)). Hence, one specific asset return can be much more volatile than others and some risk factors can vary differently over time, also moving from constant to time-varying parameters. This last setting has a similar flavour of parsimony in the sense of lopes2016parsimony.

Dynamic risk factor selection also offers an important benefit in terms of model parsimony since it imposes factor loadings of non-selected factors to be exactly equal to zero. It tends to induce much higher parsimony compared to continuous shrinkage priors applied in kastner2019sparse and many other papers nowadays. As argued in huber2020inducing and hauzenberger2020combining, continuous shrinkage priors offer a lower bound of accuracy to be achieved and, for highly parametrized models, parameter uncertainty over-inflate predictive variances.

Finally, an additional contribution of DRFDM to the literature relies on its simple and fast computation. Since we rely on a conjugate model, with closed-form solutions for posterior distributions and predictive densities, there is no need for expensive MCMC methods, making the whole process extremely fast and easily scalable for high-dimensions. As we show in the empirical section, we perform a portfolio allocation procedure using data covering both the Great Financial Crisis and the initial stage of the Covid-19 pandemic with almost 500 stock returns from the $S\&P$ 500 index that can be backtested in just few minutes. We compare results with different specification choices and the traditional Wishart Dynamic Linear Model (W-DLM) as a benchmark. The W-DLM has been a standard model in the Bayesian financial time series and in the financial industry, because of its scalability and availability of sequential filtering. Our results show that the DRFDM is able to produce not just much stronger statistical improvements but substantial increase in Sharpe Ratios and risk reduction. We also show that a mean-variance investor will be willing to pay a considerable management fee to switch from the W-DLM benchmark, from different well known estimation methods in the literature and the passive investment on the $S\&P$ index to the Dynamic Risk Factor Dependency Model.

The remainder of the paper is organized as follows. The general econometric framework is introduced in Section (ref). Section (ref) details the time-varying model and factor selection approach and how it can be applied to impose sparsity. In Section (ref) we perform our empirical analysis, providing an out-of-sample statistical and economic performance evaluation in a high-dimensional environment. Section (ref) concludes.

Econometric Framework

As mentioned in Section (ref), our work is inspired by the Cholesky-style framework in lopes2016parsimony and primiceri2005time, being closely related to the Dynamic Dependency Network Model of zhao2016dynamic. The great advantage of this framework is to model the cross-sectional contemporaneous relations among different series and customize univariate DLMs. Consider $\boldsymbol{r_{t}}$ as a $\textit{N}$-dimensional vector with asset returns time series $r_{j,t}$ and consider the following dynamic system:

equation[equation omitted — 287 chars of source]

where $\boldsymbol{\alpha}_{t}$ is a $\textit{N}$-dimensional vector of time-varying intercepts and $\boldsymbol{\Omega}_{t} = \operatorname{diag} \left(\sigma_{1 t}^{2}, \ldots, \sigma_{N t}^{2}\right)$. All contemporaneous relations among different asset returns are coming from the $\textit{N}$ $\times$ $\textit{N}$ matrix $\boldsymbol{B}_{t}$, whose off-diagonal elements $\beta_{ji t}$s (for $j \neq i $) capture the dynamic contemporaneous relationships among series $j$ and $i$ at time $t$ and $\boldsymbol{B}_{t}$ has zeroes on the main diagonal.

In the work of lopes2016parsimony, zhao2016dynamic and levy2021dynamic, $\boldsymbol{B}_{t}$ is a lower triangular matrix with zeroes in and above the main diagonal:

equation[equation omitted — 284 chars of source]

Since the error terms in $\boldsymbol{\epsilon_{t}}$ are contemporaneouly uncorrelated, the triangular contemporaneous dependencies among asset returns in Equation ((ref)) generate a fully recursive system, known as a Cholesky-style framework (west2020bayesian). Hence, each equation $j$ of the system will have its own set of $\textit{parents}$ ($\boldsymbol{r_{pa(j), t}}$), that is, will depend contemporaneously on all other asset returns above equation $j$, following the triangular format in Equation ((ref)). In words, the top asset return in the system will not have parents, the second from the top asset return will have the first time series of returns as a parent and will load on it, the third asset return will have the first two asset returns as parents and will load on them all the way to the last asset return, which will depend on all other $N-1$ returns above it.

Equation ((ref)) can be rewritten in the reduced form as

equation[equation omitted — 239 chars of source]

where $\boldsymbol{A}_t=\left(\boldsymbol{I}_N-\boldsymbol{B}_{t}\right)^{-1}$ and $\boldsymbol{u}_t = \boldsymbol{A}_t \boldsymbol{\epsilon}_{t}$. The modified Cholesky decomposition clearly appears in $\boldsymbol{\Sigma}_{t}=\boldsymbol{A}_t \boldsymbol{\Omega}_{t}\boldsymbol{A}_t^{\prime}$ which is now a full variance-covariance matrix capturing the contemporaneous relations among the $\textit{N}$ asset returns. Given the parental triangular structure of $\boldsymbol{B_{t} }$ in ((ref)), the equations will be conditionally independent, bringing the “Decoupled” aspect of the multivariate model. In other words, the multivariate model can be viewed as a set of $\textit{N}$ conditionally independent univariate DLMs that can be dealt with in a parallelizable fashion. The outputs of each equation are then used to compute $\boldsymbol{B_{t}}$ and $\boldsymbol{\Omega_{t}}$, hence recovering the full time-varying covariance matrix $\boldsymbol{\Sigma_{t}}$.

Dynamic Risk Factor Dependency Model

Although the model above shows greater flexibility, it remains highly parameterized and susceptible to producing poor out-of-sample forecasts. Notice that for a model with hundreds or thousands of equations, those asset returns at the bottom part will load on several hundreds of other assets, making out-of-sample forecasts too unstable.

Additionally, the triangular form in Equation ((ref)) makes the system dependent of the asset return ordering. As highlighted by levy2021dynamic, it is an important drawback of the Cholesky-style framework, because imposing a specific order structure can lead to inferior final decisions and harm portfolio performance. Since the "correct" series ordering is uncertain and the environment of the economy is continuously changing, levy2021dynamic propose what they call a Dynamic Ordering Learning. That is a flexible model that deals with the ordering uncertainty in a dynamic fashion, where the econometrician is able to sequentially learn the contemporaneous relations among different series over time. However, since there are $N$! possible orders to learn for each period of time, even with a flexible model it is a prohibitive task in high-dimensions.

In order to impose greater model parsimony and overcome the ordering uncertainty in high-dimensions, we propose what we call a Dynamic Risk Factor Dependency Model. Inspired by the literature on observable risk factors, in this method we augment the $N$-vector of asset returns with $K$ economically motivated risk factors, for $K << N$, and impose the restriction that all dynamic contemporaneous dependencies among asset returns are coming from them. It drastically reduces both the parameter space and the ordering uncertainty problem to a low-dimension one, being easily implemented by the Dynamic Ordering Learning approach of levy2021dynamic.

Defining a new vector of returns $\boldsymbol{\mathcal{R}}_{t} = \left(\boldsymbol{\mathcal{F}}_{t}\quad , \quad \boldsymbol{r}_{t} \right)^{\prime} $ , augmented by the $K$-dimensional vector of known risk factors, $\boldsymbol{\mathcal{F}}_{t}$, we rewrite Equation ((ref)) as:

equation[equation omitted — 355 chars of source]

where the tilde upscript represents the extended to $K+N$ dimension version of previous vectors and matrices of Equation ((ref)). Now, $\boldsymbol{\mathcal{B}}_{t}$ is a new $(K+N) \times (K+N)$ matrix containing all the dynamic contemporaneous dependencies, where both factor and asset returns are allowed to load only on the set of observable chosen factors:

equation[equation omitted — 1,003 chars of source]

Our approach can be viewed as an extension of the work of zhao2016dynamic, however, instead of following the whole triangular format as in Equation ((ref)), we impose that any asset return dependencies are coming from common factors in the economy and not on specific movements of different asset returns. Hence, for $j = 1,\dots,K$, the matrix $\boldsymbol{\mathcal{B}}_{t}$ follows the usual triangular format. However, for $j > K$, asset returns are restricted to load until series (factor) $K$. This new representation allows us to estimate a much lower number of parameters, since $\boldsymbol{\mathcal{B}}_{t}$ is filled with zeroes after the $K$-th column. The sparser $\boldsymbol{\mathcal{B}}_{t}$ is, the more stable and efficient the resulting inferences are, producing better out-of-sample predictions for decision analysis. Additionaly, it gives higher economic intuition for the variation of asset returns, since the common movements follow exposures to well know risk factors developed by the financial literature.

$\textit{K + N}$ univariate dynamic linear models

As mentioned before, given the structure of $\boldsymbol{\mathcal{B}}_{t}$, the set of $K+N$ univariate models can be represented as $K+N$ univariate recursive dynamic regressions, where we have for each $j = 1, \dots, K, \dots, K+N$: \footnote{Inspired on recent evidence of momentum on risk factors (gupta2019factor) we also have tested risk factors as a AR(1) process, but results were very similar so we decided to maintain in this paper the traditional structure of Equation ((ref)). }

equation[equation omitted — 226 chars of source]

where for $j = 1, \dots, K$, the parental set $\boldsymbol{\mathcal{R}}_{pa(j), t} = \boldsymbol{\mathcal{F}}_{pa(j), t}$ represents all risk factor series in $\boldsymbol{\mathcal{R}}_{t}$ that are above series $j$ and, for $j > K$, $\boldsymbol{\mathcal{R}}_{pa(j), t}$ possibly represents all series until series K, i.e, all risk factors in $\boldsymbol{\mathcal{F}}_{t}$.

We represent the dynamic coefficients in Equation ((ref)) evolving according to random walks:

equation[equation omitted — 323 chars of source]

and by ease of exposition, we define $x_{j, t-1}=1 $ and represent the full dynamic state and regression vectors as

$$ \boldsymbol{\theta}_{j t}=\left(

array[array omitted — 60 chars of source]

\right) \quad and \quad \mathbf{F}_{j t}=\left(

array[array omitted — 64 chars of source]

\right), $$

Hence, denoting $y_{jt}$ as the $j$-$th$ variable in $\boldsymbol{\mathcal{R}}_{t}$ we recover the traditional univariate DLM formulation as in west2006bayesian, namely

eqnarray*[eqnarray* omitted — 337 chars of source]

for $j=1,\ldots, K + N$, where again $\boldsymbol{\theta}_{j t}$ evolves over time as a simple random-walk.

\paragraph{Posterior at $t-1$.} Following the algorithmic structure of sequential learning in DLMs west2006bayesian, Chapter 4, at time $t-1$ and for each time series $j$, the joint posterior distribution of $\boldsymbol{\theta}_{j t-1} $ and $\sigma_{j t-1}$ at time $t-1$ is a multivariate Normal-Gamma\footnote{In our empirical analysis we have considered initial states centering $\theta_{j0}$ around zero ($m_{j, 0}=\boldsymbol{0}$) with $C_{j,0} = 100\boldsymbol{I}$. We set $s_{j, 0}$ as the sample variance estimate of residuals from an OLS model over the training period and we start with $n_{j,0} = 10$. }:

equation[equation omitted — 221 chars of source]

Through the random walk evolution and conjugacy, we can derive the joint prior distribution of $\boldsymbol{\theta}_{jt} $ and $\sigma_{jt}$ for time $t$ as:

equation[equation omitted — 182 chars of source]

where $r_{j t}=\kappa_{j} n_{j, t-1}$, $\boldsymbol{a}_{j t} = \boldsymbol{m}_{j, t-1}$ and $\boldsymbol{R}_{j t} = \boldsymbol{C}_{j, t-1} / \delta_{j}$. The quantities $0 < \delta_{j} \leq 1$ and $0 < \kappa_{j} \leq 1$ represent specific discount (aka forgetting) factors for $\theta_{jt}$ and $\sigma_{jt}$, respectively. Discount methods are used to induce time-variations in the evolution of parameters and have been extensively used in many applications (raftery2010online, dangl2012predictive, koop2013large, \citealp*{mcalinn2020multivariate}, amongst others) and well documented in west2006bayesian, lopes2006mcmc and prado2010time.

\paragraph{1-step ahead forecast at $t-1$.} The (prior) predictive distribution of $y_{j t}$ is a Student's $t$ distribution with $r_{j t}$ degrees of freedom:

$$ y_{j t} \mid \boldsymbol{y}_{pa(j),t}, \mathcal{D}_{t-1} \sim \mathcal{T}_{r_{j t}}\left(f_{j t} , q_{j t}\right), $$

with $f_{j t} = \mathbf{F}_{j t}^{\prime}\boldsymbol{a_{j t}}$ and $q_{j t} = s_{j, t-1}+\mathbf{F}_{j t}^{\prime} \mathbf{R}_{j t} \mathbf{F}_{j t} $. It is important to notice that in this framework we have a conjugate analysis for forward filters and one-step ahead forecasting. Therefore, we are able to compute closed-form solution for predictive densities for each equation $j$. Hence, conditional on $\textit{parents}$, it is easy to compute the joint predictive density for $\boldsymbol{y_{t}} $:

equation[equation omitted — 188 chars of source]

which simply is the product of the already computed $K + N$ different univariate Student's $t$ distributions. After the time series are decoupled for sequential analysis, they are then recoupled for multivariate forecasting. In our decision analysis at Section (ref), we divide the recoupled part in two: one related to the dynamics of the $K$ factors and the other to the dynamic of the $N$ asset returns. We will be concerned with the mean and variance of each of this parts for the portfolio allocation study:

equation[equation omitted — 236 chars of source]

for the $K$-vector of expected factor means and the $K \times K$ expected factor covariance matrix and

equation[equation omitted — 209 chars of source]

for the $N$-vector of expected asset returns and the $N \times N$ expected covariance matrix of returns. Further details about the derivations of the evolution, forecasting, updating distributions can be found in Appendices A and B.

Dynamic Model and Factor Selection

As argued by kastner2019sparse, even imposing parsimony by a factor structure, models are still rich in parameters when considering a high-dimension portfolio problem. It motivates our work to induce stronger parsimony by the use of sparsity on factor loadings in a time-varying and sequential form. This approach is convenient to improve out-of-sample predictions and introduce greater model flexibility, since for each period of time different assets can load just on a subset of risk factors.

Model uncertainty is a well-known challenge among applied researchers and industry practitioners that are interested in producing forecasts for decision-making problems. In the last decades, the Bayesian literature has addressed this question with great success. Bayesian Model Averaging and Bayesian Model Selection are well-known methodologies for static models when there is uncertainty about the predictors to include (madigan1994model and hoeting1999bayesian). More recently, raftery2010online have propose a dynamic version of Bayesian Model Selection, called Dynamic Model Selection (DMS). They suggest the use of reasonable approximations, borrowing ideas from discount (forgetting) methods, avoiding simulations of transition probabilities matrices and maintaining the conjugate form of posterior distributions, allowing analytical solutions for forward filtering and forecasting which significantly reduces the computation burden of the process. Their approach has also shown great success, being applied recently in macroeconomics and finance (dangl2012predictive, koop2013large, catania2019forecasting and levy2021dynamic).

In our portfolio allocation problem we apply DMS for each individual equation of the system in ((ref)), learning sequentially about which specifications to choose, such as the main risk factors and the best discount factors conducting time-variation in factor loadings and volatilities. Therefore, not just risk factors can have a stochastic or constant volatility, but also each individual asset return can have constant or time-varying factor loadings and volatilities, switching its behavior depending on the environment of the economy. Using DMS we are able to dynamically select the best risk factors by imposing factor loadings equal to zero for non-selected risk factors. For a highly parametrized model it can be viewed as an advantage compared to traditional shrinkage priors that set coefficients close but not equal to zero, increasing uncertainty on predictions.

It is important to highlight here that the main focus of our dynamic factor selection approach is to induce model sparsity in order to deflate the covariance structure among the universe of asset returns and to produce more stable predictions of expected returns. Although we take advantage of the fact that our DRFDM is able to predict expected returns and we use this predictions for portfolio decisions, we will not be particularly concerned on inferences about asset pricing models or best risk factors to explain stock returns. For recent advances on Bayesian model selection/averaging with focus on investigating expected returns based on factor asset pricing models, we refer to bryzgalova2019bayesian and hwang2020bayesian.

Dynamic Model Probabilities

In order to perform DMS in each equation, we introduce here the idea of dynamic model probabilities. Consider the case of a specific equation $j$ in the system discussed on the previous section. Suppose the dependent serie of this equation can load on $p$ different risk factors available. Hence, there are $2^{p}$ possible combinations of models defined by the parental set. Also, when considering $n_{\delta}$ and $n_{\kappa}$ different possible discount factors for factor loadings and volatilities, the equation $j$ will have a total of $n_{j} = 2^{p}\times n_{\delta}\times n_{\kappa} $ possible models to choose in the univariate model space.\footnote{In our portfolio study we actually consider $n_{j} = (2^{p}-1)\times n_{\delta}\times n_{\kappa} $ possible models, since we exclude the case where an asset return is not loading in any risk factor.}

The DMS approach deals with the model uncertainty by assigning probabilities for each possible model. Denote $\pi_{t-1 | t-1, i,j} = p(\mathcal{M}_{i}^{j} \mid \mathcal{D}_{t-1})$ as the posterior probability of model $i$ and equation $j$ at time $t-1$. Following raftery2010online, the predicted probability of the model $i$ given all the data available until time $t-1$ is expressed as:

align[align omitted — 119 chars of source]

where $0 \leq \alpha \leq 1$ is a forgetting factor. The main advantage of using $\alpha$ is avoiding the computational burden associated with expensive MCMC schemes to simulate the transition matrix between possible models over time. This approach has also been extensively used in the Bayesian econometric literaure in the last decade (koop2013large, zhao2016dynamic, lavine2020adaptive and beckmann2020exchange). After observing new data at time $t$, we update our model probabilities following a simple Bayes’ update:

align[align omitted — 236 chars of source]

the posterior probability of model $i$ at time $t$ where $p_{i}(y_{j t} \mid \boldsymbol{y}_{pa(j)}, \mathcal{D}_{t-1})$ is the predictive density of model $i$ evaluated at $y_{jt}$. Hence, upon the arrival of a new data point, the investor is able to measure the performance for each univariate model $i$ and to assign higher probability for those models that generate better performance.

One possible interpretation for the forgetting factor $\alpha$ is through its role to discount past performance. Combining the predicted and posterior probabilities, we can show that

equation[equation omitted — 182 chars of source]

Since $0 < \alpha \leq 1$, Equation ((ref)) can be viewed as a discounted predictive likelihood, where past performances are discounted more than recent ones. It implies that models that received higher performance in the recent past will produce higher predictive model probabilities. The recent past is controlled by $\alpha$, since a lower $\alpha$ discounts more heavily past data and generates a faster switching behavior between models over time.\footnote{The value of $\alpha$ is commonly selected to be very close to 1. Following the majority of the econometric literature, we set $\alpha = 0.99$ in our empirical study.}

The idea of DMS is to select the model with the highest model probability for each period of time. Given the fact that each equation is conditionally independent, as soon as we select the best model for each equation, the posterior model probability of the multivariate model can be viewed as a product of the $K+N$ univariate model probabilities:

$$P(\mathcal{M}_{1:(K+N)}^{*}|\mathcal{D}_{t-1}) = \prod_{j=1}^{(K+N)} P(\mathcal{M}_{j}^{*}|\mathcal{D}_{t-1}) $$

where the asterisk symbol represents the selected model. A major benefit of our our approach is that each equation in the system in ((ref)) is conditionally independent and model uncertainty can also be addressed independently for each equation $j$. Hence, if for each equation we have a possible model space of size $n_{j}$, it implies a total of $\sum_{j = 1}^{K+N} n_{j} $ possible models. In a high-dimension environment with hundreds of equations and hundreds of possible univariate models for each one, it is a massive reduction on the total model space, since in a multivariate model with no conditional independence would require a total of $\prod_{j = 1}^{K+N} n_{j} $ models.

The use of Dynamic Model Probabilities offers a substantial flexibility for our approach. The DMS approach allow us to customize individual series, proposing different models with different risk factors and variation in coefficients, switching between them in a dynamic fashion. It is a important advantage compared to the traditional Wishart-DLM, where all series are stuck to the same discount factors for state evolutions.

Dynamic Factor Ordering Learning

As highlighted in Section (ref), the Cholesky-style framework is based on the triangular structure of Equation ((ref)). Therefore, models within this framework will depend on the series ordering structure selected by the researcher, leading to different contemporaneous relations among series. The recent work of levy2021dynamic show evidences of instabilities on the "correct" ordering structure and propose a method to dynamically deal with the problem of ordering uncertainty. By the use of Dynamic Ordering Probabilities, the authors propose a Dynamic Ordering Learning (DOL) approach where the econometrician is able to select or average the outputs of different orderings in a sequential fashion. They show that taking into account the ordering uncertainty among different series improves out-of-sample predictions and portfolio performance.

In our high-dimension environment, the computations of the total number of possible orderings is prohibitive, since for $N$ financial time series, there are $N!$ possible orders to compute.\footnote{Just to give an idea of how fast the order space grows with the number of series, a multivariate model with 6 series has $6! = 720$ possible orders. Just increasing the model dimension to 10 series increases the number of possible orders to 3,628,800. It is easy to see that when we go to hundreds of series, the number of orders goes to infinity. } However, once we impose a factor dependency structure as in ((ref)), we restrict our ordering uncertainty to a low dimension, because just the top $K$ variables rely on a triangular format. Since $K<<N$, we are able to follow the DOL approach of levy2021dynamic. The idea is to assign ordering probabilities for the risk factors in a similar manner as we discussed in Section ((ref)). However, instead of selecting the best factor ordering over time, we average the outputs of each possible order, weighting by their order probabilities. For more details of the DOL procedure we refer to the paper of levy2021dynamic.

Empirical Analysis

In order to test how the DRFDM performs in a real world problem, in this section we compare our method in terms of out-of-sample statistical and portfolio performance compared to different model settings and the Wishart-DLM. Our data is based on weekly stock returns from the Standard & Poor's 500 index. A stock is included if it was traded over the full horizon 2002-2020:5 and was a constituent of the index at some period during that time, resulting in $N = 452$ stocks over 959 time periods. The data was colected from Bloomberg Terminal and log-returns were used for the analysis. In our main empirical analysis, we use as observable risk factors the 5 factors of fama2015five downloaded for the same period from the website of Kenneth French.\footnote{The data is available at \url{https://mba.tuck.dartmouth.edu/pages/faculty/ken.french/data_library.html}. We transformed daily to weekly factor returns.} The five risk factors are the market excess return (MKT), a size factor (SMB), a value factor (HML), a profitability factor (RMW) and a investment factor (CMA). In our portfolio performance below, we also show results when considering different subset of risk factors, such as the first three factors (MKT, SMB and HM), the three factors plus the Momentum factor (carhart1997persistence) and the five factors plus Momentum.

table[table omitted — 560 chars of source]

We group stocks based on eleven sectors of the Global Industry Classification Standard (GICS). It makes easier the visual interpretation of figures below. Table (ref) list the number of stocks in each sector.

First, we illustrate in Figure (ref) the mean posterior factor loadings of our DRFDM with the five factors for three specific time periods. Each row represents the factor loads of an individual stock return loading on different risk factors of each column. The first figure at the left is referred to a calm time period, while the second and third dates were intentionally chosen to consider highly turbulence periods in the stock market, the Global Financial Crisis of 2008 and the the Covid-19 pandemic. The white colour represents factor loadings equal to zero. For the three periods, there is considerable sparsity on factor loadings, being stronger at the end of 2006 and weaker for the other stressed periods. Also note that not just the sparsity is changing over time, but also the mean of factor loadings are substantially varying, with periods of higher and lower values, with signs of many asset returns changing for the same risk factor. One interesting aspect of Figure (ref) is the great importance of the market factor for asset returns, being quite rare to observe sparsity on its loadings.

figure[figure omitted — 636 chars of source]

To clarify the dynamic sparsity pattern of our approach, Figure (ref) in Appendix C focus on the time-varying movements of posterior mean factor loadings of two specific stocks of different sectors over time. In fact, Figure (ref) highlights how differently they load on risk factors, with one company inducing much more sparsity than the other. Interestingly, with the exception of the market factor, all factor loadings are set to zero at least once. Also, depending on the company, some factor loadings are set to zero for the majority of the sample period.

Although the estimation of factor loadings plays an important role, they are not the only ingredients to predict the covariance matrix of returns. As discussed in previous sections, the investor also needs to learn the time-variation on the factor covariance matrix. Figure (ref) shows the predicted mean of the time-varying factor correlation matrix over the same selected three periods. The first aspect to note is how the correlations among factors change over time. As an example is the sign changing in the correlation between the MKT and the HML factors at the end of 2006 to the stressed periods. Additionally, it seems that during bad periods the correlations tend to be stronger, specially for the three factors MKT, SMB and HML. One exception is the CMA factor, presenting very low factor correlations during these bad times.

figure[figure omitted — 712 chars of source]

The evolution of correlations among stock returns is illustrated in Figure (ref), where we display the predicted mean correlation matrix for the selected periods. Stock returns are grouped according and following the same ordering as in Table (ref). Considering the end of 2006, the sparser factor loadings in Figure (ref) and weaker factor correlations of Figure (ref) are translated into a lighter correlation matrix of returns. Again, as lighter the colours the closer to zero the correlations are. Note a sligthly stronger correlation among stocks from the Financial sector and how they are also more correlated to almost all other sectors. This patterns are repeated for the three selected periods. At the other hand, we can note how stocks from the Health Care sector are much less correlated with other stock returns. When we focus on the stressed periods of the Global Financial Crisis and the Covid-19 pandemic, the higher factor loadings and greater factor correlations are translated in higher stock return correlations. It is an expected pattern, since in bad market periods stock returns variations tend to be more related to factor movements. Interestingly, during the Covid-19 pandemic the Consumer Staples sector and some companies from the Energy sector presented a strong correlation reduction, while the Financial sector is very high correlated not just among its own companies but with other sectors.

figure[figure omitted — 517 chars of source]

Finally, Figure (ref) investigates the evolution of risk factors inclusion probabilities. It can be defined as the posterior probability of an asset return including a specific covariate variable. Hence, for a given asset return $j$, the inclusion probability of a risk factor can be computed as the sum of probabilities of all univariate models for return $j$ including this specific risk factor. Since each asset return has different inclusion probabilities for different risk factors, we display in Figure (ref) the cross-sectional average of inclusion probabilities for each period of time. We can notice that the probability of including the market factor is considerably higher than all other risk factors, which demonstrates its importance in dictating common movements in the cross-section of stock returns. Since the market factor inclusion probability always fluctuates slightly close to 100%, we conclude that models including this factor tend produce much stronger predictive densities compared to models excluding the traditional market factor. It is interesting to note a slight drop in the importance of the market factor since 2016, but with the advent of the Covid-19 pandemic, its importance returned to past values. Indeed, except the CMA factor, all other risk factors have become substantially more relevant since the beginning of the pandemic. This increase in importance was also observed during the Great Recession for SMB, HML and CMA. The general assessment after the Great Recession is that the HML factor demonstrated a higher inclusion probability than the SMB, RMW and CMA factors for the great majority of the time. Therefore, Figure (ref) clearly shows how our DRFDM is able to infer about the importance of different risk factors in a dynamic fashion, providing evidences of a time-varying impact of different risk factors on stock returns.

figure[figure omitted — 396 chars of source]

In what follows in the next subsections, we study analysis of several variants and restrictions of our DRFDM based on different discount factor specifications and risk factors to be considered. We let the first four years of data (from 2002 to 2005) as a training period and we perform statistical and portfolio out-of-sample evaluation for the next years. Therefore, we discard the first 208 data points to train our models and use the next 751 data points for evaluation.

\paragraph{Model specifications:}Before going to the statistical and portfolio analysis, we detail here several model variants and restrictions to be compared. Different model settings are based on previous tests and experience using the initial training sample and values are similar to the commonly applied by the econometric literature (gruber2017bayesian and zhao2016dynamic):

itemize• DRFDM: This model considers the 5 Fama-French factors (fama2015five) as asset returns parents, applying dynamic risk factor and discount factors selection. Hence, it learns automatically the variation in betas (factor loadings), the variation in factor covariances, the variation in return volatilities and the selection of the best risk factors for each individual asset return and for each period of time. In terms of discount factors for different degrees of variation in coefficients, we let the model choose between: $\kappa_{r} \in \{0.99, 0.995, 1 \}$ for return volatilities; $\delta \in \{0.998, 0.999, 1 \}$ for factor loadings; and $\kappa_{f} \in \{0.999, 1 \}$ for factor volatilities. Model probabilities are computed using a forgetting factor $\alpha=0.99$. • DRFDM ($\alpha=0.98$): The same as DRFDM, but considering $\alpha=0.98$ for model probabilities. • DRFDM ($\alpha=1$): The same as DRFDM, but considering $\alpha=1$ for model probabilities. In this case, it applies BMS, the static version of DMS. • DRFDM (No Spars.): A dense version of DRFDM. It does not induce sparsity by dynamic factor selection. Therefore, all 5 Fama-French factors are always considered. • 3F-DRFDM: The same as DRFDM, but restricting the set of risk factors to MKT, SMB and HML (fama1992cross). • 4F-DRFDM: The same as 3F-DRFDM, but including MOM as an additional risk factor (carhart1997persistence). • 6F-DRFDM: The same as DRFDM, but including MOM as an additional risk factor. • W-DLM: It is the standard multivariate Wishart-DLM (see prado2010time, Ch. 10). We use $\delta=0.997$ for the local level evolution discounting and $\kappa = 0.99$ for the multivariate volatility discount factor.\footnote{For both W-DLM and Factor W-DLM we set the initial prior location $S_{0}$ to a diagonal matrix whose diagonal elements are given by 0.1. We also tested using the residual variances over the training period from an OLS model, but using 0.1 instead produced a stronger benchmark. moura2020comparing have shown that using a Wishart process with a shrinkage towards a diagonal covariance matrix delivers better portfolio performance with lower turnovers.} • Factor W-DLM: It is a multivariate Wishart-DLM using the 5 Fama-French factors as observable common predictors. We have considered $\delta=0.997$ for states discounting and $\kappa = 0.99$ for the multivariate volatility discount factor.\footnote{We model risk factors in a separate W-DLM and predict asset returns and covariances conditional on risk factor predictions.}

We also show different combination of models still applying dynamic factor selection, but considering the following restrictions:

itemize• TVB: A model using time-varying betas (factor loadings), where $\delta = 0.999$. • CB: A model using constant betas (factor loadings), where $\delta = 1$. • FSV: A model using factor stochastic volatility, where $\kappa_{f} = 0.999$. • FCV: A model using factor constant volatility, where $\kappa_{f} = 1$. • SV: A model using return stochastic volatility, where $\kappa_{r} = 0.995$. • CV: A model using return constant volatility, where $\kappa_{r} = 1$.

As we explain in Appendix A, discount factors equal to one represent the case of no variation in coefficients, while discount factors lower than one induce time-variability in coefficients. Hence, the DRFDM is the most flexible model, since it allows to switch between constant and time-varying parameters and dynamically selects different subset of risk factors over time if it is empirically desirable.

Forecasting accuracy

In order to evaluate different model settings in terms of statistical accuracy, we compute density forecasts and hit rates. As we explain in details in Appendix B, for each period of time an individual asset return $j$ can load in different subsets of risk factors parents. Hence, in order to generate an asset return forecasting for the next week, the investor considers the predictions obtained from the specific risk factor parents for that asset and prior coefficients:

$$ f_{jt|t-1} = a_{j \alpha t}+\lambda_{pa(j) t|t-1}^{\prime} a_{j \beta t} $$

where $f_{jt|t-1}$ is the asset return $j$ forecasting for the next week and $\lambda_{pa(j) t|t-1}$ represents a vector of risk factor predictions. It is important to highlight here that for our Dynamic Risk Factor Dependency Model, this set of risk factor dependencies changes depending on the period of time and for different asset returns under analysis. Finally, $a_{j \alpha t}$ and $a_{j \beta t}$ are the prior means for time $t$ for, respectively, the intercept and factor loading distributions given all data available until time $t-1$.

In Table (ref) we show sign accuracy (Acc.) and the Log-Predictive Density (LPD)\footnote{We exclude from the analysis the measures related to risk factors, focusing only on Acc. and LPD for the $N$ asset returns, which is the main interest of our portfolio allocation exercise.}. The first is computed as the sum of the log of the 1-step ahead predictive densities over the evaluation period as in Equation ((ref)) for the $N$ assets available, and higher values represent better performance. The LPD gives a sense of how well a model performs out-of-sample considering its whole predictive distribution and not just its mean. Therefore, it suits quite well to our portfolio analysis, since it also takes into account the impact of the covariance structure. This metric has been applied recently in many bayesian econometric papers and has become common practice for model comparison (see koop2013large, \citealp*{zhou2014bayesian}, gruber2017bayesian and \citealp*{mcalinn2020multivariate}). The second statistical metric is simply computed as the number of corrected sign predictions averaged across all assets and over the out-of-sample evaluation period. Hence, we represent the the accuracy of model $k$ over the evaluation period as

$$ Acc_{k} = \frac{\sum_{j} \sum_{t=(T_{0}+1)}^{T} \mathbbm{1}_{\{sign(f_{jt|t-1}^k) = sign(r_{jt})\}}}{(T-T_{0})N} $$

where $T_{0}$ represents the training period and $\mathbbm{1}_{\{sign(f_{jt|t-1}^k) = sign(r_{jt})\}}$ is an indicator function, which is equal to one when the sign of forecasted asset return $j$ for period $t$ using model $k$ is the same sign of the actual observed asset return $j$ for period $t$.\footnote{A sign is equal to 1 when returns are positive and equal to -1 when are negative.}

Table (ref) provides statistical out-of-sample performances. First, what can be noticed is that all different settings are performing much better than the Wishart approach in terms of density forecast. In general, models with constant parameters produce worst density forecast as well, specially with CV. In terms of the number of factors to be considered, the density forecast measure is quite similar among different choices, with a slight worsening for models including the CMA and RMW factors. In terms of correct return sign forecasts, differences are quite small. However, the recent literature on empirical asset pricing has shown evidence that even models providing only a small improvement in return predictability are capable to generating significant impacts on portfolio performance over time (\citealp*{chinco2019sparse}, \citealp*{gu2020empirical}, \citealp*{jiang2020re} and levy2021time). For instance, one may note in Table (ref) that the model with no sparsity on factor loadings produced very similar statistical performance than those models inducing sparsity, with just a small statistical deterioration. As we show in the empirical portfolio analysis (section (ref)), achieving sparsity promotes strong mean and variance prediction stability. In a high-dimension portfolio allocation with thousands of factor loadings for each period of time, a reduction on the parameter space and small differences on predictions are able to generate great improvements on final portfolio performance. We can also conclude from Table (ref) that models with constant factor loadings tend to produce worse out-of-sample accuracy in terms of both predictive densities and hit rates.

table[table omitted — 1,227 chars of source]

It is important to notice that even though both W-DLM and Factor W-DLM provide much weaker predictive densities than other model settings, they show stronger out-of-sample accuracy in terms of hit rates, with the Factor W-DLM showing the same accuracy as the main DRFDM. However, as we show in the empirical portfolio allocation section, both Wishart approaches produce portfolios with performances considerably lower than our DRFDM. In fact, this results confirm the evidence obtained in cenesizoglu2012return, where the authors show stronger correlation between density forecast and final portfolio performance than with point forecast metrics. Since any point forecast metric is not able to incorporate any uncertainty around predictions and the covariance structure among returns, it tends to generate a lower correlation with final portfolio performances.

Finally, at the bottom of Table (ref) we provide the hit rate using the momentum signal as a return predictor. In this case, instead of using an econometric model, the investor only considers the momentum from previous 12 months as a measure to predict future returns. This type of approach has received increasing attention in the recent literature and has been applied to high-dimensional portfolio problems (\citealp*{engle2019large}, \citealp*{de2018factor} and moura2020comparing).\footnote{Since we use weekly data, we consider the momentum as the average of previous 52 returns on each stock but excluding the most recent 4 returns.}

Dynamic Portfolio Allocation

After describing the set of possible models considered in this study and their statistical performance, we discuss how our DRFDM is able to improve final investor decisions. We will take the perspective of an investor who allocates her wealth among all different stocks available in the dataset. At each period of time, the investor applies two steps. The first is to use the econometric method to generate one-week ahead forecasts of return means and covariances. In the second step, she dynamically rebalances the portfolio by finding new optimal portfolio weights. In our main analysis, we perform a Mean-Variance portfolio optimization, where the investor uses both the vector of predicted mean stock returns and its respective predicted covariance matrix, as we describe below. Although it is not the focus of our analysis, in Appendix C we also show additional results for the case where the investor only considers the predicted covariance matrix for a Global Minimum Variance portfolio. Hence, from this setup we are able to assess the economic value of the DRFDM with different settings within a dynamic framework, implementing an efficient-frontier strategy subject to return target or a minimum-risk portfolio. Below we describe in more details the portfolio strategies.

Mean-Variance Portfolio

The Mean-Variance Portfolio (MVP), also know as Efficient Frontier (EFF) portfolio or Markowitz portolio due to the seminal work of markowitz1952portfolio, solves the following investment problem in the absence of short-sales constraints:

equation[equation omitted — 325 chars of source]

where $ \mathbf{1}$ is a vector of ones, $\boldsymbol{\Sigma_{t}}$ is the covariance matrix of returns for time $t$, $\boldsymbol{\omega}_{t}$ represents portfolio weights, $\boldsymbol{\mu}$ is the expected return vector and $\tau$ is the return target. Replacing $\boldsymbol{\mu}$ by the predicted vector of returns from our econometric approach $\boldsymbol{f}_{t|t-1} = \boldsymbol{\alpha}_{t|t-1} + \boldsymbol{\beta}_{t|t-1} \boldsymbol{\lambda}_{t|t-1}$\footnote{More details about predictive moments computation can be found in Appendix B.}, and $\boldsymbol{\Sigma_{t}}$ by the respective point estimate of the predicted covariance matrix, $\widehat{\boldsymbol{\Sigma}}_{t|t-1}$, the expression for the optimal portfolio weights can be expressed as:

equation[equation omitted — 229 chars of source]

where $A= \mathbf{1}^{\prime} \widehat{\boldsymbol{\Sigma}}_{t|t-1}^{-1} \mathbf{1}$, $B= \mathbf{1}^{\prime} \widehat{\boldsymbol{\Sigma}}_{t|t-1}^{-1}\boldsymbol{f}_{t|t-1}$ and $C=\boldsymbol{f}_{t|t-1}^{\prime} \widehat{\boldsymbol{\Sigma}}_{t|t-1}^{-1} \boldsymbol{f}_{t|t-1}$ . Following engle2012dynamic, in our main results we have considered an annualized return target of $\tau = 10\%$, but we also report additional results for $\tau = 15\%$ and $20\%$.

In Appendix C we also show results for a restricted portfolio optimization, where we solve a problem similar to ((ref)), but including one additional restriction: the maximum (absolute) weights on individual stocks to be 5%.

Global Minimum Variance

In Appendix C we show as additional analysis the results for the Global Minimum Variance portfolio (GMV) with no short-sales constraints. The main investment problem of the GMV is to reduce total portfolio risk and can be expressed as:

equation[equation omitted — 213 chars of source]

where again $\boldsymbol{\Sigma_{t}}$ is the covariance matrix of returns for time $t$ and $\boldsymbol{\omega}_{t}$ represents a vector of portfolio weights. Replacing $\boldsymbol{\Sigma_{t}}$ by the predicted covariance matrix, the optimal portfolio weights for the GMP is given by:

equation[equation omitted — 190 chars of source]

Since the GMV focus is to reduce portfolio volatility, it ignores the use of the mean of returns and tend to generate lower Sharpe ratios in practice. Although it is not widely used in practice, this approach can be viewed as a way to evaluate covariance matrix estimation and is still a common practice in academic works. Besides the GMV portfolio as in Equation ((ref)), we also show results when we include an additional restriction of maximum (absolute) weights on individual stocks to be 5%, as we did on the MVP case.

Economic Performance Measures

After producing forecast outputs to dynamically build portfolios, we backtest our models in terms of portfolio performance out-of-sample. Investors face portfolio allocation problems and, at the end of the day, it is not just about out-of-sample predictability, but how predictions are translated into better final decisions, i.e., better portfolio choices. With this in mind, we evaluate portfolios based on financial metrics, such as portfolio turnovers, annualized mean excess returns (Mean), standard deviations (SD) and Sharpe ratios (SR). The latter is commonly used among practicioniers in the financial market and by academics. Despite their popularity, those portfolio metrics are unconditional measures and are not well suited for dynamic allocations with time-varying and sequential predictions (see marquering2004economic). Also, they do not take into account the investor risk aversion. In order to overcome this problems and improve our model comparisons we follow fleming2001economic and provide a measure of economic utility for investors.

We compute ex-post average utility for a mean-variance investor with a quadratic utility. As in fleming2001economic and della2009economic we can calculate the performance fee that an investor will be willing to pay to switch from the traditional Wishart Dynamic Linear Model (W-DLM) benchmark to our Dynamic Risk Factor Dependency Model. The performance fee is computed by equating the average utility of the W-DLM portfolio with the average utility of the DRFDM portfolio (or any other alternative portfolio), considering the latter with a management fee $\Phi$:

equation[equation omitted — 227 chars of source]

where $\gamma$ is the investor's degree of relative risk aversion and $R_{p,t}^{DRFDM}$ is the gross excess return of the DRFDM portfolio and $R_{p,t}^{W-DLM}$ is the gross excess return from the W-DLM portfolio. As in fleming2001economic, we report our estimates of $\Phi$ as annualized fees in basis points using $\gamma = 10$.\footnote{Appendix C provides additional results applying different values for $\gamma$ and main conclusions are maintained.}

All economic measures displayed in Section (ref) are already net of transaction costs (TC). Following marquering2004economic, we deduct transaction costs from the portfolio return ex-post. Despite the great majority of the papers related to covariance matrix estimation for portfolio allocation cited above do not take into account transaction costs in their findings, in our main results we consider TC = 5 bps of the traded volume in an effort to approximate our results to a real world example. In general, there is disagreement about which transaction cost to incorporate. In the past, many papers applied transaction costs of 50 bps, but recently this value has been substantially reduced for the most liquid stocks (french2008presidential). Hence, in the same spirit of moura2020comparing we display additional results for TC $\in \{0, 10 \}$ bps.

Additional Competing Models

Here we detail some additional covariance matrix estimation approaches included in our portfolio performance analysis. We consider some recent traditional benchmarks from the literature. We also show results for two benchmarks that do not require the use of econometric models to be estimated: the equally-weighted portfolio from our universe of stocks and the passive investment in the $S\&P$ 500 index. The latter can be viewed as a strong benchmark, since it is well known that is quite hard to beat the market.

itemize• EFM: this is a static estimator based on an exact factor model, where $\boldsymbol{\Sigma}_{f}$ is given by the sample covariance matrix of risk factors and the residual covariance matrix $\boldsymbol{\Omega}$ is a diagonal matrix filled with the sample variances estimates of residuals. We estimate residuals and factor loadings using four years rolling regressions. • DCC-NL: it is a dynamic estimator using the multivariate GARCH of engle2019large. • AFM-DCC-NL: it is the dynamic approximate factor model of de2018factor. \footnote{Following the main results from their paper, we have consider only the market factor.} • LW: it is the static linear shrinkage estimator of ledoit2004honey. • EWMA: the traditional exponentially weighted moving average estimator where $\boldsymbol{\Sigma}_{t+1}=(1-\lambda) \boldsymbol{y}_{t}^{\prime} \boldsymbol{y}_{t}+\lambda \boldsymbol{\Sigma}_{t}$. We consider two values for the decay factor, $\lambda \in \{0.97, 0.99\}$. • EW: a simple strategy considering a equal-weighted portfolio over our universe of stocks. As claimed by demiguel2009optimal, it tends to perform better than the simple unconditional covariance matrix of returns and it has been claimed to be hard to be outperformed. • $S\&P$: it represents the passive investment (buy-and-hold) on the $S\&P$ 500 index over the evaluation period.

It is important to highlight that for models EFM, DCC-NL, AFM-DCC-NL, LW and EWMA detailed above, we use the average momentum signal from the previous 12 months as vector of predicted mean returns to find optimal portfolio weights in Equation ((ref)). It follows a similar procedure as applied in engle2019large, de2018factor and moura2020comparing and can be viewed as a competitive method to forecast returns.

We recognize the existence of other recent competing models in the literature involving Bayesian analysis using MCMC methods, as we have described in Section (ref). However, the simulation schemes dramatically limit scalability and require repeat MCMC simulation analyses at each time point which makes the whole backtesting procedure computationally prohibitive. Although factor stochastic volatility models of kastner2017efficient and kastner2019sparse have been implemented in the R package factorstochvol, \footnote{See hosszejni2019modeling} the estimation is not sequential, which requires the model to be rerun at each time over the evaluation period. As argued by gruber2017bayesian, it will take several weeks to backtest a universe of stocks like ours. At the other hand, since our DRFDM estimation is sequential and does not require any simulation schemes, it is able to handle the whole estimation procedure in only few minutes.

Portfolio Performance

In order to evaluate the validity of the DRFDM out-of-sample in a real world context, we conduct a backtest analysis using the different portfolio strategies and model settings described in the previous sections. One of the great advantages of our approach is its flexibility and fast computation. For the sake of curiosity, using parallel computations among 32 cores we run our main model (DRFDM with 5 Fama-French factors) for the entire dataset and produced portfolio backtests in less than 10 minutes.\footnote{We also have repeated the exercise using almost one thousand stocks from the Russel 1000 Index and we were able to run the whole estimation procedure and portfolio backtest in less than 25 minutes. }

table[table omitted — 2,398 chars of source]

The main focus of our portfolio analysis is to find an econometric model able to satisfactorily handle the best balance between risk and return. Table (ref) present results related to the MVP portfolio with 5 bps transaction costs during the out-of-sample evaluation period (from January, 6, 2006 until May, 22, 2020). The main columns to be analyzed from this table are in terms of SR and annualized management fee ($\Phi$). What can be noticed from Table (ref) is that regardless of the selected model setting, all models outperform the EW portfolio suggested by demiguel2009optimal. It is interesting to notice that although the EW strategy presents the lowest weekly turnover, it has the worst SR because of its high volatility that is caused by the absence of a covariance structure among returns on its formulation. Also, the great majority of the models were able to outperform the passive investment on the $S\&P$500 index. In special, the SR from the DRFDM and its different subsets of risk factors counterparts dominate all other models, with DRFDM presenting an out-of-sample SR of 0.82, almost two times the SR from the $S\&P$500 index, the Factor W-DLM and the W-DLM benchmark. The latter performed quite similar to the $S\&P$500 index in terms of SR, but since it was able to produce much lower portfolio volatility, it generates utility gains for the investor compared to the market index. However, since our DRFDM was able to reduce volatilty even further and produce quite stable return predictions, it has produced a considerable utility gain for the investor. In fact, a mean-variance investor would be willing to pay an annualized management fee of 585 bps to switch fromm the traditional W-DLM benchmark to the DRFDM approach. Using different subsets of risk factors in our DRFDM setting also produced quite strong portfolio results.

We also included in our analysis what we call DRFDM (Mom signal), which represents the same model as the main DRFDM, but instead of using its predicted mean returns for portfolio optimization in Equation ((ref)), we used the same momentum signal as we did for competing models at the top of Table (ref). What we observe is that using the predictions from our approach generates more stable estimates and portfolio performance than using the momentum signal. Although the differences are small, the DRFDM (Mom signal) induces higher turnovers, what harms final portfolio performances as higher transaction costs start to be considered, as Table (ref) shows. Also, as Table (ref) had demonstrated, predictions from the DRFDM generated higher out-of-sample accuracy than the Momentum Signal. Hence, we see the mean predictions from our approach as a clear competitor to the classical Momentum Signal broadly used in the academic literature and among practitioners.

In terms of the benefits of inducing sparsity on the covariance structure of returns, we see from Table (ref) that the DRFDM (No Spars.) produced a much more volatile portfolio than when we allow to dynamically select the best risk factors for each stock return, which is translated in lower utility gains and SR for the investor. The benefits of time-varying sparsity on factor loadings are observed in all different portfolio setting of this paper, regardless of the optimization problem, amount of transaction costs or risk aversions. When the DRFDM set many factor loadings to exactly zero, it is indeed improving covariance matrix estimation by deflating its whole structure.

table[table omitted — 2,416 chars of source]

We also compare different models within our model structure but restricting their variation in coefficients to follow the same pattern for all periods of time by fixing their discount factors. For instance, models with constant volatilities (CV) tend to produce portfolios with higher volatilities and lower SR and utility gains than models with time-varying volatilities. It is interesting to notice that models with both time-varying factor volatilities (FSV) and time-varying factor loadings (TVB) were not able to considerably improve portfolio performance, giving evidence of no significant predictive power increment by allowing time-varying betas and factor volatilities for all periods of time. At the other hand, since our DRDFM was allowed to dynamically select different discount factors over time, it was able to switch between periods of low and high variation on parameters, producing better adaptation to the data and stronger portfolio improvements. In terms of the selected forgetting factor $\alpha$ for model probabilities, we see an increase in portfolio volatility when no forgetting is applied ($\alpha=1$) - which is equivalent to a static Bayesian model selection approach - and comes with a lower SR and utility gain for the investor. However, when a higher forgetting is allowed using $\alpha=0.98$, a considerable portfolio volatility reduction is observed, but worse mean excess returns and higher turnovers are produced, affecting final performance. Hence, despite models with different $\alpha$'s are still delivering robust results and are able to outperform several competitor models, it seems that applying the intermediate value $\alpha=0.99$ generates a better balance between risk and return.

table[table omitted — 2,422 chars of source]

When we focus on the competitor models at the top of Table (ref), we see lower SR and utility gains compared to our approach, jointly with much higher turnover. In special, the EWMA (0.97) produced large turnovers and volatility. Although the AFM-DCC-NL model presented lower SR and utility gains than DRDFM, it was the approach with lower annualized volatility, a pattern that is repeated throughout the different portfolio specifications in this paper. The DCC-NL and LW approaches were also able to deliver lower volatilities, but as the AFM-DCC-NL, they fail to produce a good balance between risk and return, delivering lower SR and utility gains. Since those approaches are quite unstable, they require the portfolio to be highly rebalanced over time, harming final performance. Figure ((ref)) in the Appendix compares turnovers from our DRFDM to the dynamic estimators AFM-DCC-NL and DCC-NL and it is clear the ability of the DRDFM approach to produce more stable rebalances and reducing turnovers, specially during the Great Recessions and the Covid-19 pandemic. It also can be seen as an advantage of DRFDM for portfolio managers interested to invest in low liquid markets with much higher transaction costs.

Tables (ref) and (ref) repeat the same portfolio procedure as Table (ref), using different transaction costs. The conclusion are quite similar to those described before. However, when a higher transaction cost of 10 bps is considered, we can notice from Table (ref) the considerable negative portfolio impacts on those competitor models with high turnovers, such as the DCC-NL, EWMAs, Factor W-DLM and W-DLM. Due to the characteristic of high diversification and low portfolio rebalancing changes from the DRFDM, it improves even further in terms of utility gains compared to the W-DLM benchmark, requiring the investor to pay an annualized management fee of 712 bps to change from the W-DLM to the DRDFM. When no transaction costs are considered, the W-DLM improves which reduce the utility gains of several models, including the DRFDM. However, our main model get a SR of 0.86, the same as the 6F-DRFDM and almost 23% higher than the SR obtained from the AFM-DCC-NL and DCC-NL models.

In Appendix C the reader can find several additional tables reporting results using different return targets, risk aversions, portfolio constraints and applying the Minimum Variance Portfolio optimization of Equation ((ref)). In Table (ref) we show annualized management fees when lower relative risk aversions are considered as a robustness analysis. In fact, the main conclusions remain the same for different levels of risk aversion, i.e., inducing time-varying sparsity and dynamically selecting different discount-factors for variation in coefficients by our DRFDM is able to considerably add economic value for the investor compared to the traditional W-DLM benchmark and other common competitor models regardless of the risk aversion. Tables (ref) and (ref) show mean-variance portfolio performances using higher annualized portfolio return targets of 15% and 20%, respectively. What can be observed is that not only the conclusions are the same as those using a return target of 10% from Table (ref), but are even stronger in terms of Sharpe Ratio and management fees the investor would pay to use the DRFDM. The same happens from the GMV optimization in Tables (ref), (ref) and (ref) using different transaction costs. The DRFDM and its different specifications were able to deliver volatility reduction compared to the W-DLM and produce extremely lower turnovers and economic value for the investor. In spite of the fact that the DCC-NL, AFM-DCC-NL and LW have generated lower SD than our DRFDM, they fail to produce what is the advantage from our approach and is the main investor interest at the end of the day: better returns adjusted by risk and strong utility gains. Last but not least, Tables (ref) and (ref) show mean-variance optimization and global minimum-variance portfolios with maximum weight constraints, where little impacts were observed to our approach because of the high diversification quality of the covariance matrix from the DRFDM, so the vast majority of portfolio weights were already below the 5% limit in absolute terms when the formulas from Equations ((ref)) and ((ref)) were being applied. Interestingly, competitor models with high turnovers and low portfolio diversification, such as EWMAs and Wishart models were able to considerably reduce portfolio turnovers and risks.

Summarizing, the main conclusions from different tables and results in this empirical section display similar informations: regardless of the parameters set by the investor, she is always benefited from the dynamic choices made by the DRFDM, producing lower risks, portfolio turnovers and utility improvements. Therefore, a mean-variance investor who dynamic learns about best risk factors, variation in coefficients and volatilities in an automatic and online fashion improves final decisions and utility measures compared to the benchmarks.

Conclusion

Dynamic portfolio allocation in high-dimensions require a refined estimation of the covariance matrix of returns. The goal of this paper was to introduce a fast and flexible multivariate model for returns that is able to improve predictions for bayesian portfolio decisions. Inspired by the Cholesky-style framework for multivariate inferences, we impose economically motivated risk factor dependencies that are able to solve the curse of dimensionality. Due to the low dimension of the risk factors, we are able to deal with the problem of ordering uncertainty in a sequential fashion. The conjugate format for foward filters and the use of discount factors for state evolutions and dynamic model probabilities allow our model to sequentially select the best specification choices for each asset return in parallel, such as best risk factors, degree of variation in factor loadings, volatilities and factor volatilities. We show that by the use of dynamic factor selection, we can achive higher model parsimony by time-varying sparsity on the parameter space in an online fashion.

We have found that the Dynamic Risk Factor Dependency Model is able to improve final portfolio decisions compared to the traditional W-DLM benchmark, the Equally Weighted portfolio, the passive investment on the $S\&P$ 500 index and many other competitor models in the literature. It generates not just risk reduction and better Sharpe ratios, but significant lower portfolio turnovers and utility gains for investors. We show that a mean-variance investor will be willing to pay a considerable management fee to switch from those strategies to the DRFDM approach. Also, since our approach does not require expensive MCMC schemes to draw from posterior distributions, we can backtest high-dimension portfolios in a matter of few minutes. This is good news for portfolio managers who are interested to improve investment strategies in a financial world with a large number of assets available, high model uncertainties and rapid and complex changes over time.