Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
78,259 characters · 19 sections · 66 citation commands
Factor Investing: A Bayesian Hierarchical Approach
Bayesian methods are widely applied in financial studies, ranging from stock return prediction to volatility modeling to asset allocation.\footnote{Both avramov2010bayesian and jacquier2011bayesian provide excellent academic surveys on Bayesian methods in financial econometrics and portfolio analysis.} They allow researchers to incorporate prior economic beliefs in a statistical model and evaluate uncertainty in parameter estimation or even the choice of models. Seminal Bayesian finance work, such as kandel1996predictability and barberis2000investing, point out that the estimation risk is non-negligible in a simple stock-bond allocation problem when stock returns are predictable. Successful deployment of the mean-variance efficient portfolio framework requires estimating conditional expected returns and their covariance matrix while accounting for estimation risk. By contrast, traditional frequentist statistical approaches or modern machine learning predictive methods fail to properly evaluate such prediction risk. The estimation risk issue dramatically increases when investors allocate funds among multiple assets due to the high dimension of the parameter space.
This paper introduces a Bayesian hierarchical (BH) approach to investigate the asset allocation problem when returns are predictable.\footnote{If returns are unpredictable, the mean-variance efficient portfolio is time-invariant. Therefore, the econometric interest lies in the estimation property of unconditional expected returns and the covariance matrix.} In particular, we model returns by market-timing macro predictors and assume lagged fundamental characteristics drive their predictor strength (regression coefficients). Our approach estimates regression coefficients and the covariance matrix jointly, thus making possible the creation of the mean-variance efficient portfolio. The benefits of our approach are twofold. First, Bayesian methods allow one to model expected returns and the covariance matrix jointly and learn the uncertainty of the estimation. Second, our BH approach fits multiple assets separately while sharing data information through the hyperparameters of a hierarchical prior distribution. This data sharing feature improves individual asset fitness while maintaining model heterogeneity for different assets.
The empirical exercise aims to allocate funds across multiple risky assets based on market-timing macro predictors and fundamental characteristics. The investment universe includes sector portfolios, tradable risk factors, and characteristics-sorted portfolios. We are interested in sector or factor rotation under different macroeconomic conditions. Notably, we project time-varying coefficients of each asset onto its fundamental characteristics. This setup is similar to avramov2006predicting, but they instead assume coefficients of all asset returns are driven by the same macroeconomics predictors. Consequently, the conditional expected returns and covariance matrix are driven by both macro predictors and fundamental characteristics; thus, we can study factor rotation under changing macroeconomic conditions over time. Furthermore, the mean-variance efficient portfolio created by the BH approach serves as a stochastic discount factor, and it can explain most of the known asset pricing anomalies.
The rest of the paper is organized as follows. We first introduce the econometric and financial motivation in section (ref) and (ref), and report the methodology and empirical overviews in section (ref) and (ref). Section (ref) reviews relevant literature in Bayesian econometrics and empirical asset pricing. Section (ref) introduces our BH approach setting, the corresponding Markov chain Monte Carlo (MCMC) sampling algorithm, and the portfolio construction scheme. Section (ref) documents the empirical performance for asset return prediction, portfolio evaluations, and asset pricing implications. Section (ref) concludes the paper.
Documenting asset return predictability is a popular research topic among both academic researchers and practitioners. The existing empirical literature focuses on testing the existence of return predictability for a single market index (e.g., S&P 500) using market-timing macro predictors, such as dividend yield, treasury bill rates, inflation, and so on. However, some researchers argue most of these predictors are unstable or even spurious welch2008comprehensive. The forecasting power of these predictors undoubtedly varies substantially for different assets, and the current literature struggles to find consistent predictability evidence across multiple assets huang2020time.
Particular attention should be paid to modeling multiple assets. In this case, the time series of asset return is usually short (e.g., monthly data with hundreds of observations), but the number of time series is relatively large (e.g., hundreds or even thousands of assets). Researchers either give up the potential benefit from massive data and model each asset independently (time series modeling) or ignore the heterogeneity of assets, stack all time series together, and train a single model (pool modeling).\footnote{For example, gu2020empirical use pooled modeling, whereas huang2020time use time series modeling.} Neither way properly takes advantage of information from massive data. Time series modeling has poor individual asset fitness due to the small sample size being used, whereas pooled modeling is naive and loses heterogeneous signals. Stock return data usually suffer from extremely low signal-to-noise ratios, and the predictive power of predictors changes under different macroeconomic conditions. Therefore, improving model fitness is critical.
Our BH approach sheds light on this problem from a Bayesian perspective. BH modeling provides a useful framework for studying the joint predictability for multiple assets. The hierarchical model\footnote{Chapter 5 of rossi2012bayesian is an excellent textbook reference for hierarchical modeling.} provides a way to share information across numerous assets. Hence, it can handle the cross-sectional return dependence. The sampling process of our BH approach can be interpreted as two intuitive cyclic steps. The first step (information sharing step) takes condition on the hierarchical prior and updates each asset's regression coefficients separately. Note the hierarchical prior provides information from all other assets. The second step (information grouping step) evaluates regression coefficients of all asset returns and updates parameters of the hierarchical prior to share information. Intuitively, the hierarchical prior builds an overline bridge linking multiple asset returns. The within-asset part describes the predictor-return dynamics, and the cross-asset part incorporates the heterogeneity of predictor existence and strength.
Despite the advantages above, our Bayesian framework is flexible enough to model heterogeneous time-varying coefficients driven by each asset's lagged fundamental characteristics. avramov2006predicting find these macro predictors are linked to the underlying business cycles for conditional investing or market timing, and their analysis focuses on individual stocks. Our paper pursues a similar goal: to assess the economic value of predictability and show portfolio strategies successfully rotate across different risk factor styles during changing business conditions. Furthermore, our BH approach models the covariance matrix simultaneously, thus enabling the creation of mean-variance efficient portfolios.
Thousands of stocks are traded on the secondary market in the U.S., beyond what investors would like to invest in. Academic researchers and practitioners generally perform dimension reduction on the vast individual stock universe. They usually study the asset allocation problem in a small number of tradable portfolios instead of thousands of individual stocks. kozak2020shrinking provide a Bayesian shrinkage method to estimate the efficient portfolio weights for the stochastic discount factor (SDF) in empirical asset pricing.
Sorting is arguably the most popular approach to create portfolios among financial economists. Researchers sort stocks based on values of one fundamental characteristic, such as market equity values, book-to-market ratios, and so on. Then, they split the continuum of stocks into several consecutive groups and create portfolios using stocks within each group (i.e., characteristics-sorted portfolios).\footnote{For example, some index funds or ETFs classify and group individual stocks into mega-cap, mid-cap, and small-cap portfolios, which is equivalent to sorting individual stocks by market equity values.} This sorting mechanism\footnote{feng2020deepchar discuss the security-sorting mechanism used in asset pricing.} only uses a simple relative ranking of firms in the characteristics and does not consider the scale nor build statistical models. Many papers document that stocks within the same group behave similarly, and the characteristics-sorted portfolios have relatively stable dynamic relations with macro predictors and fundamental characteristics.
Moreover, many characteristics-sorted portfolios are shown with relatively monotonic average returns, which implies return predictability through past characteristics. The long-short portfolio (long top and short bottom sorted portfolios) returns are treated as a proxy for an underlying risk factor or an anomaly fama1993common. The current trend of factor investing is driven by these intuitive innovations of dimension reduction in the vast stock universe. Investors can directly invest in a small number of risk factors or the characteristics-sorted portfolios. However, the enlarging zoo of risk factors feng2020taming creates a new challenge of selecting risk factors.
Our BH approach joins this literature branch by considering investing in many risk factors or characteristics-sorted portfolios to further reduce the dimension of the investing universe. We create mean-variance efficient portfolios based on our estimation of conditional expected returns and the covariance matrix. The cross-sectional asset pricing literature generally tests characteristics-sorted portfolios or risk factors. We follow this path to introduce the hierarchical structure for return predictability across multiple assets. Moreover, the mean-variance efficient portfolio constructed by our method can be viewed as the SDF when considering a broad cross section of assets. We demonstrate in section (ref) that it can explain many known anomalies.
Our paper bridges the gap between Bayesian econometrics and empirical finance. The highlights are as follows. First, we propose a predictive system for the cross section of asset returns using market-timing macro predictors, considering parameter uncertainty. Second, each asset's predictive model adopts heterogeneous time-varying coefficients driven by each asset's lagged fundamental characteristics. Third, the hierarchical prior structure allows the sharing of information while maintaining heterogeneity for different assets. Finally, our model estimates conditional expected returns and covariance matrix of assets jointly, thus making it possible to create mean-variance efficient portfolios.
The primary goal of our approach is to create a common predictor-return dynamic and study the optimal portfolio choice for $N$ risky assets\footnote{We mainly consider the tangency portfolio optimization in the portfolio analysis.}. Let $r_{i,t+1}$ denote the excess returns and let $x_t$ represent the vector of lagged macro predictors. Our predictive regressions provide a system for cross-sectional return prediction.
We assume $\alpha_{i,t}$ and $\beta_{i,t}$ are driven by characteristics and have a hierarchical prior distribution. We follow the recent development of machine learning in asset pricing, such as gu2020empirical and feng2020deep, and model asset returns by combining macro predictors and fundamental characteristics. The advantage of our BH approach is that each asset uses information not limited to its time series but borrows information from others. By contrast, all other prediction studies face a trade-off between sample size and model heterogeneity. A full Bayesian evaluation considers the parameter uncertainty for interval prediction and efficient portfolio optimization where machine learning fails.
We perform various empirical exercises in the U.S. equity market using monthly data from 1978 to 2018. With 10 market-timing macro predictors and 20 representative fundamental characteristics, we test three representative groups of portfolios: 10 sector portfolios (Sector 10), 20 long-short factors (Factor 20), and 100 characteristics-sorted portfolios (Char-Sort 100). A complete list of the factors and portfolios is presented in the appendix (ref). These three groups of portfolios have a diverse representation of the cross section. The purpose of the increasing number of assets is to evaluate our approach's robustness to a higher dimension asset universe.
In terms of asset return prediction, we find our BH approach outperforms many strong benchmark models, including Lasso, random forest (RF), principal component regression (PCR), and the Bayesian method in avramov2006predicting (AC2006). Although the Bayesian forecast is biased, our BH approach performs slightly weaker than the moving average, but is still robust. Furthermore, our BH approach has more accurate prediction interval coverage than alternatives because it estimates the covariance matrix jointly and considers the estimation risk of coefficients.
For efficient portfolio performance, we also find substantial positive gains using our BH approach. Our BH approach outperforms many robust benchmarks in terms of risk-adjusted evaluation, including the naive equally weighted portfolio, the passive investing portfolio (S& P 500), and the predictive methods listed above. In the past 20 years, our BH approach in sector investment delivers average monthly returns of 0.92% and a significant Jensen`s alpha of 0.32%. We find technology, energy, and manufacturing are the most important sectors in the past decade for sector rotation. For factor rotation, size, investment, and short-term reversal factors have been heavily weighted in the efficient portfolio. Finally, the SDF constructed by our BH approach explains mail anomalies that alternative models fail to.
The body of literature on Bayesian methods in asset pricing is extensive and long established.\footnote{Early studies include shanken1987bayesian, harvey1990bayesian, and mcculloch1991bayesian. Though not directly related, another body of Bayesian learning papers in asset pricing can be a future research direction for our models, such as lewellen2002learning, johannes2014sequential, and fulop2015self} pastor2000portfolio and pastor2000comparing introduce Bayesian asset pricing models to the portfolio optimization problem. avramov2004stock add time-varying alphas and time-varying risk premia into Bayesian asset pricing models, and avramov2006asset show explanations for the size and value effects through time-varying coefficients. We learn a similar structure from them and use lag characteristics and macro predictors to drive the time-varying coefficients while incorporating a novel, seemingly unrelated regression (SUR) model with hierarchical priors. Though the lens of factor investing, our paper identifies useful factors through their efficient portfolio weights.
Bayesian methods have also appeared sporadically in the literature of portfolio analysis. For the simple stock-bond allocation problem, kandel1996predictability and barberis2000investing focus on posterior predictive distribution of returns for the expected returns and covariance matrix, and avramov2006predicting extend the model by allowing time-varying coefficients, which are driven by characteristics. Unlike analyzing predictive distributions, our approach draws samples from the posterior by the MCMC approach. polson2000bayesian introduce a Bayesian seemingly unrelated regression (SUR) model with hierarchical prior and employ informative priors on the expected returns and covariance matrix. But they model the covariance matrix separately rather than learning it through the SUR model. We follow a similar SUR structure. In addition, we allow dynamic regression coefficients and learning the covariance matrix jointly with coefficients.
Finally, our paper is related to the recent literature on high-dimensional cross-sectional asset pricing models.\footnote{Other recent related publications include lettau2020PCA, demiguel2020portfolio, and freyberger2020dissecting.} gu2020empirical investigates machine learning algorithms to forecast asset returns by many firm characteristics and macro predictors. feng2020deep also find strong predictability evidence through a benchmark combination using characteristics-sorted portfolio returns. kelly2019characteristics adopt a characteristics-driven time-varying coefficient model for principal component analysis. Our study is similar to this literature in predicting returns by high-dimensional firm characteristics and macro predictors, but we propose a novel statistical model. In addition, we go one step further, from return prediction to portfolio optimization.
Section (ref) discusses the model settings. Section (ref) illustrates the reformulation of the likelihood as an SUR model. Section (ref) discusses the hierarchical prior specification. The MCMC scheme is presented in section (ref). Finally, section (ref) shows portfolio optimization based on information extracted by the MCMC algorithm.
The goal of a Bayesian investor is to maximize portfolio performance though a Bayesian predictive model of asset returns. Suppose he or she observes the historical return $R_t = (r_{1,t}, \cdots, r_{N,t})$ of $N$ assets, $P$ asset characteristics $z_t$ for each asset, and $Q$ market-timing macro predictors $x_t$. An investor updates his or her portfolio regularly, at time period $t$, models the joint predictive distribution of returns $f(R_{t+1} \mid D_t)$, and calculates asset allocation weight $W_t = (w_t^1, \cdots, w_t^N)$ accordingly. In the next time period, the realized return of the portfolio is
We denote all historical data observed up to period $t$, including asset returns $R_t$, characteristics $x_t$, and macro predictors $z_t$ as $D_t = (R_t, x_t, z_t)$. The asset allocation decision is made at the end of period $t$ based on the information in $D_t$.
If returns are unpredictable, the efficient portfolio can be constructed using the unconditional mean and covariance matrix of returns. If returns are predictable, one needs to learn the source and mechanism of return predictability. We follow the Bayesian predictive regression model in kandel1996predictability but adopt a conditional predictive formula with time-varying coefficients. For each asset $i$, its return at time period $t+1$ is assumed to be
where $x_{t}$ is the vector for $Q$ macro predictors. The residual vector $(\epsilon_{1,t+1}, \cdots, \epsilon_{N,t+1})^\intercal$ contains shocks to all asset returns and is assumed to follow a multivariate normal distribution $N(0, \Sigma)$ with a dense covariance matrix $\Sigma$.
Conditional modeling with time-varying coefficients is one solution for those “seemingly useless" predictors debated in welch2008comprehensive. The regression coefficients $\alpha_{i,t}$ and $\beta_{i,t}$ are assumed to be time-varying, driven by asset characteristics as follows:
where $\theta^{b}_{i}$ is a matrix coefficient of size $Q\times P$, and $z_{i,t}$ is the vector for $P$ portfolio characteristics. avramov2004stock gives a similar time-varying coefficients setup but assumes a common factor structure for all assets, and factor loadings and alphas are driven by lagged stock characteristics. If we plug the time-varying coefficients ((ref)) into equation ((ref)), we obtain an unconditional predictive regression in terms of $z_t$, $x_t$, and their interactions $z_t \otimes x_t$:
where $\otimes$ stands for the Kronecker product. Moreover, our linear Bayesian model allows for more intuitive results over black-box machine learning predictive models.
Notice equation ((ref)) models the $i$-th asset by macro predictors $x_t$ and its own characteristics $z_{i,t}$. Instead of modeling each asset separately, we prefer group modeling to estimate covariance and share information across different assets, which we achieve by SUR zellner1962efficient with the hierarchical prior on coefficients. The SUR models allow joint modeling of asset returns and the covariance matrix. The hierarchical prior provides a channel to share information across different assets while modeling each asset's return by its characteristics. polson2000bayesian provide a similar hierarchical SUR setting, but they model the covariance matrix by a separate shrinkage estimator. In the following two subsections, we present our SUR setup and the hierarchical prior.
Before discussing the hierarchical model, we simplify the notations of equation ((ref)) as
where $f_{i, t} = [1, z_{i,t}, x_t, (x_t\otimes z_{i,t})]$ indicates all regressors, symbol $\otimes$ denotes the Kronecker product, and $b_i = [\eta_i^\alpha, \theta_i^\alpha, \eta_i^b, \theta_i^b]$ is a vector of all corresponding regression coefficients. For each asset $i$, we stack the equations of different time periods $t$ as
where $r_i = (r_{i,2}, \cdots, r_{i,T+1})^\intercal$, $\epsilon_i = (\epsilon_{i,2}, \cdots, \epsilon_{i,T+1})^\intercal$, and $f_i$ is a matrix with $T$ rows.
The system of all $i = 1, \cdots, N$ assets is written as an SUR setup as follows. Stacking all equations, we have
where
Here, $R$ is an $NT \times 1$ vector of a stacked vector of firm returns, $F$ is an $NT \times NK$ block diagonal matrix, $B$ is an $NK \times 1$ vector, and $E$ is an $NT\times 1$ stacked vector of residuals.
Estimating the covariance matrix of many assets, $\Sigma$, is a challenging problem in general. The literature of portfolio optimization mostly develops a covariance matrix estimator independently from return predictions. Researchers either assume a low-dimensional factor structure or shrink the empirical estimator. Our SUR setup allows for joint modeling for asset returns and the covariance matrix.
We assume the covariance matrix of the stacked residual $E$ has the form $\text{Cov}(E) := \Omega = \Sigma \otimes I_N$, where $\Sigma$ is a dense matrix of cross-sectional covariance, and $I_N$ is an $N\times N$ identity matrix. This structure implies assets are correlated in the cross section but are independent across different periods, which is a standard assumption in the empirical asset pricing literature. This specification differs from the standard SUR model in zellner1962efficient as well as polson2000bayesian, who do not assume any constraint of $\Omega$ and estimate the $NT\times NT$ matrix directly. Therefore, we reduce the dimension of the parameter space dramatically from $N^2T^2$ to $N^2$.
Estimating the dense covariance matrix $\Sigma$ enables us to create the mean-variance portfolio. Many popular machine learning algorithms, such as Lasso, random forest, or deep learning, cannot model this common shock jointly with returns but assume a constant variance over time or across assets.
The current empirical literature mainly focus on the time-series predictability evidence for index returns, such as the S&P 500. Similar to the field of cross-sectional asset pricing, researchers are also interested in the common predictors across assets, such as individual stocks, sector portfolios, and sorted portfolios on market equities.
To study the joint predictability across assets, rather than putting an independent prior on $b_i$ for each asset, we assume a hierarchical structure with a common prior mean. Suppose $b_i$ has an independent and identical normal prior
where $\bar{b}$ is the prior mean for $b_i$, and $\Delta_b$ is a dense prior covariance matrix. Furthermore, the prior mean and covariance are assumed to have a second layer prior:
This setup is the standard normal-inverse-Wishart conjugate prior for $\bar{b}$ and $\Delta_b$. It assumes the average predictability of all predictors is $\bar{\bar{b}}$, and $\Delta_b$ represents the confidence of this predictability belief. We set $\bar{\bar{b}} = 0$ in practice, implying the belief a priori that the average predictability of those predictors is zero. Finally, this prior specification contains four groups of hyperparameters: $\{\bar{\bar{b}}, \Delta_{\bar{b}}, \nu_b, V_b\}$.
The normal prior on $b_i$ is equivalent to $L_2$ shrinkage of coefficients similar to the ridge regression. Because $\Delta_{\bar{b}}$ is a parameter to learn from the data, our model can adapt the shrinkage level based on the information from the data. This ability is the primary advantage of our BH approach over alternatives. When many potential and weak predictors exist, the hierarchical prior can properly shrink coefficients. From the frequentists' perspective, simply testing $H_0: \bar b_j = 0$ reveals the joint predictability implication of the jth predictor.
Notice the dimension reduction assumption $\Omega = \Sigma \otimes I_N$, where we assume the standard inverse-Wishart prior on $\Sigma$ with two hyperparameters $\nu_\Sigma$, and $V_\Sigma$ as
The MCMC sampler in section (ref) updates $\Sigma$, and $\Omega$ can be recovered by simple calculation.
The model likelihood function is multivariate normal:
Hence, the joint posterior distribution can be expressed as
Based on the above prior and likelihood specification, the next subsection describes the corresponding Gibbs sampler for model inference.
The design of our Gibbs sampler is straightforward because all priors are conjugate to the likelihood. As discussed in section (ref), we assume $\Omega = \Sigma \otimes I_N$ and update smaller $\Sigma$ instead of $\Omega$. The corresponding Gibbs sample is
Convergence is always a concern for MCMC algorithms. We confirm the convergence of the algorithm by examining the trace plot of posterior samples and calculating the effective sample size (ESS) for all parameters using the R package coda plummer2006coda. The trace plot stabilizes quickly, and we achieve an average effective sample size of 1,800 out of 2,000 posterior draws, which indicates that the sampler explores the posterior efficiently. See the appendix (ref) for details.
One can simply plug the Bayesian estimates $b_i$ and new data of $z_t, x_{i,t}$ into equation ((ref)) to make predictions for the next time period. We take the posterior average as our estimation for parameters. Specifically, for the kth draw in a total of K MCMC samples, we have
The conditional expected returns and covariance matrix proceeds similarly:
Notably, our BH approach is able to provide more information about the uncertainty of the parameters. For example, we can calculate the interval forecast for the vector of returns, or value at risk (VaR), based on the posterior variance of parameters. By contrast, most machine learning methods fail to estimate the covariance matrix and only give the point prediction of returns.
Hence, given the predicted return and covariance matrix, we create a mean-variance efficient portfolio by optimizing the standard utility function. Simply plug in the Bayesian estimates of $\mathbf{E}(R_{t+1} \mid D_t)$ and $\text{Cov}(R_{t+1} \mid D_t)$ for optimal portfolio weight calculation. The portfolio is built to maximize the mean-variance utility function:
where $R_{p,t+1} = W_t^\intercal R_{t+1}$ is the return of the portfolio with allocation weight $W_t$ and
The optimal weight is
The risk-aversion parameter $\gamma$ indicates an investor's risk preference. One convenient property for this optimal portfolio is
This formula is the same portfolio-efficiency estimate as kozak2020shrinking, who adopt a maximum a posteriori (MAP) estimator. They consider incorporating a shrinkage prior to shrink the cross section of risk anomalies: factors and characteristics-sorted portfolios. The efficient portfolio generated through a large cross section is supposed to be the stochastic discount factor that prices the cross section. Our approach can tell the same story. One major difference is that ours provides a full Bayesian estimate that allows for posterior inference, whereas the MAP estimate fails.
In the empirical exercise section, we place some constraints on the optimization of equation ((ref)). We restrict short selling, leverage and assume full position, or
This constraint eliminates leverage and short selling. We only consider the allocation of risky assets; thus, we assume the investor spends every dollar on assets. We observe macro predictors and fundamental characteristics in the current period. Given that fundamental characteristics drive time-varying predictor coefficients, our framework is a convenient way to provide or update the one-step-ahead or multi-step-ahead optimal asset allocation weights. We illustrate this convenient property in the empirical analysis.
Section (ref) introduces details of the data. The portfolio rebalancing scheme is listed in section (ref). Section (ref) illustrates the prediction performance comparison between our BH method and other workhorse benchmarks. In section (ref), we further compare the performance of the efficient portfolio and add the predictor usefulness evaluation. Finally, we show the results of asset pricing implications in section (ref).
Our data sample consists of extensive coverage of the U.S. equity market. The data are from January 1978 to December 2018. We test our approach on three major groups of assets, including 10, 20, and 100 portfolios. First, we follow the Fama-French 10-industry classifications (sector 10).\footnote{Fama and French form 10 sector portfolios at the end of June of year t with the Compustat SIC codes for the fiscal year ending in the calendar year t-1.} It classifies thousands of individual stocks into 10 portfolios. By analyzing this group of assets, we can also learn the impact of sector rotation through time. We also investigate the predictability of 20 risk factors (Factor 20) and 100 characteristics-sorted portfolios (Char-Sort 100), which are commonly used testing assets in asset pricing studies.
The number of selected portfolios ranges from 10 to 100 in the exercise. It helps us test our BH approach's robustness to different cross-section sizes for portfolio construction. Furthermore, these three asset groups are constructed in fundamentally different ways for the stock universe. It also allows us to assess the robustness of our BH approach to different types of portfolios.
We take 20 portfolio characteristics $z_{i,t}$ from the six major categories: momentum, value, investment, profitability, frictions (or size), and intangibles. Appendix (ref) lists details of all 20 characteristics. The data are from the public access library of hou2020replicating. Some minor differences exist because our work focuses on monthly asset return prediction. First, we modify the characteristics formulas, so they are updated monthly. Second, the characteristics-sorted portfolios are $5 \times 1$ univariate-sorted\footnote{We also follow the NYSE breakpoints and work on the same stock universe as Fama-French three factors.} every month. Third, the risk factors are long-short portfolios on the corresponding characteristics-sorted portfolios (top-bottom or bottom-top).
Following kelly2019characteristics, We standardize the cross section of the monthly characteristics of firms in the range of $[-1,1]$.\footnote{For example, the market equity in 2018 December is uniformly standardized to $[-1, 1]$. The firm with the lowest market equity is -1, and the firm with the highest market equity is 1. Every month, a “size" factor longs firms with size $< -0.6$ and shorts firms with size $\geq 0.6$. Therefore, this uniform standardization is a non-standard standardization that transforms the data onto $[-1, 1]$ every month. If a firm has missing values for some characteristics, the imputed values are 0, which implies the firm is not important in security sorting.} To calculate portfolio characteristics, we take the average of the sector's standardized characteristics and characteristics-sorted portfolios. For long-short risk factors, we take spread differences between the long and short portfolio characteristics.
Firm characteristics and market-timing macro predictors applied in this paper are the same as in feng2020deep. Those 10 macro predictors are listed in Appendix (ref). welch2008comprehensive study return prediction of S&P 500 returns using market-timing predictors, and amihud2002illiquidity also constructs the market illiquidity for the S&P 500 index for market timing. gu2020empirical use eight predictors of welch2008comprehensive in their analysis, all of which are covered in our list.
We estimate the model using a rolling window of the past 21 years. The prediction period is from January 1999 to December 2018. Given the constraint of computational resources, we only update the model annually (calendar year). Each year, the model is trained using a rolling window of the past 21 years (252 months). The model is fixed for the incoming year, but regression coefficients and predictors are updated monthly. Thus, we can provide an updated monthly forecast for expected returns. The dynamic asset allocation procedure is based on re-estimating and then rebalancing portfolios based on the updated monthly forecasts. We apply the same rolling-window strategy for other methods in the empirical comparison.
Below are the step-by-step details of our design of the dynamic portfolio optimization.
We compare our BH method with other commonly used machine learning methods, including Lasso, RF, and PCR. The tuning parameters of these approaches are selected by threefold cross-validation with a comprehensive list of parameter candidates. Every fold of validation contains seven consecutive years in the rolling window setting and addresses business cycles' impact on model selection. Additionally, we compare our approach with avramov2006predicting (AC2006)\footnote{AC2006 has a problem inverting matrices when the matrix of predictors is singular, due to their uninformative prior specification. Their model does not allow for heterogeneous predictors, $z_{i, t}$ for each asset. Therefore, we implement their model only with our 10 macro predictors. Specifically, we use the first five predictors in Appendix (ref) as their equity predictors, and the second five as macroeconomic predictors.}.
To gauge the performance of our Bayesian hierarchical approach, we compare it with alternative methods in terms of a few key metrics, including out-of-sample R-square, coverage, and length of the predictive interval. The measurements are defined as follows:
where $\mathbf{I}(\cdot)$ is the indicator function. We follow the literature and use moving average $\bar R_{i,t}$ as the point prediction benchmark in $R_{OOS}^2$ calculation. The out-of-sample predictive interval is at the 95% level (from 2.5% to 97.5% quantiles of the posterior distribution).
Note our BH approach offers a dynamic heterogeneous predictive interval for each asset, so we take the average interval across time and assets. Additionally, our Bayesian predictive interval is wider due to the uncertainty in parameter estimation. By contrast, most machine learning methods cannot estimate the variance or covariance of assets. Thus, we instead calculate the residual variance or covariance. Table (ref) presents results of the predictions. The table lists results of two different prior settings (tight vs. mild) \footnote{The two prior settings are mild, $\bar{\bar{b}} = 0, \Delta_{\bar{b}} = \text{diag}(0.1, K), \nu_b = 1001 + K, V_b = \text{diag}(3, K)$; tight, $ \bar{\bar{b}} = 0, \Delta_{\bar{b}} = \text{diag}(0.1, K), \nu_b = 5001 + K, V_b = \text{diag}(3, K)$. Note the mean of inverse Wishart distribution $IW(\nu,V)$ is $V / (\nu - p - 1)$ where $p$ is the dimension. Here, the tight prior implies a small prior mean of $\Delta_b$; thus, the prior standard deviation of $b$ is around 0.02, corresponding to stronger regularization on $b$. } for the BH method, indicating strong shrinkage and relatively flat prior parameter settings, respectively.
The first panel of table (ref) summarizes results for the past 20 years. We find the BH approach has a moderately higher $R_{OOS}^2$ than most other methods, including Lasso, RF, PCR, and AC2006. Admittedly, the slightly negative $R_{OOS}^2$ implies the BH approach does not outperform the moving average in the overall sample. However, if we focus only on the recent decade in the bottom panel, both the point and interval predictions of the BH approach are better than most others.
Additionally, the BH approach gives the best out-of-sample predictive interval coverage in the overall sample and subsamples. Notably, the better accurate interval coverage is due to wider intervals or considering the uncertainty of parameter estimation. The pattern is robust across all three types of portfolios: 10 sectors, 20 factors, and 100 characteristics-sorted portfolios. These are the empirical facts for the importance of accounting for estimation risk in return prediction.
Next, we proceed to dynamic asset allocation. We mainly study the risk-adjusted performance of portfolios and the asset pricing implication in this section. We compare the average monthly return, annualized Sharpe ratio, and Jensen's alpha of the monthly updated optimal portfolio using different methods. We also include the $1/N$ naive diversification strategy (equally weighted portfolio, or EW), and the buy-and-hold strategy (SPY, exchange-traded fund of S&P 500 index) as passive investment benchmarks. All of the active investing portfolios are updated monthly using the corresponding forecasts of conditional expected returns and covariance matrix.\footnote{To limit the impact of transaction fees, the monthly portfolio turnover rate is restricted to be smaller than 50%, and the maximum position of a single asset is 50% of the total portfolio.}
Table (ref) reports portfolio performance. All alternative methods fall behind passive investment (EW portfolio and S&P 500) in terms of average return, annualized Sharpe ratio, and Jensen's alpha for all three groups of assets. However, our BH approach gives the best numbers for long-only assets: 10 sector and 100 characteristics-sorted portfolios. Figure (ref) presents the evolution of \$1 since January 1998, where the BH approach enjoys the highest cumulative return after 20 years of investment.
For 10 sector portfolios, our BH approach outperforms alternatives in all categories and provides an economically and statistically significant Jensen's alpha. The capital asset pricing model (CAPM) $R^2$ or correlation with the market factor is low, indicating the BH portfolio is an excellent diversified strategy away from the market risk. Additionally, Figure (ref) exhibits the changing weights over time, which tells how the 10 sectors rotate in the past 20 years. Technology, energy, and manufacturing are the most heavily weighted sectors in the past decade. The BH approach also obtains desirable results on 100 characteristics-sorted portfolios. It achieves the highest average returns and a Sharpe ratio similar to S&P 500.
The 20 long-short factor investment case is interesting because investing in long-short factors is the practitioner's definition of factor investing. We also include passive investment benchmarks, such as an equally weighted portfolio using these 20 factors (EW) or Fama-French five factors. Note all approaches have flat returns after the 2008-2009 financial crisis (Figure (ref)). Most of these 20 factors probably do not provide significant risk premia after the crisis.
An equally weighted portfolio of 20 factors has the highest Sharpe ratio, which contradicts studies that the factor zoo should be sparse feng2020taming. Our BH approach achieves the highest monthly average return with a smaller Sharpe ratio, but the portfolio only invests a few factors in each period. Figure (ref) shows the rotation of 20 factors over time, where market equity (size), asset growth (investment), and lag monthly return (short-term reversal) factors are heavily weighted in the past 10 years.
Our BH mean-variance efficient portfolio is the stochastic discount factor (SDF) for the three different cross-sections of assets. Moreover, the efficient portfolio can also be viewed as a dimension-reduced version of the factor zoo (e.g., the first principal component for the 20 factors). Therefore, this section discusses the asset pricing implications of the efficient portfolio.
First, we evaluate our BH efficient portfolio's model-adjusted performance on commonly used factor models, including CAPM, Fama-French three factors and five factors. If the efficient portfolio cannot be priced by these asset pricing models and contains additional signals, the intercept “alpha" should be significantly positive. Table (ref) presents the results of the model-adjusted performance. Specifically, the BH efficient portfolio constructed by 10 sector portfolios generates significantly positive alphas over CAPM, whereas other methods fail. This monthly alphas' magnitude is economically significant, with 0.32% and 0.29% over CAPM and the Fama-French three-factor model. The BH efficient portfolio constructed by 20 factors has a marginally significant alpha 0.41% over CAPM, though it shrinks to an insignificant 0.18% over Fama-French five factors. Surprisingly, the equally weighted portfolio of 20 factors has superior performance, with a 70% annualized Sharpe ratio and a significant alpha over Fama-French five factors.
Second, we study whether the BH efficient portfolio can explain existing anomalies. We apply our BH efficient portfolio as the SDF model to price existing anomalies: 20 published risk factors related to our characteristics. Results are summarized in Table (ref). During the past 20 years, 8 out of 20 factors have marginally significant (10% level) alphas to CAPM and even 9 out of 20 to Fama-French three factors. By contrast, after controlling for the BH efficient portfolio constructed by those 20 factors, only two factors still have marginally significant alphas. We also check the first principal component of these 20 factors, which leaves six factors unexplained. In addition, we find the maximum alpha unexplained is 0.82% for CAPM, 0.81% for FF3, 0.48% for BH, and 0.63% for PCA. The evidence above suggests the BH SDF has the best pricing ability in this cross section of 20 long-short factors.
When returns are predictable, the mean-variance efficient portfolio framework's successful deployment requires estimating conditional expected returns and their covariance matrix while accounting for estimation risk. Bayesian methods shine in these areas. Our Bayesian hierarchical prior setting provides a way to model multiple assets other than separate time series modeling or pool modeling. It allows the sharing of information while maintaining specific heterogeneity of models for different assets. Furthermore, stock return prediction suffers from a low signal-to-noise ratio and varying predictive power over time. Our BH approach adopts heterogeneous time-varying coefficients driven by lagged fundamental characteristics to solve this problem.
Our empirical findings are also noteworthy. We study the U.S. equity market from 1978 to 2018 and find our BH approach provides better prediction performance than alternatives in terms of out-of-sample R-squared coverage of predictive intervals. The superior coverage performance is due to our BH approach accounting for the estimation risk of regression coefficients. Primarily for sector investing during the past 20 years, our BH approach gives average monthly returns of 0.92% and significant Jensen`s alpha of 0.32% . Moreover, we study the sector and factor rotations during the past decade. We find technology, energy, and manufacturing are important sectors, and size, investment, and short-term reversal are heavily weighted factors. In the end, the SDF constructed by our BH approach explains many risk anomalies that alternative models fail to.