EconBase
← Back to paper

Smoothing volatility targeting

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

140,018 characters · 17 sections · 105 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Smoothing volatility targeting

\thispagestyle{empty}

\centerline{\bf Abstract} We propose an alternative approach towards cost mitigation in volatility-managed portfolios based on smoothing the predictive density of an otherwise standard stochastic volatility model. Specifically, we develop a novel variational Bayes estimation method that flexibly encompasses different smoothness assumptions irrespective of the persistence of the underlying latent state. Using a large set of equity trading strategies, we show that smoothing volatility targeting helps to regularise the extreme leverage/turnover that results from commonly used realised variance estimates. This has important implications for both the risk-adjusted returns and the mean-variance efficiency of volatility-managed portfolios, once transaction costs are factored in. An extensive simulation study shows that our variational inference scheme compares favourably against existing state-of-the-art Bayesian estimation methods for stochastic volatility models.

Keywords: Volatility targeting, mean-variance efficiency, Bayesian methods, stochastic volatility models, variational Bayes inference.

JEL codes: G11, G12, G17, C23

{.2cm } {0.55cm}

\doublespacing

\pagenumbering{arabic}

Introduction

The widespread evidence that volatility tends to cluster over time and negatively correlates with realised returns have motivated the use of volatility targeting to dynamically adjust the notional exposure to a given portfolio. A conventional approach to volatility targeting builds upon the idea that the capital exposure to a given portfolio is levered up (scaled down) based on the inverse of the previous month's realised variance. The theoretical foundation lies in the evolution of the risk-return trade-off over time (see, e.g., moreira2017volatility).\footnote{Notice that the terms “volatility-managed”, “volatility-targeting”, “volatility-managing” are used interchangeably throughout the paper as they carry the same meaning for our purposes.} However, volatility management based on realised variance estimates is associated with a dramatic increase in portfolio turnover and significant time-varying leverage. This casts doubt on the usefulness of conventional volatility-managed portfolios, especially for large institutional investors with high {\it all-in} implementation costs (see, e.g., patton2020you).

Figure (ref) shows this case in point. The left panel shows the volatility-managed portfolio allocation based on realised variance estimates for three common portfolios; the market, and the size and momentum factors as originally proposed by fama1996multifactor and jegadeesh1993returns, respectively. Simple volatility targeting leads to a tenfold notional exposure compared to the original equity strategy. This excess leverage is pervasive across a broad set of 158 equity trading strategies which will be introduced in Section (ref). For instance, the middle panel in Figure (ref) shows that volatility targeting based on realised variance leads to a leverage between 1.8 and 4 times for more than 10%, and between 3 to 11 times for at least 1% of the original 158 equity strategies. This makes volatility-managed strategies potentially both risky and costly to implement, especially when volatility targeting is missed and/or forecasts are not sufficiently accurate (see, e.g., bongaerts2020conditional).

A simple approach towards cost mitigation is to reduce liquidity demand by slowing down the time-series variation in the factor leverage; this is often achieved by using less erratic estimates of risk, such as the realised volatility instead of the realised variance, or by introducing leverage constraints in the form of a capped notional exposure (see, e.g., moreira2017volatility,cederburg2020performance,barroso2021limits). While imposing leverage constraints may simplifies an empirical analysis, they do not regularise the often erratic monthly underlying volatility estimates and are typically set arbitrarily, absent sounded economic arguments for their optimal setup. In this respect, the economic value of leverage constraints is an indirect function of the statistical accuracy of the underlying volatility estimates.\footnote{This is akin a joint-test problem whereby leverage constraints are well-specified only to the extent that the assumptions underlying the volatility estimates are correct.}

In this paper, we propose an alternative approach towards slowing down liquidity demand in volatility-managed portfolios which is based on smoothing the predictive density of an otherwise standard stochastic volatility model. Our view is that by smoothing monthly volatility forecasts, one can regularise trading turnover and therefore mitigate the effect of transaction costs on volatility-managed portfolios. Such regularisation is achieve by a variational Bayes inference scheme which flexibly encompasses different smoothness assumptions irrespective of the underlying persistence of the latent state. Put it differently, our underlying assumption is that actual monthly returns' volatility may simply follow a conventional autoregressive latent stochastic process.\footnote{See for example, harvey_etal.1994,andersen_etal.1996,ghysels_etal.1996,gallant_etal.1997,bali2000testing,durbin_koopman.2000,jacquier_etal.2002,jacquier_etal.2004,shephard_pitt.2004,yu2005leverage,han2006asset,hansen2008consumption,bansal2010long,schorfheide2018identifying, among others. An extensive review of the use of stochastic volatility models as an alternative to ARCH-type approaches can be found in shephard2020statistical.} However, monthly volatility forecasts can be noisy, which leads to extreme portfolio turnover from volatility targeting. As a result, one could “filter out” the noise based on a posterior approximation density which embeds both non-smooth predictive densities and different types of smoothing, e.g., wavelet basis functions (see rue_held.2005).

We evaluate the economic performance of our smooth volatility prediction based on a broad sample of 158 equity trading strategies. We first consider the nine equity factors examined by moreira2017volatility. We augment the first group of test portfolios with a second group covering a broader set of trading strategies based on the list of 153 characteristic-managed portfolios, or “factors”, reported in jensen2021there. In addition to previous month's realised variance (henceforth {\tt RV}), we benchmark our smooth volatility-managed portfolios ({\tt SSV}) against several alternative implementations of volatility-targeting. The first uses the expected variance from a simple AR(1) rather than realized variance ({\tt RV AR}), which helps to reduce the extremity of the weights. Second, we consider an alternative six-month window to estimate the longer-term realised variance ({\tt RV6}) as proposed by barroso2015momentum,barroso2021limits. Third, we consider both a long-memory model for volatility forecast as proposed by corsi2009simple ({\tt HAR}), and a standard AR(1) stochastic volatility model ({\tt SV}) (see, e.g., taylor1994modeling). Finally, we consider a plain GARCH(1,1) specification ({\tt Garch}), which has been proved a challenging benchmark in volatility forecasting (see, hansen2005forecast).

Main findings

Our empirical tests evaluate the performance of alternative volatility-managed implementations of a broad set of volatility managed portfolios, each of them constructed as

align[align omitted — 96 chars of source]

where $y^\sigma_{t}$ and $y_t$ are the scaled and the original portfolio's excess returns in month $t$, respectively. Here $\widehat{\sigma}_{t-1|t}^2$ is the variance forecast of the original portfolio's returns at month $t$ based on information available up to month $t-1$. We follow cederburg2020performance and consider both an unconditional and a real-time implementation of volatility targeting. The former implies that the constant $c^*$ is chosen such that the unconditional variance of the managed $y^\sigma_{t}$ and unmanaged $y_{t}$ portfolios coincide. For the real-time implementation, $c_t^*$ is time-varying and is chosen such that the variance of the managed and unmanaged portfolios coincide only conditional on the returns up to month $t$.

Most prior studies assess the value of volatility targeting strategies by comparing the Sharpe ratios obtained by scaled factors $y_t^\sigma$ as in Eq.(ref), with the Sharpe ratios obtained from the original factors $y_t$ (see, e.g., barroso2015momentum,daniel2016momentum,moreira2017volatility,bianchi2022taming). We follow this approach and confirm the existing evidence in the literature that stand-alone investments in volatility-managed portfolios do not systematically improve upon unmanaged factors (see, e.g., cederburg2020performance,barroso2021limits). However, volatility targeting based on our smooth volatility forecasts substantially improves both upon conventional realised variance measures and a variety of competing volatility forecasting methods. Specifically, volatility-managed portfolios show a substantially lower turnover compared to alternative volatility forecasting methods. The right panel of Figure (ref) shows this case in point. The time variation of leverage for the volatility-managed market portfolio is much lower for our {\tt SSV} compared to a standard {\tt RV} or a lower-frequency {\tt RV6}.

Perhaps more importantly, we show that greater portfolio stability translates into a substantially large risk-adjusted performance. For conservative levels of transaction costs, our {\tt SSV} produces a substantially higher economic utility compared to both standard and non-standard volatility forecasting methods. For each equity strategy and volatility-targeting methodology, we estimate the spanning regression on both the scaled and unscaled returns,

align[align omitted — 82 chars of source]

The economic implication of $\alpha>0$ is that volatility scaled portfolios may expand the mean-variance frontier relative to the unscaled portfolios (see, e.g., gibbons1989test). We test this assumption by comparing the certainty equivalent return (CER) when factoring in moderate levels of notional transaction costs, with and without leverage constraints. Specifically, we compare two strategies: (i) a strategy that allocates between a given volatility-managed portfolio and its corresponding original portfolio, and (ii) a strategy constrained to invest only in the original portfolio. The baseline combination correspond to the optimal mean-variance allocation assuming a risk aversion coefficient equal to five. We show that when transaction costs are considered, our {\tt SSV} stands out as the most profitable rescaling method, on average. Perhaps more interestingly, the {\tt SSV} is the only with a positive median CER differential with respect to the unmanaged portfolio strategies. That is, the economic gain is positive for at least 50% of the equity strategies considered. Interestingly, a regularisation of the volatility targeting weights based on leverage constraints does not reduce the gap between our {\tt SSV} method and all the alternative weighting schemes we consider. Similarly, the economic gain from a mean-variance combination strategy of the unmanaged and managed portfolios is substantially in favour of our {\tt SSV} method.

In addition to the empirical analysis on a broad set of equity trading strategies, we explore the statistical underpinnings of our modeling framework through an extensive simulation exercise. We compare the estimation accuracy of our VB inference scheme against state-of-the-art Bayesian methods, such as MCMC (see, e.g., stochvol_package) and variational Bayes (see, e.g., chan_yu2022). The results show that when we do not arbitrarily impose any smoothness in the posterior estimates of the latent stochastic volatility state, our algorithm is as accurate as MCMC and existing variational Bayes methods. Yet, when we smooth the posterior estimates the accuracy deteriorates. This is expected since the wavelet basis functions mechanically tilts the posterior estimates of the parameters towards a more persistent latent state relative to the actual data generating process.

Reference literature

In addition to moreira2017volatility, our work contributes to a growing literature that seeks to understand the origins and the dynamic properties of volatility-managed portfolios (see, e.g., harvey2018impact,bongaerts2020conditional,cederburg2020performance,liu2019volatility,barroso2021limits,wang2021downside, among others). liu2019volatility shows that a real-time implementation of volatility targeting suffers from severe drawdowns, compared to unmanaged portfolios. Similarly, cederburg2020performance shows that volatility-managed portfolios do not systematically outperform the corresponding unmanaged equity strategies.

We contribute to this literature by highlighting the importance of volatility modeling for the profitability of volatility-managed portfolios. Specifically, we show that smoothing the volatility forecasts provide an intuitive regularization to volatility targeting. This translates in an economically better performance versus realised variance measures when notional trading costs are factored in. In addition, we explicitly acknowledge that the uncertainty around the volatility predictions might be pervasive. By taking a Bayesian approach we can quantify the uncertainty around the scaled portfolio returns, so that a more direct statistical comparison between scaled and unscaled factors can be made.

A second strand of literature we contribute to, relates to the estimation of stochastic volatility models. The non-linear interaction between the latent volatility state and the observed returns lead to a likelihood function that depends upon high dimensional integrals. A variety of estimation procedures have been proposed to overcome this difficulty, including the generalized method of moments (GMM) of melino_turnbull.1990, the quasi maximum likelihood (QML) approach of harvey_etal.1994 and ruiz.1994, and the efficient method of moments (EMM) of gallant_etal.1997. Within the context of Bayesian methods, the analysis of stochastic volatility models has been initially proposed by kim_etal.1998,durbin_koopman.2000,jacquier_etal.2002,jacquier_etal.2004,shephard_pitt.2004,durbin_koopman.2000. We contribute to this literature by proposing a novel variational Bayes estimation framework which allows to flexibly smooth the predictive density of the latent stochastic volatility state, irrespective of the underlying assumption about the data generating process. Our approach is general, meaning that encompasses different smoothness assumptions for the volatility forecasts without changing the underlying model structure.

Finally, this paper connects to a third strand of literature that introduces the use of variational Bayes methods for economic forecasting (see, e.g., gefang_koop_poon2019,koop_korobilis_2020,chan_yu2022). Variational approximate methods Bishop.2006 have become popular as computational feasible alternatives to Markov Chain Monte Carlo (MCMC) for approximating the posterior distributions. This type of inferential methods have been used in a wide range of applications, ranging from statistics rustagi.1976 to quantum mechanics sakurai.1994, statistical mechanics parisi.1988, machine learning hinton_vancamp.1993 and then generalized to many probabilistic models, taking advantage of the graphical models' representation jordan_etal.1999. We contribute to this literature by proposing a flexible approximation based on a Gaussian Markov random field approximation of the latent stochastic volatility state. This allows to consider both non-smooth and smooth volatility forecasts based on a simple twist in the posterior approximating density of the latent state.

Modeling framework

Let consider a standard univariate dynamic model with stochastic volatility taylor1994modeling. A general specification is based on a state-space representation of the form:

align[align omitted — 214 chars of source]

where $y_t$, $\mathbf{x}_t\in\mathbb{R}^p$, $h_t=\log\sigma_t^2$ are, respectively, the log-return, a set of covariates, and the log-volatility of an equity strategy at time $t$, for $t=1,2,\dots,n$. The error terms $\varepsilon_t$ and $u_t$ are mutually independent Gaussian white noise processes. The latent process in (ref) is a conventional autoregressive process of order one, with unconditional mean $c$, persistence $\rho$, and conditional variance $\eta^2$. We assume $|\rho|<1$, so that the initial state $h_0$ can be sampled from the marginal distribution, i.e. $h_0\sim\mathsf{N}\left(c,\frac{\eta^2}{1-\rho^2}\right)$. Notice that, for comparability with the existing literature on volatility-managed portfolios, we assume a constant mean $\mu$ in the observation equation (ref), such that there are no covariates and $\mu=\mathbf{x}_t^\intercal\beta$ with $\mathbf{x}_t$ an n-dimensional vector of ones. However, in the following, we provide the full specification of our variational Bayes inference scheme under the general model with covariates.

Variational Bayes inference

A variational Bayes approach to inference requires to minimize the Kullback-Leibler ($\mathit{KL}$) divergence between an approximating density $q(\boldsymbol{\vartheta})$ and the true posterior density $p(\boldsymbol{\vartheta}|\mathbf{y})$, Blei.2017. The $\mathit{KL}$ divergence cannot be directly minimized with respect to $\boldsymbol{\vartheta}$ because it involves the expectation with respect to the unknown true posterior distribution. ormerod2010explaining show that the problem of minimizing $\mathit{KL}$ can be equivalently stated as the maximization of the variational lower bound (ELBO) denoted by $\underline{p}\left(\mathbf{y};q\right)$:

equation[equation omitted — 357 chars of source]

where $q^*(\boldsymbol{\vartheta})\in\mathcal{Q}$ represents the optimal variational density and $\mathcal{Q}$ is a space of functions. The choice of the family of distributions $\mathcal{Q}$ is critical and leads to different algorithmic approaches. In this paper we consider two cases. The first is a mean-field variational Bayes (MFVB) approach which is based on a non-parametric restriction for the variational density, i.e. $q(\boldsymbol{\vartheta})=\prod_{i=1}^p q_i(\boldsymbol{\vartheta}_i)$ for a partition $\{ \boldsymbol{\vartheta}_1,\dots,\boldsymbol{\vartheta}_p\}$ of the parameter vector $\boldsymbol{\vartheta}$. Under the MFVB restriction, a closed form expression for the optimal variational density of each component $q(\boldsymbol{\vartheta}_j)$ is defined as:

equation[equation omitted — 353 chars of source]

where the expectation is taken with respect to the joint approximating density with the $j$-th element of the partition removed $q^\star(\boldsymbol{\vartheta}\setminus\boldsymbol{\vartheta}_j)$. This allows to implement a coordinate ascent variational inference (CAVI) algorithm to estimate the optimal density $q^*(\boldsymbol{\vartheta})$. Equation (ref) shows that the factorization of $q(\boldsymbol{\vartheta})$ plays a central role in developing a MFVB algorithm. In the following, we consider a factorization of the joint variational density of the latent log-variances $\mathbf{h}$ and the parameters $\boldsymbol{\vartheta}=\left(\mbox{\boldmath $\beta$},c,\rho,\eta^2\right)$ of the form:

equation[equation omitted — 158 chars of source]

In the following, we focus on the approximating density for the latent process $\mathbf{h}$, where the novelty of our estimation procedure lies compared to the existing literature (see, e.g., chan_yu2022). For the interested reader, in Appendix (ref) we provide the full set of derivations of the optimal variational densities for the parameters $q(\mbox{\boldmath $\beta$})$, $q(c)$, $q(\rho)$, and $q(\eta^2)$.

The marginal distribution $p\left(\mathbf{h}\right)$ of the joint vector $\mathbf{h}^\intercal=(h_0,h_1,\ldots,h_n)$ admits a Gaussian Markov random field (GMRF) representation $\mathbf{h}\sim\mathsf{N}_{n+1}(c\boldsymbol{\iota}_{n+1},\eta^2\mathbf{Q}^{-1})$ that preserves the time dependence structure implied by the autoregressive process. Specifically, the matrix $\mathbf{Q}=\mathbf{Q}(\rho)$ is a tridiagonal precision matrix with diagonal elements $q_{1,1}=q_{n+1,n+1}=1$ and $q_{i,i}=1+\rho^2$ for $i=2,\ldots,n$, and off-diagonal elements $q_{i,j}=-\rho$ if $|i-j|=1$ and $0$ elsewhere (see rue_held.2005). We exploit this representation to obtain the approximating density $q(\mathbf{h})$ as $\mathbf{h}\sim\mathsf{N}_{n+1}(\mbox{\boldmath $\mu$}_{q(h)},\mbox{\boldmath $\Omega$}_{q(h)}^{-1})$ with mean vector $\mbox{\boldmath $\mu$}_{q(h)}=\mathbf{W}\mathbf{f}_{q(h)}$ and variance-covariance matrix $\mathbf{\Sigma}_{q(h)}=\mbox{\boldmath $\Omega$}_{q(h)}^{-1}$.

Notice that the choice of $\mbox{\boldmath $\mu$}_{q(h)}$ as a linear projection $\mathbf{W}\mathbf{f}_{q(h)}$, with $\mathbf{f}_{q(h)}\in\mathbb{R}^{k}$ the projection coefficients and $\mathbf{W}$ an $(n+1)\times k$ deterministic matrix, has a direct effect on the posterior estimates of log-volatility. In Section (ref) we discuss in details how different structures of $\mathbf{W}$ leads to different posterior estimates irrespective of the underlying dynamics of the latent state. This is a key feature of our estimation strategy since it allows to customise the volatility forecasts without changing the underlying model assumptions.

In the following we focus on the more general heteroschedastic case, whereas the optimal density and the estimation details for the more restrictive homoschedastic case are discussed in Appendix (ref). The optimal parameters $\boldsymbol{\xi}=(\mathbf{f}_{q(h)},\mathbf{\Sigma}_{q(h)})$ of the approximating density $q\left(\mathbf{h}\right)$ can be found by solving the optimization problem

equation[equation omitted — 149 chars of source]

To solve the optimization we leverage on the GMRF representation of $q\left(\mathbf{h}\right)$ and exploit the results in rohde_wand2016. They provide a closed-form updating scheme for the variational parameters when the approximating density is a multivariate Gaussian. Proposition (ref) the details on the optimal updating scheme for the variational density of the latent volatility states. The proof and analytical derivations are available in Appendix (ref). A pseudo-code for the implementation of the proposed iterative estimation procedure is available in Algorithm (ref) in Appendix (ref).

propositionLet $\mbox{\boldmath $\mu$}_{q(\mathbf{s})}=(\mu_{q(s_1)},\ldots,\mu_{q(s_n)})^\intercal$ with $\mu_{q(s_t)} = (y_t-\mathbf{x}_t^\intercal\mbox{\boldmath $\mu$}_{q(\beta)})^2+\mathsf{tr}\left\{\mathbf{\Sigma}_{q(\beta)}\mathbf{x}_t\mathbf{x}_t^\intercal\right\}$, and $\mbox{\boldmath $\mu$}_{q(\beta)},\mathbf{\Sigma}_{q(\beta)}$ denote the variational mean and covariance of the regression parameters $\mbox{\boldmath $\beta$}$. Assuming a GMRF representation of $\mathbf{h}\sim\mathsf{N}_{n+1}(\mbox{\boldmath $\mu$}_{q(h)},\mbox{\boldmath $\Omega$}_{q(h)}^{-1})$, with mean vector $\mbox{\boldmath $\mu$}_{q(h)}=\mathbf{W}\mathbf{f}_{q(h)}$ and variance-covariance matrix $\mathbf{\Sigma}_{q(h)}=\mbox{\boldmath $\Omega$}_{q(h)}^{-1}$, an iterative algorithm can be set as: \begin{align} \mathbf{\Sigma}_{q(h)}^{new} &= \left[\nabla_{\boldsymbol{\mu}_{q(h)}\boldsymbol{\mu}_{q(h)}}^2 S(\boldmath $\mu$_{q(h)}^{old},\mathbf{\Sigma}_{q(h)}^{old})\right]^{-1}, \\ \mathbf{f}_{q(h)}^{new} &= \mathbf{f}_{q(h)}^{old} + \mathbf{W}^+\,\mathbf{\Sigma}_{q(h)}^{new}\nabla_{\boldsymbol{\mu}_{q(h)}} S(\boldmath $\mu$_{q(h)}^{old},\mathbf{\Sigma}_{q(h)}^{old}),\\ \boldmath $\mu$_{q(h)}^{new} &= \mathbf{W}\mathbf{f}_{q(h)}^{new}, \end{align} with $\mathbf{W}^+=(\mathbf{W}^\intercal\mathbf{W})^{-1}\mathbf{W}^\intercal$ the left Moore–Penrose pseudo-inverse of $\mathbf{W}$, and $S(\mbox{\boldmath $\mu$}_{q(h)},\mathbf{\Sigma}_{q(h)})$ equal to $\mathbb{E}_q(\log p(\mathbf{h},\mathbf{y}))$ (see Eq.(ref)), such that, \begin{align} \nabla_{\boldsymbol{\mu}_{q(h)}} S(\boldmath $\mu$_{q(h)},\mathbf{\Sigma}_{q(h)}) &= -\frac{1}{2}[0,\boldsymbol{\iota}_n^\intercal]^\intercal+\frac{1}{2}[0,\boldmath $\mu$_{q(\mathbf{s})}^\intercal]^\intercal\odot\mathrm{e}^{-\boldsymbol{\mu}_{q(h)}+\frac{1}{2}\mathsf{diag}\left(\boldsymbol{\Sigma}_{q(\mathbf{h})}\right)} \\ &\qquad-\mu_{q(1/\eta^2)}\boldmath $\mu$_{q(\mathbf{Q})}(\mbox{\boldmath $\mu$}_{q(h)}-\mu_{q(c)}\boldsymbol{\iota}_{n+1}),\\ \nabla_{\boldsymbol{\mu}_{q(h)}\boldsymbol{\mu}_{q(h)}}^2 S(\mbox{\boldmath $\mu$}_{q(h)},\mathbf{\Sigma}_{q(h)}) &=-\frac{1}{2}\mathsf{Diag}\Bigg[[0,\mbox{\boldmath $\mu$}_{q(\mathbf{s})}^\intercal]^\intercal\odot\mathrm{e}^{-\boldsymbol{\mu}_{q(h)}+\frac{1}{2}\mathsf{diag}\left(\boldsymbol{\Sigma}_{q(\mathbf{h})}\right)}\Bigg] -\mu_{q(1/\eta^2)}\mbox{\boldmath $\mu$}_{q(\mathbf{Q})}, \end{align} where $\boldsymbol{\iota}_n$ is an n-dimensional vector of ones, $\mu_{q\left(1/\eta^2\right)}$ is the variational mean of $1/\eta^2$, $\mbox{\boldmath $\mu$}_{q(\mathbf{Q})}$ is the element-wise variational mean of $\mathbf{Q}$, and $\odot$ denotes the Hadamard product.

Our approach expands the global approximation method proposed by chan_yu2022 along three main dimensions. First, we relax the assumption that the initial distribution $q(h_0)$ is independent on the trajectory of the latent state $q(\mathbf{h}_1)$, that is, we do not assume $q(\mathbf{h})=q(h_0)q(\mathbf{h}_1)$. Second, we do not make any assumption on the $\mathbf{\Sigma}_{q(h)}$, which is not fixed conditional on $\mbox{\boldmath $\mu$}_{q(h)}$, but is estimated jointly with $\mbox{\boldmath $\mu$}_{q(h)}$. Third, our latent volatility state accommodates a more general AR(1) dynamics, instead of a random walk. While the latter reduces the parameter space, it imposes a strong form of non-stationarity in the log-volatility process. In Section (ref), we show via an extensive simulation study that all these features have a significant effect on the accuracy of the variational Bayes estimates.

Smoothing the volatility estimates

The choice of $\mbox{\boldmath $\mu$}_{q(h)}$ as a linear projection $\mathbf{W}\mathbf{f}_{q(h)}$, with $\mathbf{f}_{q(h)}\in\mathbb{R}^{k}$ the projection coefficients and $\mathbf{W}$ an $(n+1)\times k$ deterministic matrix, has a direct effect on the posterior estimates of log-volatility. Figure (ref) shows examples of the shape of $\mbox{\boldmath $\mu$}_{q\left(\mathbf{h}\right)}=\mathbf{W}\mathbf{f}_{q(h)}$ for difference choices of $\mathbf{W}$ (solid line), and the corresponding confidence intervals implied by $\mathbf{\Sigma}_{q(h)}$ (dashed line). The gray trajectory represents the true simulated value of the log-stochastic volatility $\mathbf{h}^\intercal=(h_0,h_1,\ldots,h_n)$ for $n=300$. The top-left panel reports the posterior estimates obtained by setting $\mathbf{W}=\mathbf{I}_{n+1}$, with $\mathbf{I}_{n+1}$ an identity matrix of dimension $n+1$. This represents a non-smooth estimate which is akin to the output of a standard MCMC estimation scheme (see, e.g., stochvol_package).

The remaining panels of Figure (ref) highlight a key feature of our estimation strategy; that is, it allows to customise the volatility forecasts without changing the underlying model assumptions. For instance, the top-right panel shows the posterior estimates of the latent volatility state with $\mathbf{W}$ a matrix of wavelet basis functions with a fixed degree of smoothness $l=4$ wand_ormerod.2011. The fact that the matrix $\mathbf{W}$ enters both in the conditional mean and covariance of the optimal variational density $q^\ast\left(\mathbf{h}\right)$ allows to smooth not only the conditional mean of the latent volatility state, but also the corresponding confidence intervals.

The bottom panels in Figure (ref) highlight the flexibility of our approach; the left panel shows that more than one smoothing assumption can coexists in the same optimal variational density. For instance, the shape of the posterior estimates assuming $\mathbf{W}=\mathbf{I}_{n+1}$ for the first half of the sample and $\mathbf{W}$ a wavelet basis function with $l=4$ for the second half of the sample. The bottom-right panel shows that a variety of smoothing functions can be adopted; for instance, the estimates of the latent stochastic volatility can be smoothed based on $\mathbf{W}$ equal to be a B-spline basis matrix representing the family of piecewise polynomials with the pre-specified interior knots ($kn$), degree ($dg$), and boundary knots.

Figure (ref) depicts the form of $\mathbf{W}$ when B-spline and Daubechies wavelets are used. The form of $\mathbf{W}$ in case of B-spline basis functions (top) and wavelet basis functions (bottom). Right panels correspond to columns of the matrix $\mathbf{W}$. The B-spline basis functions is a sequence of piecewise polynomial functions of a given degree, in this case $dg=3$. The locations of the pieces are determined by the knots, here we assume $kn=20$ equally spaced knots. The functions that compose the wavelet basis matrix $\mathbf{W}$ are constructed over equally spaced grids on $[0,n]$ of length $R$, where $R$ is called resolution and it is equal to $2^{l-1}$, where $l$ defines the level, and as a result the degree of smoothness. The number of functions at level $l$ is then equal to $R$ and they are defined as dilatation and/or shift of a more general {\it mother} function.

Variance prediction

Consider the posterior distribution of $p(\mathbf{h},\boldsymbol{\vartheta}|\mathbf{y})$ given the information set up to time $t$, $\mathbf{y}=\left\{y_{1:t}\right\}$, and $p(h_{n+1}|\mathbf{y},\mathbf{h},\boldsymbol{\vartheta})$ the likelihood for the new latent state $h_{n+1}$. The predictive density then takes the familiar form,

equation[equation omitted — 211 chars of source]

Given a variational density $q(\mathbf{h},\boldsymbol{\vartheta})=q(\mathbf{h})q(\boldsymbol{\vartheta})$ that approximates $p(\mathbf{h},\boldsymbol{\vartheta}|\mathbf{y})$, we follow gunawan2021variational and obtain the variational predictive distribution:

align[align omitted — 324 chars of source]

where the second equality follows from Markov property. Recall that within the context of a volatility-managed portfolio our object of interest is the forecast of the variance $\sigma_t^2$, rather than the log-volatility $h_t$ for $t=n+1$. Since $h_n=\log\sigma_n^2$, the density of the conditional variance is readily available as $q\left(\sigma_{n+1}^2|\mathbf{y}\right)=\frac{\partial h_{n+1}}{\partial\sigma_{n+1}^2}q\left(h_{n+1}|\mathbf{y}\right)=\frac{1}{\sigma_{n+1}^2} q\left(h_{n+1}|\mathbf{y}\right)$. The integral in Eq.(ref) cannot be solved analytically. However, it can be approximated through Monte Carlo integration exploiting the fact that the optimal variational densities $q(h_n)$ and $q(\boldsymbol{\vartheta})$ are known and we can efficiently sample from them. A simulation-based approximated estimator for the variational predictive distribution of the conditional variance $q(\sigma^2_{n+1}|\mathbf{y})$ is therefore obtained by averaging the density $p(h_{n+1}|h_n^{(i)},\boldsymbol{\vartheta}^{(i)})$ over the draws $h_n^{(i)}\sim q(h_n)$ and $\boldsymbol{\vartheta}^{(i)}\sim q(\boldsymbol{\vartheta})$, for $i=1,\ldots,N$ from the optimal variational density, such that $\widehat{q}(\sigma^2_{n+1}|\mathbf{y}) = \frac{1}{\sigma_{n+1}^2}\frac{1}{N}\sum_{i=1}^N p(h_{n+1}|h_n^{(i)},\boldsymbol{\vartheta}^{(i)})$.

Empirical results

We now investigate the statistical and economic value of our smooth volatility forecast within the context of volatility targeting across a large set of equity strategies. We first consider the nine equity factors examined by moreira2017volatility. We collect daily and monthly data on the excess returns on the market, and the daily and monthly returns on the size, value, profitability and investment factors as originally proposed by fama2015five, in addition to the profitability and investment factors from hou2015digesting and the betting-against-beta factor from frazzini2014betting.\footnote{Data on the fama2015five factors and the jegadeesh1993returns momentum are available on the Kenneth French's website at \url{http://mba.tuck.dartmouth.edu/pages/faculty/ken.french/data_library.html}. Data on the betting-against-beta factor are available on the AQR website \url{https://www.aqr.com/Insights/Datasets/Betting-Against-Beta-Original-Paper-Data}.}

We augment the first group of test portfolios with a second group covering a broader set of trading strategies based on established asset pricing factors. We start with the list of 153 characteristic-managed portfolios, or “factors”, reported in jensen2021there. We then restrict our analysis to value-weighted strategies that can be constructed using the Center for Research in Security Prices (CRSP) monthly and daily stock files, the Compustat Fundamental annual and quarterly files, and the Institutional Broker Estimate (IBES) database. In addition, we exclude a handful of strategies for which there are missing returns. This process identifies 149 value-weighted long-short portfolios for which we collect both daily and monthly returns. For a more detailed description of the portfolio construction we refer to jensen2021there.\footnote{Data on the 153 set of characteristic-based portfolios can be found at \url{https://jkpfactors.com}. We thank Bryan Kelly for making these data available.} The combined sample consists of 158 equity trading strategies.

Construction of volatility-managed portfolios

For a given equity trading strategy, let $y_t$ be the buy-and-hold excess portfolio return in month $t$. We follow moreira2017volatility and construct the corresponding volatility-managed portfolio return $y^\sigma_{t}$ as

align[align omitted — 97 chars of source]

where $c^*$ is a constant chosen such that the unconditional variance of the managed $y^\sigma_{t}$ and unmanaged $y_{t}$ portfolios coincide, and $\widehat{\sigma}_{t-1|t}^2$ is the variance forecast of unscaled portfolio returns based on information available up to the previous month $t-1$. The objective of Eq.(ref) is to adjust the capital invested in the original equity strategy based on the inverse of the (lagged) predicted variance. Effectively, a volatility-managed portfolio is targeting a constant level of volatility, rather than a constant level of notional capital exposure. As such, the dynamics investment position in the underlying portfolio $\frac{c^*}{\widehat{\sigma}_{t-1|t}^2}$ is a measure of (de)leverage required to invest in the volatility-portfolio in month $t$. Notice that in the standard implementation in Eq.(ref) the scaling parameter $c^*$ is not know by an investor in real time as it requires to observe the full time series of the unscaled returns $y_t$ and the volatility forecasts $\widehat{\sigma}_{t|t-1}^2$.

A benchmark approach to approximate the variance forecast at month $t$, $\widehat{\sigma}_{t-1|t}^2$ is to use the previous month's realized variance (henceforth {\tt RV}) calculated based on daily portfolio returns (see, e.g., barroso2015momentum,daniel2016momentum,moreira2017volatility,cederburg2020performance,barroso2021limits),

align[align omitted — 132 chars of source]

where $y_{j,t-1}$ be the excess returns on a given portfolio in day $j=1,\ldots,\mathcal{N}_{t-1}$ for month $t-1$. In addition to the realised variance, we compare our smoothing volatility targeting approach ({\tt SSV}) against a variety of alternative rescaling approaches. The first uses the expected variance from a simple AR(1) rather than realized variance ({\tt RV AR}), which helps to reduce the extremity of the weights. Second, we follow barroso2021limits and consider an alternative six-month window to estimate the longer-term realised variance ({\tt RV6}). Third, we consider both a long-memory model for volatility forecast as proposed by corsi2009simple ({\tt HAR}), and a standard AR(1) latent stochastic volatility model ({\tt SV}) (see, e.g., taylor1994modeling). Finally, we consider a plain GARCH(1,1) specification ({\tt Garch}), which has been shown to be a challenging benchmark in volatility forecasting (see, hansen2005forecast). Throughout the empirical analysis we consider, we follow cederburg2020performance and consider both unconditional volatility targeting -- whereby $c^*$ is calibrated to match the unconditional volatility of the scaled and unscaled portfolios --, as well as real-time volatility targeting -- whereby $c_t^*$ is calibrated to match the volatility of the scaled and unscaled portfolios at each month $t$.

A simple statistical appraisal

In this section we provide a statistical appraisal of the performance of our smoothing volatility targeting approach compared to both conventional realised variance measures and benchmark volatility forecasts. This is based on the predictive density of the volatility forecasts obtained for both the non-smooth {\tt SV} and smooth {\tt SSV} stochastic volatility models. Recall that real-time volatility targeting for month $t$ takes the form $\omega_{t}=\frac{c_t^*}{\widehat{\sigma}_{t|t-1}^2}$, $t=1,\ldots,n$. As a result, given the unmanaged factors $y_{t}$ and the recursively calibrated coefficient $c_t^*$, for each month we can define the distribution of the volatility-managed returns based on the variational predictive density $q(\sigma^2_{t}|\mathbf{y})$ with $\mathbf{y}$ collecting the strategy returns up to $t-1$ (see Section (ref) for more details).

Figure (ref) shows this case in point. The top panels report the distribution of the volatility-managed portfolio returns implied by the non-smooth {\tt SV} (red area) and smooth {\tt SSV} (blue area) stochastic volatility models. For the sake of simplicity, we report the volatility-managed returns on the market portfolio over three distinct months. The returns on the unmanaged portfolio and its scaled version based on previous month's realised variance are indicated as a white and green circle, respectively. By comparing this distribution on a given month with the realised returns on a benchmark strategy for the same month, we can calculate $\text{Prob}\left(y_t^{\mathcal{M}_1}<y_t^{\mathcal{M}_0}\right)$, which is akin to the p-value on a one-side test where the null hypothesis is $\mathcal{H}_0: y_t^{\mathcal{M}_1}\leq y_t^{\mathcal{M}_0}$. For instance, a $\text{Prob}\left(y_t^{\mathcal{M}_1}<y_t^{\mathcal{M}_0}\right)<0.05$ implies that the null hypothesis $\mathcal{H}_0$ is rejected with a p-value of 0.05 in favour of the alternative $\mathcal{H}_1: y_t^{\mathcal{M}_1}>y_t^{\mathcal{M}_0}$. On the opposite, if $\text{Prob}\left(y_t^{\mathcal{M}_1}<y_t^{\mathcal{M}_0}\right)>0.05$, the null hypothesis $\mathcal{H}_0$ can not be rejected with a p-value of 0.05. Here $y_t^{\mathcal{M}_0}$ represents the returns on the benchmark volatility managing method, for e.g., {\tt RV}, whereas $y_t^{\mathcal{M}_1}$ the returns on volatility targeting based on either a non-smooth or a smooth stochastic volatility model.

The left panel shows the results for October 1995. The $\text{Prob}\left(y_t^{{\tt SSV}}<y_t^{{\tt RV}}\right)=0.07$, that is the null $\mathcal{H}_0: y_t^{{\tt SSV}}\leq y_t^{{\tt RV}}$ can not be rejected at standard significance levels. Similarly, $\text{Prob}\left(y_t^{{\tt SSV}}<y_t^{{\tt U}}\right)=0.66$, which again suggests that the returns on the {\tt SSV} volatility targeting and the unmanaged counterpart are statistically equivalent. The right panel of Figure (ref) show as another example the returns distribution on March 2009. The probability $\text{Prob}\left(y_t^{{\tt SSV}}<y_t^{{\tt RV}}\right)=0$, that is the null hypothesis $\mathcal{H}_0: y_t^{{\tt SSV}}\leq y_t^{{\tt RV}}$ is rejected with a p-value of 0.000 in favour of the alternative $\mathcal{H}_1: y_t^{{\tt SSV}}>y_t^{{\tt RV}}$. On the other hand, $\text{Prob}\left(y_t^{{\tt SV}}<y_t^{{\tt RV}}\right)=0.08$, which suggests that the {\tt SV} model produce a volatility-managed portfolio which is statistically equivalent to the one implied by a realised variance {\tt RV}. Similarly we can setup the opposite one-side test, which is for the null hypothesis $\mathcal{H}_0: y_t^{{\tt SSV}}\geq y_t^{{\tt RV}}$ against the alternative $\mathcal{H}_1: y_t^{{\tt SSV}}< y_t^{{\tt RV}}$. The bottom panel of Figure (ref) shows that the distribution of {\tt SSV} and {\tt SV} can be highly time varying. The figure shows as an example the distribution of the returns on a volatility-managed momentum portfolio. The large negative performance of the unmanaged momentum strategy in March-May 2009 coincides with the so-called “momentum crashes” (see barroso2015momentum,daniel2016momentum,bianchi2022taming).

Two interesting facts emerge. First, and perhaps not surprisingly, a non-smooth stochastic volatility model tends to produce relatively similar volatility adjusted returns with few exceptions. In this respect, a standard {\tt RV} rescaling substantially overperform (underperform) the unmanaged portfolio during periods of large negative (positive) returns. Put it differently, standard volatility targeting helps to mitigate tail risk at the expense of cutting upside opportunities. This is consistent with the abundant empirical evidence that indeed, on average, {\tt RV} targeting does not systematically outperforms unmanaged portfolios. The second interesting fact pertains our smoothing volatility targeting; the returns on the {\tt SSV} are closer to the original equity strategy.

We now take to task the intuition highlighted in Figure (ref) and compare our {\tt SSV} methodology against all of the competing volatility targeting methods, across all of the 158 equity strategies in our sample. Specifically, we calculate each month two indicator dummies $\mathbb{I}^+_{i,t},\mathbb{I}^-_{i,t}$ for each of the $t=1,\ldots,n$ and each of the $i=1,\ldots,m$ equity trading strategies,

align[align omitted — 406 chars of source]

We can then calculate $p^+_i = n^{-1}\sum_{t=1}^n\mathbb{I}^+_{i,t}$ and $p^-_i = n^{-1}\sum_{t=1}^n\mathbb{I}^-_{i,t}$, with $n$ the sample of observations, for each equity trading strategy. These indicate the frequency over the full sample with which the null hypothesis $\mathcal{H}_0$ is rejected in favour of the alternative $\mathcal{H}_1: y_t^{{\tt SSV}}>y_t^{\mathcal{M}_0}$, i.e., $p^+_i$, or the alternative $\mathcal{H}_1: y_t^{{\tt SSV}}<y_t^{\mathcal{M}_0}$, i.e., $p^-_i$.

Figure (ref) reports the difference between $p^+_i$ and $p^-_i$ for all 158 equity strategies. This indicates the imbalance between outperformance and underpeformance of our $y_{i,t}^{\tt SSV}$ compared to a benchmark $y_{i,t}^{\mathcal{M}_0}$. The left panel compares our {\tt SSV} against the original factor portfolios {\tt U} and the volatility targeting based on the realised variance {\tt RV}. The comparison against the unscaled factors confirms the results of cederburg2020performance; there is no systematic outperformance of volatility targeting versus unmanaged equity strategies over the sample under investigation. This is reflected in the fact that the difference between $p_i^+$ and $p_i^-$ is centered around zero for the cross section of equity strategies. The middle and right panel also confirms that, unconditionally over the full sample, the performance of our {\tt SSV} does not systematically dominate other competing volatility targeting methods. For instance, the spread $p_i=p_i^+-p_i^-$ is as low as -0.1 and as high as 0.05 when comparing {\tt SSV} vs {\tt RV6}. Similarly, $p_i$ ranges between -0.05 and 0.05 when comparing our {\tt SSV} against the {\tt HAR} or the {\tt Garch} methods.

The results in Figure (ref) show that the returns on volatility-managed portfolios are statistically equivalent to unscaled factors, at least unconditionally. We now look at a conditional aggregation of the indicators $\mathbb{I}^+_{i,t}$ and $\mathbb{I}^-_{i,t}$. Specifically, we construct a $p^+_t = m^{-1}\sum_{i=1}^m\mathbb{I}^+_{i,t}$ and $p^-_t = m^{-1}\sum_{i=1}^m\mathbb{I}^-_{i,t}$, with $m$ the number of equity strategies, for month $t=1,\ldots,n$. Figure (ref) reports the spread $p_t=p_t^+-p_t^-$ across the whole sample of observations. The left panel compares the performance of {\tt SSV} versus {\tt RV} and the unmanaged factors {\tt U}. Two interesting facts emerge; first, for the most part of the sample the performance of the {\tt SSV} is subpar compared to the {\tt RV}. This is primarily concentrated in the expansionary periods, whereby volatility is low and the exposure to the original unscaled portfolios is levered up (see, e.g., Figure (ref)).

Second, a smooth volatility targeting substantially improves upon {\tt RV} during the recession in the aftermath of the dot-com bubble and the great financial crisis of 2008/2009. Interestingly, most of the underperformance of {\tt SSV} versus {\tt U} is concentrated during the burst of the dot-com bubble. A possible explanation is that volatility-targeting implies a deleveraging on the original factor, in period in which high volatility did not necessarily correspond to large losses in the original equity factors. The middle and right panel in Figure (ref) shows that alternative volatility measures to {\tt RV} share a similar pattern compared to our {\tt SSV}; that is, by smoothing volatility forecasts the performance during major recessions improves at the expenses of a subpar performance during economic expansions and/or lower-volatility periods.

Economic evaluation

We begin our analysis by presenting detailed results on direct performance comparison between unscaled and scaled portfolios without considering transaction costs. Next, we build upon moreira2017volatility,cederburg2020performance and consider two distinct levels of notional transaction costs to implement volatility targeting. Finally, we compare our {\tt SSV} volatility targeting against both {\tt RV} and other competing forecasting methods when both transaction costs and leverage constraints are considered (see, e.g., barroso2021limits).

Baseline results without transaction costs

Table (ref) reports the annualised Sharpe ratio (henceforth SR) and the Sortino ratio, for both unconditional and real-time volatility targeting. For each performance measure, we report both the mean value and the 2.5th, 25th, 50th, 75th, and 97.5th percentiles across all the 158 equity trading strategies. Both the original and the volatility-managed factors yield a positive annualised Sharpe ratio, on average. The risk-adjusted performance is comparable across volatility estimates. For instance, the annualised SR from the {\tt RV} is 0.28 against 0.26 from {\tt SSV}. The dispersion of SRs across equity strategies is also quite comparable across methods. For instance, the 97.5th percentile in the distribution of SRs is 0.69 for the {\tt SSV} against 0.81 from a six-month realised variance {\tt RV6}.

To determine whether the SR from a given volatility-managed portfolio is statistically different from its unmanaged counterpart, we follow cederburg2020performance and implement a block bootstrap approach as proposed by jobson1981performance,ledoit2008robust. Table (ref) reports the percentage -- out of all 158 equity strategies -- of SR differences that are positive or negative, and are statistically significant at the 5% level. The results in Table (ref) confirms the existing evidence in the literature that volatility-managed portfolios do not systematically outperform their original counterparts (see, e.g., barroso2021limits). For instance, {\tt RV} yields a significantly larger (smaller) SR compared to unmanaged portfolios for 6% (2.5%) of the 158 equity trading strategies considered.

The percentage of higher and significant SRs slightly improves when using our {\tt SSV} method versus both {\tt RV} and all other competing volatility forecasts. Nevertheless, the percentage of significant and positive SRs tend to be quite similar across different volatility targeting estimates. Table (ref) also reveals that the gross performance across methods is quite comparable when looking at the risk-adjusted returns with a focus on downside risk only. For instance, the average Sortino ratio is 1.44, which is smaller than the 1.77 obtained from the {\tt RV}, but economically fairly close. Again, the Sortino ratios are fairly comparable across scaling methods.

Existing evidence on the performance of volatility-managed portfolios follows from a spanning regression approach of the form $y_t^\sigma=\alpha + \beta y_t+\epsilon_t$. The object of interest is the intercept $\alpha$, that is a positive $\alpha$ implies that a combination of the original unmanaged factor and its volatility-managed counterpart expands the mean-variance frontier compared to investing in the original unscaled portfolio alone (see, e.g., gibbons1989test). The top panel in Table (ref) reports the mean alpha (in %) across all the 158 equity strategies obtained from different volatility target methods. Similar to the Sharpe ratios, we report the 2.5th, 25th, 50th, 75th, and 97.5th percentile of the alphas across all rescaled portfolios, in addition to the mean value across equity strategies. Volatility targeting based on realised variance {\tt RV} achieves the highest average gross $\alpha$ (1.68%), on par with the six-month realised variance {\tt RV6}. This holds both for the unconditional and the real-time volatility implementation. The fraction of positive and significant gross alphas -- at a 5% level --, is also higher for the {\tt RV} and {\tt RV6} methods.

moreira2017volatility link their spanning test results to appraisal ratios and utility gains for investors. Both metrics can be red in the context of mean-variance portfolio choice. The appraisal ratio for a given scaled strategy is $AR=\widehat{\alpha}/\widehat{\sigma}_\varepsilon$, where $\widehat{\alpha}$ is the estimated gross alpha from the spanning regression and $\widehat{\sigma}_\varepsilon$ the root mean squared error. The squared of the appraisal ratio reflects the extent to which volatility management can be used to increase the slope of the mean-variance frontier (see, gibbons1989test). The mid panel of Table (ref) shows the results for both unconditional and real-time volatility targeting. On average, the appraisal ratio from the {\tt RV} is higher (0.05) compared to our {\tt SSV} (0.03). The cross-sectional distribution of the ARs is quite symmetric, as the mean and median estimates tend to coincide.

Perhaps more interesting is the fact that the estimates of the $\widehat{\alpha}$ from the spanning regressions can be used to quantify the utility gain from volatility management. This is achieved by comparing the certainty equivalent return (CER) for the investor who has access to both the original and the volatility-managed factor against the investor who is constrained to the original equity strategy only. We follow cederburg2020performance,barroso2021limits and define the difference in CER from the unmanaged and the scaled portfolios as

align[align omitted — 111 chars of source]

where $\text{SR}\left(y_t\right)$ is the Sharpe ratio of the unscaled portfolio and $\text{SR}\left(z^*_t\right)$ is the Sharpe ratio of the combined strategy $z_t=x_\sigma\omega_t + x$, with $\omega_t=\frac{c^*}{\widehat{\sigma}_{t|t-1}^2}$. The ex post optimal policy $\left[x_\sigma,\ x\right]^{\prime}=\frac{1}{\gamma}\widehat{\Sigma}^{-1}\widehat{\mu}$ allocates a static weight $x_\sigma$ to the volatility-managed portfolio and a static $x$ weight on the original factor, based on the sample covariance $\widehat{\Sigma}$ and the sample mean $\widehat{\mu}$ returns of the scaled and unscaled portfolios. This policy is equivalent to dynamically adjust the exposure to the original factor portfolio according to $z_t$, so that the returns on the combined strategy can be obtained as $z^*_t=z_t\cdot y_t$. The bottom panel of Table (ref) reports $\Delta CER$(%) for the unconditional and real-time volatility targeting.

We follow cederburg2020performance,wang2021downside and consider a risk aversion coefficient equal to $\gamma=5$. The $\Delta CER$ confirms that volatility targeting based on realised variance does indeed expands ex post the mean-variance frontier relative to the other volatility targeting methods, when no transaction costs or cost-mitigation strategies are considered. For instance, the $\Delta CER$ from the {\tt RV} is 18% versus 9% obtained from our {\tt SSV} smoothing volatility forecast. Interestingly, a slightly smoother estimate of realised volatility, i.e., {\tt RV6}, produces a higher $\Delta CER$(%), both unconditionally and in real time.

Turnover and leverage

A standard volatility targeting strategy is built upon scaling the original portfolio returns by $\frac{c^*}{\widehat{\sigma}_{t|t-1}^2}$. The often erratic nature of $\widehat{\sigma}_{t|t-1}^2$ based on realised volatility implies that volatility-managed portfolios are associated with high turnover and significant time-varying leverage $\omega_t$. This is likely to cast doubt on the actual usefulness of volatility targeting portfolios under common liquidity constraints (see moreira2017volatility,harvey2018impact,bongaerts2020conditional,patton2020you,barroso2021limits). Table (ref) shows the amount of portfolio turnover for different volatility targeting methods. The portfolio turnover is calculated as the average absolute change of the leverage weights $|\Delta w|$ (see moreira2017volatility). We report the mean turnover as well as the 2.5th, 25th, 50th, 75th, and 97.5th percentile across the 158 equity strategies.

Clearly, our {\tt SSV} method substantially reduces the portfolio turnover compared to all other volatility forecasting methods. For instance, the turnover from the {\tt RV} is 0.65 against a 0.05 from {\tt SSV}, on average across equity strategies. Our {\tt SSV} produces a lower turnover not only on average, but for the full cross section of equity strategies. For instance, the 2.5th (97.5th) percentile is 0.03 (0.06) for the {\tt SSV} against a 0.51 (0.91) from {\tt RV}. Perhaps not unexpectedly, the six-month realised variance implies a lower turnover compared to {\tt RV}. Nevertheless, our {\tt SSV} stands out in terms of portfolio stability, both within the context of unconditional or real-time volatility targeting.

The middle panel of Table (ref) also reports the average leverage implied by volatility targeting, i.e., $\omega_t=\frac{c^*}{\widehat{\sigma}_{t|t-1}^2}$. The real-time implementation of the {\tt RV} portfolio scaling implies a leverage that is almost twice as large as the one implied by {\tt SSV} volatility targeting (0.73). Differences across volatility methods are lower for the unconditional targeting. In addition, the bottom panel shows that our smoothing volatility forecasting method significantly reduce liquidity demand, that is increases the stability of $\omega_t$ over time. For instance, the variability of leverage from {\tt SSV} is half (0.43) compared to {\tt RV} (1.09). The leverage mitigation effect of {\tt SSV} is even more clear when looking at the real-time implementation; the standard deviation of $w_t$ is 0.27, on average across equity strategies. This compares to 1.21, 1, and 0.85 from the {\tt RV}, {\tt RV6} and {\tt RV AR}, respectively.

Main specification with transaction costs

Table (ref) shows that alternative scaling methods, such as {\tt HAR}, {\tt Garch} and {\tt RV AR} indeed helps to stabilise volatility managing compared to a standard {\tt RV}. Yet, our smoothing volatility prediction {\tt SSV} generates by the lowest and most stable liquidity demand across all methods. For each equity factor we now consider the costs of the leverage adjustment associated with volatility targeting. We follow moreira2017volatility,wang2021downside and consider two alternative levels of transaction costs of 14 basis points (bps) of the notional value traded to implement volatility targeting (see, e.g., frazzini2012trading) and a more conservative 50 basis points (see, e.g., wang2021downside).

Table (ref) reports the net-of-costs performance statistics for the managed factors. After 14 bps costs, the average SR for {\tt RV} decreases from 0.23 to 0.17. With a more conservative level of transaction costs, the average SR from {\tt RV} turns to a negative -0.11 annualised. This is in stark contrast of what we obtain by smoothing the volatility predictions; that is, our {\tt SSV} generates a remarkable stable SR of 0.25 and 0.23 after 14 and 50 basis points of notional trading costs, respectively. Perhaps more importantly, only 10% of volatility-managed portfolios produce a significantly lower SR compared to the unmanaged counterpart even with conservative 50 bps of trading costs. This is in contrast to {\tt RV}, for which 79% of Sharpe ratios are significantly lower than the unscaled portfolios. Furthermore, when we consider 50 basis points of transaction costs, the Sortino ratio from {\tt SSV} is 1.38 versus -0.69 from {\tt RV}, 0.85 from {\tt RV6} and 0.98 from a {\tt Garch} model, respectively.

Table (ref) reports the results for the spanning regression $y_t^\sigma=\alpha + \beta y_t + \epsilon_t$, with $y_t^\sigma$ the returns on the volatility managed portfolio net of transaction costs and $y_t^\sigma$ its the original equity strategy. The top panels report the estimated alphas ($\widehat{\alpha}$ in %). When considering a conservative notional trading cost of 50 basis points, our {\tt SSV} volatility forecast generates a positive alpha of 0.46% annualised. This is against a large and negative alpha from the {\tt RV}, {\tt RV AR}, {\tt HAR}, and {\tt SV} methods. Consistent with barroso2021limits, a longer-term six-month estimate of the realised variance {\tt RV6} improves the volatility-managed alphas (0.12%). Perhaps more importantly, our {\tt SSV} method generates a significantly positive alpha for 21% of the equity strategies in our sample, against, for instance, a 3%, 9%, and 14% of the strategies from the {\tt RV}, {\tt RV6} and {\tt Garch} models, respectively.

The appraisal ratio $AR$ reported in the middle panel of Table (ref) confirms that {\tt SSV} substantially improves upon realised variance measures {\tt RV}, especially when a conservative transaction cost is factored in. For instance, with 50 basis points of trading costs the {\tt SSV} is the only method that can still generate a positive appraisal ratio. By comparison, the {\tt RV}, {\tt RV6}, {\tt Garch} and {\tt RV AR} all generate significantly negative ARs. The bottom panels report the difference in the certainty equivalent return between and investor that can access both the volatility-managed and the original portfolio, and an investor constrained to invest in the original portfolio only. The utility gain $\Delta CER$(%) is highly in favour of our {\tt SSV} volatility targeting. For instance, for 14 basis points of transaction costs, the second-best performing strategy is the {\tt RV6} rescaling with a $\Delta CER$ of 9.56%, annualised, against a 14.5% from our {\tt SSV}.

Transaction costs with leverage constraints

The results in Tables (ref)-(ref) show that when conservative levels of transaction costs to implement volatility targeting are considered, the performance of standard volatility targeting methods substantially deteriorates. Standard volatility targeting strategies are not designed to mitigate transaction costs. Hence, we next evaluate whether by reducing liquidity demand via capping leverage render volatility targeting still profitable after costs. This approach does not necessarily aim at an optimal allocation from the perspective of a mean-variance investor. Rather, it is a simple, yet effective, risk-management approach that aims to regularise the capital exposure to the original equity trading strategy. We follow moreira2017volatility,cederburg2020performance,barroso2021limits,wang2021downside and consider two different levels of leverage constraint; one that cap the leverage at 1.5 times the original factor, and a second less restrictive that cap leverage at 5 times the exposure to the original factor.

Table (ref) reports the Sharpe and the Sortino ratios considering the same level of transaction costs as in Section (ref), namely 14 and 50 basis points of the notional trading exposure. Panel A shows the results for a 500% leverage constraint. For a conservative 50 basis points transaction costs our {\tt SSV} produces the highest Sharpe and Sortino ratios among the volatility targeting methods, on average across the 158 equity strategies. For instance, the {\tt SSV} generates a 0.23 Sharpe ratio on average against a dismal -0.10 annualised Sharpe ratio from the {\tt RV}. Compared to the unmanaged portfolios, the number of significantly higher SRs is also higher for the {\tt SSV} case. For instance, none of the rescaled portfolios with {\tt RV} has a positive and significant SR differential against 7% of the portfolios rescaled with {\tt SSV}.

Panel B shows the results for a more restrictive leverage constraint, which forces the exposure from volatility targeting no more than 1.5 times the original factor portfolio. Consistent with moreira2017volatility,barroso2021limits, a tighter cap does indeed regularise more the performance of volatility targeting across all competing methods. Nevertheless, the performance of our {\tt SSV} portfolio is quite stable across different levels of leverage constraints. Interestingly, unlike the case without leverage constraints, the {\tt RV6} plus leverage cap proves to be a quite competitive benchmark volatility targeting method.

Table (ref) reports the results for the spanning regressions. The top panels report the estimated alphas ($\widehat{\alpha}$ in %). When considering a conservative notional trading cost of 50 basis points, our {\tt SSV} volatility forecast generates a positive alpha of 0.46% annualised. This is against a large and negative alpha from the {\tt RV}, {\tt RV AR}, {\tt HAR}, and {\tt SV} methods. Perhaps more importantly, our {\tt SSV} method generates a significantly positive alpha for 21% of the equity strategies in our sample, against, for instance, 3%, 17%, and 14% from the {\tt RV}, {\tt RV6} and {\tt Garch} models, respectively.

The appraisal ratio $AR$ reported in the middle panel of Table (ref) confirms that our {\tt SSV} substantially improves upon standard volatility targeting based on {\tt RV}, especially when more conservative transaction costs are factored in. For instance, with 50 basis points of trading costs the {\tt SSV} is the only method that can still generate a positive appraisal ratio together with the {\tt RV6} long-term realised variance method. By comparison, the {\tt RV}, {\tt Garch} and {\tt RV AR} all generate significantly negative ARs. The bottom panels report the difference in the certainty equivalent return between and investor that can access both the volatility-managed and the original portfolio, and an investor constrained to invest in the original portfolio only. The utility gain $\Delta CER$(%) is highly in favour of our {\tt SSV} volatility targeting. For instance, for 14 (50) basis points of transaction costs, our {\tt SSV} method generates a 12% (8%) utility gain. This compares to the 7% from the {\tt HAR} with 14 basis points and 2.2% from the {\tt RV6} with 50 basis points of transaction costs.

Table (ref) reports the spanning regression results with a tighter leverage cap of 1.5. The results are largely in line with Table (ref). That is, the {\tt RV6} does indeed represents a challenging benchmark for our {\tt SSV} method when it comes to the estimated alphas. However, the $\Delta CER$(%) from the combination strategy is substantially in favour of our smoothing volatility targeting. For instance, the $\Delta CER$(%) from the {\tt SSV} is 9.52% (13.8%) with 50 (14) basis points of notional transaction costs, against a 4.5% (8/2%) from the {\tt RV6} volatility targeting.

Simulation study and inference properties

We now perform an extensive simulation study to evaluate the properties of our estimation framework in a controlled setting. We compare our variational Bayes ({\tt VB}) method against two state-of-the-art Bayesian approaches used within the context of stochastic volatility models, such as {\tt MCMC} (see, stochvol_package) and the global variational approximation recently introduced by chan_yu2022 (henceforth {\tt CY}). Since neither of the benchmark approaches entertain the possibility of arbitrarily smooth predictive densities, the baseline comparison is based on the assumption that $\mathbf{W}=\mathbf{I}_{n+1}$ and the underlying latent state follows an autoregressive dynamics. This gives a cleaner comparison of the accuracy of our variational estimates both in absolute terms and with respect to {\tt MCMC} methods.

We compare each estimation method across $N=100$ replications and for all different specifications. We consider $T=600$, consistent with the shortest time series in the empirical application, $c=0$, $\eta^2=0.1$ and both low and high persistence $\rho\in\{0.70,0.98\}$. Recall that our estimation framework is agnostic on the structure of covariance of the approximating density $\mathbf{\Sigma}_{q(h)}$ (see Proposition (ref)). However, to better understand the contribution of such generalisation compared to existing methods, we also consider the performance of a more tight parametrization with $\mathbf{\Sigma}_{q(h)}=\tau^2\mathbf{Q}^{-1}$, where $\tau^2\in\mathbb{R}^+$ and $\mathbf{Q}=\mathbf{Q}(\gamma)$ (henceforth {\tt VBH}). This provides an homoschedastic representation of the approximating density in the spirit of chan_yu2022, which further simplifies the estimation of $\mathbf{f}_{q(h)}$, $\tau^2$, and $\gamma$.

Figure (ref) reports the mean squared error and a measure of global estimation accuracy compared to the {\tt MCMC}. The mean squared error is measured as $MSE=n^{-1}\sum_{t=1}^n(h_t-\hat{h}_t)^2$, where $h_t$ and $\hat{h}$ are the simulated log-variance and its estimate, respectively. The average aggregated accuracy of variatonal Bayes with respect to the {\tt MCMC} approach is calculated as:

equation[equation omitted — 133 chars of source]

where $p(\mathbf{h}|\mathbf{y})$ is the {\tt MCMC} posterior and $q(\mathbf{h})$ is the comparing variational Bayes approximation (see wand_ormerod.2011). For the higher-persistence scenario with $\rho=0.98$ (top panels), the {\tt MCMC}, {\tt CY}, {\tt VB}, and {\tt VBH} provide statistically equivalent performances. The best approximation to the {\tt MCMC} is provided by our {\tt VB} for $\rho=0.98$.

Interestingly, for the lower-persistent scenario with $\rho=0.70$ (bottom panels), the {\tt CY} approach shows some difficulty in capturing the full extent of the dynamics of the latent stochastic volatility process. This is also reflected in a generally lower accuracy in approximating the true posterior density $p(\mathbf{h}|\mathbf{y})$ compared to the {\tt MCMC} approach. The lower accuracy of the {\tt CY} approach for $\rho=0.7$ is due to a more restrictive dynamics of the latent state imposed by their estimation setting. The approximation proposed by chan_etal.2021 is based on the computationally convenient assumption that the latent volatility state is a random walk. As a result, it shows a substantially lower accuracy when $\rho\ll1$.

Although neither the {\tt CY} nor the {\tt MCMC} approach entertain the possibility of smooth volatility forecasts, for a full comparison of the estimation accuracy of our {\tt VB} method we also evaluate the performance of two alternative smoothing approaches, with $\mathbf{W}$ either a B-spline basis matrix with knots equally spaced every 10 time points (henceforth {\tt VBS}), or a Daubechies wavelet basis matrix with $l=5$ (henceforth {\tt VBW}).\footnote{The choice of the equally spaced knots in the basis function and the $l$ for the wavelet basis matrix is such that both approaches give a similar degree of smoothness.} Notice that both these modifications of $\mathbf{W}$ represent an arbitrary intervention on the approximating density $q\left(\mathbf{h}\right)$. Compared to the baseline {\tt VB}, the smooth approximations have a lower accuracy in the estimate of the underlying AR(1) latent process. Interestingly, similar to {\tt CY} the global accuracy with respect to the {\tt MCMC} deteriorates as the persistence of the latent log-volatility process decreases.

The last column of Figure (ref) shows that our variational Bayes is less computationally expensive compared to both {\tt MCMC} and {\tt CY} methods. The gain in terms of computational cost holds for both highly persistent latent stochastic volatility (top-right panel) and lower-persistent volatility (bottom-right panel). More generally, our {\tt VB} is almost an order of magnitude faster than {\tt MCMC}, on average. This intuitively represents an advantage when implementing real-time predictions for more than a 150 equity strategies, as in our main empirical application.

Figure (ref) suggests that the accuracy of our variational Bayes estimation framework deteriorates when smoothness on the latent state is imposed via the structure in $\mathbf{W}$. We now investigate more in details why that is the case by looking at the posterior estimates of the parameters of interest $\{c,\eta^2,\rho\}$ for difference specifications of $\mathbf{W}$. Figure (ref) shows that by imposing smoothness in the form of either B-spline or a Daubechies wavelet basis forces the posterior estimates of $\rho$ to be close to one, irrespective of the actual level of persistence in the underlying latent process. Similarly, the estimates of the latent state variance $\eta^2$ are smaller for both {\tt VBS} and {\tt VBW} versus {\tt MCMC}'s, and even more so when $\rho=0.7$. Figure (ref) confirms the intuition that a lower accuracy of the posterior estimates of the latent state is due to a tight regularization of the parameters implied by smoothing. The effect on the conditional variance estimates is particularly striking.

Beside the possibility of introducing smoothness in the estimates, our variational Bayes approach relax the assumption that the initial distribution $q(h_0)$ is independent on the trajectory of the latent state $q(\mathbf{h}_1)$, that is, we do not assume $q(\mathbf{h})=q(h_0)q(\mathbf{h}_1)$. Figure (ref) shows that this generalisation has a non-negligible impact on the posterior estimate of the latent state, especially at the beginning on the sample. This is shown by comparing the global accuracy for different slices of observations. The top (bottom) panels report the global accuracy when $\rho=0.98$ ($\rho=0.7$). We report the estimation results for $t\in\left(1,10\right)$ in the left panel, $t\in\left(301,310\right)$ in the middle panel, and $t\in\left(591,600\right)$ in the right panel. The simulation results show that our variational Bayes approach maintains an optimal performance over all the timeline. On the other hand, the accuracy of {\tt CY} drops at the beginning of the time series. This is due to the restrictive independence assumption between the initial condition and the rest of the latent state trajectory $q(\mathbf{h})=q(h_0)q(\mathbf{h}_1)$.

Conclusion

Prior studies found that volatility-managed portfolios that increase leverage when volatility is low produce statistically equivalent economic value compared to the original unscaled factors. This contradicts conventional investment practice whereby risk mitigation should improve, or at least not deteriorates, portfolio returns on a risk-adjusted basis. We show that such equivalence is primarily due to the extreme leverage implied by volatility targeting. Indeed, volatility-managed portfolios based on standard realised variance tend to have extremely levered exposure to the original factors; such exposure is highly time varying. When factoring in moderate levels of notional transaction costs the benefit of volatility-managing disappears.

To regularise turnover and mitigates the effect of transaction costs on volatility-managed portfolios, we propose a novel inference scheme which allows to smooth the predictive density of an otherwise standard stochastic volatility model. Specifically, we develop a novel variational Bayes estimation method that flexibly encompasses different smoothness assumptions irrespective of the underlying persistence of the latent state. Using a large set of 158 equity strategies, we provide evidence that our smoothing volatility targeting approach has economic value when conservative levels of transaction costs are considered. This has important implications for both the risk-adjusted returns and the mean-variance efficiency of volatility-managed portfolios.

\vskip50pt \singlespacing {.05cm }

spacing{0.1}
table[table omitted — 5,137 chars of source]
table[table omitted — 6,184 chars of source]
table[table omitted — 5,408 chars of source]
table[table omitted — 5,087 chars of source]
table[table omitted — 6,172 chars of source]
table[table omitted — 9,031 chars of source]
table[table omitted — 6,210 chars of source]
table[table omitted — 6,216 chars of source]
figure[figure omitted — 1,362 chars of source]
figure[figure omitted — 625 chars of source]
figure[figure omitted — 1,487 chars of source]
figure[figure omitted — 1,205 chars of source]
figure[figure omitted — 1,234 chars of source]
figure[figure omitted — 1,178 chars of source]
figure[figure omitted — 1,402 chars of source]
figure[figure omitted — 1,342 chars of source]
figure[figure omitted — 1,491 chars of source]

\onehalfspacing