Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
29,154 characters · 16 sections · 37 citation commands
A Markowitz Approach to Managing a Dynamic Basket of Moving-Band Statistical Arbitrages
We consider the problem of managing a portfolio of statistical arbitrages (stat-arbs), where each stat-arb is a mean-reverting portfolio of a subset of an $n$-asset universe. Trading strategies based on stat-arbs have been around since the 1980s and are popular due to their documented success in practice. The strategy is based on the idea that if the price of a portfolio of assets is mean-reverting, we can profit by buying the portfolio when it is cheap and selling it when it is expensive. The literature on stat-arbs is vast, and typically focuses on one of two aspects: (i) finding a linear combination of asset prices that is mean-reverting, or (ii) finding a trading policy that profits by exploiting the mean-reversion. To the best of our knowledge, there exists no satisfying solution to the problem of managing a portfolio of multiple stat-arbs, that is independent of the process of finding them. This paper aims to fill this gap and proposes a solution based on a Markowitz optimization approach. In particular, we focus on the case of moving-band stat-arbs (MBSAs), recently introduced in johansson2023statarb, where the midpoint of the price band is allowed to move over time.
\paragraph{Stat-arbs.} Stat-arb strategies have been popular ever since their introduction by Nunzio Tartaglia in the 1980s pole2011statistical, gatev2006pairs. The first strategy was based on pairs trading, where the price difference between two assets is tracked and positions entered when this difference deviates from its mean. Its success has been demonstrated in numerous empirical studies in various markets like equities avellaneda2010statistical, commodities nakajima2019expectations, vaitonis2017statistical, and currencies fischer2019statistical; see, \eg, gatev2006pairs, avellaneda2010statistical, perlin2009evaluation, hogan2004testing, krauss2017deep, caldeira2013selection, huck2019large, dunis2010statistical. The general setting of a stat-arb is a portfolio of (possibly more than two) assets that exhibits a mean-reverting behavior feng2016signal. The literature on stat-arbs tends to split into several categories: finding stat-arbs, modeling the (mean-reverting) portfolio price, and trading stat-arbs. For a detailed overview of the literature, we refer the reader to krauss2017statistical.
\paragraph{MBSA.} MBSAs were recently introduced as an extension of the tradition stat-arb, in which the midpoint of the price band is allowed to move over time johansson2023statarb. Although the method is new, it relies on principles related to Bollinger bands that have been known for decades bollinger1992using,bollinger2002bollinger. The MBSA relies on solving a small convex optimization problem to find a portfolio of (possibly more than two) assets that vary within a moving band. The method has been shown to work well in practice, and is used in the empirical study in this paper.
\paragraph{Trading stat-arbs.} With some exceptions, the literature on trading stat-arbs is mostly limited to pairs trading or trading individual stat-arbs, rather than managing a portfolio of several stat-arbs. Several methods are based on co-integration JOHANSEN2000359, alexander2005indexing, granger1983co, engle_granger_1987, vidyamurthy2004pairs. For example, in zhao2018mean, zhao2018optimal, zhao2019optimal the authors consider a (non-convex) optimization problem for finding high variance, mean-reverting portfolios, in a co-integration space of several stat-arb spreads. Their strategy is based on finding a portfolio of spreads, defined by a co-integration subspace, and implemented using sequential convex optimization. In yamada2012model, primbs2018pairs, yamada2018model the authors model the spread of a pair of assets as an autoregressive process, a discretization of the Ornstein-Uhlenbeck process, and show how to trade a portfolio of spreads under proportional transaction costs and gross exposure constraints using model predictive control. The authors of yamada2012optimal use co-integration techniques to find pairs of assets that are mean-reverting, and show how to construct optimal mean-variance portfolios of these pairs.
Our approach differs from the above methods in that we do not assume any particular model of the price process, and we do not use statistical analysis like co-integration tests to find stat-arbs. Instead, we assume several stat-arbs (defined below as a portfolio and a price signal) are given, and the problem is to find the optimal allocation to these stat-arbs. In other words, we decouple the problem of finding stat-arbs from the problem of portfolio construction.
The rest of this paper is structured as follows. In \S(ref) we review MBSAs and set our notation. In \S(ref) we show how to manage and evaluate a dynamic basket of MBSAs. We present an empirical study of the method on recent historical data in \S(ref), and give conclusions in \S(ref).
MBSAs were recently introduced in johansson2023statarb. We start with a universe of $n$ assets, with $P_t \in \reals^n_{++}$ denoting the price (suitably adjusted) in USD, in period $t=1, 2, \ldots$.
An MBSA is defined by a vector $s \in \reals^n$ of asset holdings (in shares), with negative entries denoting short positions. The vector $s$ is typically sparse, with only a modest number of nonzero entries, but that will not affect our method. The MBSA price in period $t$ is $p_t=s^TP_t$. The MBSA midpoint price is given by the trailing $M$-period average price, \BEQ \mu_t = \frac{1}{M}\sum_{\tau=t-M+1}^t p_\tau, \EEQ where $M$ is the rolling window memory. For a good MBSA, the difference of the price and midpoint price, $p_t-\mu_t$, will oscillate in a band; we obtain a profit by buying when the price difference is low and selling when it is high. We associate with the MBSA the `alpha' value \BEQ \alpha_t = \mu_t - p_t. \EEQ If $\alpha_t$ is positive, we expect the price of the MBSA to increase, and if it is negative, we expect the price to decrease.
In johansson2023statarb we show how to discover MBSAs (\ie, the vector $s$) by approximately solving a constrained variance maximization problem, using the convex-concave procedure shen2016disciplined, lipp2016variations.
Each MBSA has a lifetime, starting at creation or discovery time $t=d$ and ending at time $t=e>d$, when we choose to decommission it. We say it is active over the periods $t=d, d+1, \ldots, e$. (See johansson2023statarb for details.)
We consider a collection of multiple active MBSAs that changes over time as new ones are discovered and some existing ones are decommissioned. We denote $K_t$ as the number of MBSAs active during period $t$, and index the MBSAs as $k=1, \ldots, K_t$. The $k$th MBSA is defined by the holdings $s^{(k)}\in\reals^n$, which gives rise to the price $p^{(k)}_t = (s^{(k)})^T P_t$, and alpha \[ \alpha^{(k)}_t = \mu^{(k)}_t - p^{(k)}_t, \] where $\mu^{(k)}_t$ is the moving midpoint of the $k$th MBSA. We collect the holdings of the $K_t$ MBSAs in a matrix \[ S_t =[ s^{(1)}\;\cdots \;s^{(K_t)}] \in\reals^{n\times K_t}, \] and introduce the $K_t$-vectors \[ p_t=(p^{(1)}_t,\ldots,p^{(K_t)}_t), \qquad \mu_t=(\mu^{(1)}_t,\ldots,\mu^{(K_t)}_t), \qquad \alpha_t=(\alpha^{(1)}_t,\ldots,\alpha^{(K_t)}_t), \] \ie, we extend the notation introduced in the previous paragraph to the setting of multiple MBSAs.
\paragraph{Arb-level and asset-level portfolio holdings.} The (MBSA) holdings during period $t$ are denoted by $q_t \in \reals^{K_t}$. (The units of $q_t$ are `shares' of the MBSAs.) The corresponding asset-level holdings are $h_t=S_tq_t \in \reals^n$ (in shares), or, in USD, $P_t \circ (S_t q_t)$, where $\circ$ denotes the elementwise (Hadamard) product.
The cash account is denoted by $c_t \in \reals$ (in USD). Since we will use the cash account as part of our collateral for short positions, we will assume $c_t>0$ for all $t$, \ie, we do not borrow cash. The total portfolio value at time $t$ is then $p_t^T q_t + c_t$ in USD.
\paragraph{Trading cost.} We denote the (asset level) trades vector at time $t$ as $z_t = h_t- h_{t-1}$, \ie, the change in asset-level holdings from time $t-1$ to time $t$, in shares. The trading cost at time $t$ is given by \[ (\kappa_t^{\text{trade}})^T |z_t|, \] where $\kappa_t^{\text{trade}} \in \reals^n_+$ is the vector of one half the bid-ask spread for each asset at time $t$, in USD per share, and the absolute value is elementwise.
\paragraph{Holding cost.} The holding cost is given by \[ (\kappa_t^{\text{short}})^T (-h)_+, \] where $\kappa_t^{\text{short}} \in \reals^n$ is the vector of shorting rates for each asset over period $t$, in USD per share per trading period. Here $(u)_+=\max\{u,0\}$ is the nonnegative part of $u$, and in the expression above it is applied elementwise.
We manage a dynamic basket of MBSAs using a Markowitz optimization approach markowitz_1952, markowitz1959portfolio, boyd2023markowitz, maximizing the expected return adjusted for transaction and holding costs, subject to constraints on the portfolio holdings which include a risk limit. We assume that we know the current asset prices $P_t$, the arb-to-asset transformation matrix $S_t$, and the previous portfolio holdings $h_{t-1}$ (and $q_{t-1}$) and cash value $c_{t-1}$. Our goal is to find the new portfolio holdings $h_t$ (and $q_t$ and $c_t$). We denote the candidate values, which are optimization variables, as $h$, $q$, and $c$, respectively.
\paragraph{Objective.} The objective is to maximize the alpha exposure, adjusted for transaction and shorting costs, \[ \alpha_t^T q - \gamma^{\text{trade}} (\kappa_t^{\text{trade}})^T(h-h_{t-1}) - \gamma^{\text{short}} (\kappa_t^{\text{short}})^T (-h)_+, \] where $\gamma^{\text{trade}}$ and $\gamma^{\text{short}}$ are positive parameters trading off the alpha exposure, transaction cost, and holding cost. The true values of $\kappa_t^{\text{trade}}$ and $\kappa_t^{\text{short}}$ are not known at time $t$, but are estimated from historical data. This objective is a concave function of the variables $h$ and $q$.
\paragraph{Cash-neutrality.} We assume that the portfolio is cash-neutral boyd2023markowitz,grinold2000active, \ie, \BEQ p_t^T q = 0. \EEQ In other words, the total portfolio value is equal to the value of the cash account. (We can also add a limit on the market exposure (or beta), but we found empirically that this makes a small difference on top of the cash neutrality constraint.) This constraint is often used in long-short trading strategies grinold2000active. The cash neutral constraint is a linear equality constraint on the variable $q$.
\paragraph{Collateral constraint.} We require that \[ P_t^T(h)_+ + (c)_+ \geq \eta (P_t^T(h)_- + (c)_-), \] where $\eta \geq 1$ is a parameter. Here $(u)_- = (-u)_+=\max\{-u, 0\}$ is the nonpositive part of $u$, applied elementwise to $h$ above. This constraint ensures that the long position is at least $\eta$ times the short position. In the form above, it is not a convex constraint, but subtracting $(P_t^T(h_t)_- + (c_t)_-)$ from both sides yields the equivalent convex constraint \[ P_t^Th + c \geq (\eta-1) (P_t^T(h)_- + (c)_-). \] The cash-neutrality constraint implies $P_t^T h=0$, so this simplifies to \[ c \geq (\eta-1) P_t^T(h)_-, \] \ie, the cash account is at least $(\eta-1)$ times the short position. For example, we can set $\eta=2.02$ to keep a collateral of 102% of the short position, as has typically been required by regulation d2002market, geczy2002stocks.
\paragraph{MBSA size limit constraints.} We add the constraints \[ |q|_k(p_t)_k \leq \xi_t^{(k)} c, \quad k=1,\ldots,K_t, \] \ie, the absolute value of the position in the $k$th MBSA is at most $\xi_t^{(k)}$ times the total portfolio value (which is the same as the cash account), where $\xi_t^{(k)}$ is a parameter. These constraints serve two purposes. First, they act as a form of regularization in case there are more MBSAs than assets, in which case the covariance matrix in the risk constraint (defined below) will be ill-conditioned or singular. Second, they limit extreme positions in any single MBSA. During the active period $t=d, d+1, \ldots, e$ of an MBSA, we set $\xi_t^{(k)}=\xi$ for $t=d, d+1, \ldots, e-l$, and then linearly decrease $\xi_t^{(k)}$ to zero over the next $l$ periods, where $l$ is a parameter, say $l=21$ periods, and $\xi$ is a parameter common to all MBSAs. In other words we use the size limit constraints to gradually decommission MBSAs. These constraints are linear inequality constraints on the variables $q$ and $c$, and so are convex.
\paragraph{Risk constraint.} The traditional definition of risk of a portfolio is the variance of the portfolio return, expressed as a quadratic form of the asset weights with an estimated asset return covariance matrix. Taking the squareroot we obtain the standard deviation of portfolio return, \ie, its volatility, which we can convert to USD by multiplying by the portfolio value.
Here we propose a different risk measure, which we have found empirically to work better, and is more in line with the spirit of MBSAs. Our risk is based on fluctuation of the portfolio value around its expected midpoint, and is directly expressed in USD. Recall that the portfolio value is $p_t^Tq_t$, which we expect to fluctuate around the midpoint value $\mu_t^Tq_t$, both of these in USD. We take the risk to be an estimate of the short term mean square value of the difference of the portfolio value and the midpoint value, $(p_t-\mu_t)^T q_t$. We express this mean-square value (now using $q$, not $q_t$) as \[ q^T \Sigma_t q , \] where $\Sigma_t$ is an average of the recent values of $(p_\tau - \mu_\tau) (p_\tau - \mu_\tau)^T$. We limit our risk using the constraint \[ \|\Sigma_t^{1/2} q\|_2 \leq \sigma c, \] where $\sigma > 0$ is a parameter setting the target risk (which is unitless) and $c$ is the portfolio value. This risk limit is a convex constraint in the variables $q$ and $c$, specifically a second-order cone (SOC) constraint BoV:04.
The covariance matrix $\Sigma_t$ can be estimated in many ways; see, \eg, johansson2023covariance. We express it in terms of asset prices as \[ \Sigma_t = S_t^T \Sigma^{P}_t S_t, \] where $\Sigma^{P}_t\in\reals^{n\times n}$ is an estimate of the short-term covariance matrix of the asset prices. We estimate $\Sigma^{P}_t$ as follows. First, define the centered price vectors \[ \tilde P_t = P_t - \bar P_t, \quad t = 1, \ldots, T, \] where $\bar P_t$ is the $M$-period rolling window mean of the asset prices. (This is the same $M$ used to define the MBSA midpoints in (ref).) Then, $\Sigma^{P}_t$ is an average of the recent values of $\tilde P_\tau \tilde P_\tau^T$. If a linear estimator, like a rolling window or exponentially weighted moving average (EWMA), is used to estimate this expectation, then this is equivalent to estimating $\Sigma_t$ directly using the same method, \ie, to center the MBSA prices and then compute an average of the outer product of the centered prices. We will use the iterated EWMA predictor johansson2023covariance,cov_barrat_2022, to estimate the covariance matrix of the centered prices $\tilde P_t$. (This is not equivalent to estimating $\Sigma_t$ directly with an iterated EWMA, but in practice very similar.)
We assemble the objective and constraints described above. We take $h_t$, $q_t$, and $c_t$, as optimal values of $h$, $q$, and $c$ in the optimization problem \BEQ
\EEQ with variables $h \in \reals^{n}$, $q \in \reals^{K_t}$, and $c$. The problem data are: \BIT • $h_{t-1}$, $q_{t-1}$, $c_{t-1}$, the previous portfolio holdings at the asset-level and arb-level, and cash; • $\kappa_t^{\text{trade}}$ and $\kappa_t^{\text{short}}$, (predictions of) the trading and holding costs; • $S_t$, the arb-to-asset transformation matrix; • $\eta$, the collateral parameter; • $P_t$ and $p_t$, the asset and MBSA prices; • $\xi_t^{(k)}$, the MBSA size limits; • $\Sigma_t$, a prediction of the short-term covariance matrix of the MBSA prices; • $\sigma$, a target risk, expressed as a fraction of the portfolio value. \EIT The problem (ref) is a convex optimization problem, more specifically one that can be transformed to a second-order cone program (SOCP). The solution to this optimization problem gives the new MBSA portfolio $q_t$, the asset-level portfolio holdings $h_t$, and the cash account $c_t$ at time $t$.
\paragraph{Extensions and variations.} We can add additional constraints, such as a leverage constraint, asset position limits, maximum market exposure, \etc boyd2023markowitz. We can also soften some constraints, if it is not critical that they are satisfied exactly and softening improves performance boyd2023markowitz. For example, to soften the arb-to-asset constraint, we remove the constraint $h=S_tq$ and add a penalty term $\gamma^{\text{hold}}\|P_t\circ(h - S_tq)\|_1$ to the objective. (Note that softening the arb-to-asset constraint implicitly also softens the risk and cash-neutrality constraints if these are expressed in terms of the arb-level portfolio $q$.) Softening some constraints allows the optimizer to choose values that violate the original hard constraints when necessary, which can reduce unnecessary trading, and therefore transaction cost.
We illustrate the MBSA portfolio management strategy on recent historical data. Code to replicate the experiments is available at
\paragraph{Data set.} We gather daily price data from the CRSP US Stock Databases using the Wharton Research Data Services (WRDS) portal WRDS. The data set consists of adjusted asset prices of 15405 stocks from January 4, 2010, to December 30, 2023, for a total of 3282 trading days.
\paragraph{Monthly search for MBSAs.} We search for MBSAs every 21 trading days, with the same setup and parameters as described in detail in johansson2023statarb.
\paragraph{Dynamic management of the MBSAs.} Every day we solve the Markowitz optimization problem (ref) to rebalance our portfolio. Each MBSA is kept in the portfolio for 500 trading days. After that $\xi$ is reduced linearly to zero over the next 21 trading days.
\paragraph{Risk model.} As described in \S(ref), we decompose the covariance matrix as \[ \Sigma_t = S_t^T \Sigma^{P}_t S_t, \] where $\Sigma^{P}_t$ is the short-term covariance matrix of the asset prices. The covariance matrix $\Sigma^{P}_t$ is estimated using the iterated EWMA (IEWMA) predictor (see DCC and johansson2023covariance) on the centered prices $\tilde P_t = P_t - \bar P_t$, where $\bar P_t$ is the 21-day rolling window mean of the asset prices. For the IEWMA predictor, we use a 125-day half-life for volatility estimation, and a 250-day half-life for correlation estimation. To reduce trading induced by the risk-model, we smooth the covariances with a 250-day half-life EWMA johansson2023covariance.
\paragraph{Parameters.} For the MBSA problem we use the same parameters as in johansson2023statarb. For the Markowitz optimization problem we use $\gamma^{\text{trade}}=1$, $\eta=1$, $\xi=1$, and $\sigma^{\text{tar}}=10\%$ annualized.
\paragraph{Trading and shorting costs.} Our numerical experiments take into account transaction costs, \ie, we buy assets at the ask price, which is the (midpoint) price plus one-half the bid-ask spread, and we sell assets at the bid price, which is the price minus one-half the bid-ask spread. We use 0.5% as a proxy for the annual shorting cost of stocks, which is well above what is typically observed in practice for liquid stocks d2002market, geczy2002stocks, kim2023shorting, drechsler2014shorting. We also note that we have tested the method with shorting costs upwards of 10% annually, and it remains profitable.
This section explains how we simulate the portfolio, and the metrics we use to evaluate the performance.
\paragraph{Cash account.} We initialize the portfolio with a cash account at time $t=0$, $c_0 = 1$. The cash account evolves as \[ c_{t+1} = c_t - (q_{t+1}-q_t)^Tp_{t+1} - \phi_t = c_t + q_t^Tp_{t+1} - \phi_t, \quad t=0,\ldots, T, \] where $\phi_t$ is the transaction and holding cost, consisting of the trading cost at time $t+1$ and the holding cost over period $t$. (The last equality follows from the cash-neutrality constraint.)
\paragraph{Portfolio NAV.} The net asset value (NAV) of the portfolio, including the cash account, at time $t$ is \[ V_{t} = c_t + q^T_t p_t = c_t. \] Due to the market neutrality constraint in (ref) we have $V_t=c_t$, \ie, the NAV is equal to the cash account.
\paragraph{Return.} The return at time $t$ is \[ r_t = \frac{V_t - V_{t-1}}{V_{t-1}}, \quad t = 1, \ldots, T. \] We report several standard metrics based on the returns $r_t$. The average return is \[ \overline r = \frac{1}{T} {\sum_{t=1}^{T}{r_t}}, \] which we multiply by $250$ to annualize. The return volatility is \[ \left(\frac{1}{T} \sum_{t=1}^{T}\left(r_t-\overline r\right)^2\right)^{1/2}, \] which we multiply by $\sqrt{250}$ to annualize. The annualized Sharpe ratio is the ratio of the annualized average return to the annualized volatility. Finally, the maximum drawdown is \[ \max_{1 \leq t_1 < t_2\leq T} \left(1- \frac{V_{t_2}}{V_{t_1}}\right), \] the maximum fractional drop in value form a previous high.
\paragraph{Turnover.} The turnover at time $t$ is \[ T = \frac{1}{2} \sum_{i=1}^{n} |((P_t\circ h_t)_i - (P_t\circ h_{t-1})_i)/V_t| = \frac{1}{2} \|P_t\circ (h_t - h_{t-1})\|_1 / V_t, \] which we multiply by $250$ to annualize. It measures the amount of trading in the portfolio grinold2000active. For example, a turnover of $0.01$ means that the average of total amount bought and total amount sold is 1% of the total portfolio value.
\paragraph{Active return and risk.} We define the active return as the return of the portfolio minus the return of the market, which we take to be the S&P 500. The active risk is the standard deviation of the active return.
\paragraph{Residual return and risk.} Given the portfolio returns $r_1,r_2,\ldots, r_T$, and the corresponding market returns $r^m_1,r^m_2,\ldots, r^m_T$, we construct the linear model \[ r_t = \beta r^m_t + \theta_t, \quad t=1,\ldots,T, \] where $\beta r^m_t$ is the return explained by the market (with $\beta\in\reals^n$), and $\theta_t\in\reals^n$ is the residual return at time $t$. The residual risk is the standard deviation of the residual return. The mean residual return is often referred to as the alpha of the portfolio grinold2000active (not to be confused with the alpha of an MBSA). The information ratio is the ratio of the portfolio alpha to the residual risk.
\paragraph{Portfolio performance.} The portfolio performance is summarized in table (ref).
We attain an average annual return of 19% at an annual volatility of 12%, corresponding to a Sharpe ratio of 1.61, with a maximum drawdown of 15% over the roughly 10-year period. The average turnover is 136, which corresponds to a daily turnover of roughly 50%. In comparison to other stat-arb strategies in the literature, this seems like a reasonable level of turnover guijarro2021deep.
\paragraph{Active and residual return and risk.} Here we compare the MBSA strategy to the market, represented by the S&P 500 with an initial investment of \$1 and diluted with cash to attain the same annualized risk as the MBSA strategy. Table (ref) gives a numerical summary. The MBSA strategy attains an annualized Sharpe ratio of 1.61 (compared to 0.66 for the market) with a residual return (alpha) of 18% and a market beta of 11%. The information ratio is 1.53.
\paragraph{NAV evolution.} Figure (ref) shows the NAV of the MBSA strategy and the market, and the 250-day half-life EWMA correlation between the two.
As seen, the MBSA strategy outperforms the market, with very low correlation, 15% over the whole period. This suggests mixing the MBSA strategy with the market. For example, mixing 90% of the MBSA strategy and 10% of the market yields an annualized Sharpe ratio of 1.66, slightly better than the MBSA strategy alone.
\paragraph{Annual performance.} Here we break down the metrics reported above over the 11-year period into the performance metrics for each year. Figure (ref) shows the annual performance of the MBSA strategy and the market. The MBSA strategy has positive return in each of the 11 years, whereas the market has negative return in 2 of the 11 years. The MBSA strategy outperforms the market in 8 of the 11 years, and has more stable performance.
Finally, figure (ref) shows the annual residual return and risk, and market beta. As seen, most of the MBSA success is unexplained by the market.
We have shown how to manage a dynamic basket of moving-band stat-arbs, based on a long-short Markowitz optimization strategy. We presented an empirical study of the method on recent historical data, showing that it can to outperform the market with low correlation.