Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
142,160 characters · 14 sections · 0 citation commands
Sparse spanning portfolios and under-diversification with second-order stochastic dominance
We know for decades that the diversification benefits measured by the volatility of portfolio returns are limited when we invest beyond 10 to 20 assets; see e.g.\ Evans and Archer (1968), Klemkosky and Martin (1975), Elton and Gruber (1977). Practitioners coin the term over-diversification. At the opposite end of the spectrum, we often observe under-diversification among households (Campbell (2006), Calvet, Campbell and Sodini (2007)). It might be caused by information acquisition costs (Van Nieuwerburgh and Veldkamp (2010)), overconfidence (Anderson (2013)), solvency requirements (Liu (2014)), or overweighting low probability events (Dimmock et al.\ (2021)). A characteristic-based demand system might also explain why institutions and households hold a small set of stocks (Koijen and Yogo (2019)).
Possible over-diversification contributes to motivating the recent literature on sparse construction of mean-variance (MV) portfolios within the Modern Portfolio Theory (Markowitz (1952)) through imposing constraints on the portfolio weights; see e.g.\ Jagannathan and Ma (2003), DeMiguel et al.\ (2009), Brodie et al.\ (2009), Fan, Zhang and Yu (2012), Ao, Li and Zheng (2019), and Caner, Medeiros and Vasconscelos (2023). Such a construction limits the impact of transaction costs, and eases monitoring and risk management. It also achieves statistical regularisation of the investment portfolio in the presence of ill-conditioned large covariance matrices. Whether limitations of diversification benefits beyond a given small number of assets still hold true when we leave the MV paradigm is an open problem. This paper targets the following questions: Is it possible to build a sparse portfolio of dimension $q$ from a large set of assets of dimension $p$ so that we cannot get further improvement from considering additional assets in a second-order stochastic dominance (SSD) paradigm? If not, how much do we lose by limiting ourselves to this sparse portfolio in terms of expected utilities compatible with SSD? Can we design an optimization algorithm to compute this sparse portfolio from available data? Do we have the asymptotic statistical guarantee that we cannot improve on the estimated expected utility loss due to under-diversification by considering another sparse portfolio of the same fixed dimension?
The theory of stochastic dominance (SD) gives a systematic framework for analyzing investor behavior under uncertainty (see Chapter 4 of Danthine and Donaldson (2014) for an introduction oriented towards finance). Stochastic dominance ranks portfolios based on general regularity conditions for decision making under risk (Hadar and Russell (1969), Hanoch and Levy (1969), and Rothschild and Stiglitz (1970)). SD uses a distribution-free assumption framework which allows for nonparametric statistical estimation and inference methods. We can see SD as a flexible model-free alternative to MV dominance of Modern Portfolio Theory (Markowitz (1952)). The MV criterion is consistent with Expected Utility for elliptical distributions such as the normal distribution (Chamberlain (1983), Owen and Rabinovitch (1983), Berk (1997)) but has limited economic meaning when we cannot completely characterize the probability distribution by its location and scale. Simaan (1993), Athayde and Flores (2004), and Mencia and Sentana (2009) develop a mean-variance-skewness framework based on generalizations of elliptical distributions that are fully characterized by their first three moments. SD presents a further generalization that accounts for all moments of the return distributions without necessarily assuming a particular family of distributions.
Second-order SD (SSD) spanning (Arvanitis et al.\ (2019)) is a model-free alternative to MV spanning of Huberman and Kandel (1987) (see also Jobson and Korkie (1989), De Roon, Nijman, and Werker (2001), Ardia, Laurent, and Sessinou (2024)). Spanning occurs if introducing new securities or relaxing investment constraints does not improve the investment possibility set for a given class of investors. MV spanning checks if the MV frontier of a set of assets is identical to the MV frontier of a larger set made of those assets plus additional assets (Kan and Zhou (2012), Penaranda and Sentana (2012)). Here we investigate such a problem for investors with risk-averse preferences which are interested in the whole return distributions generated by two sets of assets, a sparse subset of dimension $q$ (10 assets) with a limited number of assets coming from a much larger set of dimension $p$ (500 assets).
The first contribution of the paper is to introduce the concept of sparse SSD spanning. We propose a theoretical measure for sparse spanning based on second-order stochastic dominance. For economic interpretation, we provide a representation based on a class of concave utility functions without assuming differentiability. When sparse SSD spanning occurs, a risk-averse investor will not improve her expected utility by shifting from the sparse subset to the larger investment opportunity set. On the contrary, if it does not occur, the risk-averse investor suffers an expected utility loss since we work with a subset instead of the full set of assets. Hence we further provide a lower bound that takes the interpretation of an optimal expected utility loss that cannot be improved upon by any sparse subset made of $q$ assets. We know that we suffer a loss because of the sparsity constraint but we cannot do better though investing optimally in only $q$ assets under an SSD criterion. To check sparse SSD spanning on data, we develop consistent and feasible estimation procedures based on Linear Programming (LP) and a greedy algorithm, namely the Forward Stepwise Selection (FSS) algorithm. We use a finite set of increasing piecewise-linear functions, restricted to the bounded empirical supports, that are constructed as convex mixtures of appropriate “ramp functions” (in the spirit of Russel and Seo (1989)) in our representation as in Arvanitis, Scaillet and Topaloglou (2020a,b). For every such utility function, we solve two embedded linear maximization problems. It is an improvement over the implementation in Arvanitis and Topaloglou (2017) and Arvanitis, Scaillet and Topaloglou (2020b) where they formulate the empirical counterpart in terms of Mixed-Integer Programming (MIP) problems. MIP problems are NP-complete, and far more difficult to solve. Our numerical approximations are simple and fast since they are based on standard LP. They suit better computationally intensive optimisation methods, which otherwise become quickly computationally demanding in empirical work on large data sets. Those formulations are reminiscent of the LP programs developed in the early papers of testing for SSD efficiency of a given portfolio by Post (2003) and Kuosmanen (2004) (see also Scaillet and Topaloglou (2010)).
Since we aim at a sparse solution computed from a large dimensional problem, we rely on a greedy optimisation algorithm. We use a discrete combinatorial algorithm for maximizing a function subject to a cardinality constraint. It starts with the empty set, and then adds elements to it in $r$ iterations. In each iteration, the algorithm adds to its current solution the single element increasing the value of this solution by the most, i.e., the element with the largest marginal value with respect to the current solution. In the context of submodular maximization (see Buchbinder and Feldman (2018) for a survey), this simple FSS algorithm checking for incremental gain at each step using nested models is usually referred to simply as “the greedy algorithm”. In the case of submodular functions, it returns a solution that is provably within a constant factor of the optimum (Nemhauser, Wolsey and Fisher (1978)), and it turns out to be the best approximation ratio possible for the problem (Nemhauser and Wolsey (1978)). Submodular functions have a natural diminishing return property: adding an element to a larger set results in smaller marginal increase in the value of the function compared to adding the element to a smaller set. Submodular functions share also a natural sub-additivity property: for two disjoint sets, submodularity implies sub-additivity (but the converse is not true). In the context of risk measures, sub-additivity is related to the notion of diversification, a desirable property to be classified as a coherent risk measure (Artzner at al.\ (1998)). The lower partial moments that we use below to construct our test statistics have a type of put option pay-off and are measure of downside risk (Bawa (1975), Fishburn (1977)). They are known to be coherent risk measures and, as such, benefit from the sub-additivity property. Bian et al.\ (2017) extend guarantee results of the greedy algorithm for cardinality constrained maximization of non-submodular nondecreasing set functions, in particular nondecreasing standard LP problems with non-degenerate basic feasible solution (Bertsimas and Tsitsiklis (1997), Ch.\ 3) that we implement in our empirics.
We choose that approach over penalization methods currently used for building sparse MV portfolios for two reasons. First, we wish to bound the relative error without any assumptions on the underlying sparsity for the true parameter. It is useful to show the consistency of our empirical strategy irrespective of sparse spanning being present or absent. Our proof relies on the recent work of Elenberg et al.\ (2018) (see Das and Kempe (2011) for the linear regression case). Contrary to prior work in the MV setting, we require neither assumptions on the sparsity of the underlying problem nor i.i.d.\ returns. We establish multiplicative approximation guarantees from the best-case sparse solution. Our results improve over previous work by providing bounds on a solution that is guaranteed to match the desired sparsity and cannot be further decreased. Convex methods for linear regressions such as the standard LASSO objective (Tibshirani (1996)) require strong assumptions on the model and the data, such as the unrepresentable condition on the parameter vector and i.i.d.\ data (Zhao and Yiu (2006), Meinshausen and Buhlmann (2006)), in order to provide exact sparsity guarantees on the recovered solution (see Zhang (2009) for use of these assumptions in greedy least squares regression). More specifically, when the number $r$ of iterations is equal to $r=q\ln T$, $T$ being the time-series sample size, we show that the algorithm provides a consistent estimate of the bound of the expected utility loss computed from financial returns satisfying a mixing condition. Mixing holds true for many time series models such as ARMA models as well as several GARCH and stochastic volatility processes (see Francq and Zakoian (2011) for several examples). It allows us to build a path of the estimated bound as a function of the sparsity constraint $q$, and verify when we have a sufficiently large $q$ to get sparse SSD spanning, namely when the bound vanishes. Second, the only input we need is the sparsity number $q$ of assets. Hence, we avoid the selection problem of a tuning parameter, namely the regularization parameter in penalization methods. As discussed in Brodie et al.\ (2009), a portfolio selection with a LASSO approach regulates the amount of shorting. In our setting, we use short-sales constraints which corresponds to using an implicit large regularisation parameter for the LASSO penalty. Our numerical approach based on a greedy algorithm however does not require the true portfolio to be sparse, and a large regularisation parameter is not required for developing valid statistical inference. As a by-product, our approach also provides a selection algorithm for sparse MV spanning under multivariate normality using the equivalence with sparse SSD spanning for elliptical distributions. It allows to bypass the regularization of ill-conditioned estimates of large covariance matrices (see e.g.\ Fan, Liao, and Shi (2015), Ledoit and Wolf (2017)).
The second contribution of the paper aims at checking on large datasets of equity returns whether sparse SSD holds or not. We find that there is no benefit from expanding a sparse opportunity set beyond 45 assets. The optimal sparse portfolio invests in 10 industry sectors and cuts tail risk when compared to a sparse MV portfolio. On a rolling-window basis, the number of assets shrinks to 25 assets in crisis periods, while standard factor models cannot explain the performance of the sparse portfolios. .
The paper is organized as follows. In Section 2, we establish our probabilistic framework, and review the definition of SSD. In Section 3, we define the relevant concept of sparse SSD spanning and provide convenient functional representations. In Section 4, we construct an estimate of the bound for sparse SSD spanning by using empirical analogues. We exploit the limiting distribution of the empirical process underlying the estimator which has the form of a Gaussian process. Our estimation strategy builds on LP and an FSS algorithm. We show the asymptotic optimal recovery of the sparse solution, namely statistical approximation guarantee of the greedy algorithm output for a given $q$ when $T$ becomes large. In Section 5, we describe the numerical implementation aspects of our empirical procedures. In Section 6, we analyze large datasets of equity returns to study whether sparse SSD holds or not and compare with results given by the construction of sparse MV portfolios with the MAXSER approach of Ao, Li and Zheng (2019). We provide concluding remarks in Section 7. We provide our proofs and the list of factors used in the empirical application in the Appendix. The Online Appendix discusses the concept of approximate sparse spanning. Given a fixed support dimension, it specifies the low dimensional portfolio set that comes closer (in an appropriate sense defined later on in the paper) spanning the high dimensional one. It also gathers Monte Carlo experiments to assess the finite sample properties of our procedure for sparse SSD spanning.
We describe our limiting economy for a large number of financial assets. We denote the financial returns by a process $X^{\infty}$ living in $\ell^{\infty}\left(\mathbb{N},\mathbb{R}\right)$, which is the space of bounded real valued sequences equipped with the uniform metric. $X^{(i)}$ denotes the $i^{\text{th}},\:i\in\mathbb{N}$ coordinate, $X$ denotes the projection of $X^{\infty}$ in the first $p$ coordinates, and $\mathbb{P}$ denotes the distribution of $X^{\infty}$. We suppress dependence on $p$ for brevity.
We introduce the associated portfolio weights with short-sales constraints. Short-sales constraints on the asset allocation promote sparsity (Brodie et al.\ (2009)); our approach can be used to trace further patterns of (desired) sparsity. The set $\Lambda_{\infty}$ is a non-empty subset of the $\mathbb{N}$-simplex $\left\{ \lambda\in\mathbb{R}^{\mathbb{N}}:\lambda_{i}\geq0,i\in\mathbb{N},\sum_{i=0}^{\infty}\lambda_{i}=1\right\}$, and for $p\in\mathbb{N}$, $\Lambda=\left\{ \lambda\in\Lambda_{\infty},\sum_{i=0}^{p-1}\lambda_{i}=1\right\} $ denotes the $p-1$ dimensional unit sub-simplex of $\Lambda_{\infty}$ and $K$ is a non-empty closed subset of $\Lambda$.\footnote{Positive portfolio weights summing to one induce compactness of the parameter space $\Lambda$ which facilitates proofs. If short sales are allowed, we can alternatively assume that portfolio weights lie inside compact sets because of lending restrictions. The unit ball $\ell^{1}$-type restrictions imposed by the simplex consideration along with the Lipschitz continuity of $(x)_{+}$ imply distributional robustness: the expectations are equivalent to expectations w.r.t.\ the worst case distribution in a Wasserstein neighbourhood of the underlying distribution; see Theorem 1 of Gao, Chen, and Kleywert (2017).} In the present context, $X$ is a random vector of financial returns for $p$ base assets, while $\Lambda$ represents a set of portfolios formed on $X$. The process $X^{\infty}$ idealizes the high dimensional situation in the limiting case where $p\rightarrow +\infty$. Our first assumption specifies probabilistic properties for $X^{\infty}$. It requires mild moment existence conditions (bounded sequence of first order moments), and a lower bound on the associated supports consistent with non-logarithmic returns. There, $\mathrm{supp}$ denotes support of the distribution of the random variable involved, $\bar{\text{co}}$ denotes the closure of the convex hull, and $(x)_{+}:=\max(0,x)$. Given the restrictions that its elements satisfy, $\Lambda_{\infty}$ is considered topologized by the $l_{1}$ norm and $\lambda,\:\kappa$ denote generic elements of $\Lambda_{\infty}$.
The assumption implies that for any $\lambda\in\Lambda_{\infty}$, $\sum_{i=0}^{\infty}\lambda_{i}X^{(i)}$ is a well defined random variable and that the Lower Partial Moment Differential (LPMD) defined by
is also bounded (and hence exists) and continuous in $z,\lambda,\kappa$.\footnote{This is due to the monotonicity of the integral, $\mathbb{E}\left[\left|\sum_{i=0}^{\infty}\lambda_{i}X^{(i)}\right|\right]\leq\max_{i}\mathbb{E}\left[\left|X^{(i)}\right|\right]<+\infty$. The assumption implies that the partial moment $\mathbb{E}\left[\left(z-\sum_{i=0}^{\infty}\lambda_{i}X^{(i)}\right)_{+}\right]$ is continuous in $z,\:\lambda$ via dominated convergence, and that it is also bounded in $\lambda$ for any $z$, even though $\Lambda_{\infty}$ is not ($l_{1}$-) totally bounded. Along with the Lipschitz continuity property of $\left(\cdot\right)_{+}$, it also implies that for any $\lambda\neq\kappa$,
} The LPMD ((ref)) takes the interpretation of the difference between two coherent risk measures evaluated on the portfolios at hand. Assumption (ref) facilitates the definition of SSD for the constructed portfolios:
The definition is simply an adaptation of the usual SSD relation in our high dimensional framework. Using the classical Russell and Seo (1989) utility representations, we obtain the well known result that $\kappa\underset{\text{SSD}}{\succeq}\lambda$ iff the former is preferred to the latter by every increasing and concave utility. Thus, SSD exemplifies universal choices w.r.t.\ every insatiable and risk averse investor.
Arvanitis et al.\ (2018) define the notion of SSD Spanning as an extension of the MV analogue. It involves comparison of portfolio sets that are not necessarily singletons.
If the sets are not related by inclusion and $K\underset{\text{SSD}}{\succeq}\Lambda$, then we have necessarily $K\underset{\text{SSD}}{\succeq}K\cup\Lambda$. Furthermore, spanning would be trivial if $K\supseteq\Lambda$ were allowed. Hence, we can always consider that $K$ lies inside $\Lambda$. Spanning admits an economic interpretation when $K\subseteq\Lambda$; it means that extension of the investment opportunity set from $K$ to $\Lambda$ does not improve investment possibilities for any risk averter. Hence, no spanning means that the extension contains a non dominated element. It is formalized as follows: $K\underset{\text{SSD}}{\nsucceq}\Lambda$ iff $\exists\lambda\in\Lambda:\forall\kappa\in K,\:\kappa\underset{\text{SSD}}{\nsucceq}\lambda$, i.e., $\lambda$ is maximal (efficient) w.r.t.\ $K$.
Under some further structure on $K$, SSD spanning admits an empirically useful characterization involving a saddle-type point of the LPMDs.
We extend the notion in the high dimensional setting, by also allowing\textcolor{black}{ a potentially unknown} low dimensional investment opportunity set to SSD span a high dimensional superset. In order to formally define this and extend it to the limiting case where $p\rightarrow +\infty$, we introduce the following notation for the support of a portfolio set: $\mathrm{csupp}\left(K\right):=\#\left\{ i:\kappa_{i}\neq0,\kappa\in K\right\} $. By construction $\mathrm{csupp}\left(\Lambda\right)=p$. We suppose that as $p\rightarrow+\infty$ and $\lim_{p\rightarrow +\infty}\Lambda=\Lambda_{\infty}$, where the limit is interpreted in the Painleve-Kuratowski convergence mode. The sequence $\left(\Lambda\right)_{p}$ is by construction monotone increasing.
\textcolor{black}{Definition (ref) generalizes Definition (ref) in a twofold manner. First, it allows for a limiting high dimensional setting thus providing the proper framework for addressing the empirical questions listed in the introduction. Second, it only prescribes the existence of a “low-dimensional” spanning subset of $\Lambda$, whereas for the original definition the spanning subset is exogenously given. It implies that any procedure designed to test whether SS-SSD holds would have to search for a spanning set inside the collection of “low-dimensional” subsets of $\Lambda$. It is useful even in the case where SS-SSD does not hold. As the following paragraph suggests, such a procedure, if consistent, would end up with a sparse portfolio set that “comes as close as possible” to SSD span its high dimensional universe of portfolios. }
As in Lemma (ref), we obtain a useful characterization of SS-SSD by considering the collection $\mathcal{L}_{p,q}=\left\{ K\subset\Lambda:K\text{ closed},0<\mathrm{csupp}\left(K\right)\leq q\right\} $:\footnote{ When $\Lambda$ is itself a simplicial complex, then $\mathcal{L}_{p,q}$ is also a simplicial complex of dimension $q-1$. Then and if $p\geq2q$, $\mathcal{L}_{p,q}$ has a geometric realization as a sub-simplex of the standard $p-1$ simplex (see the Geometric Realization Theorem in Edelsbrunner (2014)).}
The possibility of interchanging the order of appearance of the optimization operators in the characterization $\inf_{\mathcal{L}_{p,q}}\sup_{\Lambda}\inf_{K}\sup_{z\in Z}D\left(z,\kappa,\lambda,\mathbb{P}\right)\leq0$ to $\sup_{z\in Z}\sup_{\Lambda}\inf_{\mathcal{L}_{p,q}}\inf_{K}$ will greatly facilitate numerical aspects as well as the derivations of limiting properties for the empirical procedures. It actually holds via the use of appropriate minimax theorems and the extension of our assumption framework.
We have $\sup_{\Lambda}\inf_{\mathcal{L}_{p,q}}\inf_{K}D\left(z,\kappa,\lambda,\mathbb{P}\right) =\inf_{\mathcal{L}_{p,q}}\inf_{K}\mathbb{E}\left[\left(z-\sum_{i=0}^{\infty}\kappa_{i}X_{t}^{(i)}\right)_{+}\right] $ \break $-\inf_{\Lambda}\mathbb{E}\left[\left(z-\sum_{i=0}^{\infty}\lambda_{i}X_{t}^{(i)}\right)_{+}\right]$, given an arbitrary threshold $z$, so that we can separate the optimizations w.r.t.\ the “parameter sets” $\Lambda$ and $\mathcal{L}_{p,q}\times K$. It is useful especially in the case where we approximate the outer optimization over $Z$ by some discretization, as in our empirical numerical implementations.
In this section, and given the latency of $\mathbb{P}$, we are interested in the empirical approximation of the element of $\mathcal{L}_{p,q}$ that approximately spans $\Lambda$ for a fixed $q$, and the subsequent estimation of the associated diversification loss $M\left(\Lambda,\mathcal{L}_{p,q},\mathbb{P}\right)$. We employ the empirical analogues of the functionals that characterize spanning, and design the sparse optimization involved via a greedy algorithm. We establish consistency using the results on statistical guarantees by Elenberg et al.\ (2018). We derive the usual parametric $\sqrt{T}$ rate and the limiting distribution, and construct a conservative inferential procedure based on fast subsampling.
Consider the sequence $\left(X_{t}^{\infty}\right)_{t\in\mathbb{Z}}$ where for all $t$, $X_{t}^{\infty}\underset{\text{d}}{=}X^{\infty}$ and $\underset{\text{d}}{=}$ denotes equality in distribution. Suppose, that for some $p$, a sample of $\left(X_{t}\right)_{t=1,\dots,T}$ is available from the sequence $\left(X_{t}\right)_{t\in\mathbb{Z}}$. Denote with $\mathbb{P}_{T}$ its empirical distribution function in $\mathbb{R}^{p}$ (in what follows $\mathbb{P}$ also identifies the distribution of $X_{0}$ in $\mathbb{R}^{p}$ without inconsistency due to the Daniel-Kolmogorov Theorem). We approximate $D\left(z,\kappa,\lambda,\mathbb{P}\right)$ by $D\left(z,\kappa,\lambda,\mathbb{P}_{T}\right)$ and design a procedure that evaluates $\inf_{\mathcal{L}_{p,q}}\sup_{\Lambda}\inf_{K}\sup_{z\in Z}D\left(z,\kappa,\lambda,\mathbb{P}_{T}\right)$.
Given Lemma (ref), we design our empirical procedure as follows: for fixed $q$, formulate the empirical optimization problem $\sup_{z\in Z}\sup_{\Lambda}\inf_{\mathcal{L}_{p,q}}\inf_{K}D\left(z,\kappa,\lambda,\mathbb{P}_{T}\right)$, as
Given $\mathcal{K}\left(\Lambda,\mathcal{L}_{p,q},z,\mathbb{P}_{T}\right)$ the numerical technology evaluating $\mathcal{J}\left(\Lambda,z,\mathbb{P}_{T}\right)$ and $M\left(\Lambda,\mathcal{L}_{p,q},\mathbb{P}_{T}\right)$ is the same as the one employed in the SD literature in the low dimensional settings. The main issue here is to design a procedure that evaluates $\mathcal{K}\left(\Lambda,\mathcal{L}_{p,q},z,\mathbb{P}_{T}\right)$. The outer optimization there involves searching over low dimensional subsets of $\Lambda$. As explained in the introduction, we favor a procedure based on a greedy algorithm that approximates $\mathcal{K}\left(\Lambda,\mathcal{L}_{p,q},z,\mathbb{P}_{T}\right)$ over procedures based on penalization. We use the FSS Algorithm (Algorithm 2 in Elenberg et al.\ (2018)). Let us denote $r_{T}\left(q\right)$ the number of iterations performed.
The last step (d), for $i=r_{T}\left(q\right)$, returns $\mathcal{K}^{\mathrm{FS}}\left(\Lambda,\mathcal{L}_{p,q},z,\mathbb{P}_{T},r_{T}\left(q\right)\right)$, namely the numerical approximation of $\mathcal{K}\left(\Lambda,\mathcal{L}_{p,q},z,\mathbb{P}_{T}\right)$ in ((ref)) by the greedy algorithm. The next section provides details on the numerical aspects of the three optimizations appearing in ((ref)), including the implementation of FSS.
We examine the issues of consistency, rates of convergence and limiting distribution of $M\left(\Lambda,\mathcal{L}_{p,q},\mathbb{P}_{T}\right)$ given $\mathcal{K}^{\mathrm{FS}}\left(\Lambda,\mathcal{L}_{p,q},z,\mathbb{P}_{T},r_{T}\left(q\right)\right)$. Our analysis depends on the asymptotic behavior of the empirical process $\sqrt{T}D\left(z,\kappa,\lambda,\mathbb{P}_{T}-\mathbb{P}\right)$, and of the process $G_{T}\left(z,\kappa,\lambda\right):=\sqrt{T}\left[g\left(z,\lambda,\mathbb{P}_{T}\right)-g\left(z,\lambda,\mathbb{P}\right)\right]^{T}\left(\kappa-\lambda\right)$, as well as of the empirical moment process $\frac{1}{T}\sum_{t=0}^{T}\left(z_{T}-\sum_{i=0}^{\infty}\lambda_{i}X_{t}^{(i)}\right)_{+}$, where the subdifferential $g\left(z,\lambda,\mathbb{Q}\right):=\mathbb{E}_{\mathbb{Q}}\left[X\mathbb{I}\left\{ z\geq\sum_{i=0}^{\infty}\lambda_{i}X^{(i)}\right\} \right]$ and $\mathbb{E}_{\mathbb{Q}}$ denotes integration w.r.t.\ the measure $\mathbb{Q}$. Specifically, consistency is facilitated if the first and the second processes are asymptotically tight over appropriate subsets of parameters, and the third process (locally) uniformly converges to its population counterpart. This behavior depends on stationarity and mixing rates for the returns process involved as well as a stricter moment existence condition compared to Assumption (ref).
The stationarity, ergodicity and mixing rates conditions as well as the moment existence condition hold for several geometrically ergodic (finite dimensional), linear as well as GARCH type models with values in Euclidean spaces. Those are frequently employed in empirical finance with data consistent parameter restrictions; see Francq and Zakoian (2011). Using the Daniell-Kolmogorov Theorem we have that stationarity and mixing rates hold for the $\left(X_{t}^{\infty}\right)_{t\in\mathbb{Z}}$ process whenever they hold uniformly over the collection of finite dimensional parts of the process. Thereby, they hold whenever the finite dimensional parts of the process are consistent with the aforementioned models with uniform parameter restrictions.
We obtain the following limit theory; let $\ell^{\infty}\left(Z\times\Lambda^{\infty}\times\Lambda^{\infty}\right)$ denote the space of real valued bounded functions on $Z\times\Lambda^{\infty}\times\Lambda^{\infty}$ equipped with the $\sup$ norm. We use $\rightsquigarrow$ to denote weak convergence.
The condition $\frac{\ln p}{\sqrt{T}}\rightarrow 0$ that appears in the final pair of results of the theorem is somewhat stricter than the usual $\frac{\ln p}{T}\rightarrow 0$ that appears in the literature, it however facilitates standard rates and limiting Gaussianity for the empirical processes involved and thus the results that go beyond consistency. It ensures that the bracketing entropy of $\Lambda$ grows at an appropriate rate in order for tightness to hold in the limit. For the notion of the bracketing entropy numbers of a metric space, see Section 5 of Andrews (1994) and Ch.\ 2 of van der Vaart and Wellner (1996). In our context, it corresponds to the mapping that keeps track of the logarithm of the minimal number of $\delta$-brackets (w.r.t.\ the $l_{1}$ norm) of real sequences with absolutely convergent series needed to cover the particular neighborhood, for each $\delta>0$.
We furthermore use an assumption concerning a property of restricted strong convexity (see Ch.\ 9 of Wainwright (2019) for reviewing restricted strong convexity in high-dimensional statistics) for the LPMs as $p\rightarrow +\infty$. In this section, we also denote $\Lambda$ with $\Lambda_{p}$ whenever it is important to keep track of the dimension of the portfolio space. For $p\gg m\in\mathbb{N}$, we denote the set $\left\{ \left(\lambda,\lambda^{\star}\right)\in\Lambda_{p}\times\Lambda_{p}:\mathrm{csupp}\left(\lambda\right)\leq m,\mathrm{csupp}\left(\lambda^{\star}\right)\leq m,\mathrm{csupp}\left(\lambda-\lambda^{\star}\right)\leq m\right\} $ with $\Lambda_{\left(m\right)}$. $\tilde{\Lambda}_{\left(m\right)}$ denotes the set obtained by keeping the first component $\lambda$ of the pairs $\left(\lambda,\lambda^{\star}\right)$ that define the elements of $\Lambda_{\left(m\right)}$.
Due to Assumption (ref), the existence of the continuous density $f$, Theorem 1 of Savare (1996) and given the distributional derivative of $\left(x\right)_{+}$ (see p.\ 1 in Savare (1996)), we obtain that $\mathbb{E}\left(z-\sum_{i=0}^{\infty}\kappa_{i}X_{0,i}\right)_{+}$ is twice differentiable and the Hessian assumes the form $\int_{\mathbb{R}^{q}}XX^{T}\delta\left(z-\kappa^{T}X\right)f\left(X\right)dX$, where $\delta$ denotes the Dirac Delta function. Using Example 27 in Estrada and Kanwal (2012), the latter equals $\int_{z=\kappa^{T}X}XX^{T}f\left(X\right)dX$. The eigenvalue restriction part of Assumption (ref) thus follows whenever the minimum eigenvalue of $V$, the second moment matrix of $X$, is dominated by $\ln T$ as $T\rightarrow\infty$, due to the Cauchy's eigenvalue interlacing theorem (see for example Hwang (2004)). It means that we can also accommodate second moment matrices that become asymptotically singular at a slowly varying rate. Controlling the asymptotic singularity of the Hessian is related to the geometric interpretation provided in Remark 4 of Elenberg et al.\ (1998). By doing so, we control the way the set function values diminish via adding individual features through a stepwise procedure compared to adding multiple features at once as in a LASSO penalization, approximating the diminishing return property of a submodular function. The analysis in Par.\ 5 of Kim and Pollard (1990) implies that the same control suffices when $f$ is continuously differentiable, which is more in line with the bounded support framework of our applications. When $V$ has a Kac-Murdock-Szego type Toeplitzian structure (Trench (1999)), where $V_{i,j}=v^{|i-j|},\:i,j=1,\dots,p$ for $v\in [0,1)$, Assumption (ref) holds trivially since the minimum eigenvalue of the matrix is then uniformly bounded in $p$ (Trench (1999), p.\ 182). Such matrices appear in zero mean normalised autoregressive progresses. In the case of the zero mean spiked identity model (Example 7.18 in Wainwright (2019)), where $V=\mathbf{Id}+\mu \mathbf{1}\mathbf{1}'$ for some $\mu\in [0,1)$, the results in the aforementioned example imply that Assumption (ref) holds for every fixed value of $\mu$ as well as when $\mu$ converges to one with $T$ at a slower than logarithmic rate. It fails to hold whenever $\mu$ becomes asymptotically null with faster rates. When $X$ is not necessarily zero mean, then the variational representation of the minimum eigenvalue $\lambda_{\min}(A)$, of a p.d.\ matrix $A$, implies that Assumption (ref) holds whenever $\min\left\{\lambda_{\min}(\mathbb{E}(X)\mathbb{E}(X)'),\lambda_{\min}(\mathrm{Var}(X))\right\}\ln T\rightarrow\infty$. Due to $\Lambda$ being simplicial, no restricted smoothness conditions for the maximum eigenvalue of the Hessian are required. Hence, the assumption can accommodate without such restrictions covariance matrices for models arising in strong factor settings where the largest eigenvalue is of the same order as the number of assets.
The RSC assumption along with Assumption (ref) imply analogous RSC properties for the empirical LPMs $\frac{1}{T}\sum_{t=0}^{T}\left(z-\sum_{i=0}^{\infty}\kappa_{i}X_{t}^{(i)}\right)_{+}$ with probability converging to one (w.h.p.). It enables the use of the results of Elenberg et al.\ (2018) on statistical guarantees for the FSS Algorithm. Using the above, we first obtain the following consistency result.
Whenever Assumption (ref) holds for some $q^{\star}\in\mathbb{N}$, Theorem (ref) implies then that the mapping $q\rightarrow M^{\mathrm{FS}}\left(\Lambda,\mathcal{K}_{p,q},\mathbb{P}_{T},q\ln T\right)$ converges in probability to $q\rightarrow M\left(\Lambda_{\infty},\mathcal{K}_{\infty,q},\mathbb{P}\right)$ uniformly in $q\leq q^{\star}$.
Theorem 2 holds whether we have sparse spanning or not at the limit. We do not need to assume sparsity in the population. The statistical guarantee result of Theorem 2 is a strong advantage of the greedy algorithm over penalization methods.
We are further occupied with the determination of the rates of convergence and the distributional limit for the deviation $M^{\mathrm{FS}}\left(\Lambda,\mathcal{\mathcal{L}}_{p,q},\mathbb{P}_{T},q\left(\ln T\right)^{\epsilon}\right)-M\left(\Lambda^{\infty},\mathcal{\mathcal{L}}_{\infty,q},\mathbb{P}\right)$, that gauges the gap between $M^{\mathrm{FS}}\left(\Lambda,\mathcal{\mathcal{L}}_{p,q},\mathbb{P}_{T},q\left(\ln T\right)^{\epsilon}\right)$, which is returned by the greedy algorithm on the data, and the limit $M\left(\Lambda^{\infty},\mathcal{\mathcal{L}}_{\infty,q},\mathbb{P}\right)$. To this end, we augment $r_{T}$ to $q\left(\ln T\right)^{\epsilon}$ for some arbitrary $\epsilon>1$, in order to facilitate arguments that estimate the rate of the approximation of the infimum of $\frac{1}{T}\sum_{t=0}^{T}\left(z-\sum_{i=0}^{\infty}\kappa_{i}X_{t}^{(i)}\right)_{+}$ over the empirical solution in $\mathcal{\mathcal{L}}_{p,q}$, by the infimum of $\mathbb{E}\left[\left(z-\sum_{i=0}^{\infty}\kappa_{i}X_{t}^{(i)}\right)_{+}\right]$ over the population solution. Given the second result of Theorem (ref), we obtain standard rates and a distributional limit defined as a saddle type point of a zero mean Gaussian process, using among others the generalized Delta method applicable due to the Hadamard directional differentiability of the optimization functionals that appear in the definition of spanning (see C\'arcamo et al.\ (2020)).
In practice, for fixed $p$, and as long as $p>q\ln T$, $\epsilon$ can be chosen conveniently close to 1, so that $p>q(\ln T)^{\epsilon}$ and the method is applicable. Theorem (ref) allows for the construction of a feasible inferential procedure based on subsampling in the spirit of Linton et al.\ (2014) (see also Linton et al.\ (2005)) that approximates the asymptotic quantiles of the limit in ((ref)). To get a viable numerical strategy, we design the subsampling technique to avoid the costly numerical search of the FSS algorithm inside each subsample. To this end, let $\kappa_{z,T}$ denote the solution of $\inf_{\mathrm{csupp}\left(\kappa\right)\leq q}\frac{1}{T}\sum_{t=0}^{T}\left(z_{t}-\sum_{i=0}^{\infty}\kappa_{i}X_{t}^{(i)}\right)_{+}$ over $\mathcal{L}_{p,q}$. Denote with $\Gamma^{\star}$ the subset of $\Gamma$ that contains the triplets at which some accumulation point of $\kappa_{z,T}$ appears. Let $0<b_{T}\leq T$, and consider the subsamples from the original observations $(X_{j})_{j=t,\ldots t+b_{T}-1}$ for all $t=1,2,\ldots,T-b_{T}+1$. For $\alpha\in\left(0,1\right)$, denote with $q_{T,B_{T}}\left(1-\alpha\right)$ the $1-\alpha$ quantile of the subsample empirical distribution of $\left(\sqrt{b_{T}}\left(\sup_{Z\times\Lambda_{p}}D\left(z,\kappa_{z,T},\lambda,\mathbb{P}_{t,b_{T}}\right)-M^{\mathrm{FS}}\left(\Lambda,\mathcal{\mathcal{L}}_{p,q},\mathbb{P}_{T},q\left(\ln T\right)^{\epsilon}\right)\right)\right)_{t=1,\dots,T-b_{T}+1}$, where $\mathbb{P}_{t,b_{T}}$ denotes the empirical distribution of $(X_{j})_{j=t,\ldots t+b_{T}-1}$ and we use the same $\kappa_{z,T}$ across subsamples. Hence, we get a fast subsampling method (Hong and Scaillet (2006)) that we use in our empirics to build confidence intervals for the estimated diversification loss.
Our final result depends on a condition on the elements of $\Gamma^{\star}$ that avoids limiting degeneracies (Condition ND below). They would imply poor higher order properties for the conservative inference that we consider in Proposition (ref). We say that a triplet in $\Gamma^{\star}$ is trivial if the variance of $\mathcal{G}$ there is zero. We have triviality when the first element of the triplet is $\inf Z$. It is also the case when $\lambda$ coincides with the $\text{\ensuremath{\kappa}}$ appearing in the triplet. Then, $\lambda$ is by construction an efficient element of $\Lambda_{\infty}$ that is also $q$-sparse. Whenever the elements of $X_{p}$ are linearly independent for $p$ larger than the maximum desired value of $q$ for the analysis at hand, trivialities can occur only if SS-SSD holds. This linear independence holds for Gaussian returns for example.
When spanning does not hold, then a form of non-degeneracy of $\mathbb{P}$ suffices for ND; for any optimal $K\in\mathcal{L}_{\infty,q}$, it suffices that no random variables in $X$ exist, corresponding to coordinates outside the support of $K$, that are obtainable as linear combinations of elements of $X$ that correspond to coordinates in the support of $K$. When spanning holds, the same non-degeneracy condition suffices, as long as there are triples in $\Gamma^{\star}$ that correspond to $z\neq\inf Z$. The latter can be achieved by trimming $z$ to be greater than or equal to an arbitrarily close number above the infimum of the minima of the empirical supports of the elements of $X$. It comes at the cost of weakening the SD relation. The non-degeneracy condition can be empirically tested via augmenting the vector of the elements of $X$ that correspond to coordinates in the support of the FS solution for the optimal $K$, with a single at a time element of $X$ outside $K$ and performing rank tests for the relevant $(q+1)\times (q+1)$ empirical covariance matrices-see for example Robin and Smith (2000).
The evaluation of the subsample quantile has small computational burden since we avoid the costly sparse optimization w.r.t.\ $\kappa$ inside each subsample. Usually, $Z$ is approximated by some finite discretization and optimization w.r.t.\ $\lambda$ is performed via linearization of the SD conditions and the use of LP methods. Then, the computational cost of sparse optimization is avoided and the asymptotic results in ((ref))-((ref)) hold as long as the discretized set converges to a dense subset of $Z$.
In the special case where the problem $M\left(\Lambda_{\infty},\mathcal{\mathcal{L}}_{\infty,q},\mathbb{P}\right)$ has a unique optimizer - possible only if SS-SSD does not hold, ((ref)) implies asymptotic normality. It occurs whenever the maximal expected utility difference between an efficient element of $\Lambda_{\infty}$ and its approximate counterpart of dimension $q$ occurs at a unique Russell-Seo utility for a unique pair of efficient-approximate efficient portfolios. In such a case, we can exploit normality to obtain a result like ((ref)). A feasible normality result requires a consistent estimator for the limiting variance. It can be obtained via a subsampling methodology that does not involve subsample optimizations, as long as stricter moment conditions hold for $X_{0}$, and a non-degeneracy condition for the covariance kernel of $\mathcal{G}$ holds in some neighborhood of the optimizer.
Proposition (ref) also implies an obvious conservative Kolmogorov-Smirnov testing procedure for the null hypothesis of $q$ sparse spanning based on subsampling. The null hypothesis is rejected iff zero does not lie inside the confidence interval. Given a finite set $Q\subset \mathbb{N}^{\star}$, it is also easy to use the result above in order to test more complicated hypotheses; e.g.\ the hypothesis that sparse spanning holds for at least some $q\in Q$ would be rejected iff zero lies outside the associated confidence interval for $\max\left\{q\in Q\right\}$. An interesting extension would be the construction of a test for the null of SS-SSD spanning when $q$ is allowed to diverge with rates dominated by the logarithm of the sample size.
For $q<p$, we consider the following empirical optimization problem
The utility class interpretation of Arvanitis, Scaillet and Topaloglou (2020a,b) implies that we can represent ((ref)) in terms of expected utility as:
with $\mathcal{U}:=\left\{ u\in\mathcal{C}^{0}:u(y)=\text{\ensuremath{\int}}_{\underline{x}}^{\overline{x}}v(x)r(y;x)dx\;v\in\mathcal{V}\right\}$, $\mathcal{V}:=\left\{ v:\mathcal{X}\rightarrow\mathbb{R}_{+}:\int_{\mathcal{X}}v\left(x\right)=1\right\}$, and $ r(y;x):=(y-x)1(y\leq x),\;(x,y)\in\mathcal{X}^{2}$.
The set $\mathcal{U}$ is comprised of normalized, increasing, and concave utility functions that are constructed as convex mixtures of elementary Russell et Seo (1989) ramp functions $r(y;x),\:x\in\mathcal{X}$. This representation is used in the numerical implementation via
We approximate every element of $\mathcal{U}$ with arbitrary prescribed accuracy using a finite set of increasing and concave piecewise-linear functions in the following way:
For $N_{1},N_{2}$ integers greater than or equal to $2$, first, $\mathcal{X}$ is partitioned into $N_{1}$ equally spaced values as $\underline{x}=z_{1}<\cdots<z_{N_{1}}=\overline{x}$, where $z_{n}:=\underline{x}+\frac{n-1}{N_{1}-1}(\overline{x}-\underline{x})$, $n=1,\cdots,N_{1}$. Second, $[0,1]$ is partitioned as $0<\frac{1}{N_{2}-1}<\cdots<\frac{N_{2}-2}{N_{2}-1}<1$. Using these partitions, an approximate optimization problem is considered:
where $\qquad$ $\underline{\mathcal{U}}:=\left\{ u\in\mathcal{C}^{0}:u(y)=\sum_{n=1}^{N_{1}}v_{n}r(y;z_{n})\:v_{n}\text{\ensuremath{\in}}V\right\},$ and the set of allowable weights $V:=\left\{ v\in\left\{ 0,\frac{1}{N_{2}-1},\cdots,\frac{N_{2}-2}{N_{2}-1},1\right\} ^{N_{1}}:\sum_{n=1}^{N_{1}}v_{n}=1\right\}$.
By construction, every $u\in\underline{\mathcal{\mathcal{U}}}$ consists of at most $N_{2}$ linear line segments with endpoints at $N_{1}$ possible outcome levels. Furthermore, $\underline{\mathcal{\mathcal{U}}}\subset\mathcal{U}$, which is finite as it has $N_{3}:=\frac{1}{(N_{1}-1)!}\prod_{i=1}^{N_{1}-1}(N_{2}+i-1)$ elements and subsequently (ref) approximates (ref) from below as the partitioning scheme is refined; $(N_{1},N_{2}\rightarrow\infty)$. Then, for every $u\in\underline{\mathcal{\mathcal{U}}}$, the two embedded utility maximization problems in ((ref)) can be solved using LP. Consider $ c_{0,n}:=\sum_{m=n}^{N_{1}}\left(c_{1,m+1}-c_{1,m}\right)z_{m}$, $c_{1,n}:=\sum_{m=n}^{N_{1}}w_{m}$, and $ \mathcal{N}:=\left\{ n=1,\cdots,N_{1}:v_{n}>0\right\} \bigcup\left\{ N_{1}\right\}$. Then, for any given $u\in\underline{\mathcal{\mathcal{U}}}$, $\mathbb{\sup_{\boldsymbol{\lambda}\in\mathrm{\Lambda}}}\mathbb{E}_{\mathbb{P}_{T}}\left[u\left(X^{\mathrm{T}}\boldsymbol{\lambda}\right)\right]$ is the optimal value of the objective function of the following LP problem in canonical form: $ \max T^{-1}\sum_{t=1}^{T}y_{t}$ s.t.\ $y_{t}-c_{1,n}X_{t}^{\mathrm{T}}\boldsymbol{\lambda}\leq c_{0,n}$, $t=1,\cdots,T$, $n\in\mathcal{N}$, $\sum_{i=1}^{M}\lambda_{i}=1$, $\lambda_{i}\geq0,\:i=1,\cdots,M$, and $y_{t}$ free, $t=1,\cdots,T$. The LP problem always has a feasible solution and has $\mathcal{O}(T+N)$ variables and constraints. In the empirical application, we take $N_{1}=10$ and $N_{2}=5$. Thus, we end up with $N_{3}=\frac{1}{9!}\prod_{i=1}^{9}(4+i)=715$ distinct utility functions and $2N_{3}=1\,430$ small LP problems, which is time manageable with modern-day computer hardware and solver software. We use a desktop PC with a 3.6 GHz, 24-core Intel i7 processor, with 128 GB of RAM, using MATLAB and GAMS with the Gurobi optimization solver. We start with an empty set, and then we gradually increase the number of assets adding 1 asset at a time until we find a set $K\subset\Lambda$ with $\mathrm{csupp}\left(K\right)\leq q$ and such that $K\underset{\text{SSD}}{\succeq}\Lambda$. In each iteration, we search for the asset that increases ((ref)) the most.
The overall procedure consists of the following steps: \\ For $w=1$ to $q$:
Given the output of the last step of the procedure above, and since in the empirical applications $p$ is fixed, the optimal $q$, i.e., the one that provides the portfolio that comes closest in eliminating the empirical utility loss, can be readily estimated. To do so, and if the output of step 4 does not already imply zero optimal empirical utility loss, we may continue for $w>q $ up to $p$.
We analyze large datasets of equity returns. We investigate the performance of our strategy based on the S&P 500 index constituents, and we compare the results with the sparse mean-variance efficient portfolios (MAXSER) of Ao, Li, and Zheng (2019). We consider the period from January 1981 to December 2020, namely a total of 480 monthly return observations.
Starting with the empty set, we implement our sparse dominance methodology as described above, adding one element at a time in $r$ iterations. In each iteration, the algorithm adds to its current solution the single element decreasing the value of this solution by the most, i.e., the element with the largest marginal value with respect to the current solution. The target is to get the optimal portfolio with size $q$ that yields the minimal empirical diversification loss. Whenever the latter is zero, a sparse portfolio of support $q$ is built from a large set of assets of support $p$ which cannot be improved in terms of expected utility from the consideration of additional assets (full diversification).
In Figure (ref), we observe that the number of assets that yield zero diversification loss is 45 (upper panel)\footnote{Analogous analysis has been done for the FTSE100 constituents as well as the 49 Industry portfolios of Kenneth French. For the FTSE100, we get a subset $K$ with size $q=25$ that yields zero diversification loss, while, for the 49 Industry portfolios, the size is 13 assets.}. In the same figure (lower panel), we also observe that the MAXSER portfolio of Ao, Li, and Zheng (2019) consists of 32 assets.\footnote{The tuning parameter $\lambda$ in MAXSER is the regularisation parameter in the LASSO penalization. To determine it, we use the 10-fold cross-validation procedure, which is described in Section 1.5.1.\ of their paper.} We evaluate the diversification loss, namely the estimated expected utility loss, of the optimal MAXSER portfolio with respect to the SSD portfolio with the smallest number of stocks reaching the zero bound. In the same graph, the upper bound of a 95% as well as 90% one-sided confidence intervals (CI) corresponding to the portfolio reaching the zero bound (45 assets) are additionally reported. We build them with the fast subsampling method of Section (ref) whose validity is established in Proposition (ref). We observe that the diversification loss of the MAXSER portfolio is between the loss of the SS-SSD portfolio for $q=32$ and the 90% confidence interval.
We compare the in-sample performance of the MAXSER and the SS-SSD optimal portfolios as well as the $1/N$ (equally-weighted) portfolio with $N=p=500$. We compute the first four moments of portfolio returns (Average, Standard Deviation, Skewness and Kurtosis), as well as a number of commonly used parametric performance measures for portfolios: the Sharpe Ratio, the Downside Sharpe Ratio of Ziemba (2005), the 95% Value-at-Risk (with a positive sign for a loss), the 95% Expected Shortfall (with a positive sign for a loss), the Upside Potential and Downside Risk (UP) ratio of Sortino and van den Meer (1991), the Opportunity Cost, and the Certainty Equivalent return (CEQ).
The definition of Downside Sharpe Ratio uses the downside variance (or more precisely the downside risk) defined as $ \sigma_{P_{-}}^{2}=\frac{\sum_{t=1}^{T}(R_{t})_{-}^{2}}{T-1}, $ where $(R_{t})_{-}$ is the return of portfolio $P$ at day $t$ which is below zero (i.e., those with losses). Given that the total variance equals twice the downside variance $2\sigma_{P_{-}}^{2}$, the Downside Sharpe Ratio is given by $ \mathrm{S}_{P}=\frac{\bar{R}_{P}-\bar{R}_{f}}{\sqrt{2}\sigma_{P-}}, $ where $\bar{R}_{P}$ is the average period return of portfolio $P$ and $\bar{R_{f}}$ is the average risk free rate.
The UP ratio compares the upside potential to the shortfall risk over a specific target (benchmark): $ \textrm{UP ratio}=\frac{\frac{1}{T}\sum_{t=1}^{T}(R_{P,t}-R_{f,t})_+}{\sqrt{{\frac{1}{T}\sum_{t=1}^{T}((R_{f,t}-R_{P,t})_+)^{2}}}}, $ where $R_{P,t}$ is the realized monthly return of the portfolio $P$ for the out-of-sample period, $T$ is the number of experiments performed, and $R_{f,t}$ is the monthly return of the benchmark (the riskless asset). The numerator equals the average excess return over the benchmark reflecting the upside potential while the denominator measures the downside risk (i.e., shortfall risk over the benchmark).
Both the Downside Sharpe and UP Ratios are viewed to be more appropriate measures of performance than the typical Sharpe Ratio given the asymmetric return distribution of assets.
Moreover, we compute a certainty equivalent return (CEQ) for the three portfolios based on the exponential and power utility functions: $ \mathbb{E}_{\mathbb{P}_{T}}[u(1+R_{\mathrm{P}})]=u(1+ CEQ). $ For its calculation, exponential and power utility functions are used, consistent with second degree stochastic dominance. For the coefficient of risk aversion alternative values are employed.
Finally, the Opportunity Cost $\theta$ of Simaan (2013) is used, which is a useful measure for the economic significance of the performance difference of two portfolios. It is defined as the return that needs to be added to (or subtracted from) the MAXSER portfolio return $R_{\mathrm{MAXSER}}$, so that the investor is indifferent (in utility terms) between the the two different portfolios: $ \mathbb{E}_{\mathbb{P}_{T}}[u(1+R_{\mathrm{MAXSER}}+\theta)]=\mathbb{E}_{\mathbb{P}_{T}}[u(1+R_{\mathrm{SS-SSD}})]. $ A positive (negative) Opportunity Cost implies that the investor is better (worse) off if he invests in the SS-SSD over the MAXSER portfolio. We use the same type of definition for the Opportunity Cost for the $1/N$ portfolio. Given that this measure takes into account the entire probability distribution of asset returns, it is suitable to evaluate strategies even when the asset return distribution is not normal. Again we use exponential and power utility functions under alternative values for the coefficient of risk aversion.
Table (ref) reports the performance and risk measures of the in-sample performance of the three portfolios. They allow to finer distinguish the differences between the portfolios. We observe that the Average as well as the Standard Deviation for the SS-SSD portfolio are higher that those of the MAXSER portfolio, while the Sharpe Ratio is slightly lower. It is expected, since the Sharpe Ratio is the maximization target in the construction of the MAXSER portfolio. The Skewness is less negative and the Kurtosis is higher. Although the Sharpe Ratio of the SS-SSD portfolio is slightly lower, the Downside Sharpe Ratio as well as the UP Ratio are higher. The VaR and the Expected Shortfall (with a positive sign for a loss) are lower as expected when investors want to mitigate the impact of large losses. The SS-SSD portfolio targets and achieves a transfer of probability mass from the left to the right tail of the return distribution when compared to the MAXSER portfolio. We also observe that both the SS-SSD and the MAXSER portfolios outperform the $1/N$ portfolio in all performance and risk measures. The CEQ is higher for the SS-SSD portfolio, and the Opportunity Cost is always positive, indicating that a positive return should be added in the MAXSER or the $1/N$ portfolio to achieve the same expected return with the SS-SSD portfolio. The CEQ for the optimal SS-SSD with $q = 45$ range from 0.958% to 1.183% for the exponential utility and from 1.184% to 6.214% for the power utility. For the optimal MAXSER portfolio with $q = 32$, we get a CEQ between 0.855% and 1.060% for the exponential utility and an Opportunity Cost between 1.060% and 5.536% for the power utility. For the $1/N$ portfolio, we get a CEQ between 1.016% and 0.670% for the exponential utility and between 1.016% and 5.243% for the power utility.
In order to economically quantify the diversification loss in Figure (ref), we also report in Table (ref) the performance measures for $P(5)$ and $P(10)$ i.e., optimal portfolios $P(q)$ with cardinality constraints $q=5$ and $q=10$ assets. We observe that both portfolios exhibit significantly worse performance than the SS-SSD and MAXSER portfolios. Moreover, we get a very negative Skewness and positive Kurtosis, while the VaR and Expected Shortfall are huge. We also calculate the CEQ as well as the Opportunity Cost $\theta$: $ \mathbb{E}_{\mathbb{P}_{T}}[u(1+R_{\mathrm{P(q)}}+\theta)]=\mathbb{E}_{\mathbb{P}_{T}}[u(1+R_{\mathrm{SS-SSD}})]. $ For $q=5$, the CEQ ranges from 0.496% to 0.815% for the exponential utility and from 0.916% to 4.257% for the power utility, while the Opportunity Cost ranges from 0.042% to 0.062% for the exponential utility and from 0.051% to 0.122% for the power utility. For $q=10$, the CEQ ranges from 0.629% to 0.928% for the exponential utility and from 1.019% to 4.910% for the power utility, while the Opportunity Cost for the diversification loss ranges from 0.047% to 0.068% for the exponential utility and from 0.069% to 0.145% for the power utility. We observe differences in the CEQ between the various portfolios around 0.1%, 1%, or 1.5% depending on the chosen utility function. They correspond to around 1.2%, 12%, or 18% on an annual basis.
\begingroup {6pt}
\endgroup
Finally, Table (ref) reports the average and standard deviation of asset weights of the major Industries selected by each one of the two portfolios. We observe that both portfolios are well diversified and invest in almost the same Industries, with different overall weights. The table also exhibits the % of overlap, between assets selected by SS-SSD and MAXSER. The overlap is high, since both strategies select assets from the same industries. Moreover, Table (ref) shows the average skewness and kurtosis of the assets selected by both strategies. We observe that the SS-SSD strategy picks assets with higher skewness and kurtosis, as an attempt to increase the right tail and diminish the left one.
\begingroup {6pt}
\endgroup
\begingroup {6pt}
\endgroup
We conduct out-of-sample backtesting experiments and we evaluate the optimal SS-SSD portfolios achieving a zero diversification loss in a rolling-window scheme. We use a window width of 240 monthly return observations. A stock is excluded from the asset pool if it has missing data in the 240-month training period; the number of stocks varies over time and can be smaller than the total number of constituents of the S$\&$P 500. Each month the portfolios are constructed using the monthly returns during the prior 240 months. The clock is advanced and the realized returns of the optimal portfolios are determined from the actual returns of the various assets. The same procedure is then repeated for the next time period and the ex post realized returns over the period from 01/2001 to 12/2020 (240 months) are computed. The out-of-sample test is a real-time exercise avoiding a potential look-ahead bias and mimicking the way that a real-time investor acts in practice.
We again compare the performance of the optimal SS-SSD portfolios with that of the MAXSER portfolios of Ao, Li, and Zheng (2019). The upper panel of Figure (ref) plots the number of stocks of the optimal SS-SSD portfolios through time that eliminate the diversification loss, as well as the number of stocks of the efficient MAXSER portfolio. The lower panel plots the estimated expected loss of the optimal MAXSER corresponding to the inefficient SSD portfolios with the same number of stocks as MAXSER. The diversification loss is zero for the efficient SS-SSD portfolios corresponding to the upper panel by construction. On a rolling-window basis, the number of assets in the SS-SSD portfolios is always higher than in the MAXSER portfolios. It shrinks to around 25 assets in the crisis periods of 2008-2009 and at the beginning of the Covid-19 period. Otherwise the number of assets in the SS-SSD portfolios is stable between 30 and 35. The number of assets in the MAXSER portfolios is more volatile.
Figure (ref) illustrates the out-of-sample cumulative returns of the SS-SSD, the MAXSER, the $1/N$ and the S$\&$P 500 portfolios during the period (January 2001 to December 2020). The grey areas are the NBER recession periods. We observe that the SS-SSD optimal portfolio has a 19.3 times higher value at the end of the holding period compared to the beginning, while the MAXSER portfolio has a 17.1 times higher value. The $1/N$ portfolio follows with 14.3 higher value than at the beginning of the period. Finally, the S$\&$P 500 portfolio exhibits the worst performance, with 4.2 higher value than the initial.
Next, we compare the performance of the SS-SSD optimal portfolio with the performance of the MAXSER optimal portfolio using both nonparametric tests as well as parametric performance measures.
We use the pairwise (non-)dominance test of Anyfantaki et al.\ (2022), for a risk-adjusted comparison of the out-of-sample performance of the SS-SSD and MAXSER portfolios.
The definition for second order stochastic non-dominance is the following:
Strict second order stochastic non-dominance holds iff $\kappa$ achieves a higher expected utility for some non-decreasing and concave utility function or achieves the same expected utility for every non-decreasing and concave utility function. Equivalently, strict stochastic non-dominance holds iff $\kappa$ is strictly preferred to $\lambda$ by some risk averter, or every risk averter is indifferent between them.
We test the null hypothesis ${H'_0}$ vis-\`{a}-vis the alternative hypothesis :
For the pairwise test of the two portfolios, the test statistic is $ \xi_T = \sup_{z\in\mathcal{Z}}D\left(z,\kappa,\lambda,F\right) $. To calculate the $p$-value, we use block-boostrapping. The $p$-value is approximated by $\tilde{p}_{j}=\frac{1}{R}\sum_{r=1}^{R}\{\xi_{T,r}^{\star}>\xi_{T}\},$ where $\xi_{T}$ is the test statistic, $\xi_{T,r}^{\star}$ is the bootstrap test statistic, averaging over $R=1000$ replications. We reject the null hypothesis of non-dominance if the $p$-value is lower than $5\%$. The test statistic $\xi_{T}$ is -0.0012, and the $p$-value is estimated at $4.4\%$. We thus reject the null hypothesis of non-dominance of portfolio SS-SSD over MAXSER.
We also compute a number of parametric performance measures to compare the out-of-sample performance of the optimal portfolios. Apart from the performance measures we used in the in-sample analysis, we additionally compute the Portfolio Turnover (PT), which measures the degree of rebalancing required to implement each one of the two strategies. For any portfolio strategy $P$, the portfolio turnover is defined as the average of the absolute change of weights over the $T$ rebalancing points in time and across the $N$ available assets: $ \mathrm{PT}=\frac{1}{T}\sum_{t=1}^{T}\sum_{i=1}^{N}(|w_{P,i,t+1}-w_{P,i,t}|), $ where $w_{P,i,t+1}$ and $w_{P,i,t}$ are the optimal weights of asset $i$ under strategy $P$ (SS-SSD or MAXSER) at time $t$ and $t+1$, respectively.
The performance of the portfolios is also assessed under the risk-adjusted (net of transaction costs) returns measure of DeMiguel et al.\ (2009) which is an indicator of how the proportional transaction cost generated by the portfolio turnover affects the portfolio returns. We use a transaction cost of 35 bps, which is typical in the literature. For this, the change in the net of transaction cost wealth $\mathrm{NW}_{P}$ of portfolio $P$ through time is defined as $ \mathrm{NW}_{P,t+1}=\mathrm{NW}_{P,t}(1+R_{P,t+1})[1-\mathrm{trc}\times\sum_{i=1}^{N}(|w_{P,i,t+1}-w_{P,i,t}|), $ where $\mathrm{trc}$ is the proportional transaction cost and $R_{P,t+1}$ is the realized return of portfolio $P$ at time $t+1$. Then, the portfolio return net of transaction costs is defined as $ \mathrm{RTC}_{P,t+1}=\frac{\mathrm{NW}_{P,t+1}}{\mathrm{NW}_{P,t}}-1. $
The return-loss measures the additional return needed so that the MAXSER optimal portfolio performs equally well with the SS-SSD portfolio is defined as $ \mathrm{R}_{Loss}=\frac{\mu_{\mathrm{SSD}}}{\sigma_{\mathrm{SSD}}}\times\sigma_{\mathrm{MAXSER}}-\mu_{\mathrm{MAXSER}}, $ where $\mu_{\mathrm{MAXSER}}$ and $\mu_{\mathrm{SSD}}$ are the out-of-sample mean of monthly $\mathrm{RTC}$ for the MAXSER and the SS-SSD opportunity set, and $\sigma_{\mathrm{MAXSER}}$ and $\sigma_{\mathrm{SSD}}$ are the corresponding standard deviations.
\begingroup {6pt}
\endgroup Table (ref) reports the parametric performance measures (monthly) for the MAXSER, the SS-SSD optimal portfolios and the $1/N$ portfolio for the sample period. The higher the value of each one of these measures, the greater the investment opportunities for the relative portfolio.
We observe that the Average, the Sharpe Ratios and the Downside Sharpe Ratios of the SS-SSD optimal portfolios are higher than those of the MAXSER optimal portfolios. It reflects an increase in the risk-adjusted performance (i.e., an increase in the expected return per unit of risk) and hence expands the investment opportunities for risk-averse investors. The same is true for the UP Ratio. The Value-at-Risk and the Expected Shortfall (with a positive sign for a loss) of the SS-SSD portfolios are lower, indicating lower downside losses. Furthermore, the MAXSER portfolios induce slightly less portfolio turnover than the SS-SSD portfolios. The SS-SSD strategy may have more frequent rebalancing and incur higher transaction costs, but the additional performance justifies the additional cost; see Carroll et al.\ (2017). The return-loss measure that takes into account transaction costs, is positive. The CEQ of the SS-SSD optimal portfolios is the highest in all cases. Finally, the Opportunity Cost is always positive, indicating that a positive return should be added in the MAXSER or in the $1/N$ portfolio to achieve the same expected return with the SS-SSD portfolio. The $1/N$ portfolio exhibits again the worst performance, dominated by both the SS-SSD as well as the MAXSER portfolios.
Let us now analyze the composition of the SS-SSD and the MAXSER portfolios through time. Figure (ref) reports the optimal average weights of the major Industries selected by each one of the two portfolios during the out-of-sample period. We observe that both portfolios are well diversified and invest in almost the same Industries, with different overall weights.
We further analyze the characteristics of the SS-SSD and the MAXSER portfolios through time, by estimating the Jensen Alpha and Beta coefficients of the individual stocks of these portfolios, each month. Figures (ref) and (ref) exhibit the range (max, min, and quartiles) of the Jensen Alpha and Beta coefficients, estimated from the CAPM, of the individual stocks of these portfolios during the out-of-sample period. For the estimation of the Alpha and Beta coefficients, the previous 5 years of individual monthly returns have been used (60 monthly returns). We can observe that the Beta coefficients in the SS-SSD portfolios have a more defensive profile. The heterogeneity of Alpha and Beta coefficients is lower in the SS-SSD portfolios.
Finally, we investigate which factors explain the returns of the active investors with SSD preferences. To do so, we start with the classical single factor model (CAPM), and we additionally use five asset pricing models that are popular in the literature. First, we use, the Fama-French 6-factor model (2016), which is the Fama and French 3-factor model augmented by profitability (RMW - robust minus weak), investment (CMA - conservative minus aggressive), and momentum (UMD - up minus down). Second, the q-6-factor model of Hou, Xue and Zhang (2015), including the original market and size factors of Fama-French model, augmented by a profitability (ROE - return on equity) and investment factor (I/A - investment to assets). Third, the M4 4-factor model of Stambaugh and Yuan (2017) including the standard market and size factors along with two composite factors for profitability (PERF - performance) and investment (MGMT - management). Fourth, the Barillas and Shanken 6-factor model (2018), who use a Bayesian approach, suggesting the model of six factors including market, I/A, ROE, SMB, the value factor HMLm from Asness and Frazzini (2013), and UMD. Finally, the 3-factor model of Daniel, Hirshleifer, and Sun (2020) introducing behavioral-related factors such as the market factor augmented by long- and short-term mispricing factors (FIN and PEAD, respectively). The last is included to give an economic insight on behavioral influence. A brief description of the factors is given in the Appendix.
We consider linear regression models of the following form: $R_{P,t}-R_{f,t}=a +\sum_{i}b_{i}R_{i,t}+e_{t}, $ where $R_{P,t} -R_{f,t}$ is the excess return of either the SS-SSD or MAXSER optimal portfolio at period $t$, $R_{i,t}$ is the return on the $i$th factor and $e_t$ is the error term. If the exposures $b_{i}$ to the various factors capture all variation in expected returns, the intercept $a$ is zero since the factors are tradable.
Tables (ref) - (ref) report the coefficient estimates of the factor models, as well as their respective $t$-statistics and $p$-values. The results indicate that none of the factor models could explain the performance of the two strategies. In particular, a close to zero market loading indicates a market neutral exposure. The intercept $a$ is statistically different from zero in all cases.
For all factor models, we observe that the beta market is smaller than one (defensive) for both portfolios as expected. When the Fama and French 6-factor model is used, the negative sign for the SMB factor loading and positive sign for the HML factor loading correspond to an additional defensive tilt of the SS-SSD portfolio returns. Defensive strategies overweight large value stocks and underweight small growth stocks (Novy-Marx (2016)).
We also observe that the only factors that are significant for the MAXSER returns are the FIN factor of the 3-factor model of Daniel, Hirshleifer, and Sun (2020), and the MGMT factor of the Stambaugh and Yuan(2017), four-factor model. The FIN factor (long-horizon financing factor) exploits the information in manager decisions to issue or repurchase equity in response to persistent mispricing, while the MGMT, or Management factor, is the excess returns of stocks with high ranking on management-related anomalies over the return of those with low ranking. On the other hand, there is no statistically significant factor that explains the returns of the SS-SSD portfolios.
Finally, in order to understand whether the results are explained by the long-short nature of the factors, we construct long-only factors for the Fama and French 5-factor model. Table (ref) confirms that the performance of the SS-SSD and MAXSER optimal portfolios is not explained by traditional factors even if we consider their long-only legs.
The results seem to indicate that other factors drive the performance of these portfolios. We also observe that most of the factors in all factor models are significant in the case of the $1/N$ portfolio (apart from the long-only factors), with a positive significant market loading. This observation exemplifies that the two dynamic strategies are markedly different from the naive $1/N$ one.
Our new methodology designed to target sparse spanning portfolios shows that we can often limit ourselves to a subset of a large investment opportunity set without sacrificing expected utility because of under-diversification. It also reveals that a sparse mean-variance portfolio selection (MAXSER) yields under-diversification w.r.t.\ an optimal sparse spanning portfolio. This paper focuses on second-order stochastic dominance but could be modified to accommodate higher-order stochastic dominance. We could then check whether the empirical findings extend in such settings as well.
The methodology avoids the use of LASSO-type regularizations on the stochastic dominance inequalities. It does not require fine tuning regularization parameters. Its asymptotic behavior is known whether sparse spanning holds of not. Importantly, it enables the investigation of the relation between under-diversification loss and the sparsity (cardinality) constraint.
The FSS greedy algorithm technology can be felt as time consuming especially if it is employed in resampling frameworks to get suitable statistical inference. Even though our fast subsampling methodology avoids this, it could however be of interest to alleviate its associated numerical cost, and provide paths for further research. One example is the possibility of exploiting the geometric realization of $\mathcal{L}_{p,q}$ as a sub-simplex of $\Lambda$, when the latter is a simplex for large enough $p$ - see Edelsbrunner (2014). Another example concerns the investigation of the existence of suitable smooth approximations of the Russell-Seo utilities, for the subsequent use of greedy algorithms that exploit smoothness; e.g.\ the Orthogonal Matching Pursuit in Elenberg et al.\ (2018).
In relation to the cost of resampling, it could be of interest to approximate the upper tail behavior of the limiting distribution of the under-diversification loss. The latter could be related to extensions of approximations of the analogous probabilities for the supremum of Gaussian random fields via topological features of the underlying parameter space like its Euler characteristic-see for example Takemura and Kuriki (2003). Another approach, especially when testing for sparse spanning, is via the combined use of Empirical Likelihood Ratio statistics with conservative chi-squared based rejection regions formed by moment selection; see for example Arvanitis and Post (2024).
The sparse spanning methodology could be used as an alternative selection framework to identify the factors out of a large set of factors and anomalies (for example, the set of 153 factors of Jensen, Kelly and Pedersen (2023), or the set of 193 factors of Hou, Chen, and Zhang (2020, 2021) that explain the returns of funds. Chen et al.\ (2023) impose a sparsity assumption via a regularized regression approach, the adaptive LASSO estimator, from the machine learning literature, to select a model of 9 factors.