EconBase
← Back to paper

Spanning Tests for Markowitz Stochastic Dominance

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

105,566 characters · 14 sections · 0 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Spanning Tests for Markowitz Stochastic Dominance

abstractWe derive properties of the cdf of random variables defined as saddle-type points of real valued continuous stochastic processes. This facilitates the derivation of the first-order asymptotic properties of tests for stochastic spanning given some stochastic dominance relation. We define the concept of Markowitz stochastic dominance spanning, and develop an analytical representation of the spanning property. We construct a non-parametric test for spanning based on subsampling, and derive its asymptotic exactness and consistency. The spanning methodology determines whether introducing new securities or relaxing investment constraints improves the investment opportunity set of investors driven by Markowitz stochastic dominance. In an application to standard data sets of historical stock market returns, we reject market portfolio Markowitz efficiency as well as two-fund separation. Hence, we find evidence that equity management through base assets can outperform the market, for investors with Markowitz type preferences. \\ Key words and phrases: Saddle-Type Point, Markowitz Stochastic Dominance, Spanning Test, Linear and Mixed integer programming, reverse S-shaped utility.\\[3mm] JEL Classification: C12, C14, C44, C58, D81, G11. \\ Acknowledgements: We would like to thank the Editor, Co-Editors and the two referees for constructive criticism and numerous suggestions which have led to substantial improvements over the previous version. We thank the participants at the SFI research days 2018 for helpful comments.

Introduction

An essential feature of any model trying to understand asset prices or trading behavior is an assumption about investor preferences, or about how investors evaluate portfolios. The vast majority of models assume that investors evaluate portfolios according to the expected utility framework. Investors are assumed to act as non -atiable and risk averse agents, and their preferences are represented by increasing and globally concave utility functions.

Empirical evidence suggests that investors do not always act as risk averters. Instead, under certain circumstances, they behave in a much more complex fashion exhibiting characteristics of both risk loving and risk averting. They seem to evaluate wealth changes of assets w.r.t.\ benchmark cases rather than final wealth positions. They behave differently on gains and losses, and they are more sensitive to losses than to gains (loss aversion). The relevant utility function can be either concave for gains and convex for losses (S-Shaped) or convex for gains and concave for losses (reverse S-Shaped). They seem to transform the objective probability measures to subjective ones using transformations that potentially increase the probabilities of negligible (and possibly averted) events, which, in some cases, share similar analytical characteristics to the aforementioned utility functions. Examples of risk orderings that (partially) reflect such findings are the dominance rules of behavioral finance (see Friedman and Savage (1948), Baucells and Heukamp (2006), Edwards (1996), and the references therein).

Accordingly, stochastic dominance has been used over the last decades in this framework, having more generally evolved into an important concept in the fields of economics, finance and statistics/econometrics (see inter alia Kroll and Levy (1980), McFadden (1989), Levy (1992), Mosler and Scarsini (1993), and Levy (2005)), since it enables inference on the issue of optimal choice in a non-parametric setting. Several statistical tools have been developed to test whether, given some fixed notion of stochastic dominance, a probability distribution of interest (or some random element that represents it) dominates any other similar distribution in a given set, i.e., the former is super-efficient over the latter set (see Arvanitis et al.\ (2018)). Analogous procedures have been developed to test whether this distribution is not dominated by any other member of the given set, i.e., whether it is an efficient element of it (see Linton Post and Wang (2014)). We can find some illustrative examples in the application sections of Horvath, Kokoszka, and Zitikis (2006), where interest lies in distributions of income, or Post and Levy (2005), Scaillet and Topaloglou (2010), Linton, Post and Whang (2014), where interest lies in financial portfolios.

There is a large evolving literature on the first (FSD) and on the second (SSD) order stochastic dominance. We can characterize FSD via the choice under uncertainty of every non-satiable investor, while we can characterize SSD by the analogous choice of every risk averse and non-satiable investor (see Hadar and Russell (1969), Hanoch and Levy (1969), and Rothschild and Stiglitz (1970)). Higher order stochastic dominance relations impose more restrictions on the underlying utilities of the set of investors while retaining non-satiety and risk aversion. Dropping global risk aversion, Levy and Levy (2002) formulate the notions of prospect stochastic dominance (PSD) (see also Levy and Wiener (1998), Levy and Levy (2004)) and Markowitz stochastic dominance (MSD). Those notions investigate choices by investors who have S-shaped utility functions and reverse S-shaped utility functions. Arvanitis and Topaloglou (2017) accordingly develop consistent statistical tests for PSD and MSD super-efficiency.

Given a stochastic dominance relation, the concept of stochastic spanning subsumes the aforementioned notion of super-efficiency. It is an idea of Thierry Post, influenced by Mean-Variance spanning in Huberman and Kandell (1987), that was formulated in the context of second order stochastic dominance in Arvanitis et al.\ (2018). It is yet generalizable to arbitrary stochastic dominance relations. Given such a relation, and if the underlying set of efficient elements, i.e., the efficient set, is non-empty, a spanning set is simply any superset of the efficient set. As such, we can use a spanning set to provide an ”outer approximation” of the underlying efficient set, and/or, when small enough, to provide with a desirable reduction of the initial set of distributions upon which the stochastic dominance ordering is defined, and which could be complicated. In such a case, we can reduce the examination of the optimal choice problem, to a potentially easier and more parsimonious one. Both issues are of interest to financial economics since the underlying distributions often represent the return behaviour of financial assets and the dominance orderings reflect classes of investor preferences (e.g.\ for the FSD and SSD, as well as the PSD and MSD rules and their relations to classes of utility functions, see Levy and Levy (2002)). Those notions could also be of potential interest in any field of economic theory or decision science that examines optimal choice under uncertainty.

For example, if a strict subset of a universe of available assets is known to be spanning w.r.t.\ a stochastic dominance relation that reflects all preferences with some sort of combination of local risk aversion with local risk seeking behavior (see for example the MSD preorder defined in Section 3.1), any investor with such a disposition towards risk can safely restrict her choice to the spanning set. On the contrary, if it is not spanning, there must exist investors with suchlike preferences that benefit from the enlargement of the investment opportunities from the subset to the superset. This implies that stochastic spanning can be useful in extracting important properties of financial markets for investment decisions taylor made for particular shapes of utility functions.

Hence the following question naturally arises: for some fixed stochastic dominance relation, is a given set of assets spanned by a (possibly economically relevant) subset? When the two sets are not equal, spanning occurs if and only if a functional defined by a complex recursion of optimizations is zero (see for example the discussion in page 6 of Arvanitis et al.\ (2018) for the case of SSD, or Proposition (ref) below for the case of MSD). Its empirical verification is usually analytically intractable due to the dependence of the functional on the generally unknown underlying distributions and/or due to the complexity of the optimizations involved. Hence, this is not of direct practical use. However, we can design non-parametric tests of the null hypothesis of spanning given the existence of empirical information. The limit theory of tests for stochastic spanning\footnote{Spanning tests subsume as special cases tests of super-efficiency w.r.t.\ the underlying preorder. For example, procedures developed in Scaillet and Topaloglou (2010), Arvanitis and Topaloglou (2017), and Linton et al.\ (2014), can be considered as spanning tests for singleton spanning sets.} usually involves null weak limits represented as a finite recursion of optimization functionals applied on some relevant Gaussian process that could have the form of a saddle-type functional. The possibility of the existence of atoms in their distribution affects the issue of asymptotic exactness of the aforementioned tests which are usually based on resampling procedures such as bootstrap and subsampling (Linton et al.\ (2005)). In order to obtain exactness, we cannot thus rely on standard probabilistic results used in the previous work on tests of super-efficiency, due to the complexity of the aforementioned functional.

Hence, our first contribution is the theoretical study of continuity properties of the cdf of random variables defined as saddle type points of real valued stochastic processes. Section 2 of the paper sets up the probabilistic framework, and derives new properties of the law of a random variable defined by a finite number of nested optimizations on a continuous process w.r.t.\ possibly interdependent parameter spaces. Beside its usefulness for the limit theory of spanning tests developed in this paper, this result is also a non-trivial extension to results concerning suprema of other stochastic processes and can be useful in other econometric settings (see Section 2 for references and examples).

Our second contribution is the following. The results in Arvanitis et al. (2018) concern the concept of stochastic spanning w.r.t.\ the SSD relation, which essentially represents all preferences with global risk aversion, and are derived in a context of bounded support for the underlying distributions. We expect that analogous, yet possibly more complex results, on the properties of spanning sets, their representation by relevant functionals, the construction of testing procedures, and the derivation of their limit theory hold if we extend to local risk aversion and general supports. Statistical tests concerning the issue of super-efficiency w.r.t.\ stochastic dominance rules representing local attitudes towards risk have already appeared in the literature (see for example Post and Levy (2005), or Arvanitis and Topaloglou (2017)), but to our knowledge the concept of spanning has not been studied yet for such dominance relations.

Section 3 investigates the concept of stochastic spanning w.r.t.\ the MSD preorder in the context of financial portfolios formation. We define the notion and provide with an original characterization of spanning by the zero of a functional. Using the principle of analogy, we define the non-parametric test statistic, derive its limit distribution under the null hypothesis, and define a subsampling algorithm for the approximation of the asymptotic critical values. Among others, we use the new probabilistic results of Section 2 and a novel combinatorial argument, for the derivation of asymptotic exactness when the relevant limit distribution is non-degenerate and a restriction on the significance level holds. In particular, we derive consistency of the subsampling procedure. In contrast to the results in Arvanitis et al.\ (2018), we allow for unbounded supports for the return distributions, and we suppose that the relevant parameter spaces are simplicial complexes. We explain in Section 3 why those extensions are useful and how we have to modify the theoretical arguments to accommodate them.

Section 4 provides with a numerical implementation consisting of a finite set of Linear Programming (LP) and Mixed Integer Programming (MIP) problems, the latter being highly non linear optimization problems to solve.

Inspired by Arvanitis and Topaloglou (2017), who show that the market portfolio is not MSD efficient, we test in an empirical application in Section 5, whether investors with MSD preferences could beat the market through equity management, according to Markowitz preferences. We use equity portfolios as base assets. We show that the market portfolio is not Markowitz efficient, and the two-fund separation theorem does not hold for MSD investors. Thus, combinations of the market and the riskless asset do not span the portfolios created according to the MSD criterion. We also show that equity managers with MSD preferences could generate portfolios that yield 30 times higher cumulative return than the market over the last 50 years. Standard performance and risk measures show that the optimal MSD portfolios better suit the MSD investors that are risk averse for losses and risk lovers for gains. It achieves a transfer of probability mass from the left to the right tail of the return distribution when compared to the market portfolio. Its return distribution exhibits less negative skewness, less kurtosis, and less negative tail risk. Finally, using the four-factor model of Carhart (1997) and the five-factor model of Fama and French (2015), we investigate which factors explain these returns. We find that a defensive tilt explains part of the performance of the optimal MSD portfolios, while momentum and profitability do not.

In the final section, we conclude. We present the proofs of the main and the auxiliary results in the Appendix.

Probabilistic Results

Suppose that $\Lambda_{1},\Lambda_{2},\ldots,\Lambda_{s}$ are separable metric spaces, and let $\Lambda:=\prod_{i=1}^{s}\Lambda_{i}$ be equipped with the product topology. Consider the functional $\mathcal{\text{oper}}:=\mbox{opt}_{1}\circ\mbox{opt}_{2}\circ\cdots\circ\mbox{opt}_{s}$ where $\mbox{opt}_{i}=\sup\:\text{or}\:\inf$ w.r.t.\ to some non-empty compact $\Lambda_{i}^{\star}\subseteq\Lambda_{i}$, for $i=1,\ldots,s$. When $i>1$, $\Lambda_{i}^{\star}$ is allowed to depend on the elements of $\prod_{j=1}^{i-1}\Lambda_{i-j}^{\star}$.

The probabilistic framework follows closely Chapter 2 of Nualart (2006). It consists of a complete probability space $\left(\Omega,\mathcal{F},\mathbb{P}\right)$, where $\mathcal{F}$ is generated by some isonormal Gaussian process $W=\left\{ W\left(h\right),h\in H\right\} $ and $H$ is an appropriate Hilbert space. $X$ is some vector valued stochastic process on $\Lambda$ with sample paths in the space of continuous functions $\Lambda\rightarrow\mathbb{\mathbb{R}}^{q}$ equipped with the uniform metric. In many applications, $X$ is a Gaussian weak limit of some net of processes. We denote the Malliavin derivative operator (see Nualart (2006)) by $D$ and by $\mathbb{D}^{1,2}$ the completion of the family of Malliavin differentiable random variables w.r.t.\ the norm $\sqrt{\mathbb{E}\left[z^{2}+\left(Dz\right)^{2}\right]}$.

We are interested in the form of the support and the continuity properties of the cdf of the law of the random variable $\xi:=\mathcal{\text{oper}}X_{\lambda}$. The following assumption describes sufficient conditions for the aforementioned law to have a countable number of atoms while being absolutely continuous when restricted between their successive pairs. Given this, the result to be established below, allows first for the random variable at hand to be defined by saddle-type functionals,\footnote{The term ”saddle-type” is used here in a somewhat abusive manner, since commutativity between the successive optimization operations does not hold in general.} and second for discontinuities of the resulting cdf. Hence, it generalizes known results concerning the absolute continuity of the distribution of suprema of stochastic processes. For an excellent treatment of those see inter alia, Propositions 2.1.7 and 2.1.10 of Nualart (2006), and for the discontinuities related literature on the fibering method and its probabilistic applications, see Lifshits (1983).

assumptionFor the process $X$ suppose that: \begin{enumerate} • $\mathbb{E}\left[\sup_{\Lambda}\left(X_{\lambda}^{2}\right)\right]<+\infty$. • For all $\lambda\in\Lambda$, $X\left(\lambda\right)\in\mathbb{D}^{1,2}$, and the $H$ -valued process $DX$ has a continuous version and $\mathbb{E}\left[\sup_{\Lambda}\|DX_{\lambda}\|^{2}\right]<+\infty$. • For some countable $\mathcal{T}\subset\mathbb{R}$, $\mathbb{P}\left(\left\{ \xi=\tau\right\} \cap\Omega_{\tau}\right)\geq0$ holds if and only if $\tau\in\mathcal{T}$, where $\text{\ensuremath{\Omega}}_{\tau}$ denotes $\left\{ \omega\in\Omega:DX_{\lambda}\left(\omega\right)=0\mbox{ for some \ensuremath{\lambda}}\mbox{ such that }\tau=X_{\lambda}\left(\omega\right)\right\} $. \end{enumerate}

In the usual case where $X$ is zero-mean Gaussian, we can establish the first condition by strong results that imply the subexponentiality of the distribution of $\sup_{\Lambda}X_{\lambda}$, like Proposition A.2.7 of van der Vaart and Wellner (1996). Its validity follows from conditions that restrict the packing numbers of $\Lambda\times\mathbb{R}$ metrized as a totally bounded metric space by the use of the covariance function of $X$, to be polynomially bounded, something that is easily established if the $\Lambda_{i}$ are subsets of Euclidean spaces for all $i$. In the same respect, the second condition is easily established as in Nualart (2006) (see page 110). More specifically, if $K\left(\lambda_{1},\lambda_{2}\right)$ is the aforementioned covariance function, then $H$ is the closed span of $\left\{ h_{\lambda}\left(\cdot\right)=K\left(\lambda,\cdot\right),\lambda\in\Lambda\right\} $, with inner product $\left\langle h_{\lambda_{1}},h_{\lambda_{2}}\right\rangle _{H}=K\left(\lambda_{1},\lambda_{2}\right)$, whence $DX_{\lambda}=K\left(\lambda,\lambda\right)$. In this case, the previous along with dominated convergence implies the existence of $\mathbb{E}\left[\sup_{\Lambda}\|DX_{\lambda}\|^{2}\right]$. The third condition is the most difficult to establish. In the cases that we have in mind, we can derive ”outer approximations” of $\mathcal{T}$ by analogous, as well as easier to establish, properties of random variables that are stochastically dominated by $\xi$, see for example the corollary below.

We are now able to state and prove the main probabilistic result.

thmUnder Assumption (ref), the law of $\xi$ has connected support, say $\textnormal{supp}\left(\xi\right)$, that contains $\mathcal{T}$. If $\tau\in\mathcal{T}$, the cdf of the law evaluated at $\tau$ has a jump discontinuity of size at most $\mathbb{P}\left(\text{\ensuremath{\Omega}}_{\tau}\right)$. If $\tau_{1},\tau_{2}$ are successive elements of $\mathcal{T}$, the law restricted to $\left(\tau_{1},\tau_{2}\right)$ is absolutely continuous w.r.t.\ the Lebesgue measure. If $\mathcal{T}$ is bounded from below then the law restricted to $\left(-\infty,\inf\mathcal{T}\right)$ is absolutely continuous w.r.t.\ the Lebesgue measure. Dually, if $\mathcal{T}$ is bounded from above then the law restricted to $\left(\sup\mathcal{T},+\infty\right)$ is absolutely continuous w.r.t.\ the Lebesgue measure.

Theorem (ref) encompasses the standard absolute continuity results in the aforementioned literature that hold when $\text{oper}$ is a composition of suprema (or dually infima), the parameter spaces $\Lambda$ are not dependent, and $\mathbb{P}\left(\Omega_{\tau}\right)=0$, for all $\tau\in\mathcal{T}$. Furthermore, even in the special case where $\mathcal{T}$ is a singleton, the result is a generalization of Theorem 2 of Lifshits (1983) since it allows for non-Gaussianity, dependence between the domains of the optimization operators, as well as saddle-type optimizations. The following corollary focuses on this particular case and estimates the size of the potential jump discontinuity by assuming the existence of an auxiliary random variable that is stochastically dominated by $\xi$.

corSuppose that Assumption (ref) is satisfied. Furthermore, suppose that $\mathcal{T}=\left\{ c\right\} $, $\xi\geq\eta$, $\mathbb{P}$ a.s., and that $\text{supp}\left(\eta\right)=\left[c,+\infty\right)$. Then, $\text{supp}\left(\xi\right)=\left[c,+\infty\right)$, its cdf is absolutely continuous on $\left(c,+\infty\right)$, and it may have a jump discontinuity of size at most $\mathbb{P}\left(\eta=c\right)$ at $c$.

The results above, and especially the previous corollary, are useful for the derivation of the limit theory for our test of stochastic spanning (see Arvanitis et al.\ (2018) for the case of SSD based on other arguments). For a given pair of sets of probability distributions driven by sets of portfolio allocations, the null hypothesis of spanning posits that, for any distribution in the first set, there exists some in the other one that dominates it. Below, such a hypothesis is represented by a functional of the form $\sup\sup\inf$ of an appropriate set of moment conditions parameterized by such a $\Lambda$. We can obtain a test statistic through a scaled empirical version of this functional. Under the null limit theory for the test statistic, the results above are useful for the construction of an asymptotically exact decision procedure based on a resampling scheme. They do so by providing with restrictions on the asymptotic significance level that guarantee the convergence of the critical values to continuity points of the null limiting cdf. In such frameworks, $X$ is usually zero-mean Gaussian, while $\xi$ is conveniently defined as a difference between infima of $X$ defined on different regions of $\Lambda$ with given properties (see the following sections for explicit derivations of those properties in the case of MSD).

We can meet similar probabilistic structures in other econometric applications. An example concerns the null hypothesis of nesting of a given statistical model by a set identified model represented by moment inequalities. More specifically, suppose that given a random matrix $Y$, a statistical model is comprised by a set of probability distributions conditional on $Y$ and parameterized by a Euclidean parameter $\varphi\in\Phi$. A second statistical model is comprised by the set of probability distributions conditional on $Y$ that satisfy the conditional moment inequalities $\mathbb{E}\left[\mathbf{g}\left(\theta\right)\vert Y\right]\leq\mathbf{0}_{d}$, for some $\theta\in\Theta$, where $\Theta$ is again a subset of some Euclidean space, the moment function $\mathbf{g}=\left(g_{1},g_{2},\dots,g_{d}\right)$ is finite dimensional and the inequality sign is interpreted pointwisely. We are interested in testing the hypothesis that the first model is nested in the second model, i.e., that, for any $\varphi\in\Phi$, there exists some $\theta\in\Theta$ such that $\mathbb{E}_{\varphi}\left[\mathbf{g}\left(\theta\right)\vert Y\right]\leq\mathbf{0}_{d}$, where $\mathbb{E}_{\varphi}$ denotes expectation w.r.t.\ the distributions corresponding to $\varphi$. When $\Phi$ is a singleton, we obtain specification hypotheses similar to the ones in Guggenberger, Hahn and Kim (2008). Under some further conditions on the properties of $\Phi,\Theta$ and $\mathbf{g}$, the null hypothesis of nesting is equivalent to that $\sup_{\varphi\in\Phi}\sup_{j=1,2,\dots,d}\inf_{\theta\in\Theta}\mathbb{E}_{\varphi}\left[g_{j}\left(\theta\right)/Y\right]\leq0$. If sampling is available for any $\varphi\in\Phi$ in the first model (this would be trivial in the specification related to the singleton case for $\Phi$), we can form test statistics via empirical counterparts of the functional $\sup_{\varphi\in\Phi}\sup_{j=1,2,\dots,d}\inf_{\theta\in\Theta}\mathbb{E}_{\varphi}\left[g_{j}\left(\theta\right)\vert Y\right]$. Then, the results above are also useful for the construction of asymptotically exact decision procedures in such a context.

A Spanning Test for MSD

We now introduce the concept of stochastic spanning for the MSD relation. We initially provide some order theoretical characterization of the concept, and derive an analytical representation using a functional defined by recursive optimizations. We then construct a testing procedure using a scaled empirical counterpart of that functional and subsampling. We derive its first order limit theory mainly thanks to Corollary (ref).

MSD and Stochastic Spanning

Given $\left(\Omega,\mathcal{F},\mathbb{P}\right)$, suppose that $F$ denotes the cdf of some probability measure on $\mathbb{R}^{n}$ with finite first moment.\footnote{In comparison to the spanning test for the SSD relation of Arvanitis et al.\ (2018), we do not assume that the random variables have compact supports.} Let $G(z,\lambda,F)$ be $\int_{\mathbb{R}^{n}}\mathbb{I}\{\lambda^{Tr}$$u\leq z\}dF(u)$, i.e., the cdf of the linear transformation $\mathbb{R}^{n}\ni x\rightarrow\lambda^{Tr}x$ where $\lambda$ assumes its values in $\mathbb{L}$ which is a closed non-empty subset of $\mathbb{S}=\{\lambda\in\mathbb{R}_{+}^{n}:$$\boldsymbol{1}^{Tr}\lambda$$=1,\}$. Analogously, let $\mathbb{K}$ denote some distinguished subcollection of $\mathbb{L}$. In the context of financial econometrics, $F$ usually represents the joint distribution of $n$ base asset returns, and $\mathbb{S}$ the set of linear portfolios that can be constructed upon the previous.\footnote{The base assets are not restricted to be individual securities but are defined simply as the extreme points of the maximal portfolio set $\mathbb{S}$.} The parameter set $\mathbb{L}$ represents the collection of feasible portfolios formed by economic, legal, and/or other investment restrictions. We denote generic elements of $\mathbb{L}$ by $\lambda,\kappa$, etc. In order to define the concepts of MSD and subsequently of spanning, we consider \[ \mathcal{J}(z_{1},z_{2},\lambda;F):=\int_{z_{1}}^{z_{2}}G\left(u,\lambda,F\right)du. \]

defn$\kappa$ weakly Markowitz-dominates $\lambda$, denoted by $\kappa\succcurlyeq_{M}\lambda$, iff \begin{equation} \begin{array}{c} \Delta_{1}\left(z,\lambda,\kappa,F\right):=\mathcal{J}\left(-\infty,z,\kappa,F\right)-\mathcal{J}\left(-\infty,z,\lambda,F\right)\leq0,\,\forall z\in\mathbb{R}_{-},\:and\\ \Delta_{2}\left(z,\lambda,\kappa,F\right):=\mathcal{J}\left(z,+\infty,\kappa,F\right)-\mathcal{J}\left(z,+\infty,\lambda,F\right)\leq0,\,\forall z\in\mathbb{R}_{++}. \end{array} \end{equation}

The existence of the mean of the underlying distribution implies that we can allow the limits of integration above to assume extended values, hence the integral differences $\Delta_{1}$ and $\Delta_{2}$ in ((ref)) are well defined. Levy and Levy (2002) show that $\kappa\succcurlyeq_{M}\lambda$ iff the expected utility of $\kappa$ is greater than or equal to the expected utility of $\lambda$ for any utility function in the set of increasing and, concave on the negative part and convex on the positive part real functions (termed as reverse S-shaped (at zero) utility functions). Such utility functions represent preferences towards risk that are associated with risk aversion for losses and risk loving for gains. Hence, in financial economics, Markowitz-dominance is the case iff portfolio $\kappa$ is weakly prefered to portfolio $\lambda$ by every reverse S-shaped individual investor.

The uncountable system of inequalities in ((ref)) defines an order on $\mathbb{L}$. If those are satisfied as equalities, the pair $\left(\kappa,\lambda\right)$ belongs to the possibly non-trivial equivalence part of the order. Strict dominance $\kappa\succ_{M}\lambda$ corresponds to the irreflexive part of the order and it holds iff at least one of the previous inequalities holds strictly for some $z\in\mathbb{R}$, i.e., portfolio $\kappa$ is strictly prefered to portfolio $\lambda$ by some reverse S-shaped individual investor. Finally, given the possibility that $\Delta_{1}$ and/or $\Delta_{2}$ can change sign as functions of $z$, the relation is not generally total. When this is the case, we cannot compare $\kappa$ and $\lambda$ w.r.t.\ $\succcurlyeq_{M}$.

As in the Mean-Variance case, we can define the efficient set of $\mathbb{L}$ w.r.t.\ $\succcurlyeq_{M}$, as the set of maximal elements of the preorder. This means that $\kappa$ lies in the efficient set iff for any $\lambda\in\mathbb{L}$, either $\kappa\succcurlyeq_{M}\lambda$ or $\kappa$ is incomparable to $\lambda$. The efficient set has the property that, for any $\lambda\in\mathbb{L}$, there exists some $\kappa$ in the former for which $\kappa\succcurlyeq_{M}\lambda$. Any superset of the efficient set has also the same property, but the efficient set is minimal (if we ignore equivalencies) w.r.t.\ this property. This observation motivates the definition of MSD spanning. This is analogous to the concept of Mean-Variance spanning introduced by Huberman and Kandel (1987), and extended to the SSD case by Arvanitis et al.\ (2018).

defn$\mathbb{K}$ Markowitz-spans $\mathbb{L}$ (say $\mathbb{K}\succcurlyeq_{M}\mathbb{L}$) iff for any $\lambda\in\mathbb{L}$, $\exists\kappa\in\mathbb{K}:\kappa\succcurlyeq_{M}\lambda$. If $\mathbb{K=\left\{ \kappa\right\} }$, $\kappa$ is termed as Markowitz super-efficient.

Spanning sets always exist since by construction $\mathbb{L}\succcurlyeq_{M}\mathbb{L}$. The efficient set minimally (ignoring equivalencies) spans $\mathbb{L}$, in the sense that any other spanning set must be a superset of it. Hence, we can view any spanning subset of $\mathbb{L}$ as an “outer approximation” of the efficient set. Due to the complexity of ((ref)) w.r.t.\ the Mean-Variance case, the mathematical properties of the efficient set are generally difficult to derive, but fortunately, they are approximable by properties of sequences of spanning sets that converge to it (see below).

Furthermore, if $\mathbb{K}\succcurlyeq_{M}\mathbb{L}$, the optimal choice of every reverse S-shaped investor function lies necessarily inside $\mathbb{K}$. Hence, if $\mathbb{K\subset\mathbb{L}}$ and spanning occurs, we can reduce the problem of optimal choice within $\mathbb{L}$ to the analogous problem within $\mathbb{K}$, and the latter is more parsiminious than the former. Dually, if $\mathbb{K}$ does not span $\mathbb{L}$, there must exist optimal choices, and thereby investment opportunities, in the increment $\mathbb{L}-\mathbb{K}$ for some MSD investors. Therefore we can motivate the interest in the verification of spanning by tractability reasons related to optimal portfolio choice, or by detection of new investment opportunities.

Super-efficiency (Arvanitis and Topaloglou (2017)) corresponds to the existence of a greatest element for $\succcurlyeq_{M}$, i.e., of a unique (excluding equivalencies) element that weakly Markowitz-dominates every element of $\mathbb{L}$. Given the complexity of ((ref)), greatest elements do not generally exist. This implies that the notion of spanning not only encompasses that of super-efficiency but it is also a property of the order that will more often hold.

The above raise the following question: given $\mathbb{\ensuremath{K}}$, a non empty subset of $\mathbb{L}$,\footnote{We do not look at the issue of the selection of $\mathbb{\ensuremath{K}}$. Here, the latter is considered as given. In some cases, we can select $\mathbb{\ensuremath{K}}$ by economically relevant information, see for example the application in Arvanitis et al. (2018) for SSD. We leave the issue of the selection of a candidate spanning set, especially when this selection is related to the approximation of the efficient set, for future research.} is $\mathbb{K}\succcurlyeq_{M}\mathbb{L}$? The following proposition provides with an analytical characterization by means of nested optimizations.

propSuppose that $\mathbb{K}$ is closed. Then $\mathbb{K}\succcurlyeq_{M}\mathbb{L}$ iff \begin{equation} \xi\left(F\right):=\max_{i=1,2}\sup_{\lambda\in\mathbb{L}}\sup_{z\in A_{i}}\inf_{\kappa\in\mathbb{K}}\Delta_{i}\left(z,\lambda,\kappa,F\right)=0, \end{equation} where $A_{1}=\mathbb{R}_{-},\:A_{2}=\mathbb{R}_{++}$. Spanning does not occur iff $\xi\left(F\right)>0$.

The case of super-efficiency is then trivially obtained.

corUnder the scope of the previous lemma, $\kappa$ is Markowitz super-efficient iff \[ \max_{i=1,2}\sup_{\lambda\in\mathbb{L}}\sup_{z\in A_{i}}\Delta_{i}\left(z,\lambda,\kappa,F\right)=0. \]

Given $\mathbb{\ensuremath{K}}$, it is generally difficult to directly use the previous proposition since $F$ is usually unknown and/or the optimizations involved are infeasible. However, given the availability of a sample containing information for $F$ and in conjunction with the principle of analogy, it provides the backbone for the construction of inferential procedures that address MSD spanning.

A Consistent Non-parametric Test

Hypotheses Structure and Test Statistic

We employ Lemma (ref) to construct a non-parametric test for MSD spanning. If $\mathbb{K}\succcurlyeq_{M}\mathbb{L}$ is chosen as the null hypothesis, the hypothesis structure takes the form:\footnote{Corollary (ref) implies that the hypotheses are in the special case of super-efficiency as in Arvanitis and Topaloglou (2017).} \[

array[array omitted — 89 chars of source]

\]

To design the decision rule, we extend our framework as follows. Consider a process $\left(Y_{t}\right)_{t\in{\mathbb{Z}}}$ taking values in $\mathbb{R}^{n}$. $Y_{t_{i}}$ denotes the $i^{th}$ element of $Y_{t}$. The sample of size $T$ is the random element $\left(Y_{t}\right)_{t=1,\ldots,T}$. In our portfolio framework, it represents the observable returns of the $n$ financial base assets. We denote the unknown cdf of $Y_{t}$ by $F$, and the empirical cdf by $F_{T}$. We consider the test statistic \[ \xi_{T}:=\xi\left(\sqrt{T}F_{T}\right)=\max_{i=1,2}\sup_{\lambda\in\mathbb{L}}\sup_{z\in A_{i}}\inf_{\kappa\in\mathbb{K}}\Delta_{i}\left(z,\lambda,\kappa,\sqrt{T}F_{T}\right), \] which is the $\sqrt{T}$-scaled empirical analog of $\xi\left(F\right)$. We can equivalently express $\xi_{T}$ as a usual scaled empirical average:

equation[equation omitted — 194 chars of source]

where \[ q_{i}\left(z,\lambda,\tau,Y_{t}\right):=

casesK\left(z,\lambda,\kappa,Y_{t}\right), & i=1\\ \left[\left(\lambda^{\prime}Y_{t}\right)_{+}-\left(\kappa^{\prime}Y_{t}\right)_{+}-v\left(z,\lambda,\kappa,Y\right)\right], & i=2

, \] with $v\left(z,\lambda,\kappa,Y\right):= K\left(z,\lambda,\kappa,Y\right)-K\left(0,\lambda,\kappa,Y\right)$, and $K\left(z,\lambda,\kappa,Y\right):= \left(z-\kappa^{\prime}Y\right)_{+}-\left(z-\lambda^{\prime}Y\right)_{+}$. This is instrumental in the numerical implementation of ((ref)) in Section 4. When $\mathbb{K}$ is a singleton, the test statistic coincides with the one used in Arvanitis and Topaloglou (2017).

Null Limit Distribution

In order to show that our testing procedure is asymptotically meaningful, we need a limit theory for $\xi_{T}$ under the null hypothesis. We derive it using the following assumption.

assumptionFor some $0<\delta$, $\mathbb{E}\left[\left\Vert Y_{0}\right\Vert ^{2+\delta}\right]<+\infty$. $\left(Y_{t}\right)_{t\in{\mathbb{Z}}}$ is $a$-mixing with mixing coefficients $a_{T}=O(T^{-a})$ for some $a>1+\frac{2}{\eta},\:0<\eta<2$, as $T\rightarrow\infty$. Furthermore, \[ \mathbb{\mathbb{\mathbb{V}=E}}\left[\left(Y_{0}-\mathbb{\mathbb{E}}Y_{0}\right)\left(Y_{0}-\mathbb{\mathbb{E}}Y_{0}\right)^{T}\right]+2\sum_{t=1}^{\infty}\mathbb{\mathbb{E}}\left[\left(Y_{0}-\mathbb{\mathbb{E}}Y_{0}\right)\left(Y_{t}-\mathbb{\mathbb{E}}Y_{t}\right)^{T}\right] \] is positive definite.

The mixing rates condition is implied by stationarity and geometric ergodicity. The latter holds for many stationary models used in the context of financial econometrics, like ARMA, GARCH-type, and stochastic volatility models (see Francq and Zakoian (2011) for several examples). The moment existence condition enables the validity of a mixing CLT. A CLT typically holds under stricter restrictions. The positive definiteness of the long run covariance matrix is for instance satisfied, if $\left(Y_{t}\right)_{t\in{\mathbb{Z}}}$ is a vector martingale difference process and the elements of $Y_{0}$ are linearly independent random variables. From the compactness of $\mathbb{L}$, the previous implies that $\sup_{\lambda\in\mathbb{L}}\int_{-\infty}^{+\infty}\sqrt{G\left(u,\lambda,F\right)\left(1-G\left(u,\lambda,F\right)\right)}du<+\infty,$ which is a uniform version of the analogous condition used in Horvath et al.\ (2006).

We establish the limit theory below via the use of the concept of Skorokhod representations along with an iterative consideration of the dual notions of epi/hypo-convergence. The result depends on the contact sets \[ \Gamma_{i}=\left\{ \lambda\in\mathbb{L},\kappa\in\mathbb{K},z\in A_{i}:\Delta_{i}\left(z,\lambda,\kappa,F\right)=0\right\} . \] For any $i$, $\Gamma_{i}$ is non empty since $\Gamma_{i}^{\star}\equiv\left\{ \left(\kappa,\kappa,z\right),\kappa\in\mathbb{K},z\in A_{i}\right\} \subseteq\Gamma_{i}$. Furthermore, if the support of $F$ is bounded, for any $\lambda\in\mathbb{L},\kappa\in\mathbb{K},\:\exists z\in A_{i}:\left(\lambda,z\right)\in\Gamma_{i}$, for all $i=1,2$,.\footnote{For example, since the support is bounded, we can cover it by some hypercube of the form $\left[z_{l},z_{u}\right]^{n}$ where we can choose $z_{l}$ as negative. Obviously, $\left(\lambda,z_{l}\right)\in\Gamma_{1},$ for any $\lambda\in\mathbb{L}$. } Hence, $\Gamma_{i}^{\star}\subset\Gamma_{i}$.

In what follows, we denote convergence in distribution by $\rightsquigarrow$.

propSuppose that $\mathbb{K}$ is closed, Assumption (ref) holds, and $\mathbf{H_{0}}$ is true. Then as $T\rightarrow\infty$, $\xi_{T}\rightsquigarrow\xi_{\infty},$ where $\xi_{\infty}:=\max_{i=1,2}\sup_{z\in A_{i}}\sup_{\lambda}\inf_{\kappa}\Delta_{i}\left(z,\lambda,\kappa,\mathcal{G}_{F}\right),\:\left(\lambda,z,\kappa\right)\in\Gamma_{i},$ and $\mathcal{G}_{F}$ is a centered Gaussian process with covariance kernel given by $\text{Cov}(\mathcal{G}_{F}(x),\mathcal{G}_{F}(y))=\sum_{t\in\mathbb{Z}}\text{Cov}\left(\mathbb{I}_{Y_{0}\leq x},\mathbb{I}_{Y_{t}\leq y}\right)$ and $\mathbb{P}$ almost surely uniformly continuous sample paths defined on $\mathbb{R}^{n}$.\footnote{See Theorem 7.3 of Rio (2013).}

The covariance kernel above, and thereby $\mathcal{G}_{F}$, are well defined due to the mixing condition and the existence of $\sup_{\lambda\in\mathbb{L}}\int_{-\infty}^{+\infty}\sqrt{G\left(u,\lambda,F\right)\left(1-G\left(u,\lambda,F\right)\right)}du$ implied by Assumption (ref) (see Remark 1 in Arvanitis and Topaloglou (2017)).

A Subsampling Based Testing Procedure: Limit Theory and Combinatorial Considerations

We cannot directly use the result in Proposition (ref) for the construction of an asymptotic decision rule since the distribution of $\xi_{\infty}$ depends on the unknown covariance kernel of $\mathcal{G}_{F}$. We can establish a feasible decision rule by the use of a resampling procedure. We consider subsampling, as in Linton et al.\ (2014)-see also Linton et al.\ (2005). This resampling is of a non-parametric nature since we do not want to specify parametric conditional distributions for the multivariate return dynamics.

lyxalgorithm*The testing procedure consists of the following steps: \begin{enumerate} • Evaluate $\xi_{T}$ at the original sample value. • For $0<b_{T}\leq T$ , generate subsample values from the original observations $(Y_{i})_{i=t,\ldots t+b_{T}-1}$ for all $t=1,2,\ldots,T-b_{T}+1$. • Evaluate the test statistic on each subsample value, obtaining $\xi_{T,b_{T},t}$ for all $t=1,2,\ldots,T-b_{T}+1$. • Approximate the cdf of the asymptotic distribution of $\xi_{T}$ by\\ $s_{T,b}(y)=\frac{1}{T-b_{T}+1}\sum_{t=1}^{T-b_{T}+1}1\left(\xi_{T,b_{T},t}\leq y\right)$ and evaluate its $1-\alpha$ quantile $q_{T,b_{T}}\left(1-\alpha\right)$. • $\text{{Reject}\:}\mathbf{H_{0}}\:\text{{iff}\:}\xi_{T}>q_{T,b_{T}}(1-\alpha).$ \end{enumerate}

We derive the first order limit theory via the use of Proposition (ref) and of relevant results from the theory of subsampling. We first make the following standard assumption in the subsampling methodology.

assumptionSuppose that $\left(b_{T}\right)$, possibly depending on $\left(Y_{t}\right)_{t=1,\ldots,T}$, satisfies \[ \mathbb{P}\left(l_{T}\leq b_{T}\leq u_{T}\right)\rightarrow1, \] where $(l_{T})$ and $(u_{T})$ are real sequences such that $1\leq l_{T}\leq u_{T}$ for all $T$, $l_{T}\rightarrow\infty$ and $\frac{u_{T}}{T}\rightarrow0$ as $T\rightarrow\infty$.

The assumption does not provide with much information on the practical choice of the subsampling rate for fixed $T$. It is designed to handle issues like asymptotic exactness and consistency. In the following section, along with the numerical implementation for $\xi_{T}$, we discuss a method of fixed $T$ correction for the algorithm above, in the spirit of Arvanitis et al. (2018), that involves the use of several subsampling rates.

Asymptotic exactness is derivable by results like Theorem 3.5.1 in Politis, Romano and Wolf (1999). The latter requires continuity of the limit cdf at the quantile corresponding to the significance level $\alpha$. Even when the distribution of $\xi_{\infty}$ is non-degenerate, it is possible that it has a cdf with a unique discontinuity at zero (see the proof of Lemma (ref) in the Appendix). If there exists a lower bound for $\xi_{\infty}$, Corollary (ref) provides with an estimate for the cdf jump size at zero. Then the use of the aforementioned theorem becomes possible by properly restricting $\alpha$. This is where the new probabilistic results of Section 2 become useful in our context. It turns out (see the proof of Lemma (ref) in the Appendix) that we can obtain such a bound in the form of a non-negative random variable defined as the difference between the suprema at $\mathbb{L}$ and $\mathbb{K}$ respectively, of a linear Gaussian process. Hence, we get the needed estimate of the jump size as the probability that the latter random variable attains the value zero.

In order to evaluate this, we essentially use some combinatorial notions that allow the estimation of the proportion of the linear functions for which their unique maximizer over $\mathbb{L}$ is a common extreme point of both the parameter spaces.

defnSuppose that $\mathbb{M},\mathbb{N}$ are simplicial complexes inside $\mathbb{S}$ and $\mathbb{M}\supseteq\mathbb{N}$. The set of effective extreme points of $\mathbb{N}$ w.r.t.\ $\mathbb{M}$ is \[ e_{\mathbb{M}}\left(\mathbb{N}\right):=\left\{ \lambda\:\text{is an extreme point of }\mathbb{N}:\exists\text{ extreme point \ensuremath{s} of \ensuremath{\mathbb{S}}:}\left\Vert \lambda-s\right\Vert \leq\inf_{\kappa\in\mathbb{M}}\left\Vert \kappa-s\right\Vert \right\} . \] Furthermore, if $\lambda\in e_{\mathbb{M}}\left(\mathbb{N}\right)$ then the set of the adjoint to $\lambda$ extreme points of $\mathbb{S}$ is \[ c\left(\lambda\right):=\left\{ s\:\text{is an extreme point of }\mathbb{S}:\left\Vert \lambda-s\right\Vert \leq\inf_{\kappa\in\mathbb{M}}\left\Vert \kappa-s\right\Vert \right\} . \]

Given the non-linear simplicial complex forms of $\mathbb{M},\mathbb{N}$, the notion of an effective extreme point essentially picks the extreme points of $\mathbb{N}$ that can be restricted to $\mathbb{M}$ maximizers of linear real functions defined on $\mathbb{S}$. Given any such extreme point, its adjoint set essentially picks up the extreme points of the incorporating simplex $\ensuremath{\mathbb{S}}$ that are closer to it than any other extreme point of $\mathbb{M}$.

defnThe $\mathbb{M}$-character of $\lambda\in e_{\mathbb{M}}\left(\mathbb{N}\right)$ w.r.t.\ $s\in c\left(\lambda\right)$ is \[ ch_{\mathbb{M}}\left(s,\lambda\right):=\#\left\{ \kappa\in e_{\mathbb{M}}\left(\mathbb{N}\right):\left\Vert \lambda-s\right\Vert =\left\Vert \kappa-s\right\Vert \right\} . \] Furthermore, the $\mathbb{M}$-character of $\mathbb{N}$ is \begin{equation} ch_{\mathbb{M}}\left(\mathbb{N}\right):=\sum_{\lambda\in e_{\mathbb{M}}\left(\mathbb{N}\right)}\sum_{s\in c\left(\lambda\right)}\frac{\left(n-ch_{\mathbb{M}}\left(s,\lambda\right)\right)!}{n!}. \end{equation}

The ratio $\frac{\left(n-ch_{\mathbb{M}}\left(s,\lambda\right)\right)!}{n!}$ counts the proportion of linear real functions with unique maximizer $s$ over $\mathbb{S}$, and unique maximizer $\lambda$ over the restricted $\mathbb{M}$. Hence, $ch_{\mathbb{M}}\left(\mathbb{N}\right)$ counts the proportion of such functions for which the maximizer over $\mathbb{M}$ is an extreme point of $\mathbb{N}.$ Suppose now that $Z$ follows a non-degenerate, zero mean, $n$-dimensional Normal distribution. The characterization ((ref)) of the $\mathbb{M}$-character of $\mathbb{N}$ allows the bounding from above (see the proof of Lemma (ref) in the Appendix) of the probability of the event $\sup_{\lambda\in\mathbb{M}}\lambda'Z=\sup_{\lambda\in\mathbb{N}}\lambda'Z$ by $ch_{\mathbb{M}}\left(\mathbb{N}\right)$, and this is directly related to the estimation of the potential jump size discontinuity of the cdf of $\xi_{\infty}$ at zero. Thereby, if we assume that $\mathbb{L}$ and $\mathbb{K}$ are simplicial complexes, and if $ch_{\mathbb{L}}\left(\mathbb{K}\right)$ is easy to evaluate, the previous definitions greatly facilitate and are key for the derivations of the asymptotic exactness of our test.

assumption$\mathbb{L}$ and $\mathbb{K}$ are simplicial complexes inside the standard simplex $\mathbb{S}=\{\lambda\in\mathbb{R}_{+}^{n}:1^{\prime}$$\lambda$$=1,\}$ and $e_{\mathbb{L}}\left(\mathbb{K}\right)\subset e_{\mathbb{L}}\left(\mathbb{L}\right)$.

The simplicial form of $\mathbb{L}$ and $\mathbb{K}$ generalizes considerably the setting of Arvanitis et al.\ (2018). There, those spaces are restricted as convex polytopes. Here, they need not be convex, and they can be disconnected, discrete, etc. This could be useful when the investment categories are constrained because of SRI screening, restrictions on foreign investment, restrictions on available type of shares, etc. This generalization allows for the establishment of the asymptotic validity and thereby the applicability of our test in more complicated scenarios. For example, suppose that $\mathbb{K}=\mathbb{K}_{1}\cup\mathbb{K}_{2}$ which are disjoint simplices. If $\mathbb{K}\succcurlyeq_{M}\mathbb{L}$, but neither $\mathbb{K}_{1}\succcurlyeq_{M}\mathbb{L}$, nor $\mathbb{K}_{2}\succcurlyeq_{M}\mathbb{L}$, then this directly implies that the efficient set is disconnected. If $e_{\mathbb{L}}\left(\mathbb{K}\right)\subset e_{\mathbb{L}}\left(\mathbb{L}\right)$ holds, Assumption (ref) holds for the pairs $\left(\mathbb{L},\mathbb{K}\right)$, $\left(\mathbb{L},\mathbb{K}_{1}\right)$, $\left(\mathbb{L},\mathbb{K}_{2}\right)$. Then, we can use the test to determine the spanning relations inside each pair and thereby determine the disconnectedness of the efficient set. If the convex polytope case is not generalized as in this paper, the determination of the spanning relation between the elements of the first pair is not feasible.

Assumption (ref) implies that $e_{\mathbb{L}}\left(\mathbb{L}\right)$ is finite. The $e_{\mathbb{L}}\left(\mathbb{K}\right)\subset e_{\mathbb{L}}\left(\mathbb{L}\right)$ part implies $ch_{\mathbb{L}}\left(\mathbb{K}\right)\leq1$. Indicative examples are the following. First, consider the trivial case where $\mathbb{K}$ is interior to $\mathbb{L}.$ Then, it is obvious that $ch_{\mathbb{L}}\left(\mathbb{K}\right)=0$. Second, consider the case where $\mathbb{L=S}$, and $e_{\mathbb{L}}\left(\mathbb{K}\right)\neq\emptyset$ and Assumption (ref) holds. Then, $ch_{\mathbb{L}}\left(\mathbb{K}\right)=\frac{\#e{}_{\mathbb{L}}\left(\mathbb{K}\right)}{n}<1.$ Finally, suppose that $n=3$, $\mathbb{L}$ is a line in the interior of the triangle, such that each boundary point of the line has a minimal distance from a unique triangle vertex and that both boundary points have the same distance from the remaining vertex. Furthermore, suppose that $\mathbb{K}$ is some half of that line. Then, $e_{\mathbb{L}}\left(\mathbb{L}\right)$ consists of both the line boundary points, and $e_{\mathbb{L}}\left(\mathbb{K}\right)$ consists of the boundary point that lies in the chosen half. If $\lambda$ is an effective extreme point in either set, the cardinality of $c\left(\lambda\right)$ equals two. Moreover, $ch_{\mathbb{L}}\left(s,\lambda\right)$ equals 1 if $s$ lies closer to $\lambda$ than to the other boundary point of the line, and equals 2 in the other case. Hence, $ch_{\mathbb{L}}\left(\mathbb{K}\right)=\frac{1}{2}$.

Given our assumptions and the new combinatorial arguments not used previously in the literature, we prove in the Appendix (see Lemma (ref)) that the probability that the aforementioned bounding random variable attains the value zero is less than or equal to $ch_{\mathbb{L}}\left(\mathbb{K}\right)$. Then, via the use of Corollary (ref), we establish that, when $\xi_{\infty}$ is non degenerate, the $1-\alpha$ quantile is a continuity point for its cdf when $\alpha<1-ch_{\mathbb{L}}\left(\mathbb{K}\right)$. Hence, we immediately obtain the following first order limit theory for the subsampling testing procedure described above via Theorem 3.5.1 in Politis, Romano and Wolf (1999).

thmSuppose that Assumptions (ref), (ref) and (ref) hold. For the testing procedure described in Algorithm (ref), we have that
enumerate• If $\mathbf{H_{0}}$ is true and $\xi_{\infty}$ is constant, then, \[ \lim_{T\rightarrow\infty}\mathbb{P}\left(\xi_{T}>q_{T,b_{T}}\left(1-\alpha\right)\right)=0. \] • If $\mathbf{H_{0}}$ is true, $\xi_{\infty}$ is non-constant, and $\alpha<1-ch_{\mathbb{L}}\left(\mathbb{K}\right)$, then, \[ \lim_{T\rightarrow\infty}\mathbb{P}\left(\xi_{T}>q_{T,b_{T}}\left(1-\alpha\right)\right)=\alpha. \] • If $\mathbf{H_{a}}$ is true, then, \[ \lim_{T\rightarrow\infty}\mathbb{P}\left(\xi_{T}>q_{T,b_{T}}\left(1-\alpha\right)\right)=1. \]

When the distribution of $\xi_{\infty}$ is degenerate, the procedure is asymptotically conservative even if the restriction $\alpha<1-ch_{\mathbb{L}}\left(\mathbb{K}\right)$ does not hold. This is reminiscent of the results in Linton et al.\ (2005) concerning testing procedures for superefficiency w.r.t.\ several stochastic dominance relations. The non-degeneracy of the aforementioned limit distribution is not easy to establish except for cases such as the one about bounded supports which was discussed above.

When the distribution of $\xi_{\infty}$ is non-degenerate, the procedure is asymptotically exact if the restriction $\alpha<1-ch_{\mathbb{L}}\left(\mathbb{K}\right)$ holds. The restriction on the significance level is non-binding in usual applications. For example, when $\mathbb{L}=\mathbb{S}$ and $\mathbb{K}$ is a singleton, i.e., when the test is applied for super-efficiency, it implies at worst that $\alpha<1/2$, something that is usually satisfied. The closer to binding the restriction becomes, the more extreme points of $\mathbb{L}=\mathbb{S}$ exist inside $\mathbb{K}$. An extreme case is when $n$ is large, $\mathbb{K}$ is finite, and contains $n-1$ extreme points. In such a case, the result leads to subsampling tests that tend to asymptotically favor the null hypothesis of spanning. We could handle that by breaking up $\mathbb{K}$ is ”smaller pieces” and iterating the testing procedure w.r.t.\ them. For example, we can apply the procedure for any subset of $\mathbb{K}$ that contains $m$ points, for $m$ sufficiently small in order to obtain a meaningful significance level. If for some subset, we cannot reject spanning, we can infer that we cannot reject spanning for the initial $\mathbb{K}$, since supersets of spanning sets are spanning sets from Definition (ref). It is also possible that the structure of the efficient set prohibits such a $\mathbb{K}$ to be a spanning set. We leave the study of such questions for future work. In any case, the testing procedure is consistent.

Under some assumptions, we can prove, using again among others the main result, that an analogous testing procedure based on block bootstrap is generally asymptotically conservative and consistent.

A Numerical Implementation and Bias Correction

We first describe a potential numerical implementation via the use of a testing procedure asymptotically equivalent to the one of Subsection (ref), and obtained by finite approximations of the $A_{i},\:i=1,2$, as well as applications of mixed integer and linear programming. For each $T$, let $A_{i}^{\left(T\right)}$ denote a finite subset of $A_{i}$ for each $i$. Then consider the test statistic defined by \[ \xi_{T}^{\star}:=\max_{i=1,2}\sup_{\lambda\in\mathbb{L}}\sup_{z\in A_{i}^{\left(T\right)}}\inf_{\kappa\in\mathbb{K}}\Delta_{i}\left(z,\lambda,\kappa,\sqrt{T}F_{T}\right), \] and modify the algorithm of Subsection (ref) by using $\xi_{T}^{\star}$ in place of $\xi_{T}$. Under the previous assumption framework if, as $T\rightarrow+\infty$, $A_{i}^{\left(T\right)}$ appropriately approximates $A_{i}$, the modified procedure has the same first order limit theory with the original one.

thmSuppose that Assumptions (ref), (ref) and (ref) hold. If, as $T\rightarrow+\infty$, $A_{i}^{\left(T\right)}$ converges to some dense subset of $A_{i}$ in Painleve-Kuratowski sense for all $i=1,2$, the results of Theorem (ref) hold also for the modified procedure.

Now, the integration by parts formula for Lebesgue-Stieljes integrals and the commutativity of suprema imply that

equation[equation omitted — 219 chars of source]

where the $q_{i}$ are defined in (ref). From the finiteness of $A_{i}^{\left(T\right)},\:i=1,2$, the non trivial parts of the optimizations involved concern the $n_{i,T}:=\sup_{\lambda\in\mathbb{L}}\inf_{\kappa\in\mathbb{K}}\frac{1}{\sqrt{T}}\sum_{t=1}^{T}q_{i}\left(z,\lambda,\kappa,Y_{t}\right)$. Furthermore, \[ n_{1,T}=\inf_{\kappa\in\mathbb{K}}\frac{1}{\sqrt{T}}\sum_{t=1}^{T}\left(z-\kappa^{\prime}Y\right)_{+}-\inf_{\lambda\in\mathbb{L}}\frac{1}{\sqrt{T}}\sum_{t=1}^{T}\left(z-\lambda^{\prime}Y\right)_{+}, \] and we can reduce each of the minimizations involved to the solution of linear programming problems.

There is a set of at most $T$ values, say ${\cal R}=\{r_{1},r_{2},...,r_{T}\}$, containing the optimal value of the variable $z$ (see Scaillet and Topaloglou (2010) for the proof). Thus, we solve smaller problems $P(r)$, $r\in{\cal R}$, in which $z$ is fixed to $r$. Now, each of the above minimization problems boils down to a linear problem. Without loss of generality, the first optimization problem is the following:

subequations\begin{eqnarray} \mathtt{\min} & & \sum_{t=1}^{T}W_{t}\nonumber \\ \mathtt{s.t.} & & W_{t}\geq r-\kappa^{\prime}Y_{t},\quad\forall t\in T\nonumber \\ & & e^{\prime}\kappa=1,\nonumber \\ & & \kappa\geq0,\nonumber \\ & & W_{t}\geq0,\quad\forall t\in T. \end{eqnarray}

Furthermore, and via the results in the first Appendix of Arvanitis and Topaloglou (2017), we have that \[ n_{2,T}=\sup_{\lambda\in\mathbb{L}}\frac{1}{\sqrt{T}}\sum_{t=1}^{T}\max\left(\lambda'Y_{t},z\right)-\sup_{\kappa\in\mathbb{K}}\frac{1}{\sqrt{T}}\sum_{t=1}^{T}\max\left(\kappa'Y_{t},z\right). \] Hence, we need to solve both optimization problems appearing above. We do so via representing them as MIP programs. Again, there is a set of $T$ values, say ${\cal R^{\prime}}=\{r_{1}^{\prime},r_{2}^{\prime},...,r_{T}^{\prime}\}$, containing the optimal value of the variable $z$ (see Arvanitis and Topaloglou (2017) for the proof). Thus, we solve smaller problems $P(r)$, $r\in{\cal R^{\prime}}$, in which $z$ is fixed to $r$. Consider without loss of generality the first optimization problem:

eqnarray[eqnarray omitted — 427 chars of source]

Hence, the computational cost of the implementation above consists of $\text{card}A_{1}$ linear programming problems, $\text{card}A_{2}$ mixed integer programming problems, and three trivial optimizations.

Secondly, and although the tests above have asymptotically correct size, it is expected that the quantile estimates $q_{T,b_{T}}(1-\alpha)$ may be biased and sensitive to the subsample size $b_{T}$ in finite samples of realistic dimensions for $n$ and $T$. To correct for small-sample bias and reduce the sensitivity to the choice of $b_{T}$, we follow Arvanitis et al. (2018). For a given significance level $\alpha$, we compute the quantiles $q_{T,b_{T}}(1-\alpha)$ for a range of values for the subsample size $b_{T}$. We then estimate the intercept and slope of the following regression line using OLS regression analysis:

equation[equation omitted — 147 chars of source]

We then estimate the bias-corrected $(1-\alpha)$-quantile as the OLS predicted value for $b_{T}=T$:

equation[equation omitted — 137 chars of source]

Since $q_{T,b_{T}}(1-\alpha)$ converges in probability to $q(\xi_{\infty},1-\alpha)$ and $(b_{T})^{-1}$ converges to zero as $T\rightarrow0$, $\hat{\gamma}_{0;T,1-\alpha}$ converges in probability to $q(\xi_{\infty},1-\alpha)$, and the asymptotic properties are not affected.

Monte Carlo Study

We now design Monte Carlo experiments to evaluate the size and power of our testing procedure in finite samples. We allow for conditional heteroskedasticity consistent with empirical findings on returns of financial data as observed in the empirical application below. The multivariate return process $\left(Y_{t}\right)_{t\in\mathbb{Z}}$ is a vector GARCH(1,1) process, which is transformed to accommodate both spanning (size) and non spanning cases (power) for $K$ given assets. Such a process permits both temporal and cross sectional dependence between the random variables stacked in the vector process.

Suppose that $(z_{t}), t\in\mathbb{Z}$, are i.i.d. with mean zero, unit variance, and $\mathbb{E}\left[\vert z_{t}\vert^{2+\epsilon}\right]<\infty$, for some $\epsilon>0$. We assume that the cdf of $z_{t}$ is strictly increasing. Furthermore, we define the components of the return process for $i=1,...,K-1$ as

eqnarray*[eqnarray* omitted — 150 chars of source]

with $\mathbb{E}\left[a_{i}z_{t}^{2}+\beta_{i}\right]^{1+\epsilon}<1$, for some $\epsilon>0$, and $\omega_{i},a_{i},\beta_{i}\in\mathbb{R}{}_{++}$, $\mu_{i}\in\mathbb{R}{}_{+}$. For asset $i=K$, we define \[ y_{K,t}=v_{1}\left(z_{t}h_{K-1,t}^{1/2}\right)_{+}+v_{2}\left(z_{t}h_{K-1,t}^{1/2}\right)_{-}, \] with $v_{1},v_{2}\in\mathbb{R}$.

Let $\mathbf{\tau=}\left(0,0,...,1,0\right)$, $\mathbf{\tau}^{\star}\mathbf{=}\left(0,0,0,...,1\right)$, and $\mathbb{L}:=\left\{ \left(\lambda,0,0,\right)^{Tr},\mathbf{\tau,\tau}^{\star}\right\}$, with $\lambda\in\mathbb{R}{}_{+}^{K-2}$ and $1^{Tr}\lambda=1$. Using this portfolio space, we obtain the following result on Markowitz-spanning.

propIf $\mu_{i}=0$ for $i=1,...,K-1$, $\left\vert v_{1}\right\vert >\sqrt{\frac{\max\left\{ \omega_{i},a_{i},\beta_{i}\text{, }i=1,...,K-1\right\} }{\min\left\{ \omega_{i},a_{i},\beta_{i}\text{, }i=1,...,K-1\right\} }}$ and $\left\vert v_{2}\right\vert <\sqrt{\frac{\min\left\{ \omega_{i},a_{i},\beta_{i}\text{, }i=1,...,K-1\right\} }{\max\left\{ \omega_{i},a_{i},\beta_{i}\text{, }i=1,...,K-1\right\} }}$, then for $M \leq K-2$, the subset $\mathbb{\mathbb{K}}:=\left\{ \left(\lambda,0,0\right)^{Tr},\tau^{\star}\right\}$ with $\lambda\in\mathbb{R}{}_{+}^{M}$ and $1^{Tr}\lambda=1$, Markowitz-spans $\mathbb{L}$, while $\mathbb{\mathbb{K}}\setminus \left\{ \tau^{\star}\right\} $ does not Markowitz-span $\mathbb{L}$.

The statement of Proposition (ref) extends Proposition 4 of Arvanitis and Topaloglou (2017) to allow for $K$ assets with any subset of $M$ spanning assets, as well as non Gaussian innovations. Its proof follows the same arguments as in Arvanitis and Topaloglou (2017), and is thus omitted. It depends on $\mathbf{\tau}^{\star}$ being a Markowitz super-efficient portfolio w.r.t.\ the portfolio space. The design of Monte Carlo experiments in a dynamic setting is not easy for our testing procedure since we need to work with stationary distributions and different assets. The properties of those distributions required to show spanning and no spanning results are often difficult to characterize.\footnote{Another example is a process with different positive means and no serial dependence such that

eqnarray*[eqnarray* omitted — 146 chars of source]

with $\mu_i>0$, $i =1,...,K-1$. Then, if $\displaystyle v_2 > \frac{\max\{\mu_i, i= 1,...,K-1\}}{\min\{\mu_i, i= 1,...,K-1\}}$ and if $\displaystyle 0 < v_1 <\frac{\min\{\mu_i, i= 1,...,K-1\}}{\max\{\mu_i, i= 1,...,K-1\}},$ the spanning results stated in Proposition (ref) also hold. We have checked in unreported simulation results that the spanning test behaves also well in such an example including the case of student innovations with infinite variance. }

We present our Monte Carlo results in Table (ref). The number of replications to compute the empirical size and power is 1000 runs. We use either a combination of 2 assets ($M=2$) plus portfolios $\mathbf{\tau}$ and $\mathbf{\tau}^{\star}$ (Panel A for $K=4$), or a combination of 10 assets ($M=10$) plus portfolios $\mathbf{\tau}$ and $\mathbf{\tau}^{\star}$ (Panel B for $K=12$). We do so to gauge the testing performance both in a small and a larger number of assets to accommodate the empirical setting where we investigate spanning with up to 10 base assets. To meet the conditions of Proposition (ref), we set the parameters of the multivariate GARCH process as $\mu_{i}=0$, for $i=1,...K-1$, while we choose $(a_i) = (0.4, 0.45, 0.5)$, $(\beta_i) = (0.5, 0.45, 0.4)$, $(\omega_{i})=(0.5,0.5,0.5)$, for $i=1,2,3$ (Panel A), and similarly $(a_i) = (0.4, 0.41..., 0.5)$, $(\beta_i) = (0.5, 0.49,..., 0.4)$, $(\omega_{i})=(0.5,...,0.5)$, for $i=1,..,11$ (Panel B). We set $v_1 = 1.5$ and $v_2 = 0.5$, so that $\left\vert v_{1}\right\vert >\sqrt{\frac{\max\left\{ \omega_{i},a_{i},\beta_{i}\text{, }i=1,...,K-1\right\} }{\min\left\{ \omega_{i},a_{i},\beta_{i}\text{, }i=1,...,K-1\right\} }}$ and $\left\vert v_{2}\right\vert <\sqrt{\frac{\min\left\{ \omega_{i},a_{i},\beta_{i}\text{, }i=1,...,K-1\right\} }{\max\left\{ \omega_{i},a_{i},\beta_{i}\text{, }i=1,...,K-1\right\} }}$. We use innovations generated by a Student distribution with 5 degrees of freedom.\footnote{Unreported simulation results for Gaussian innovations are similar.}

We use three different sample sizes. For $T=300$, we get the subsampling distribution of the test statistic for subsample sizes $b_{T}\in\{50,100,150,200\}$. We set $b_{T}\in\{100,200,300,400\}$ for $T=500$, and $b_{T}\in\{120,240,360,480\}$ for $T=1000$. We present the results using the original subsampling critical values (without bias correction) as well as the ones obtained using the bias correction method. The comparison shows that the bias correction improves a lot the inference in finite samples. The bias correction method eliminates the size distortion and delivers excellent properties under the alternative hypothesis with empirical powers above 90% for a nominal size of 5%.

In our simulations, the computational time is only marginally increasing with the number of assets, and is mainly increasing with the number of observations. For example, we have roughly 5 minutes for $T=300$ and the double for $T=500$ per run. Therefore we believe that the procedure can scale up to a couple of hundred assets.

table[table omitted — 1,684 chars of source]

Empirical Applications

In this application, $\mathbb{L}$ consists of all convex combinations of the market portfolio, the T-bill, and a set of base assets. There is no need to explicitly allow for short selling in this application, because the market portfolio has no binding short-sales restrictions; non-binding constraints do not affect the efficiency classification.

Thanks to our spanning testing procedure, we want to check whether the two-fund separation theorem holds: can all MSD investors combine the T-bill and the market portfolio to span the whole set of their efficient portfolios?

If not, there is indication that active management for MSD investors according to their preferences could outperform any combination of the market portfolio and the riskless asset. This is studied in our second empirical application.

We use as base assets either the Fama and French (FF) size and book to market portfolios, a set of momentum portfolios, a set of industry portfolios, or a set of beta or size decile portfolios as described below, along with the market portfolio and the T-bill. If the number of base assets equals $n$, $\mathbb{L}$ is essentially the union of the relevant $n-2$ subsimplex of the standard $n-1$ simplex with $\left\{ \left(0,\cdots,1\right)\right\} $, where the latter signifies the market portfolio. The base assets, aside the market portfolio and the T-bill are the following portfolios:

itemize• The 6 FF benchmark portfolios: They are constructed at the end of each June, and correspond to the intersections of 2 portfolios formed on size (market equity, ME) and 3 portfolios formed on the ratio of book equity to market equity (BE/ME). • The 10 momentum portfolios: They are constructed monthly using NYSE prior (2-12) return decile breakpoints. The portfolios include NYSE, AMEX, and NASDAQ stocks with prior return data. To be included in a portfolio for month $t$ (formed at the end of month $t-1$), a stock must have a price for the end of month $t-13$ and a good return for $t-2$. • The 10 industry portfolios: They are constructed by assigning each NYSE, AMEX, and NASDAQ stock to an industry portfolio at the end of June of year $t$ based on its four-digit SIC code at that time. The industries are defined with the goal of having a manageable number of distinct industries that cover all NYSE, AMEX, and NASDAQ stocks. • The 10 size decile portfolios: We use a standard set of ten active US stock portfolios that are formed, and annually rebalanced, based on individual stock market capitalization of equity (ME or size), each representing a decile of the cross-section of NYSE, AMEX and NASDAQ stocks in a given year. • The 10 beta decile portfolios: We use a set of ten active US stock portfolios that are formed, and annually rebalanced, based on individual stock beta, each representing a decile of the cross-section of NYSE, AMEX and NASDAQ stocks in a given year.

For each dataset, we use data on monthly returns (month-end to month-end) from January 1930 to December 2016 (1044 monthly observations) obtained from the data library on the homepage\footnote{ http://mba.turc.dartmouth.edu/pages/faculty/ken.french} of Kenneth French. The test portfolio is the Fama and French market portfolio, which is the value-weighted average of all non-financial common stocks listed on NYSE, AMEX, and Nasdaq, and covered by CRSP and COMPUSTAT.

The portfolios used as base assets are of particular interest, because a wealth of empirical research, starting with Banz (1981), Basu (1983), and Fama and French (1993, 1997), suggests that the historical return spread between small value stocks and small growth stocks defies rational explanations based on investment risk. Moreover, book-to-market based sorts are the basis for the factor model examined in Fama and French (1993). Additionally, academics and practitioners show strong interest in momentum portfolios. Empirical evidence indicates that common stocks exhibit high returns on a period of 3-12 months will overperform on subsequent periods. This momentum phenomenon is an important challenge for the concept of market efficiency. Finally, industry sorted portfolios have posed a particularly challenging feature from the perspective of systematic risk measurement (see Fama and French (1997)). Beta-sorted portfolios have been used extensively to test the Sharpe-Lintner Mossin Capital Asset Pricing Model (CAPM) (see Black, Jensen, and Scholes (1972), Blume and Friend (1973), Fama and MacBeth (1973), Reinganum (1981), and Fama and French (1992), among others). Equity portfolios have also been at the center of the empirical literature in the stochastic dominance framework, see for example Post (2003), Kuosmanen (2004), Post and Levy (2005), Scaillet and Topaloglou (2010), Post and Kopa (2013), Gonzalo and Olmo (2014), among others.

To focus on the role of preferences and beliefs, we adhere to the assumptions of a single-period, portfolio-oriented model of a competitive capital market. The model-free nature of SD tests seems an advantage in this application area, because financial economists disagree about the relevant shape of utility functions of investors and the probability distribution of stock returns.

Results of the MSD Spanning Test

Arvanitis and Topaloglou (2017) report evidence against the market portfolio being MSD efficient for data up to December 2012. We corroborate their findings in unreported results for our whole period up to December 2016 as well as two sub-periods, the first one from January 1930 to June 1975, a total of 522 monthly observations, and the second one from July 1975 to December 2016, 522 monthly observations. Thus, we find evidence that passive investment is suboptimal for investors with MSD preferences. Equity management, instead of a standard buy-and-hold strategy on the market portfolio, seems more appealing for investors with reverse S-shaped utility functions. The MSD inefficiency of the market portfolio is not affected by transformations that are increasing and convex over gains and increasing and concave over losses, i.e., reverse S-shaped transformations.

Since the market is MSD inefficient, our next research hypothesis is whether two-fund separation holds, i.e., whether all MSD investors can satisfy themselves with combining the T-bill and the market portfolio only. The test of MSD efficiency for a given portfolio developped by Arvanitis and Topaloglou (2017) cannot answer that question since their approach is limited to the simple case of a spanning test for $\mathbb{K}$ being a singleton, and not any linear combination of two assets.

For non-normal distributions, two-fund separation generally does not occur, unless one assumes that preferences are sufficiently similar across investors (see, for example, Cass and Stiglitz (1970)). Our MSD spanning test can analyze two-fund separation without assuming a particular form for the return distribution or utility functions.

We get the subsampling distribution of the test statistic for subsample size $b_{T}\in[120,240,360,480]$. Using OLS regression on the empirical quantiles $q_{T,b_{T}}(1-\alpha)$, for significance level $\alpha=0.05$, we get the estimate $q_{T}$ for the critical value. We reject the MSD spanning if the test statistic $\xi_{T}$ is higher than the regression estimate $q_{T}$. In all the considered cases, $\mathbb{L}=\mathbb{S}$, and $\alpha<\frac{3}{4}\leq1-ch_{\mathbb{L}}\left(\mathbb{K}\right)$ holds. Hence, if our assumption framework is valid, we expect that asymptotic exactness holds. We find that:

itemize• The 6 FF benchmark portfolios: The regression estimate $q_{T}=15.74$ is lower than the value of the test statistic $\xi_{T}=26.78$. • The 10 momentum portfolios: The regression estimate $q_{T}=19.42$ is lower than the value of the test statistic $\xi_{T}=41.55$. • The 10 industry portfolios: The regression estimate $q_{T}=22.46$ is lower than the value of the test statistic $\xi_{T}=31.74$. • The 10 size decile portfolios: The regression estimate $q_{T}=19.62$ is lower than the value of the test statistic $\xi_{T}=32.34$. • The 10 beta decile portfolios: The regression estimate $q_{T}=31.48$ is lower than the value of the test statistic $\xi_{T}=44.76$.

The results suggest the rejection of MSD spanning and thus of the two-fund separation theorem for MSD investors. We get similar findings (unreported results) for the two subperiods 01/1930-06/1975 and 07/1975-12/ 2016.

As a final step in this analysis, we test for two-fund separation using the Mean-Variance criterion rather than the MSD criterion. We use the same methodology as for the above prospect spanning test, but we restrict the utility functions to take a quadratic shape. We solve the embedded expected-utility optimization problems (for every given quadratic utility function) using quadratic programming. In contrast to MSD spanning, we cannot reject the Mean-Variance spanning at conventional significant levels.

The combined results of the market MSD efficiency and market MSD spanning tests suggest that combining the T-bill and market portfolio is not optimal for some MSD investors. Investors with reverse S-shaped utility functions are investors that could outperform the market by staying away from a buy-and-hold strategy on the market. Active investors often take concentrated positions in assets with high upside potential or follow dynamic strategies like momentum. They can also prefer looking at defensive strategies. That can produce opportunities with positively skewed returns, or at least less negatively skewed, which are attractive for MSD investors.

Performance Summary of the MSD portfolios

The rejection of the spanning hypothesis implies that there exists at least one portfolio in $\mathbb{L}$ which is weakly prefered to every portfolio in $\mathbb{K}$ by at least one reverse S-shaped utility function (see Definition (ref)). Such a portfolio is by construction efficient w.r.t.\ $\mathbb{K}$ (see Definition 2.1 in Linton et al.\ (2014) for the SSD case which can be easily generalized to our MSD case). The empirical version of such a portfolio is the optimal portfolio $\lambda$ that maximizes $\xi_{T}$ for the particular sample value. In what follows, and given this characterization, we analyze the performance of such empirically optimal MSD portfolios through time, compared to the performance of the market portfolio (buy-and-hold strategy).

We resort to backtesting experiments on a rolling window basis. The rolling horizon computations cover the 642-month period from 07/1963 to 12/2016. At each month, we use the data from the previous 30 years (360 monthly observations) to calibrate the procedure. We solve the resulting optimization model for the MSD spanning test and record the optimal portfolio made of the base assets as well as the market portfolio and the T-bill. We determine the realized return of the chosen MSD optimal portfolio from the actual returns of the asset weight allocation picked by the optimizer for that month. Then, we repeat the same procedure for the next one-month rolling window and compute the ex-post realized returns for the period from 07/1963 to 12/2016. Therefore, the MSD optimal portfolios are outcomes of the testing procedure based on an unconditional distribution updated for each rolling window and performance is realized out of the optimization sample (no look-ahead bias).

Let us first compute the cumulative performance of the MSD optimal portfolios as well as the market portfolio for the entire sample period from July 1963 to December 2016 based on the optimal portfolio weights obtained for each one-month rolling window. The value for the MSD optimal portfolios is 426 times higher at the end of the holding period compared to the initial value, while the market portfolio is only 13.9 times higher. Hence, the relative performance of MSD type investors is 30 times higher than the performance of the market in the evaluated period. Such an increase of 3000% is significant at any significance level (unreported results).

To get further insights of the differences between two investment strategies, we report the first four moments of the realized returns and the Value-at-Risk in Table 2. We further compute a number of commonly used performance measures: the Sharpe ratio, the downside Sharpe ratio, the return loss and the opportunity cost.

The downside Sharpe ratio based on the semi-variance (Ziemba (2005)) is considered to be a more appropriate measure of performance than the typical Sharpe ratio given the asymmetric return distribution of the assets.

To account for transaction costs, we use the proposal of DeMiguel et al. (2009). This indicates the way that the proportional transaction costs, generated by the portfolio turnover, affect the portfolio returns. Let $trc$ be the proportional transaction cost, and $R_{P,t+1}$ the realized return of portfolio $P$ at time $t+1$. The change in the net of transaction cost wealth $NW_{P}$ of portfolio $P$ through time is,

equation[equation omitted — 100 chars of source]

The portfolio return, net of transaction costs is defined as

equation[equation omitted — 70 chars of source]

Let $\mu_{M}$ and $\mu_{MSD}$ be the out-of-sample mean of ((ref)) for the market portfolio and the MSD optimal portfolios, and $\sigma_{M}$ and $\sigma_{MSD}$ be the corresponding standard deviations. Then, the return-loss measure is,

equation[equation omitted — 80 chars of source]

i.e., the additional return needed so that the market performs equally well with the MSD optimal portfolios. We follow the literature and use 35 bps for the transaction costs of stocks and bonds.

Finally, the opportunity cost presented in Simaan (2013) gauges the economic significance of the performance difference between two portfolios. Let $R_{MSD}$ and $R_{M}$ be the realized returns of the MSD optimal portfolios and the market portfolio, respectively. Then, the opportunity cost $\theta$ is defined as the return that needs to be added to (or subtracted from) the market return $R_{M}$, so that the investor is indifferent (in utility terms) between the strategies imposed by the two different investment opportunity sets, i.e.,

equation[equation omitted — 53 chars of source]

A positive (negative) opportunity cost implies that the investor is better (worse) off if the investment opportunity set allows for MSD type investing. The opportunity cost takes into account the entire probability density function of asset returns and hence it is suitable to evaluate strategies even when the distribution is not normal. For the calculation of the opportunity cost, we use the following utility function which satisfies the curvature of Markowitz theory (reverse-S-shaped):

equation[equation omitted — 139 chars of source]

where $c$ is the coefficient of loss aversion (usually $c=2.25$) and $a, b > 1$. We use several values of $a,b$ in Table 2 to drive the curvature of the utility functions.

table[table omitted — 2,249 chars of source]

Table (ref) reports the performance and risk measures for the MSD optimal portfolios and the market portfolio. These measures allow us to better figure out the differences between the market portfolio and the MSD strategy. The mean is higher for the MSD optimal portfolio and the variance is lower, which results in a higher Sharpe ratio. The skewness is less negative as expected for a portfolio built for investors with preferences towards risk that are associated with risk aversion for losses and risk loving for gains. The kurtosis and VaR are lower as expected when investors want to mitigate the impact of large losses. The MSD portfolio targets and achieves a transfer of probability mass from the left to the right tail of the return distribution when compared to the market portfolio. The opportunity cost is above 70 bps and increases with the curvatures of the gain and loss parts of the utility function.

Table (ref) reports the descriptive statistics regarding the weight allocation of the MSD optimal portfolios. They load mainly on big size FF portfolios (FF portfolios), several momentum portfolios (momentum portfolios), telecommunications, health, energy and utilities (industry portfolios), small caps (size portfolios ), and low and medium beta (beta sorted portfolios), in addition to the market portfolio and the T-Bill.

table[table omitted — 1,727 chars of source]

We also investigate which factors explain the returns of the active investors with MSD preferences. To do so, we use the four-factor model of Carhart (1997) which adds momentum in the three-factor model of Fama and French (1992, 1993), as well as the Fama and French five-factor model (2015). Our empirical test examines whether these models explain the returns on MSD portfolios that dominate any combination of the market and the riskless asset, namely whether standard factors used in the empirical asset pricing literature are potential drivers of returns of MSD optimal portfolios.

First, we consider the following linear regression (Carhart four-factor model):

equation[equation omitted — 113 chars of source]

where $R_{it}$ is the return of the MSD optimal portfolio at period $t$, $R_{Ft}$ is the riskless rate, $R_{Mt}$ is the return on the value-weight (VW) market portfolio, $SMB_{t}$ is the return on a diversified portfolio of small stocks minus the return on a diversified portfolio of big stocks, $HML_{t}$ is the difference between the returns on diversified portfolios of high and low BE/ME stocks, $MOM_{t}$ is the average return on the two high prior return portfolios minus the average return on the two low prior return portfolios, and $e_{it}$ is a zero-mean residual. If the exposures $b_{i}$, $s_{i}$, $h_{i}$, and $r_{i}$ to the market, size, value, and momentum factors capture all variation in expected returns, the intercept $a_{i}$ is zero.

table[table omitted — 850 chars of source]

Table (ref) reports the coefficient estimates of the four factors, as well as their respective $t$-statistics and $p$-values. The results indicate that apart from the momentum ($MOM$), all the other three factors explain part of the performance of the optimal MSD portfolios. The intercept is not zero, which indicates that perhaps other factors drive the performance of the MSD portfolios as well.

We additionally consider the following linear regression (five-factor model):

equation[equation omitted — 126 chars of source]

where $R_{it}$ is the return of the MSD optimal portfolio at period $t$, $R_{Ft}$, $R_{Mt}$, $SMB_{t}$ and $HML_{t}$ as before, $RMW_{t}$ is the difference between the returns on diversified portfolios of stocks with robust and weak profitability, $CMA_{t}$ is the difference between the returns on diversified portfolios of the stocks of low and high investment firms, which are called conservative and aggressive, and $e_{it}$ is a zero-mean residual. If the exposures $b_{i}$, $s_{i}$, $h_{i}$, $r_{i}$, and $c_{i}$ to the market, size, value, profitability and investment factors capture all variation in expected returns, the intercept $a_{i}$ is zero.

table[table omitted — 887 chars of source]

Table (ref) reports the coefficient estimates of the five-factor model, as well as their respective $t$-statistics and $p$-values. The results indicate that, apart from the profitability ($RMW$), all the other four factors explain part of the performance of the optimal MSD portfolios. The intercept clearly not being zero indicates that other factors possibly drive the performance of the MSD portfolios as well.

In both factor models, we observe that the beta market is slightly smaller than one (defensive) for the MSD portfolios as expected. The negative sign for the SMB factor loading and positive sign for the HML factor loading correspond to an additional defensive tilt. Defensive strategies overweight large value stocks and underweight small growth stocks (see Novy-Marx (2016)).

Conclusions

We have derived properties of the cdf of a random variable defined by recursive optimizations applied on a continuous stochastic process w.r.t.\ possibly dependent parameter spaces. Those properties extend previous results and can be useful for the derivation of the limit theory of tests for stochastic spanning w.r.t.\ stochastic dominance relations.

As a theoretical application, we have defined the concept of spanning, constructed an analogous test based on subsampling, and derived the first-order limit theory and a numerical implementation for the case of the MSD relation.

We have used the non-parametric test in an empirical application, inspired by Arvanitis and Topaloglou (2017), who show that the market portfolio is not MSD efficient. The spanning test enables us to explore whether MSD equity managers could outperform the market portfolio. First, we test whether the market portfolio is MSD efficient, and then whether the two-fund separation theorem holds for investors with MSD preferences. We use as base assets either the FF size and book to market portfolios, a set of momentum portfolios, a set of industry portfolios, or a set of beta or size decile portfolios. Empirical results indicate that the market portfolio is not MSD efficient, and the two-fund separation theorem does not hold for MSD investors. Thus, the combination of the market and the riskless asset do not span the portfolios created according to the MSD criterion. Hence, there exist MSD investors that could benefit from investment opportunities that involve assets beyond portfolios constructed solely by the market portfolio and the safe asset. We verify this by showing that equity managers with MSD preferences could generate portfolios that yield 30 times higher cumulative return than the market over the last 50 years. The return distribution of the MSD optimal portfolio is less negatively skewed, less leptokurtic, and thiner left-tailed, when compared to the market portfolio. Finally, using the four-factor model of Carhart (1997) and the five-factor model of Fama and French (2015), we investigate which factors explain these returns. We find that a defensive tilt explains part of the performance of the optimal MSD portfolios, while momentum and profitability do not.

The derivations and methodology used above can also be explored for other forms of stochastic dominance relations, such as the first- or the third-order, or Prospect stochastic dominance. We leave such issues for future research.

thebibliography{10} \bibitem{AHP} Arvanitis, S., Hallam, M. S., Post, T. and N. Topaloglou. 2018. Stochastic spanning. Forthcoming in the Journal of Business and Economic Statistics (http://dx.doi.org/10.1080/07350015.2017.1391099). \bibitem{at}Arvanitis, S., and N. Topaloglou. 2017. Testing for prospect and Markowitz stochastic dominance efficiency. Journal of Econometrics 198(2), 253-270. \bibitem{bar} Banz, Rolf W. 1981. The relationship between return and market value of common stocks. Journal of Financial Economics 9(1), 3-18. \bibitem{Baucells} Baucells, M., and F.H. Heukamp. 2006. Stochastic dominance and cumulative prospect theory. Management Science 52, 1409-1423. \bibitem{black} Black, F., Jensen, M. and M. Scholes. 1972. The capital asset pricing model: some empirical tests, in M.C. Jensen (ed.). Studies in the Theory of Capital Markets, Praeger: New York, 79-124. \bibitem{blum} M.E. Blume and I. Friend. 1973. A new look at the Capital Asset Pricing Model. Journal of Finance 28(1), 19-34. \bibitem{ca} Carhart, M. 1997. On persistence in Mutual Fund Performance. Journal of Finance 52(1), 57-82. \bibitem{cs} Cass, D., and J. E. Stiglitz. 1970. The structure of investor preferences and asset returns, and separability in portfolio allocation: A contribution to the pure theory of mutual funds. Journal of Economic Theory, 2(2), 122-160. \bibitem{cort}Cortissoz, J. 2007. On the Skorokhod representation theorem. Proceedings of the American Mathematical Society 135(12), 3995-4007. \bibitem{key-1} DeMiguel, V., L. Garlappi and R. Uppal, 2009, Optimal versus naive diversification: How inefficient is the 1/n portfolio strategy?. Review of Financial Studies 22, 1915-1953. \bibitem{Edwards} Edwards, K.D. 1996. Prospect theory: A literature review, International Review of Financial Analysis 5, 18-38 \bibitem{fama3} Fama, E. and K. French. 1992. The Cross-Section of Expected Stock Returns. Journal of Finance 47(2), 427-465. \bibitem{fama} Fama, E. and K. French. 1993. Common Risk Factors in the Returns on Stocks and Bonds. Journal of Financial Economics 33, 3-56. \bibitem{fama2} Fama, E. and K. French. 1997. Industry costs of equity. Journal of Financial Economics 43, 153-193. \bibitem{fama4} Fama, E. and K. French. 2015. A five-factor asset pricing model. Journal of Financial Economics 116, 1-22. \bibitem{famaM} Fama, E.F. and J.D. MacBeth. 1973. Risk, return and equilibrium: empirical tests. The Journal of Political Economy 81, 607-636. \bibitem{fz}Francq, C., and J. M. Zakoian. 2011. GARCH models: structure, statistical inference and financial applications. John Wiley & Sons. \bibitem{friedman} Friedman, M., and L. J. Savage. 1948. The utility analysis of choices involving risk. Journal of Political Economy 56, 279-304. \bibitem{Gonzalo} Gonzalo, J. and J. Olmo. 2014. Conditional Stochastic Ddominance Tests in Dynamic Settings. International Economic Review 55(3), 819-838. \bibitem{GHK} Guggenberger, P., Hahn, J., & Kim, K. (2008). Specification testing under moment inequalities. Economics Letters, 99(2), 375-378. \bibitem{key-4} Hadar, J. and W.R. Russell. 1969. Rules for ordering uncertain prospects. American Economic Review 59, 2-34. \bibitem{key-5} Hanoch, G., and H. Levy. 1969. The efficiency analysis of choices involving risk. Review of Economic Studies 36, 335-346. \bibitem{Horvath} Horvath, L. and Kokoszka, P. and R. Zitikis. 2006. Testing for Stochastic Dominance Using the Weighted McFadden-type Statistic. Journal of Econometrics 133, 191-205. \bibitem{hub}Huberman, G. and S. Kandel. 1987. Mean-Variance Spanning. Journal of Finance 42, 873-888. \bibitem{knight}Knight, K. 1999. Epi-convergence in distribution and stochastic equi-semicontinuity. Working Paper, Department of Statistics, University of Toronto. \bibitem{kroll=00003D00003D000026Levy}Kroll, Y., and H. Levy. 1980. Stochastic Dominance Criteria: A Review and Some New Evidence, in Research in Finance, Vol. II, Greenwich: JAI Press, pp. 263-277 \bibitem{kuosman} Kuosmanen, T. (2004). Efficient diversification according to stochastic dominance criteria. Management Science 50(10), 1390-1406. \bibitem{Levyp}Levy, H. 1992. Stochastic Dominance and Expected Utility: Survey and Analysis. Management Science 38, 555-593. \bibitem{Levyb}Levy, H. 2015. Stochastic dominance: Investment decision making under uncertainty. Springer. \bibitem{LL} M. Levy and H. Levy. 2002. Prospect Theory: Much Ado about Nothing?. Management Science 48, 1334-1349. \bibitem{LL2} Levy, H., M. Levy. 2004. Prospect theory and mean-variance analysis. Review of Financial Studies 17(4), 1015-1041. \bibitem{Lif}Lifshits, M. A. 1983. On the absolute continuity of distributions of functionals of random processes. Theory of Probability and Its Applications 27(3), 600-607. \bibitem{linton1} Linton, O., Maasoumi, E. and Y.-J. Whang. 2005. Consistent Testing for Stochastic Dominance under General Sampling Schemes. Review of Economic Studies 72, 735-765. \bibitem{lpw}Linton, O., Post, T., and Whang, Y. J. 2014. Testing for the stochastic dominance efficiency of a given portfolio. The Econometrics Journal 17(2), 59-74. \bibitem{McFadden}McFadden, D. 1989. Testing for Stochastic Dominance, in Studies in the Economics of Uncertainty, eds. T. Fomby and T. Seo, New York: Springer- Verlag, pp. 113\textendash 134. \bibitem{Molch}Molchanov, I. 2006. Theory of random sets. Springer Science and Business Media. \bibitem{Mosler=00003D00003D000026Scarcini}Mosler, K., and M. Scarsini. 1993. Stochastic Orders and Applications, a Classified Bibliography, Berlin: Springer-Verlag. \bibitem{nar}Narici, L., and E. Beckenstein. 2010. Topological vector spaces. CRC Press. \bibitem{nov}Novy-Marx, R. 2016. Understanding defensive equity. Working Paper, Simon Graduate School of Business, University of Rochester and NBER. \bibitem{nual}Nualart, D. 2006. The Malliavin calculus and related topics. Berlin: Springer. \bibitem{pol}Politis, D. N., J. P. Romano and M. Wolf. 1999. Subsampling. Springer New York. \bibitem{post}Post, T. 2003. Empirical Tests for Stochastic Dominance Efficiency. The Journal of Finance 58: 1905-1931. \bibitem{pk}Post, T., and M. Kopa. 2013. General linear formulations of stochastic dominance criteria. European Journal of Operational Research 230(2), 321-332. \bibitem{pl}Post, T., and H. Levy. 2005. Does risk seeking drive stock prices? A stochastic dominance analysis of aggregate investor preferences and beliefs. Review of Financial Studies 18(3), 925-953. \bibitem{Rein} M.R. Reinganum. 1981. A new empirical perspective on the CAPM. Journal of Financial and Quantitative Analysis 16(4), 439-462. \bibitem{Rio}Rio, E. 2013. Inequalities and limit theorems for weakly dependent sequences. 3`eme cycle. pp.170. \textlesscel-00867106\textgreater. \bibitem{key-6} Rothschild, M. and J.E. Stiglitz. 1970. Increasing Risk: I. A definition. Journal of Economic Theory 2(3), 225-243. \bibitem{st}Scaillet, O., and N. Topaloglou. 2010. Testing for stochastic dominance efficiency. Journal of Business and Economic Statistics 28(1), 169-180. \bibitem{key-33}Simaan, Y., 1993, Portfolio selection and asset pricing-three-parameter framework. Management Science 39, 568-577. \bibitem{rank}Sidak, Z., P. K. Sen, and J. Hajek. 1999. Theory of rank tests. Academic Press. \bibitem{tob}Tobin, J. 1958. Liquidity Preference as Behavior Towards Risk. Review of Economic Studies 25, 65-86. \bibitem{van}van der Vaart, A. W., and J. A. Wellner. 1996. Weak Convergence. Springer New York. \bibitem{key-3} Ziemba, W., 2005, The symmetric downside risk sharpe ratio. Journal of Portfolio Management 32, 108-122.