The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
161,416 characters
Sharpe Ratio Analysis in High Dimensions: Residual-Based Nodewise Regression in Factor Models
\title{ Sharpe Ratio Analysis in High Dimensions:\\Residual-Based Nodewise Regression in Factor Models}
\author{\textsc{Mehmet Caner\thanks{
North Carolina State University, Nelson Hall, Department of Economics, NC 27695. Email: [email removed]. }}
\and \textsc {Marcelo Medeiros
\thanks{Department of Economics, Pontifical Catholic University of Rio de Janeiro - Brazil. Email: [email removed]}}
\and \textsc{Gabriel F. R. Vasconcelos
\thanks{Head of Quantitative Research, BOCOM BBM Bank. Av. Barão de Tefé, 34 - 20º e 21º floors - Rio de Janeiro - RJ, 20220-460. Email: [email removed].
We are very grateful to the co-editor, Torben Andersen, the Associate Editor and an anonymous referee for very insightful comments and suggestions which led to a much improved version of the manuscript. We thank Vanderbilt Economics Department seminar guests and the participants of the World Congress of the Econometric Society for useful comments. Finally, we are thankful for the comments by Harold Chiang, Maurizio Daniele, Anders Kock, Srini Krishnamurthy, and Michael Wolf. Medeiros acknowledges the partial financial \textnormal{sup}port from CNPq and CAPES.
}}}
\date{\today}
\maketitle
\begin{abstract}
We provide a new theory for nodewise regression when the residuals from a fitted factor model are used. We apply our results to the analysis of the consistency of Sharpe Ratio estimators when there are many assets in a portfolio. We allow for an increasing number of assets as well as time observations of the portfolio. Since the nodewise regression is not feasible due to the unknown nature of idiosyncratic errors, we provide a feasible-residual-based nodewise regression to estimate the precision matrix of errors which is consistent even when number of assets, $p$, exceeds the time span of the portfolio, $n$. In another new development, we also show that the precision matrix of returns can be estimated consistently, even with an increasing number of factors and $p>n$. We show that: (1) with $p>n$, the Sharpe Ratio estimators are consistent in global minimum-variance and mean-variance portfolios; and (2) with $p>n$, the maximum Sharpe Ratio estimator is consistent when the portfolio weights sum to one; and (3) with $p<<n$, the maximum-out-of-sample Sharpe Ratio estimator is consistent.
\end{abstract}
\section{Introduction}
One of the key issues in finance, especially in empirical asset pricing, is the trade-off between the returns and the risk of a portfolio. One important way to quantify such trade-off is via the Sharpe Ratio.
We contribute to this literature by studying the case when the number of assets, namely $p$, grows with the time span of the portfolio, $n$. To obtain the Sharpe Ratio, and also its maximum, we make use of the asset return's precision matrix. In order to get an estimate of the precision matrix for asset returns in a large portfolio, we propose that an approximate factor model governs the dynamics of excess returns. Hence, asset returns (excess returns over a risk-free asset) can be explained by an increasing but known number of factors with unknown idiosyncratic errors entering the linear relation in an additive way. One major difference with the previous literature is that, in our case, the precision matrix has to be sparse. Therefore, this is a hybrid method that combines factor models with high-dimensional econometrics.
The first step in getting the Sharpe Ratio and its maximum involves the estimation of the precision matrix of the idiosyncratic terms (errors). Estimating the such precision matrix is not an easy task, and the simple nodewise regression idea as in \cite{mein2006} is not feasible. Therefore, we provide a simple, feasible residual-based nodewise regression method to estimate the precision matrix of errors in a factor model setup even if $p>n$. This feasible residual-based nodewise regression is a new idea, and it is shown to be consistently estimating the precision matrix of the errors which is our first contribution. Next, we obtain consistent estimators to the precision matrix of asset returns, even if $p>n$, which is our second technical contribution. Although, we focus on factor models in asset pricing, our methodology can be applied to any situation where the interest is the precision matrix of the errors of a linear regression model.
Next, by using the precision matrix estimator for returns we can link our technical analysis to the financial econometrics literature. We make three contributions towards Sharpe Ratio analysis. First, we consider the Sharpe Ratios in the global minimum-variance portfolio and Markowitz mean-variance portfolio. We develop consistent estimators even if $p>n$, and both dimensions diverge. Second, we consider the rate of convergence and consistency of the maximum Sharpe Ratio when the portfolio weights are normalized to one. Recently, \cite{maller2002}, and \cite{maller2016} analyze the limit with a fixed number of assets and extend that approach to a large number of assets, but a number less than the time span of the portfolio. Their papers make a key discovery: in the case of weight constraints (summing to one), the formula for the maximum Sharpe Ratio depends on a technical term, unlike the unconstrained maximum Sharpe Ratio case. Practitioners could obtain the minimum Sharpe Ratio instead of the maximum if they are using the unconstrained formula. Our paper extends their paper by analyzing two issues. First, the case if $p>n$, with both quantities growing to infinity, and second, by handling the uncertainty created by this technical term, which we can estimate and use to obtain a new constrained and consistent maximum Sharpe Ratio. The assumption of constant loadings in the factor model is clearly a constraint for portfolio analysis over longer horizons. However, the setup where $p>n$ provides the statistical tools for us to analyze portfolios in short horizons and small samples as high-dimensional asymptotics can be seen as a good approximation for situations when $n$ is small but $p$ is large compared to $n$. Third, only in the case of $p<<n$, we obtain the consistency of our nodewise-based maximum-out-of-sample Sharpe Ratio estimate, with both $p,n$ growing to infinity and $p/n \to 0$. We also provide an analysis of the Sharpe Ratio with only portfolio weights estimated in the formula. In that way, we can see the effect of estimated portfolio on getting the optimal Sharpe Ratio. Our analysis shows this is possible when $p<n$ only.
\subsection{The Sparsity of the Precision Matrix}
There are several reasons motivating the assumption of sparsity of the precision matrix of the errors from the factor model. In technical terms, this is a convenient and widely used asymptotic tool when we want to consider high dimensional problems when $p>n$. The sparsity assumption on the precision matrix of errors gives rise to a direct way of estimating the precision matrix for the returns via Sherman-Morrison-Woodbury formula. We solve two technical issues with this assumption. First, consistent estimation of the precision matrix of returns is possible, yielding consistent estimation of the Sharpe Ratio and it's maximum, even in constrained case. Also, as far as we know, in the case of $p>n$, we do not know any other consistent estimation results for global minimum variance and Markowitz portfolios, as well as the constrained maximum Sharpe Ratio in the literature.
The sparsity assumption on the precision matrix of the errors from a factor model can be also justified in situations of interest in the empirical finance literature. First, even though we do not assume normality of the errors here, in this particular case the conditional independence of two errors given all the other errors, is represented by a zero entry in the precision matrix of errors. This is explained in p.1436-1439 of \cite{mein2006}. So, in the case of normally distributed data, sparsity can be thought as a conditional independence restriction. When the errors follow an elliptical distribution, conditional uncorrelatedness of two errors amount to a zero cell in the precision matrix as discussed in Section 2.4 of \cite{flw2018}. The authors claim that sparse precision matrix may be more useful when we estimate a network of stocks, by taking out common factors from returns and analyzing the conditional independence among idiosyncratic components (errors). Finally, there are a number of recent papers in the literature showing that after removing common factors, the covariance matrix of the errors is ``almost'' block diagonal, yielding a sparse precision matrix; see, for example, \citet{fan2016factors} and \citet{dBmMrR2018}. When the covariance matrix is block-diagonal, the precision matrix can be computed by inverting the estimated covariance matrix, which in turn can be consistently estimated by several different methods. However, even in this case, there are potential benefits of estimating the precision matrix directly as shown in our simulations and empirical exercise; see also \citet{ieee}.
\subsection{A Brief Review of the Literature and Main Takeaways}
In terms of the literature on nodewise regression and related methods, the most relevant papers are as follows. \cite{mein2006} establish the nodewise regression approach and provide an optimality result when data are normally distributed. \cite{chang2019} extend the nodewise regression method to time-series data and build confidence intervals for the elements in the precision matrix. However, the goal of \cite{chang2019} only centers on the elements of the precision matrix, and there is no connection to factor models. Furthermore, their results are based on the precision matrix of observed data and not on the residuals of a first-stage estimator. Finally, the authors do not consider the case of maximum Sharpe Ratio, and it is not clear if their results are directly applicable to financial applications. \cite{canerkock2018} establish uniform confidence intervals in the case of high-dimensional parameters in heteroskedastic setups using nodewise regression, but, as in the previous paper, there is no connection to factor models in empirical finance. \cite{caner2019} provide the variance, the risk, and the weight estimation of a portfolio via nodewise regression. They take the nodewise regression directly from \cite{mein2006} and apply it to returns. However, they assume that the precision matrix of returns is sparse. Hence, it is more restrictive and less realistic than the method we propose. We combine factor models with the sparsity of the precision matrix of errors. As a consequence, our method is much more connected to typical empirical asset pricing models. Furthermore, we do not impose any sparsity on the precision matrix of returns. \cite{caner2019} also has no proofs about the estimation of the Sharpe Ratio.
In terms of recent contributions to the literature on factor models and sparse regression, we highlight \cite{jFrMmM2021}. The authors consider the combination of factor models and sparse regression in a very general setting. More specifically, they analyze a panel data model with a factor structure and idiosyncratic terms that are sparsely related. They also provide an inference procedure designed to test hypotheses on the entries of the covariance matrix of the residuals of pre-estimated models, including principal component regressions. Our paper differs from theirs in several directions. First, \cite{jFrMmM2021} considers only the covariance matrix and not the precision matrix. Second, their approach is not based on nodewise regressions. Finally, Sharpe Ratio estimation and portfolio allocation are not considered. A seminal paper is by \cite{gagl2016}, where they analyze time-varying risk premia in large portfolios with factor models. They develop a structural model, and can tie that to factor models, and after that, they can estimate time-varying risk-premia. One of their main assumptions is that the maximum eigenvalue of covariance matrix of errors in the factor structure can diverge. Also, they assume sparsity of covariance matrix of errors and observed factors in the factor model. We also use diverging eigenvalue assumption in Assumption \ref{as7}(i) in our paper, as well as an increasing number of factors here, but with the assumption of sparsity on the precision matrix of errors. \cite{gos2019} develop a diagnostic test for omitted factors in factor models. They rely on residuals rather than errors for their tests. As clear in their analysis, working with residuals pose major difficulties. We also face the similar difficulty in our paper. Then, \cite{gos2020} analyze large conditional factor models. They analyze conditional risk premia even when the number of assets dominate the time span of the portfolio.
In a recent paper, \cite{flw2018} use sparse precision matrix estimation with hidden factors. Their approach uses a Dantzig based constrained estimator for precision matrix. The main differences are that the type of estimator depends on magnitude of coefficients in the precision matrix, with larger coefficients, and that the rate of estimation slows down considerably as seen in their equation (2.12)-result 2. Also, they assume bounded-finite $l_{\infty}$ matrix norm, which is restrictive. We allow diverging matrix $l_{\infty}$ norm. Also they do not apply their results to Sharpe Ratio analysis in high dimensions as we do.
Recently, important contributions have been obtained in this area by using shrinkage and factor models. \cite{lw2017} propose a nonlinear shrinkage estimator in which small eigenvalues of the sample covariance matrix are increased and large eigenvalues are decreased by a shrinkage formula. Their main contribution is the optimal shrinkage function, which they find by minimizing a loss function. The maximum out-of-sample Sharpe Ratio is an inverse function of this loss. Their results cover the independent and identically distributed case and when $p/n \to (0,1) \cup (1,+\infty)$. For the analysis of mean-variance efficiency, \cite{ao2019} make a novel contribution in which they take a constrained optimization, maximize returns subject to the risk of the portfolio, and show that it is equivalent to an unconstrained objective function, where they minimize a scaled return of the portfolio error by choosing optimal weights. To obtain these weights, they use lasso regression and assume a sparse number of nonzero weights of the portfolio, and they analyze $p/n \to (0,1)$. They show that their method maximizes the expected return of the portfolio and satisfies the risk constraint. Their paper is an important result on its own. One key paper in the literature is by \cite{fan2011} which assumes an approximate factor model, but, on the other hand, the authors assume conditional sparsity-diagonality of the covariance matrix of errors. \cite{fan2011} show for the first time how to build a precision matrix of returns in a large portfolio via factor models. Therefore, it is a key paper in the high-dimensional econometrics literature.
Regarding other papers, Ledoit and Wolf (2003,2004)\nocite{lw2003,lw2004} propose a linear shrinkage estimator of the covariance matrix and apply it to portfolio optimization. \cite{lw2017} shows that nonlinear shrinkage performs better in out-of-sample forecasts. \cite{lai2011}, and \cite{garlappi2007} approach the same problem from a Bayesian perspective by aiming to maximize a utility function tied to portfolio optimization. Another avenue of the literature improves the performance of the portfolios by introducing constraints on the weights. This type of literature is in the case of the global minimum-variance portfolio. Examples of works investigating this problem include \cite{jag2003} and \cite{fanli2012}. We also see a combination of different portfolios proposed by \cite{kan2007}, and \cite{tu2011}. Very recently, \cite{dlz2020} extended factor models to assumptions that are more consistent with principal components analysis. They provide consistent estimation of the risk of the portfolio under the sparsity of covariance of errors with a fixed number of factors. \cite{bgs2018}, \cite{bddgl2009}, \cite{cr1983}, \cite{dgnu2009}, \cite{fls2015} analyze the mutual fund industry, sparsely constructed Markowitz portfolio, arbitrage and factor models in large portfolios, sparsely constructed mean-variance portfolios, and risks of large portfolios, respectively.
\subsection{Organization of the Paper}
This paper is organized as follows. Section \ref{fnwe} considers our assumptions and feasible precision matrix estimation for errors. Section \ref{fnwr} provides the feasible precision matrix estimate for asset returns.
Section \ref{sratio} analyzes consistency of the Sharpe Ratio in a portfolio with large number of assets in three different scenarios.
Section \ref{sims} provides simulations that compare several methods. Section \ref{emp} presents an out-of-sample forecasting exercise. The main proofs are in the Supplement A, common proofs used for Theorems 3-8 are in Supplement B, Supplement C contains proofs related to section 4.4, and the Supplement D has a proof of mean-variance efficiency of a large portfolio in case of out-of-sample context, and some extra simulation results.
\subsection{Notation}
Let $\| \boldsymbol\nu \|_{l_1}, \| \boldsymbol\nu \|_{l_2}, \| \boldsymbol\nu \|_{\infty}$ be the $l_1, l_2, l_{\infty},$ norms of a generic vector $\boldsymbol\nu$. Let
$\| \boldsymbol v \|_n^2:= n^{-1} \sum_{t=1}^n v_t^2$ which is the prediction norm for an $n \times 1$ vector $\boldsymbol v$. Let $\textnormal{Eigmin}(\boldsymbol A)$ represents the minimum eigenvalue of a matrix $\boldsymbol A$, and $\textnormal{Eigmax}(\boldsymbol A)$ represent the maximum eigenvalue of the matrix $\boldsymbol A$. For a generic matrix
$\boldsymbol A$, let $\| \boldsymbol A \|_{l_1}, \| \boldsymbol A \|_{l_{\infty}}, \| \boldsymbol A \|_{l_2}$, be the $l_1$ induced matrix norm (i.e. maximum absolute column sum norm), $l_{\infty}$ induced matrix norm (i.e. maximum absolute row sum norm), spectral matrix norm, respectively. $\| \boldsymbol A \|_{\infty}$ is maximum absolute value of element of a matrix, and also a norm (but not a matrix norm). Matrix norms have the additional desirable feature of submultiplicativity property. For further information on matrix norms, see p.341 of \cite{hj2013}.
\section{Factor Model and Feasible Nodewise Regression}\label{fnwe}
We start with the following model for the $j$th asset return (excess asset return) at time $t$, $y_{j,t}$, for $j=1,\cdots, p$, and time periods $t=1,\cdots,n$, such that
\begin{equation}
y_{j,t} = \boldsymbol b_j' \boldsymbol f_t + u_{j,t}.\label{rv1}
\end{equation}
where $\boldsymbol b_j$ is a $K \times 1$ vector of factor loadings, $\boldsymbol f_t$ is the $K \times 1$ vector of common factors to all assets' returns, and $u_{j,t}$ is the scalar error (idiosyncratic) term for asset return $j$ at time $t$. All the factors are assumed to be observed. This model is used by \cite{fan2011}. From this point on, when asset return is mentioned, it should be understood as excess asset return.
For the $j$th asset return we can rewrite (\ref{rv1}) in the vector form, for $j=1,\cdots, p$:
\begin{equation}
\boldsymbol y_j = \boldsymbol X' \boldsymbol b_j + \boldsymbol u_j,\label{rv2}
\end{equation}
where $\boldsymbol X = (\boldsymbol f_1, \cdots, \boldsymbol f_n)$ is a $K \times n$ matrix, and $\boldsymbol y_j = (y_{j,1}, \cdots, y_{j,n})'$ is a $n \times 1$ vector of returns of the $j$th asset. We can also express the same relation in a matrix form as follows:
\begin{equation}
\boldsymbol Y = \boldsymbol B \boldsymbol X + \boldsymbol U,\label{rv3}
\end{equation}
where $\boldsymbol Y$ is a $p\times n$ matrix, $\boldsymbol B$ is a $p \times K$ matrix, and $\boldsymbol U$ is a $p \times n$ matrix.~\footnote{
We can also write the returns for each period in time, $t=1,\cdots, n$
\[
\boldsymbol y_t = \boldsymbol B \boldsymbol f_t + \boldsymbol u_t,
\]
where $\boldsymbol y_t =(y_{1,t}, \cdots, y_{j,t}, \cdots, y_{p,t})'$ is a $p \times 1$ vector.}
Define the covariance matrix of the $p \times 1$ vector of errors $\boldsymbol u_t:=(u_{1,t}, \cdots, u_{j,t},\cdots, u_{p,t})'$ as $\boldsymbol\Sigma_n:= \mathbb{E}\left[\boldsymbol u_t \boldsymbol u_t'\right]$.
We take $\{(\boldsymbol f_t,\boldsymbol u_t)\}_{t=1}^n$ to be a strictly stationary, ergodic, and strong mixing sequence of random variables.
Also, let ${\cal F}_{-\infty}^0, {\cal F}_{n}^{\infty}$ be the $\boldsymbol\Sigma$- algebras generated by $\{(\boldsymbol f_t, \boldsymbol u_t)\}$, for $-\infty < t \le 0$, and $n \le t <\infty$, respectively. Denote the strong mixing coefficient as $\alpha(n):= \textnormal{sup}_{ {\cal A} \in {\cal F}_{-\infty}^0, {\cal B} \in {\cal F}_{n}^{\infty}}
| \mathbb{P} ({\cal A}) P ({\cal B}) - \mathbb{P} ( {\cal A} \cap {\cal B} )|.$
In Assumption \ref{as7} below, we assume that maximum eigenvalue of $ \boldsymbol\Sigma_n$ can grow with sample size, this is due to $\boldsymbol\Sigma_n$ being a $p \times p$ matrix where $p$ may grow with $n$. We will assume sparsity for the precision matrix of errors $\boldsymbol\Omega:= \boldsymbol\Sigma_n^{-1}$, but we do not subscript $\boldsymbol\Omega$ with $n$ to avoid cumbersome notation. Each row of $\boldsymbol\Omega$ will be denoted as a $1\times p$ vector $\boldsymbol\Omega_j'$. We represent the indices of nonzero cells in $\boldsymbol\Omega_j'$ as $S_j$, for $l=1,\cdots, p$,
\[
S_j:= \{ j: \Omega_{j,l} \neq 0 \},
\]
where $\Omega_{j,l}$ represents the $l$th element in the $j$th row of $\boldsymbol\Omega$. Let $S_j^c$ represents the index set of all zero elements in the $j$th row of $\boldsymbol\Omega$. Define the cardinality of the non-zero cells in the $j$th row of the precision matrix as $s_j:= |S_j|$, which can be nondecreasing in $n$, but we do not subscript that with $n$.
Denote the maximum number of nonzero elements across all rows $j=1,\cdots,p$ of the precision matrix $\boldsymbol\Omega$ as $\bar{s}:= \max_{1 \le j \le p} s_j$, which is nondecreasing in $n$.
This last definition plays a key role in analysis of the rate of convergence of estimation errors. Note that, just to be clear, when $n \to \infty$, we allow $p \to \infty$, $K \to \infty$, and $\bar{s} \to \infty$. As in the literature, we do not subscript them by $n$. Also, we allow for $p>n$, when $n \to \infty, p \to \infty$, and $p/n \to \infty$ in our analysis in Theorems 1-7, which can be considered ultra-high dimensional portfolio analysis.
For future references, we denote all of the asset returns except the $j$th one as
\begin{equation}
\boldsymbol Y_{-j} = \boldsymbol B_{-j} \boldsymbol X + \boldsymbol U_{-j},\label{-yj}
\end{equation}
where $\boldsymbol Y_{-j}$, of dimension $(p-1) \times n$, is the $\boldsymbol Y$ matrix without the $j$th row, $\boldsymbol B_{-j}$ is the $(p-1)\times K$ matrix which is $\boldsymbol B$ without the $j$th row, and $\boldsymbol U_{-j}$ is the $(p-1) \times n$ matrix given by $\boldsymbol U$ matrix without the $j$th row.
It has been well established in the literature that in case of known $U_j$, $\gamma_j$, which is essential input in nodewise regression, can be recovered with the following lasso problem, with a sequence $\lambda_n >0$, for all $j=1,\cdots, p$,
\begin{equation}
\tilde{\boldsymbol\gamma}_j= \operatorname*{\arg\min}_{\boldsymbol\gamma_j \in \mathbb{R}^{p-1}} \left[ \| \boldsymbol u_j - \boldsymbol U_{-j}' \boldsymbol\gamma_j \|_n^2 + 2 \lambda_n \|\boldsymbol \gamma_j \|_1\right].\label{rv8}
\end{equation}
The main issue with (\ref{rv8}) is, unlike nodewise regression in \cite{canerkock2018}, it is infeasible due to error terms regressed on each other. We now show how to turn this to feasible regression and still consistently estimate $\boldsymbol\gamma_j$.
To get estimates for $\boldsymbol b_j$ and $\boldsymbol B_{-j}'$, \cite{fan2011} use the Ordinary Least Squares (OLS) and show that \footnote{See p.3347 of \cite{fan2011}.}
\begin{equation}
\widehat{\boldsymbol b}_j - \boldsymbol b_j = (\boldsymbol X \boldsymbol X')^{-1} \boldsymbol X \boldsymbol u_j.\label{rv11}
\end{equation}
By equation (\ref{rv2}) we can define the OLS residual as
\begin{eqnarray}
\widehat{\boldsymbol u}_j &=& \boldsymbol y_j- \boldsymbol X' \widehat{\boldsymbol b}_j
= \boldsymbol u_j - \boldsymbol X' (\boldsymbol X\boldsymbol X')^{-1} \boldsymbol X \boldsymbol u_j \nonumber \\
& = & \boldsymbol M_X \boldsymbol u_j,\label{rv12}
\end{eqnarray}
where $\boldsymbol X$ is a $K \times n$ matrix and
\begin{equation}
\boldsymbol M_X := \boldsymbol I_n - \boldsymbol X'(\boldsymbol X \boldsymbol X')^{-1} \boldsymbol X.\label{rv12a}
\end{equation}
Then, by OLS, with $\widehat{\boldsymbol B}_{-j}'$ and $\boldsymbol B_{-j}'$ being $K \times (p-1)$ matrices such that $\widehat{\boldsymbol B}_{-j}' - \boldsymbol B_{-j}' = (\boldsymbol X \boldsymbol X')^{-1} \boldsymbol X \boldsymbol U_{-j}'$.
Define the residuals by transposing (\ref{-yj}) such that
\begin{eqnarray}
\widehat{\boldsymbol U}_{-j}' &= &\boldsymbol Y_{-j}' - \boldsymbol X' \widehat{\boldsymbol B}_{-j}'
= \boldsymbol U_{-j}' - \boldsymbol X' (\widehat{\boldsymbol B}_{-j}' - \boldsymbol B_{-j}') \nonumber \\
& = & \boldsymbol U_{-j }' - \boldsymbol X' (\boldsymbol X \boldsymbol X')^{-1} \boldsymbol X \boldsymbol U_{-j}' \nonumber \\
& = & \boldsymbol M_X \boldsymbol U_{-j}'.\label{rv13}
\end{eqnarray}
Note that $\widehat{\boldsymbol U}_{-j}'$ is a $n \times (p-1)$ matrix $\boldsymbol M_X$ is a $n \times n$ matrix, and $\boldsymbol U_{-j}'$ is a $n \times (p-1)$ matrix.
Next,
use (\ref{rv12}) and (\ref{rv13}):
\begin{equation}
\widehat{\boldsymbol u}_j = \widehat{\boldsymbol U}_{-j}'\boldsymbol\gamma_j + \boldsymbol\eta_{xj},\label{rv14}
\end{equation}
where
\begin{equation}
\boldsymbol\eta_{xj}:= \boldsymbol M_X \boldsymbol \eta_j,\label{rv14a}
\end{equation}
is a $n \times 1$ vector, with $\boldsymbol \eta_j:= \boldsymbol u_j - \boldsymbol U_{-j}' \boldsymbol \gamma_j$. Of course, the key difficulties are how the new $\boldsymbol\eta_{xj}$ and the usage of the residuals affect the consistent estimation of $\boldsymbol\gamma_j$. We define a feasible nodewise estimator
\begin{equation}
\widehat{\boldsymbol\gamma}_j= \operatorname*{\arg\min}_{\boldsymbol\gamma_j \in \mathbb{R}^{p-1}} \left [\| \widehat{\boldsymbol u}_j - \widehat{\boldsymbol U}_{-j}'\boldsymbol\gamma_j \|_n^2 + 2 \lambda_n \|\boldsymbol\gamma_j \|_1 \right].\label{rv15}
\end{equation}
Then, to define $\widehat{\boldsymbol\Omega}_j'$, which is the $j$th row of the precision matrix estimator, we need
\begin{equation}
\widehat{\tau}_j^2:= \widehat{\boldsymbol u}_j'(\widehat{\boldsymbol u}_j - \widehat{\boldsymbol U}_{-j}'\widehat{\boldsymbol\gamma}_j)/n.\label{rv16}
\end{equation}
Now, to form the $j$th row of $\widehat{\boldsymbol\Omega}$, set the $j$th element in the $j$th row as
\begin{equation}
\widehat{\Omega}_{j,j}:=1/\widehat{\tau}_j^2.\label{ojj}
\end{equation}
\begin{equation}
\widehat{\boldsymbol\Omega}_{j,-j}':= -\frac{1}{\widehat{\tau}_j^2}\widehat{\boldsymbol\gamma}_j'.\label{o-jj}
\end{equation}
We want to show that for each $j=1,\cdots, p$, $\widehat{\boldsymbol\Omega}_j'$ is consistent. We can write $\widehat{\boldsymbol\Omega}_j'= \widehat{\boldsymbol C}_j'/\widehat{\tau}_j^2$ with
$\widehat{\boldsymbol C}_j'$ being an $1 \times p$ matrix of ones in $j$th cell and $-\widehat{\boldsymbol\gamma}_j'$ in the other cells.
\subsection{Assumptions and a Key Result}
In this part, we provide the assumptions that will be needed for consistency for the $j$th row of the precision matrix estimator. Let $u_{j,t}$ be the $j$ the element of the $p\times 1$ vector $\boldsymbol u_t$. Similarly, $\boldsymbol u_{-j,t}$ is the $(p-1)\times 1$ vector of errors in $t$th time period, except the $j$th term in $\boldsymbol u_t$. Define $\eta_{j,t}:= u_{j,t}- \boldsymbol u_{-j,t}'\boldsymbol\gamma_j$.
\begin{assum}\label{as1}
(i). $\{\boldsymbol u_t \}_{t=1}^n, \{ \boldsymbol f_t \}_{t=1}^n$ are sequences of (strictly) stationary and ergodic random variables. Furthermore, $\{\boldsymbol u_t \}_{t=1}^n, \{\boldsymbol f_t \}_{t=1}^n$ are independent. $\boldsymbol u_t$ is a ($p \times 1$) zero mean random vector with covariance matrix $\boldsymbol\Sigma_n$ ($p \times p$). $\textnormal{Eigmin} (\boldsymbol\Sigma_n) \ge c > 0$, with $c$ a positive constant, and $\max_{1 \le j \le p} \mathbb{E}\left[u_{j,t}^2\right] \le C < \infty$.
(ii). For the strong mixing variables $\boldsymbol f_t, \boldsymbol u_t$:
$\alpha (t) \le \exp (-C t^{r_0})$, for a positive constant $r_0>0$.
\end{assum}
\begin{assum}\label{as2}
There exists positive constants $r_1$, $r_2$, $r_3 >0$ and another set of positive constants $B_1$, $b_2$, $b_3$, $s_1$, $s_2$, $s_3 >0$, and for $t=1,\cdots, n$, and
$j=1,\cdots, p$, with $k=1,\cdots, K$
(i).
\[
\mathbb{P}\left[|u_{j,t} | > s_1\right] \le \exp[-(s_1/B_1)^{r_1}].
\]
(ii). \[
\mathbb{P}\left[| \eta_{j,t} | > s_2 \right]\le\exp [-(s_2/b_2)^{r_2}].
\]
(iii). \[
\mathbb{P}\left[|f_{k,t} | > s_3\right]\le \exp [-(s_3/b_3)^{r_3}].
\]
(iv). There exists $0< \gamma_1 < 1$ such that $\gamma_1^{-1} = 3 r_1^{-1} + r_0^{-1}$, and we also assume
$3 r_2^{-1} + r_0^{-1} > 1$, and $3 r_3^{-1} + r_0^{-1} >1$.
\end{assum}
Define $\gamma_2^{-1}:= 1.5 r_1^{-1} + 1.5 r_2^{-1} + r_0^{-1}$, and $\gamma_3^{-1}:= 1.5 r_1^{-1} + 1.5 r_3^{-1} + r_0^{-1}$, let
$\gamma_{\min}:=\min (\gamma_1, \gamma_2, \gamma_3)$.
\begin{assum}\label{as3}
(i). $[\ln (p)]^{(2/\gamma_{min}) -1 } = o(n)$, and (ii). $K^2 = o(n)$, (iii). $K = o(p)$.
\end{assum}
\begin{assum}\label{as4}
(i). $\textnormal{Eigmin}[\textnormal{cov}(\boldsymbol f_t)] \ge c > 0$, with $\textnormal{cov}(\boldsymbol f_t)$ being the covariance matrix of the factors $\boldsymbol f_t$, $t=1,\cdots, n$.
(ii). $ \max_{1 \le k \le K } \mathbb{E}\left[f_{kt}^2\right] \le C < \infty$, $\min_{1 \le k \le k} \mathbb{E}\left[f_{kt}^2\right] \ge c > 0$.
(iii). $\max_{1 \le j \le p}\mathbb{E}\left[\eta_{j,t}^2\right] \le C < \infty$.
\end{assum}
\begin{assum}\label{as5} $\bar{s}$, $K$, $p$, and $n$ are such that
(i). $ K^2 \bar{s}^{3/2} \frac{\ln(p)}{n} \to 0.$
(ii). $ \bar{s} \sqrt{\frac{\ln(p)}{n}} \to 0.$
\end{assum}
Note that Assumptions \ref{as1}-\ref{as3} are standard assumptions and are used in \cite{fan2011} as well. Also, we get $0 < \gamma_2 < 1, 0 < \gamma_3 < 1$ given Assumption \ref{as2}(iv). Furthermore, by Assumption 3, $\sqrt{\frac{\ln(p)}{n}} = o(1)$. Note that, Stationary GARCH models with finite second moments and continuous error distributions, as well as causal ARMA processes with continuous error distributions, and a certain class of stationary Markov chains satisfy our Assumptions \ref{as1}-\ref{as2} and are discussed in p.61 of \cite{chang2019}. \cite{chang2019} also uses similar assumptions.
Assumption \ref{as4}(i)-(ii) is also used in \cite{fan2011}, and the nodewise error assumption \ref{as4}(iii) is used in \cite{canerkock2018}. Assumption \ref{as5} shows the interaction of sparsity of the precision matrix with factors. They both contribute negatively to biases that our analysis will show below.
Before the next theorem, we define $\lambda_n$ formally. Let $C>0$ be a generic positive constant, then
\begin{equation}
\lambda_n:= C \left[ K^2 \bar{s}^{1/2} \frac{ln p}{n} + \sqrt{\frac{ln p}{n}}\right]=o(1),\label{lar}
\end{equation}
where we specify tuning parameter in Lemma A.5 in Supplement, and the asymptotic negligibility is by Assumption 5. Note that in tuning parameter $\lambda_n$, the first term involving $K^2$ is due to nodewise regression via factor models. In \cite{caner2019}, without factor models, they have the second term only $\sqrt{\frac{lnp}{n}}$. We now provide one of the main Theorems in the paper. Theorem provides consistent estimates for the rows of the precision matrix of errors.
\begin{thm}\label{rthm1}
Under Assumptions \ref{as1}-\ref{as5}
\[ \| \widehat{\boldsymbol\Omega} - \boldsymbol\Omega \|_{l_{\infty}}:= \max_{1 \le j \le p} \| \widehat{\boldsymbol\Omega}_j' - \boldsymbol\Omega_j' \|_1=
\max_{1 \le j \le p} \| \widehat{\boldsymbol\Omega}_j - \boldsymbol\Omega_j \|_1 = O_p\left( \bar{s} \lambda_n\right) = o_p (1).\]
\end{thm}
\textbf{Remarks:}
\begin{enumerate}
\item
Note that $\hat{\boldsymbol\Omega}_j, \boldsymbol\Omega_j$ are not columns of $\hat{\boldsymbol\Omega}, \boldsymbol\Omega$ respectively. $\hat{\boldsymbol\Omega}_j, \boldsymbol\Omega_j$ are column representation of row vectors $\hat{\boldsymbol\Omega}_j', \boldsymbol\Omega_j'$, respectively.
\item
As long as Assumption \ref{as5} is maintained, the rate of approximation error in Theorem \ref{rthm1} matches the
case where $B=0$ in the factor model (i.e., there is no factor structure). This is the case considered in \cite{caner2019}. The number of factors $K$ increases the approximation error through $\lambda_n$.
\end{enumerate}
\section{Precision Matrix Estimate for The Returns}\label{fnwr}
Assuming orthogonality between factors and the idiosyncratic errors, the $(p\times p)$ covariance matrix of the asset returns is defined as:
\begin{equation}
\boldsymbol\Sigma_y = \boldsymbol B\textnormal{cov}(\boldsymbol f_t)\boldsymbol B' + \boldsymbol\Sigma_n.\label{sigmay}
\end{equation}
We start with the precision matrix formula for the asset returns, based on factor model that we used. Using Sherman-Morrison-Woodbury formula, as in p.13 of \cite{hj2013}, $\boldsymbol\Gamma:=\boldsymbol\Sigma_y^{-1}$ is defined as:
\begin{equation}
\boldsymbol\Gamma:= \boldsymbol\Omega - \boldsymbol\Omega \boldsymbol B\left[ \{\textnormal{cov} (\boldsymbol f_t)\}^{-1} + \boldsymbol B' \boldsymbol\Omega \boldsymbol B\right]^{-1} \boldsymbol B' \boldsymbol\Omega,\label{smw1}
\end{equation}
and the precision matrix estimator for the returns is
\begin{equation}
\widehat{\boldsymbol\Gamma}:= \widehat{\boldsymbol\Omega} - \widehat{\boldsymbol\Omega} \widehat{\boldsymbol B}\left[\{\widehat{\textnormal{cov}(\boldsymbol f_t)}\}^{-1} + \widehat{\boldsymbol B}' \widehat{\boldsymbol\Omega}_{sym} \widehat{\boldsymbol B}\right]^{-1} \widehat{\boldsymbol B}' \widehat{\boldsymbol\Omega},\label{smw2}
\end{equation}
where $\widehat{\boldsymbol\Omega}_{sym}:= \frac{\widehat{\boldsymbol\Omega} + \widehat{\boldsymbol\Omega}'}{2}$ is the symmetrized version of our feasible nodewise regression estimator for the precision matrix for errors.
$\widehat{\textnormal{cov}(\boldsymbol f_t)}= n^{-1} \boldsymbol X \boldsymbol X' - n^{-2} \boldsymbol X \boldsymbol 1_n \boldsymbol 1_n' \boldsymbol X'$ is the estimator for the covariance matrix of returns, and it is given in
p.3327 of \cite{fan2011} with $\boldsymbol 1_n$ representing a $(n \times 1)$ vector of ones. Also, $\widehat{\boldsymbol B}= (\boldsymbol Y \boldsymbol X')(\boldsymbol X \boldsymbol X')^{-1}$ is the least-squares estimator for the factor model in (\ref{rv3}). In addition, $\widehat{\boldsymbol B}$ is a $(p \times K)$ matrix, and $\widehat{\textnormal{cov}(\boldsymbol f_t)}$ is a $K \times K$ matrix. Note that we use a symmetric version of our precision matrix estimator for errors in the term in square brackets in equation (\ref{smw2}). There is a technical reason behind that. The proofs depend on the symmetry of the matrix in the square brackets in (\ref{smw2}), but the other parts in the proof do not need symmetry of the precision matrix estimator. Hence, we use both symmetrized, $\widehat{\boldsymbol\Omega}_{sym}$ and standard (non-symmetric version) of the precision matrix estimator, $\widehat{\boldsymbol\Omega}$.
We want to rewrite the precision matrix and it's estimator so that it's convenient to analyze them technically. In this respect, define
\[
\boldsymbol L:= \boldsymbol B\left[ \{\textnormal{cov}(\boldsymbol f_t)\}^{-1} + \boldsymbol B' \boldsymbol\Omega \boldsymbol B\right]^{-1} \boldsymbol B',
\]
and
\[
\widehat{\boldsymbol L}:= \widehat{\boldsymbol B}\left[ \{\widehat{\textnormal{cov} (\boldsymbol f_t)}\}^{-1} + \widehat{\boldsymbol B}' \widehat{\boldsymbol\Omega}_{sym} \widehat{\boldsymbol B}\right]^{-1} \widehat{\boldsymbol B}'.
\]
As a consequence,
\begin{equation}
\boldsymbol\Gamma= \boldsymbol\Omega - \boldsymbol\Omega \boldsymbol L \boldsymbol\Omega, \quad \widehat{\boldsymbol\Gamma}= \widehat{\boldsymbol\Omega} - \widehat{\boldsymbol\Omega} \widehat{\boldsymbol L} \widehat{\boldsymbol\Omega}.\label{sw3}
\end{equation}
We need to find $ \max_{1 \le j \le p} \| \widehat{\boldsymbol\Gamma}_j - \boldsymbol\Gamma_j \|_1$ where $\boldsymbol\Gamma_j'$ and $\widehat{\boldsymbol\Gamma}_j'$ are the $1\times p$ dimensional rows of the precision matrix of the returns and its estimator, respectively. $\boldsymbol\Gamma_j$ and $\widehat{\boldsymbol \Gamma}_j$ are simply transposes of these rows which are $p \times 1$. In this respect, using (\ref{sw3}) we have that
\begin{equation}
\max_{1 \le j \le p} \| \widehat{\boldsymbol\Gamma}_j - \boldsymbol\Gamma_j \|_1=
\max_{1 \le j \le p } \| \widehat{\boldsymbol\Gamma}_j' - \boldsymbol\Gamma_j' \|_1= \max_{1 \le j \le p} \|( \widehat{\boldsymbol\Omega}_j' - \boldsymbol\Omega_j')-(\widehat{\boldsymbol\Omega}_j' \widehat{\boldsymbol L} \widehat{\boldsymbol\Omega} - \boldsymbol\Omega_j' \boldsymbol L \boldsymbol\Omega)\|_1.\label{sw4}
\end{equation}
Our aim is to simplify and get rates of convergence for the right side term in (\ref{sw4}).
To get consistency and rate of convergence results for the precision matrix for returns, rather than the errors as in Theorem \ref{rthm1} above, we need the following assumption on factor loadings.
\begin{assum}\label{as6} The factor loadings are such that:
(i). $ \max_{1 \le j \le p} \max_{1 \le k \le K} |b_{jk}| \le C <\infty$.
(ii). $\| p^{-1} \boldsymbol B' \boldsymbol B - \boldsymbol \Delta \|_{l_2} = o(1)$ for some $K \times K$ symmetric positive definite matrix $\boldsymbol\Delta$ such that $\textnormal{Eigmin} (\Delta)$ is bounded away from zero.
\end{assum}
Also, a strengthened assumption on sparsity compared to Assumption \ref{as5} is provided.
\begin{assum}\label{as7}
Assume that
(i). $\textnormal{Eigmax}(\boldsymbol\Sigma_n) \le C r_n$, with $C>0$ a positive constant, and $r_n \to \infty$ as $n \to \infty$, and
$r_n/p \to 0$, and $r_n$ is a positive sequence.
(ii).
\[
\bar{s} l_n \to 0,
\]
where
\begin{equation}
l_n:= r_n^2 K^{5/2} \max\left(\bar{s} \lambda_n, \bar{s}^{1/2} K^{1/2} \sqrt{\frac{\max[\ln(p), \ln(n)]}{n}}\right) .\label{dn}
\end{equation}
\end{assum}
Specifically, the rate $l_n$ is the rate of estimation error for $\|\widehat{\boldsymbol L} - \boldsymbol L \|_{l_{\infty}}$ as in Lemma A.13 in Supplement A. Note that Assumption \ref{as6} is used in \cite{fan2011}. Assumption \ref{as7}(i) is used in \cite{gagl2016}. Assumption \ref{as7}(i) allows for the maximal eigenvalue of $\boldsymbol\Sigma_n$ to grow with $n$. In the special case of a diagonal $\boldsymbol\Sigma_n$, due to Assumption \ref{as1}(i), the maximum eigenvalue of a diagonal $\boldsymbol\Sigma_n$ matrix is finite. However, a diagonal matrix of variance of errors case is empirically less relevant and less realistic. We expect the errors to be correlated across assets. For an example of where the maximum eigenvalue of $\boldsymbol\Sigma_n$ may diverge, we show that this may be the case for block diagonal matrix structure for $\boldsymbol\Sigma_n$ in (\ref{suffc}). Note that \cite{s1992} criticizes standard Arbitrage Pricing Theory since eigenvalue of the residual covariances must be bounded even when the number of assets diverge. Our Assumption \ref{as7}(i) moves away from maximum bounded eigenvalue assumption. Our residual covariances approximate error covariances very well and this can be seen in (\ref{pla63}) and (\ref{pla64}) in Supplement A.
Assumption \ref{as7}(ii) is a sparsity assumption which tradeoffs between maximal eigenvalue and the sparsity of the precision matrix. This assumption is needed to analyze the precision matrix for the asset returns. To give an example, ignoring constants, we can have $\bar{s}= \ln(n), K =\ln(n), p=2n$, and $r_n= n^{1/5}$, $\lambda_n = O[\max(\ln(n)^{7/2}/n, \sqrt{\ln(n)/n}]$. Then, Assumption \ref{as7}(ii) is satisfied
\[
n^{2/5} (ln n)^{7/2} max((lnn)^{9/2}/n, (ln n)^{3/2}/n^{1/2}) \to 0.
\]
Next, we define sample mean of the asset returns and the population mean of asset returns.
Let $\widehat{\boldsymbol\mu}:=\frac{1}{n}\sum_{t=1}^n\boldsymbol y_t$, where $\boldsymbol y_t$ is a $p \times 1$ vector of asset returns. Let $\boldsymbol\mu:=\mathbb{E}[ y_t]$. Next theorem provides one of our main results, which is the consistent estimation of the precision matrix for asset returns. Since the precision matrix of asset returns is in the formula of the Sharpe Ratio, as will be shown in Section 4, this theorem is crucial for subsequent analysis.
\begin{thm}\label{rthm2}
(i). Under Assumptions \ref{as1}-\ref{as4}, and \ref{as6}-\ref{as7}
\[ \max_{1 \le j \le p} \| \widehat{\boldsymbol\Gamma}_j - \boldsymbol\Gamma_j \|_1 = O_p \left(\bar{s}l_n\right)= o_p (1).\]
(ii). Under Assumptions \ref{as1}-\ref{as6}
\[ \| \widehat{\boldsymbol\mu} - \boldsymbol\mu \|_{\infty} = O_p \left[\max\left(K \sqrt{\frac{\ln(n)}{n}}, \sqrt{\frac{\ln(p)}{n}}\right)\right] = o_p (1).\]
\end{thm}
\textbf{Remarks:}
\begin{enumerate}
\item
This theorem merges two key concepts: factor models and nodewise regression in high dimensional models.
Theorem \ref{rthm2} clearly shows that there is a tradeoff between the maximal eigenvalue of the errors, the number of factors, and the sparsity of the precision matrix.
Increasing the number of factors in our model badly affect the rate of estimation of the precision matrix of the returns.
\item
Although we focus on factor models in empirical asset pricing, the vector $\boldsymbol f_t$ can be seen as any set of random variables satisfying Assumptions 1-5.
\end{enumerate}
\subsection{Two examples relating precision matrix restrictions to covariance matrix}
We now illustrate how specific structures of the covariance matrix are compatible with the sparsity assumption for the precision matrix. We provide two examples for errors, one block-diagonal covariance matrix for errors, and the other one is the Toeplitz form for the covariance matrix of errors. Then, we provide how they affect the precision matrix and Assumption \ref{as7}(i).
\subsubsection{Block Diagonal Covariance Matrix for Errors} Suppose that there are $m=1,\cdots,M$ blocks in a $p \times p$ covariance matrix of the errors.
\[
\boldsymbol\Sigma_n:=
\begin{bmatrix}
\boldsymbol\Sigma_{n,p_1} & \cdots & \boldsymbol{0} & \cdots & \boldsymbol{0}\\
\vdots & \ddots & \cdots & \cdots & \vdots\\
\boldsymbol{0} & \cdots & \boldsymbol\Sigma_{n,p_m} & \cdots & \boldsymbol{0} \\
\vdots &\cdots&\cdots&\ddots&\vdots\\
\boldsymbol{0} & \cdots & \boldsymbol{0}& \cdots & \boldsymbol\Sigma_{n,p_M}
\end{bmatrix}.
\]
Each block $\boldsymbol\Sigma_{n,p_m}$ is of dimension $p_m \times p_m$ and $\sum_{m=1}^M p_m = p$. Clearly, the inverse is sparse as well:
\[
\boldsymbol\Omega:=\boldsymbol\Sigma_n^{-1}=
\begin{bmatrix}
\boldsymbol\Sigma_{n,p_1}^{-1} & \cdots & \boldsymbol{0} & \cdots & \boldsymbol{0}\\
\vdots & \ddots & \cdots & \cdots & \vdots\\
\boldsymbol{0} & \cdots & \boldsymbol\Sigma_{n,p_m}^{-1} & \cdots & \boldsymbol{0} \\
\vdots &\cdots&\cdots&\ddots&\vdots\\
\boldsymbol{0} & \cdots & \boldsymbol{0}& \cdots & \boldsymbol\Sigma_{n,p_M}^{-1}
\end{bmatrix}.
\]
The sparsity assumption -- Assumption \ref{as1} -- for $\boldsymbol\Omega$ can be translated into $\boldsymbol\Sigma_n$ as $\max_{1 \le m \le M} \max_{1 \le j \le p_m} s_{p_m} = \bar{s}$, where this is the maximum number of nonzero cells in a given row of a block, across all blocks. For Assumption \ref{as7} we need the following inequality from Corollary 6.1.5 of \cite{hj2013}, by seeing that spectral radius of a matrix is larger than or equal to absolute value of any eigenvalue for any square matrix $\boldsymbol A$. Therefore,
\begin{equation}
\textnormal{Eigmax} (\boldsymbol A) \le \min\left(\| \boldsymbol A \|_{l_1}, \| \boldsymbol A \|_{l_{\infty}}\right).\label{evin}
\end{equation}
For the same inequality also see Theorem 5.6.9a of \cite{hj2013}. Relating to Assumption \ref{as7}(i)
\[ \textnormal{Eigmax} (\boldsymbol\Sigma_n) \le \max_{1 \le j_1 \le p} \sum_{j_2=1}^p | \Sigma_{n,j_1, j_2}|=
\max_{1 \le j_1 \le p} \sum_{j_2=1}^p\left|\mathbb{E} [u_{j_1,t} u_{j_2,t}]\right|,\]
where $\Sigma_{n,j_1, j_2}$ is the $j_1, j_2$ element of covariance matrix of errors.
By (\ref{evin}), this last inequality becomes
\[
\textnormal{Eigmax} (\boldsymbol\Sigma_n)\le \max_{1 \le m \le M} \max_{1 \le j_1 \le p_m} \sum_{j_2=1}^{p_m}\left|\mathbb{E}[u_{j_1,t} u_{j_2,t}]\right|.
\]
It is easy to see that using Assumption \ref{as1}(i), and under sufficient conditions for Assumption \ref{as7}(i), with $p_m \to \infty$ as $n \to \infty$
\begin{equation}
\max_{1 \le m \le M} \max_{1 \le j_1 \le p_{m}} \max_{1 \le j_2 \le p_m}\left|\mathbb{E}[u_{ j_1,t} u_{j_2,t}]\right| \le C < \infty, \,
\max_{1 \le m \le M} \frac{p_m}{p} \to 0, \, r_n:= \max_{1 \le m \le M} \,p_m,\label{suffc}
\end{equation}
we get $\textnormal{Eigmax} (\boldsymbol\Sigma_n) \le C r_n, r_n/p \to 0$. This allows the size of the blocks to be increasing with $p$, but the ratio of the maximum block size to total number of parameters should be small.
\subsubsection{Toeplitz Analysis}
In this case, the correlation among errors are $\mathbb{E}[u_{j,t} u_{i,t}]=\rho^{|i-j|}$, with $|\rho| < 1$. Then.
\[
\boldsymbol\Sigma_n=
\begin{bmatrix}
1 & \rho & \rho^2 & \cdots & \rho^{p-1} \\
\rho & 1 & \cdots & \cdots & \cdots \\
\rho^2 & \cdots & 1 & \cdots & \cdots \\
\vdots & \cdots & \cdots & \ddots & \vdots\\
\rho^{p-1} & \cdots & \rho^2 & \rho & 1\\
\end{bmatrix}.
\]
We have the tri-diagonal inverse, with all other cells being zero except the main and two adjacent diagonals.
\[
\boldsymbol\Sigma_n^{-1}= \frac{1}{1-\rho^2}
\begin{bmatrix}
1 & -\rho & 0 & 0 & \cdots & 0\\
-\rho & 1+\rho^2 & -\rho & 0 & \cdots & 0 \\
0 & -\rho & 1+\rho^2 & -\rho & \cdots & 0 \\
\vdots & \cdots& \ddots& \ddots & \ddots & \vdots\\
0 &\cdots & 0 & -\rho & 1+\rho^2 & -\rho\\
0 & \cdots & 0& & -\rho & 1
\end{bmatrix}
\]
Clearly $\bar{s}=3$, and the covariance matrix for errors is not sparse. For Assumption \ref{as7}(i), using (\ref{evin})
\[
\textnormal{Eigmax} (\boldsymbol\Sigma_n) \le \| \boldsymbol\Sigma_n \|_{l_{\infty}} = \max_{ 1 \le j_1 \le p} \sum_{j_2=1}^p\left| \rho ^{|j_2-j_1|}\right|.
\]
Clearly Assumption \ref{as7}(i) is satisfied since the sum on the right side converges to a constant.
\subsection{Algorithm For Asset Return Based Precision Matrix Estimation}
Here we provide a practical algorithm to get the precision matrix estimator for asset returns, $\widehat{\boldsymbol\Gamma}$, and it will depend on the residual-based nodewise regression estimator $\widehat{\boldsymbol\Omega}$, and its symmetric version $\widehat{\boldsymbol\Omega}_{sym}$.
\begin{enumerate}
\item
Use equation (\ref{rv12}) to set up the residual from a least squares based regression via known factors with $\boldsymbol y_j$ as the $j$th asset returns ($n \times 1$)
\[
\widehat{\boldsymbol u}_j = \boldsymbol y_j - \boldsymbol X' \widehat{\boldsymbol b}_j,
\]
with $\widehat{\boldsymbol b}_j= (\boldsymbol X\boldsymbol X')^{-1}\boldsymbol X \boldsymbol y_j$, and $\boldsymbol X=(\boldsymbol f_1, \cdots, \boldsymbol f_t, \cdots, \boldsymbol f_n): K \times n$ matrix with $\boldsymbol f_t: K \times 1$ known factor vector.
\item
Form the transpose matrix of residuals for all asset returns except $j$th one, $\widehat{\boldsymbol U}_{-j}'$, which is a $n \times p-1$ matrix as in (\ref{rv13})
\[
\widehat{\boldsymbol U}_{-j}' = \boldsymbol Y_{-j}' - \boldsymbol X' \widehat{\boldsymbol B}_{-j}',
\]
where $\widehat{\boldsymbol B}_{-j}'= (\boldsymbol X \boldsymbol X')^{-1} \boldsymbol X \boldsymbol Y_{-j}'$, $(K \times p-1$ matrix), which is the transpose of factor loading estimates, $\boldsymbol Y_{-j}': (n \times p-1$) is the transpose matrix of asset returns except the $j$th asset.
\item
Run (\ref{rv15}), nodewise regression of $\widehat{\boldsymbol u}_{j}$ on $\widehat{\boldsymbol U}_{-j}'$ via lasso, and get $\lambda_n$ from Cross-validation or Generalized Information Criterion as in Section 5.1.
\item
Use equation (\ref{rv16}) to get $\widehat{\tau}_j^2$.
\item
Now form $\widehat{\boldsymbol\Omega}_j'$ which is a row in the precision matrix estimate for the errors with $1/\widehat{\tau}_j^2$ as $j$th element of that $j$th row, and put all other elements of the $j$th row, as $-\widehat{\boldsymbol \Gamma}_j'/\widehat{\tau}_j^2$.
\item
Run steps 1-5 for all $j=1,\cdots, p$. Stack all rows $j=1,\cdots, p$ to form $p \times p$ matrix: $\widehat{\boldsymbol\Omega}$. Form symmetric version by
$\widehat{\boldsymbol\Omega}_{sym}:= \frac{\widehat{\boldsymbol\Omega} + \widehat{\boldsymbol\Omega}'}{2}$.
\item
Form
\[
\widehat{\boldsymbol B}= (\boldsymbol Y \boldsymbol X')(\boldsymbol X \boldsymbol X')^{-1},
\]
which is a $p \times K$ matrix of OLS estimates, where $\boldsymbol Y: p\times n$ matrix of all asset returns, where $j=1,\cdots, p$ represent all column-assets, and rows $t=1,\cdots, n$ time periods. Also form the covariance matrix estimate for factors
\[
\widehat{\textnormal{cov}(\boldsymbol f_t)} = n^{-1} \boldsymbol X \boldsymbol X' - n^{-2} \boldsymbol X \boldsymbol 1_n \boldsymbol 1_n' \boldsymbol X',
\]
where $\boldsymbol 1_n$ is $n \times 1$ column vector of ones.
\item
Now form the precision matrix estimate for all asset returns by (\ref{smw2}) and steps 6-7:
\[
\widehat{\boldsymbol\Gamma}= \widehat{\boldsymbol\Omega} - \widehat{\boldsymbol\Omega} \widehat{\boldsymbol B}\{ [\widehat{\textnormal{cov}(\boldsymbol{f}_t)}]^{-1} + \widehat{\boldsymbol B}'\widehat{\boldsymbol\Omega}_{sym} \widehat{\boldsymbol B}\}^{-1} \widehat{\boldsymbol B}'\widehat{\boldsymbol\Omega}.
\]
We use $\widehat{\boldsymbol\Omega}_{sym}$ in the inverse in square brackets, so that we can use specific inequalities for the inverse in our proof. $\widehat{\boldsymbol\Omega} $ is the nodewise regression estimator, and $\widehat{\boldsymbol\Omega}_{sym}$ is the symmetrized version.
\end{enumerate}
\section{Sharpe Ratio Analysis with Large Number of Assets}\label{sratio}
In this section, we apply the results, mainly the estimation of precision matrix of returns, to the analysis of the Sharpe Ratio with large number of assets. Specifically, we allow $p \to \infty$, when $n \to \infty$. There will be four themes in each subsection below. But all of these themes relate to the analysis of consistency of the Sharpe Ratio in portfolios with a large number of assets. All our theoretical analysis is without transaction costs, however in simulations and also in empirical exercise we consider the presence of transaction costs.
The first subsection analyzes the Sharpe Ratio of Global Minimum Variance (GMV) portfolio, and Markowitz Mean-Variance (MMV) portfolio. In the GMV portfolio, we choose the weights to minimize the variance of the portfolio and restricted to sum one. Short-sales are allowed. The Sharpe Ratio is then constructed by dividing the mean portfolio returns by its standard deviation. In MMV portfolio, weights are chosen exactly as GMV but we also impose a target for the portfolio mean return.
The second subsection considers choosing the weights of the portfolio in such a way to maximize the Sharpe Ratio, subject to weights of the portfolio adding up to one. Short sales are allowed. The main difference between GMV in Section \ref{gmv2}, and the Constrained Maximum Sharpe Ratio in Section \ref{model}, is that weights are chosen to minimize the variance in GMV portfolio and then the Sharpe Ratio is computed and, in case of the Constrained Maximum Sharpe Ratio, weights are chosen to maximize the Sharpe Ratio directly. Both methods use the same constraint that the weights of the portfolio should add up to one. In case of the MMV portfolio in Section \ref{mwtzmv} weights are chosen first to minimize the portfolio variance under the conditions described earlier and then, the Sharpe Ratio is computed. The constraint of weights adding up to one is helpful in visualizing assets in percentage terms.
In the third subsection, we analyze the maximum out-of-sample Sharpe Ratio. Here, we do not have a constraint that all weights of the portfolio should add up to one as in Sections \ref{gmv2}, \ref{mwtzmv}, and \ref{model}. The analysis is out-sample unlike the GMV, MMV, and Constrained Maximum Sharpe Ratio portfolios. Weights are chosen to maximize the portfolio returns subject to a constraint of a given variance. But the maximum out-of-sample Sharpe Ratio use estimated weights, with population out-sample mean return vector and the out-sample covariance matrix of returns in the formula. Since the maximum eigenvalue of out-sample covariance matrix of returns is growing, this affects the estimation error rate. Specifically, Sections \ref{gmv2}, \ref{mwtzmv}, and \ref{model} allow $p>n$ and we still get consistency, when $n \to \infty, p \to \infty$. With the maximum-out-of-sample Sharpe Ratio we get consistency only when $p<n$ and $n \to \infty, p \to \infty$.
In the fourth subsection, we consider the effect of estimated portfolio weights on obtaining the optimal Sharpe Ratio in large samples. Specifically, we estimate the weights and substitute this into the Sharpe Ratio formula, with keeping $\boldsymbol\mu, \boldsymbol\Sigma_y$ intact, and then try to show that this estimate is consistent. We show that it is possible only in the case of $p<n$, and this includes diverging number of assets and time span.
Before we state the theorems, we need the following sparsity assumption. Assumption \ref{as8}(i) below replaces Assumption \ref{as7}(ii). In Assumption \ref{as8}(ii), the first term shows square of the maximum Sharpe Ratio is lower bounded, (scaled by $p$), to be positive. Scaling by $p$ is needed since the numerator is summed over $p$ terms. In a similar way, the second term in Assumption \ref{as8}(ii) imposes that the variance of the GMV portfolio (scaled) to be finite. The variance of the GMV portfolio is $\left[\frac{\boldsymbol 1_{p}'\boldsymbol\Gamma \boldsymbol 1_p}{p}\right]^{-1}$. Let $c>0$ be a positive constant.
\begin{assum}\label{as8} Assume that
(i). \[ K^{3} \bar{s} l_n = o(1),\]
(ii). \[ \frac{\boldsymbol\mu'\boldsymbol\Gamma \boldsymbol\mu}{p} \ge c > 0, \quad \frac{\boldsymbol 1_{p}'\boldsymbol\Gamma \boldsymbol 1_p}{p} \ge c > 0.\]
\end{assum}
\subsection{Commonly Used Portfolios with a Large Number of Assets}\label{gmv1}
Here, we provide consistent estimates of the Sharpe Ratio of the GMV and MMV portfolios when $p>n$.
\subsubsection{Global Minimum-Variance (GMV) Portfolio}\label{gmv2}
In this part, we analyze the Sharpe Ratio that we can infer from the GMV portfolio. This is the portfolio in which weights are chosen to minimize the variance of the portfolio subject to the weights summing to one. Specifically,
\begin{equation}
\boldsymbol w_{nw} = \operatorname*{\arg\min}_{\boldsymbol w \in \mathbb{R}^p} \boldsymbol w' \boldsymbol\Sigma_y \boldsymbol w, \quad \mbox{\textnormal{subject to}} \quad \boldsymbol w' \boldsymbol 1_p =1.\label{35a}
\end{equation}
The solution to the above problem is well known and is given by
\[
\boldsymbol w_{nw} = \frac{\boldsymbol\Sigma_y^{-1} \boldsymbol 1_p}{\boldsymbol 1_p' \boldsymbol\Sigma_y^{-1} \boldsymbol 1_p}.
\]
Next, substitute these weights into the Sharpe Ratio formula, normalized by the number of assets
\begin{equation}
SR= \frac{\boldsymbol w'_{nw} \boldsymbol \mu}{\sqrt{\boldsymbol w_{nw}' \boldsymbol\Sigma_y \boldsymbol w_{nw}}} = \sqrt{p}\left(\frac{\boldsymbol 1_p' \boldsymbol\Sigma_y^{-1} \boldsymbol\mu}{p}\right)\left(\frac{\boldsymbol 1_p' \boldsymbol\Sigma_y^{-1} 1_p}{p}\right)^{-1/2}.\label{5.1}
\end{equation}
We estimate (\ref{5.1}) by nodewise regression, noting that $\boldsymbol\Gamma:=\boldsymbol\Sigma_y^{-1}$,
\begin{equation}
\widehat{SR}_{nw} = \sqrt{p}\left(\frac{\boldsymbol 1_p' \widehat{\boldsymbol\Gamma} \widehat{\boldsymbol \mu}}{p}\right)\left(\frac{\boldsymbol 1_p' \widehat{\boldsymbol\Gamma} \boldsymbol 1_p}{p}\right)^{-1/2}.\label{5.2}
\end{equation}
The following theorem is also valid when $p>n$ and establishes both consistency and rate of convergence in the case of the Sharpe Ratio in the global minimum-variance portfolio.
\begin{thm}\label{gmv}
Under Assumptions \ref{as1}--\ref{as4}, \ref{as6}, \ref{as7}(i), and \ref{as8} with $\left|\boldsymbol1_p'\boldsymbol\Gamma\boldsymbol\mu\right|/p \ge C > 0$,
\[
\left| \frac{\widehat{SR}_{nw}^2}{SR^2} - 1
\right| = O_p\left(K^{3/2}\bar{s}l_n\right)=o_p (1).
\]
\end{thm}
\textbf{Remarks:}
\begin{enumerate}
\item
We see that a large $p$ only affects the error by a logarithmic factor as in the definition of $l_n$ in (\ref{dn}). The estimation error increases with the non-sparsity of the precision matrix.
\item
In the case of non-sparse precision matrix, we can only get consistency when $p<<n$. To show this, in case of non-sparse precision matrix $\bar{s}=p$, where all the rows of precision matrix consists of non-zero cells. Then, using (\ref{lar})(\ref{dn}) and Assumption 3, after simplifying expressions, we must have that
\[
K^{3/2} \bar{s} l_n = K^{3/2} p l_n = r_n^2 K^4 p \max[p^{3/2} K^2 \ln(p)/n + p \sqrt{\ln(p)/n}, p^{1/2} K^{1/2} \sqrt{\ln/n}] \to 0,
\]
to get consistency.
\item
Condition $|\boldsymbol 1_p'\boldsymbol\Gamma\boldsymbol\mu |/p \ge C > 0$ is discussed in detail in Remark 3 of Theorem 7.
\end{enumerate}
\subsubsection{Markowitz Mean-Variance (MMV) Portfolio}\label{mwtzmv}
\cite{mar52} portfolio selection is defined as finding the smallest variance given a desired expected return $\rho_1$. The decision problem is
\[
\boldsymbol w_{mv} = \operatorname*{\arg\min}_{\boldsymbol w \in \mathbb{R}^p} (\boldsymbol w' \boldsymbol\Sigma_y \boldsymbol w) \quad \mbox{\textnormal{such that}} \quad \boldsymbol w'\boldsymbol 1_p=1, \quad \textnormal{and}\quad \boldsymbol w'\boldsymbol \mu=\rho_1.
\]
The formula for optimal weight is
\begin{equation}
\boldsymbol w_{mv} = \left[\frac{D - \rho_1 F}{AD - F^2} \right] (\boldsymbol\Sigma_y^{-1} \boldsymbol 1_p/p) + \left[\frac{\rho_1 A - F}{AD - F^2} \right] (\boldsymbol\Sigma_y^{-1} \boldsymbol\mu/p),\label{38}
\end{equation}
where we use $A,F,D$ formulas $A:= \boldsymbol 1_p'\boldsymbol\Gamma \boldsymbol 1_p/p, F := \boldsymbol 1_p' \boldsymbol\Gamma \boldsymbol\mu/p, D:= \boldsymbol\mu' \boldsymbol\Gamma \boldsymbol\mu/p$, with $\boldsymbol\Gamma:= \boldsymbol\Sigma_y^{-1}$. We define the estimators of these terms as $\widehat{A}:=\boldsymbol 1_p' \widehat{\boldsymbol \Gamma}\boldsymbol 1_p/p, \widehat{F}:= \boldsymbol 1_p' \widehat{\boldsymbol \Gamma}\widehat{\boldsymbol \mu}/p, \widehat{D}:=\widehat{\boldsymbol \mu}' \widehat{\boldsymbol \Gamma} \widehat{\boldsymbol \mu}/p$.
The optimal variance of the portfolio in this scenario is normalized by the number of assets
\begin{equation}
V=\frac{1}{p} \left[\frac{A \rho_1^2 - 2 F \rho_1 + D }{AD-F^2} \right].\label{38aa}
\end{equation}
The estimate of that variance is
\[ \widehat{V} = \frac{1}{p} \left[ \frac{\widehat{A} \rho_1^2 - 2 \widehat{F} \rho_1 + \widehat{D}}{\widehat{A} \widehat{D} - \widehat{F}^2} \right].
\]
By our constraint, we obtain
\begin{equation}
\boldsymbol w_{mv}'\boldsymbol\mu = \rho_1.\label{38a}
\end{equation}
Using the variance $V$ above
\begin{equation}
SR_{mv} = \rho_1 \sqrt{p \left( \frac{AD - F^2}{A \rho_1^2 - 2 F \rho_1 + D}
\right)}.\label{30a}
\end{equation}
The estimate of the Sharpe Ratio under the MMV portfolio is
\[
\widehat{SR}_{mv} = \rho_1 \sqrt{ p \left( \frac{\widehat{A} \widehat{D} - \widehat{F}^2}{\widehat{A} \rho_1^2 - 2 \widehat{F} \rho_1 + \widehat{D} } \right) }.
\]
We provide the maximum Sharpe Ratio (squared) consistency in this framework when the number of assets is larger than the sample size. This is a novel result in the literature.
\begin{thm}\label{mmv}
Under Assumptions \ref{as1}-\ref{as4}, \ref{as6},\ref{as7}(i), and \ref{as8} with condition $\left|\boldsymbol 1_p'\boldsymbol\Gamma\boldsymbol\mu /p\right| \ge C > 0$ and $AD - F^2 \ge C_1 > 0$, $A \rho_1^2 - 2 F \rho_1 + D \ge C_1 > 0$, with $\rho_1$ uniformly bounded away from zero and infinity, we have that
\[
\left|\frac{\widehat{SR}_{mv}^2}{SR_{mv}^2} - 1
\right| = O_p\left( K^3 \bar{s} l_n\right)=o_p(1).\]
\end{thm}
\textbf{Remarks:}
\begin{enumerate}
\item
Condition $A D - F^2 \ge C_1 > 0$ shows that the variance is bounded away from infinity, and $A \rho_1^2 - 2 F \rho_1 - D \ge C_1 > 0$ restricts the variance to be positive and bounded away from zero.
\item
We provide the rate of convergence of the estimators, which increases with $p$ in a logarithmic way as in $l_n$ definition in (\ref{dn}), and the non-sparsity of the precision matrix linearly affects affects the error.
\item
To get consistency when there is non-sparse precision matrix, the same analysis in Remark 2 of Theorem \ref{gmv} applies, with $\bar{s}=p$, we need
$p<n$.
\item
Number of factors slows the rate of convergence of estimation error to zero here. This is due to the fact that we have an extra constraint that is affected by number of factors compared with GMV Portfolio.
\end{enumerate}
\subsection{Maximum Sharpe Ratio: Portfolio Weights Normalized to One}\label{model}
In this section, we define the maximum Sharpe Ratio when the portfolio weights are normalized to one. This, in turn will depend on a critical term that will determine the formula below.
The maximum Sharpe Ratio is defined as follows, with $\boldsymbol w$ as the $p \times 1$ vector of portfolio weights:
\[
\max_{\boldsymbol w} \frac{\boldsymbol w'\boldsymbol\mu}{\sqrt{\boldsymbol w' \boldsymbol\Sigma_y \boldsymbol w}},\quad \textnormal{subject to} \quad \boldsymbol 1_p' \boldsymbol w =1,
\]
where $\boldsymbol 1_p$ is a vector of ones. This maximum Sharpe Ratio is constrained to have portfolio weights that sum to one. \cite{maller2016} shows that depending on a scalar, it has two solutions. When $\boldsymbol 1_p' \boldsymbol\Sigma_y^{-1} \boldsymbol\mu > 0$, with $\boldsymbol\Gamma:= \boldsymbol\Sigma_y^{-1}$, we have the square of the maximum Sharpe Ratio:
\begin{equation}
MSR^2= \boldsymbol\mu' \boldsymbol\Sigma_y^{-1} \boldsymbol\mu.\label{1}
\end{equation}
\noindent When $\boldsymbol 1_p' \boldsymbol \Sigma_y^{-1} \boldsymbol \mu >0$, \cite{maller2002} show
\[ {\boldsymbol w}_{c,1}:= \frac{\boldsymbol \Sigma_y^{-1} \boldsymbol \mu}{\boldsymbol 1_p' \boldsymbol \Sigma_y^{-1} \boldsymbol \mu}.\]
\noindent On the other hand, when $\boldsymbol 1_p' \boldsymbol\Sigma_y^{-1}\boldsymbol\mu \le 0$, we have
\begin{equation}
MSR_c^2 = \boldsymbol\mu' \boldsymbol\Sigma_y^{-1}\boldsymbol\mu - (\boldsymbol 1_p' \boldsymbol\Sigma_y^{-1} \boldsymbol\mu)^2/(\boldsymbol 1_p' \boldsymbol\Sigma_y^{-1}\boldsymbol 1_p).\label{e1}
\end{equation}
This is equation (6.1) of \cite{maller2016}. Equation (\ref{1}) is used in the literature, and this is the formula when the weights do not necessarily sum to one given a return constraint as in \cite{ao2019}. In case of $\boldsymbol 1_p' \boldsymbol \Sigma_y^{-1} \boldsymbol \mu \le 0$, in equations (2.7)-(2.10) of \cite{maller2002}, there is an approximation to optimal portfolio weights. To be specific, with a positive $\delta >0$, optimal portfolio weights, which is $(p \times 1)$ vector:
\[ \boldsymbol w_{c,2}:= (\delta \boldsymbol u_{max}', 1 - \delta \boldsymbol 1_{p-1}' \boldsymbol u_{\max})',\]
where
\[ \boldsymbol u_{max}:= \frac{(\boldsymbol A_{p-1}' \boldsymbol A_{p-1})^{-1} \boldsymbol A_{p-1}' \boldsymbol z_{\max}}{\sqrt{\boldsymbol z_{max}' \boldsymbol A_{p-1} (\boldsymbol A_{p-1}' \boldsymbol A_{p-1})^{-2} \boldsymbol A_{p-1}' \boldsymbol z_{\max}}}\]
is a $(p-1) \times 1$ matrix with $\boldsymbol A_{p-1}:= (I_{p-1},-1_{p-1})': p \times p-1$ matrix, with $1_{p-1}$ a $(p-1)$ column vector of ones, and
\[ z_{\max}:= \boldsymbol \Sigma_y^{-1} \left( \boldsymbol I_p - \frac{\boldsymbol 1_p \boldsymbol 1_p' \boldsymbol \Sigma_y^{-1}}{\boldsymbol 1_p' \boldsymbol \Sigma_y^{-1} \boldsymbol 1_p}
\right) \frac{\boldsymbol \mu}{MSR_c}\]
is of dimension $p \times 1$.
When $\delta \to \infty$, the weights can provide the maximum Sharpe Ratio: $MSR_c$, as discussed in p.504 of \cite{maller2002}.
These equations can be estimated by their sample counterparts, but in the case of $p>n$, $\widehat{\boldsymbol\Sigma}_n$ is not invertible, so we need to use new tools from high-dimensional statistics. We use the nodewise regression precision matrix estimate of \cite{mein2006}. This estimate is denoted by $\widehat{\boldsymbol\Omega}$. $\widehat{\boldsymbol \Omega}$ is incorporated into the precision matrix of returns $\hat{\boldsymbol \Gamma}$.
We will also introduce the maximum Sharpe Ratio, which addresses the uncertainty regarding whether we should analyze $MSR$ or $MSR_c$. This is
\[
(MSR^*)^2 = MSR^2 1_{ \{\boldsymbol 1_p' \boldsymbol\Sigma_y^{-1}\boldsymbol\mu > 0 \}} + MSR_c^2 1_{ \{\boldsymbol1_p' \boldsymbol\Sigma_y^{-1}\boldsymbol\mu \le 0 \}}.
\]
Note also that with $\boldsymbol 1_p' \boldsymbol \Sigma_y^{-1} \boldsymbol \mu =0$, $MSR=MSR_c$.
The estimators for $MSR, MSR_c, MSR^*$ will be introduced in the next subsection.
\subsubsection{Consistency and Rate of Convergence of Constrained Maximum Sharpe Ratio Estimators}\label{cmsr}
First, when $\boldsymbol 1_p' \boldsymbol\Sigma_y^{-1} \boldsymbol\mu > 0$, we have the square of the maximum Sharpe Ratio as in (\ref{1}). Namely, the estimate of the square of the maximum Sharpe Ratio is:
\begin{equation}
\widehat{MSR}^2= \widehat{\boldsymbol\mu}' \widehat{\boldsymbol\Gamma} \widehat{\boldsymbol\mu}.\label{est1}
\end{equation}
\begin{thm}\label{nwsr}
Under Assumptions \ref{as1}-\ref{as4}, \ref{as6},\ref{as7}(i), \ref{as8} with $\boldsymbol 1_p'\boldsymbol\Gamma\boldsymbol\mu > 0$,
\[
\left| \frac{\widehat{MSR}^2}{MSR^2} -1 \right| =O_p\left( K^2 \bar{s} l_n\right)= o_p(1).
\]
\end{thm}
\textbf{Remarks:}
\begin{enumerate}
\item
We allow $p>n$ and $p$ can grow exponentially in $n$. We also allow for time-series data and establish a rate of convergence. The number of assets, on the other hand, can also increase the error on a logarithmic scale, as can be seen in (\ref{dn}). So assumption on sparsity of the precision matrix helps us derive this result.
\item
When there is no sparsity of the precision matrix, i.e. $\bar{s}=p$, we can still get consistency but for $p<<n$. To see this, consider the error rate in Theorem \ref{nwsr} above, with $l_n$ definition in (\ref{dn})
\[K^2 \bar{s} l_n = K^2 p l_n \to 0.\]
This implies that to get consistency we need $p<n$.
\end{enumerate}
If $\boldsymbol 1_p' \boldsymbol\Sigma_y^{-1}\boldsymbol\mu \le 0 $, the Sharpe Ratio is minimized, as shown on p.503 of \cite{maller2002}. The new maximum Sharpe Ratio in the case when $\boldsymbol 1_p' \boldsymbol\Sigma_y^{-1}\boldsymbol \mu \le 0$ is in Theorem 2.1 of \cite{maller2002}.
The square of the maximum Sharpe Ratio when $\boldsymbol 1_p' \boldsymbol\Sigma_y^{-1}\boldsymbol\mu \le 0$ is given in (33).
An estimator in this case is
\begin{equation}
\widehat{MSR}_c^2= \widehat{\boldsymbol\mu}' \widehat{\boldsymbol\Gamma} \widehat{\boldsymbol\mu} - (\boldsymbol 1_p' \widehat{\boldsymbol\Gamma} \widehat{\boldsymbol\mu})^2/(\boldsymbol 1_p' \widehat{\boldsymbol\Gamma}\boldsymbol 1_p).\label{e2}
\end{equation}
The optimal portfolio allocation for such a case is given in (2.10) of \cite{maller2002}, and shown in $\boldsymbol w_{c,2}$ here in Section 4.2. The limit for such estimators when the number of assets is fixed ($p$ fixed) is given in Theorems 3.1b-c of \cite{maller2016}.
\begin{thm}\label{tme1}
If $\boldsymbol 1_p' \boldsymbol\Gamma\boldsymbol\mu \le 0$, and under Assumptions \ref{as1}-\ref{as4},\ref{as6},\ref{as7}(i), \ref{as8} with $AD - F^2 \ge C_1 > 0$, where $C_1$ is a positive constant,
\[
\left|
\frac{\widehat{MSR}_c^2}{MSR_c^2} - 1
\right| = O_p\left(K^2 \bar{s} l_n\right)= o_p (1).
\]
\end{thm}
\textbf{Remarks:}
\begin{enumerate}
\item
In Theorem \ref{tme1}, we allow $p>n$, and time-series data are allowed, unlike the iid or normal return cases in the literature when dealing with large $p,n$.
\item
Case of non-sparse precision matrix proceeds in the same way as Remark 2 of Theorem \ref{nwsr}. To have consistency, we need $p<n$, with non-sparse case $\bar{s}=p$.
\end{enumerate}
We provide an estimate that takes into account uncertainties about the term $\boldsymbol 1_p'\boldsymbol\Sigma_y^{-1}\boldsymbol\mu$. Note that the term can be consistently estimated, as shown in Lemma \ref{tl3} in Supplement B. A practical estimate for a maximum Sharpe Ratio that will be consistent is:
\[
\widehat{MSR}^* = \widehat{MSR} 1_{ \{\boldsymbol 1_p'\widehat{\boldsymbol\Gamma} \widehat{\boldsymbol\mu} > 0 \} } + \widehat{MSR}_c 1_{ \{\boldsymbol1_p' \widehat{\boldsymbol\Gamma} \widehat{\boldsymbol\mu} < 0 \} },
\]
where we excluded the case of $\boldsymbol 1_p' \widehat{\boldsymbol\Gamma} \widehat{\boldsymbol\mu} =0$ in the estimator. That specific scenario is very restrictive in terms of returns and variance.
Note that under a mild assumption, when $\boldsymbol 1_p'\boldsymbol\Gamma\boldsymbol\mu > 0$, we have $\boldsymbol 1_p'\widehat{\boldsymbol\Gamma}\widehat{\boldsymbol\mu} >0$, and when $\boldsymbol 1_p'\boldsymbol\Gamma\boldsymbol\mu < 0$, we have $\boldsymbol 1_p' \widehat{\boldsymbol\Gamma} \widehat{\boldsymbol\mu}<0$ with probability approaching one in the proof of Theorem \ref{tme2}. Note that $\boldsymbol\Gamma:= \boldsymbol\Sigma_y^{-1}$.
\begin{thm}\label{tme2}
Under Assumptions \ref{as1}-\ref{as4},\ref{as6},\ref{as7}(i), \ref{as8}, with $AD - F^2 \ge C_1 > 0$, where $C_1$ is a positive constant, and assuming
$|\boldsymbol 1_p' \boldsymbol\Gamma \boldsymbol\mu |/p \ge C > 2 \epsilon > 0$, with a sufficiently small positive $\epsilon > 0$, and $C$ being a positive constant,
\[ \left|
\frac{(\widehat{MSR}^*)^2}{(MSR^*)^2} - 1
\right| = O_p ( K^2 \bar{s} l_n)= o_p (1).\]
\end{thm}
\textbf{Remarks:}
\begin{enumerate}
\item
In the case of $p>n$, we only consider consistency since standard central limit theorems (apart from those in rectangles or sparse convex sets) do not apply, and ideas such as multiplier bootstrap and empirical bootstrap with self-normalized moderate deviation results do not extend to this specific Sharpe Ratio formulation.
\item
The case of non-sparse precision matrix with $\bar{s}=p$ proceeds in the same way as in Remark 2 after Theorem \ref{nwsr}.
\item
Condition $|\boldsymbol 1_p'\boldsymbol\Gamma\boldsymbol\mu |/p \ge C> 2 \epsilon > 0$ shows that apart from a small region around 0, we include all cases. This is similar to the $\beta-\min$ condition in high-dimensional statistics used to achieve model selection. Note
\[ \left|\boldsymbol 1_p'\boldsymbol\Gamma\boldsymbol\mu /p \right| = \left|\sum_{j=1}^p \sum_{k=1}^p \Gamma_{j,k} \mu_k/p\right|,\]
which is a sum measure of roughly theoretical mean divided by standard deviations. It is difficult to see how this double sum in $p$ will be a small number, unless
the terms in the sum cancel out one another. Therefore, we exclude that type of case with our assumption. Additionally, $\epsilon$ is not arbitrary, from the
proof this is the upper bound on the $| \widehat{F} - F|$ in Lemma \ref{tl3} in Supplement B, and it is of order
\[
\epsilon= O(K\bar{s}l_n)=o(1),
\]
where the asymptotically small term follows Assumption \ref{as8}.
\end{enumerate}
\subsection{Maximum Out-of-Sample Sharpe Ratio}\label{mos}
This section analyzes the maximum out of Sharpe Ratio that is considered in \cite{ao2019}. To obtain that formula, we need the optimal calculation of the weights of the portfolio. The optimization of the portfolio weights is formulated as
\begin{equation}
\boldsymbol w_{mos} = \operatorname*{\arg\max}_{\boldsymbol w\in\mathbb{R}^p} \boldsymbol w'\boldsymbol \mu \quad \textnormal{subject to} \quad \boldsymbol w'\boldsymbol\Sigma_y \boldsymbol w \le \sigma^2, \label{www}
\end{equation}
where we maximize the return subject to a specified positive and finite risk constraint, $\sigma^2>0$. Equation (A.2) of \cite{ao2019}
defines the estimated maximum out-of-sample ratio when $p<n$, with the inverse of the sample covariance matrix, $\widehat{\boldsymbol\Sigma}_y^{-1} = [\frac{1}{n} \sum_{t=1}^n\boldsymbol y_t \boldsymbol y_t']^{-1}$ used as an estimator for the precision matrix estimate:
\[
\widehat{SR}_{moscov} := \frac{\boldsymbol\mu' \widehat{\boldsymbol\Sigma}_y^{-1} \widehat{\boldsymbol\mu}}{\sqrt{\widehat{\boldsymbol\mu}' \widehat{\boldsymbol\Sigma}_y^{-1} \boldsymbol\Sigma_y \widehat{\boldsymbol\Sigma}_y^{-1} \widehat{\boldsymbol\mu}}}.
\]
The theoretical version is written as, by definition of $\boldsymbol\Gamma:= \boldsymbol\Sigma_y^{-1}$,
\[
SR^* := \sqrt{\boldsymbol\mu'\boldsymbol\Gamma\boldsymbol\mu}.
\]
Then, equation (1.1) of \cite{ao2019} shows that when $p/n \to r_1 \in (0,1)$, the above plug-in maximum out-of-sample ratio cannot consistently estimate the theoretical version.
The optimal weights of a portfolio are given in (2.3)
of \cite{ao2019} in an out-of-sample context given a risk level. This comes from maximizing the expected portfolio return subject to its variance being constrained by the square of the risk, where this is shown in (\ref{www}).
Since $\boldsymbol\Gamma:= \boldsymbol\Sigma_y^{-1}$, the formula for weights is
\[
\boldsymbol w_{mos} = \frac{\sigma \boldsymbol\Gamma \boldsymbol\mu}{\sqrt{\boldsymbol\mu' \boldsymbol\Gamma\boldsymbol\mu}}.
\]
The estimates that we will use
\[ \widehat{\boldsymbol w}_{mos} = \frac{\sigma \widehat{\boldsymbol \Gamma} \widehat{\boldsymbol \mu}}{\sqrt{\widehat{\boldsymbol \mu}' \widehat{\boldsymbol \Gamma} \widehat{\boldsymbol \mu}}}.\]
Our maximum out-of-sample Sharpe Ratio estimate using the nodewise estimate $\widehat{\boldsymbol\Gamma}$ is:
\[
\widehat{SR}_{mos} := \frac{\widehat{\boldsymbol w}_{mos}' \mu}{\sqrt{\widehat{\boldsymbol w}_{mos}' \Sigma_y \widehat{\boldsymbol w}_{mos}}}=
\frac{\boldsymbol\mu' \widehat{\boldsymbol\Gamma} \widehat{\boldsymbol\mu}}{\sqrt{ \widehat{\boldsymbol\mu}' \widehat{\boldsymbol\Gamma}' \boldsymbol\Sigma_y \widehat{\boldsymbol\Gamma} \widehat{\boldsymbol\mu}}}.
\]
Below we provide a sparsity assumption for the case of maximum out of sample Sharpe Ratio.
\begin{assum}\label{as9}
\[
p \bar{s} l_n = o(1).
\]
\end{assum}
\begin{thm}\label{msros}
Under Assumptions \ref{as1}-\ref{as4},\ref{as6}, \ref{as7}(i), \ref{as8}, \ref{as9}
\[ \left| \left[\frac{\widehat{SR}_{mos}}{SR^*}\right]^2 - 1\right| = O_p ( K^2 \bar{s} l_n)=o_p(1).\]
\end{thm}
\textbf{Remarks:}
\begin{enumerate}
\item
Note that p.4353 of \cite{lw2017} shows that the maximum out-of-sample Sharpe Ratio is equivalent to minimizing a certain loss function of the portfolio. The limit of the
loss function is derived under an optimal shrinkage function in Theorem 1. After that, they provide a shrinkage function even in the cases of $p/n \to r_1 \in (0,1) \cup (1,+\infty)$. Their proofs allow for iid data, which is restrictive since it does not allow for correlation in returns across time.
\item
We cannot have $p>n$ in this theorem, due to Assumption \ref{as9}, this shows the difficulty of maximum out of sample estimation. Mainly, $\textnormal{Eigmax} (\boldsymbol\Sigma_y) = O(p)$ caused this problem in the proofs and provide the need for Assumption \ref{as9}.
\item
$p \bar{s} l_n = o(1)$ can be also obtained in non-sparse precision matrix, although the conditions will be more restrictive. To see this, we now have $\bar{s}=p$ in non-sparse case. So
\[
p \bar{s} l_n = p^2 r_n^2 K^{5/2} \max[p \lambda_n, p^{1/2} K^{1/2} \sqrt{\ln(n)/n}] \to 0,
\]
by (\ref{lar})(\ref{dn}). This implies that we need $p<n$.
\item
The case of large non-negative weights can be handled with our analysis. This is the case of growing exposure, where the weights are depending on growing sparsity of the precision matrix, hence taking large values. For this, in Supplement D, our proof of Theorem D.1-analyzing mean of the portfolio- provides insight into this issue. Our Assumption 1 allows $\bar{s}$ to be nondecreasing in $n$.
\end{enumerate}
\subsection{Portfolio Estimation Based Sharpe Ratio Analysis}
In this section for the scenarios we considered in Sections 4.1-4.2, we form the estimate of the portfolio weights and substitute that into the Sharpe Ratio. To understand the effects of only portfolio estimation for consistent estimation of Sharpe Ratio, we keep $\boldsymbol\mu, \boldsymbol\Sigma_y$ as constants in Sharpe Ratio estimates. We start with GMV portfolio. The estimated portfolio weights are
\[
\hat{\boldsymbol w}_{nw}:= \frac{\hat{\boldsymbol\Gamma} \boldsymbol1_p}{\boldsymbol1_p' \hat{\boldsymbol\Gamma} \boldsymbol1_p}.
\]
The Sharpe Ratio estimate of this portfolio is:
\[ \widehat{SR}_{nw,p}:= \frac{\hat{\boldsymbol w}_{nw}' \boldsymbol\mu}{\sqrt{\hat{\boldsymbol w}_{nw}' \boldsymbol\Sigma_y \hat{\boldsymbol w}_{nw}}} = \frac{p^{1/2} \left(\boldsymbol 1_p' \hat{\boldsymbol \Gamma}' \boldsymbol \mu/p\right)}{\sqrt{\boldsymbol 1_p' \hat{\boldsymbol \Gamma}' \boldsymbol \Sigma_y
\hat{\boldsymbol \Gamma} \boldsymbol 1_p/p}}.\]
The optimized-target population Sharpe Ratio is given in (\ref{5.1}).
\begin{corollary}\label{cor1}
Under Assumptions \ref{as1}-\ref{as4},\ref{as6}, \ref{as7}(i), \ref{as8}, \ref{as9} with $|\boldsymbol 1_p' \boldsymbol \Gamma \boldsymbol \mu|/p \ge C > 0$
\[ \left| \left[\frac{\widehat{SR}_{nw,p}}{SR}\right]^2 - 1\right| = O_p ( K \bar{s} l_n)=o_p(1).\]
\end{corollary}
Now we consider the Sharpe Ratio based on Markowitz portfolio. The estimated portfolio weights are
\[
\hat{\boldsymbol w}_{mv}= \frac{\hat{D} - \rho_1 \hat{F}}{\hat{A} \hat{D} - \hat{F}^2} \left( \hat{\boldsymbol \Gamma} \boldsymbol 1_p/p
\right) + \frac{\rho_1 \hat{A} - \hat{F}}{\hat{A} \hat{D} - \hat{F}^2} \left( \hat{\boldsymbol \Gamma} \hat{\boldsymbol \mu}/p
\right).
\]
These are estimates by plugging in terms in equation (\ref{38}). Denote the Sharpe Ratio based on portfolio weight estimates
\[ \widehat{SR}_{mv,p} = \frac{\hat{\boldsymbol w}_{mv}' \mu }{\sqrt{\hat{\boldsymbol w}_{mv}' \boldsymbol \Sigma_y \hat{\boldsymbol w}_{mv}}}.\]
The optimal Sharpe Ratio is in (\ref{30a}) in this case.
\begin{corollary}\label{cor2}
Under Assumptions \ref{as1}-\ref{as4},\ref{as6}, \ref{as7}(i), \ref{as8}, \ref{as9} with $A \rho_1^2 - 2 F \rho_1 + D \ge C_1>0,
AD -F^2 \ge C_1 > 0$, and $|\boldsymbol 1_p'\boldsymbol \Gamma \boldsymbol \mu | \ge C >0$, with $\rho_1$ bounded away from zero and infinity,
\[ \left| \left[\frac{\widehat{SR}_{mv,p}}{SR_{mv}}\right]^2 - 1\right| = O_p ( K^{5/2} \bar{s} l_n)=o_p(1).\]
\end{corollary}
In case of constrained maximum Sharpe Ratio in section 4.2, when $\boldsymbol 1_p' \boldsymbol \Sigma_y^{-1} \boldsymbol \mu >0$, we can establish the portfolio weight estimates
\[ \hat{\boldsymbol w}_{c,1}= \frac{\hat{\boldsymbol \Gamma} \hat{\boldsymbol \mu}}{\boldsymbol 1_p' \hat{\boldsymbol \Gamma} \hat{\boldsymbol \mu}}.\]
Constrained maximum Sharpe Ratio estimate when $\boldsymbol 1_p' \boldsymbol \Sigma_y^{-1} \boldsymbol \mu >0$ is:
\[ \widehat{MSR}_{p}:= \frac{ \hat{\boldsymbol w}_{c,1}' \boldsymbol \mu}{\sqrt{\hat{\boldsymbol w}_{c,1}' \boldsymbol \Sigma_y \hat{\boldsymbol w}_{c,1}}}
= \frac{\hat{\boldsymbol \mu}' \hat{\boldsymbol \Gamma}' \boldsymbol \mu}{\sqrt{\hat{\boldsymbol \mu}' \hat{\boldsymbol \Gamma}' \boldsymbol \Sigma_y \hat{\boldsymbol \Gamma} \hat{\boldsymbol \mu}}}.\]
The optimal Sharpe Ratio in this case is in (\ref{1}).
\begin{corollary}\label{cor3}
Under Assumptions \ref{as1}-\ref{as4},\ref{as6}, \ref{as7}(i), \ref{as8}, \ref{as9} with $\boldsymbol 1_p' \boldsymbol \Sigma_y^{-1} \boldsymbol \mu \ge C >0$
\[ \left| \left[\frac{\widehat{MSR}_{p}}{MSR}\right]^2 - 1\right| = O_p ( K^{2} \bar{s} l_n)=o_p(1).\]
\end{corollary}
The constrained maximum Sharpe Ratio weights when $\boldsymbol 1_p' \boldsymbol \Sigma_y^{-1} \boldsymbol \mu \le 0$ are more complicated as seen in $\boldsymbol w_c$ in Section 4.2. The estimate is:
\[ \hat{\boldsymbol w}_{c,2}:= (\delta \hat{\boldsymbol u}_{max}, 1 - \boldsymbol 1_{p-1}' \hat{\boldsymbol u}_{\max})', \]
with
\[ \hat{\boldsymbol u}_{\max}:= \frac{(\boldsymbol A_{p-1}' \boldsymbol A_{p-1})^{-1} (\boldsymbol A_{p-1}' \hat{\boldsymbol z}_{max})}{\sqrt{\hat{\boldsymbol z}_{max}' \boldsymbol A_{p-1} (\boldsymbol A_{p-1}' \boldsymbol A_{p-1})^{-2} \boldsymbol A_{p-1}' \hat{\boldsymbol z}_{max}}}.\]
\[ \hat{\boldsymbol z}_{max}:= \hat{\boldsymbol \Gamma} \left( I_p - \frac{\boldsymbol 1_p \boldsymbol 1_p' \hat{\boldsymbol \Gamma}}{\boldsymbol 1_p' \hat{\boldsymbol \Gamma} \boldsymbol 1_p}
\right) \frac{\hat{\boldsymbol \mu}}{\widehat{MSR}_c}.\]
Note that maximum Sharpe Ratio in this second constrained case is:
\[ \widehat{MSR}_{c,p}= \frac{\hat{\boldsymbol w}_{c,2}' \boldsymbol \mu}{\sqrt{\hat{\boldsymbol w}_{c,2}' \boldsymbol \Sigma_y \hat{\boldsymbol w}_{c,2}}}.\]
Using $\hat{\boldsymbol w}_{c,2}$ poses several challenges.
Taking $\delta \to \infty$ to reach the optimal Sharpe Ratio is key but the rate may play a role and also the weights depend on $\hat{\boldsymbol u}_{\max}$ term which depends on $\hat{\boldsymbol z}_{max}$ that depends on precision matrix estimate $\hat{\boldsymbol \Gamma}$, mean estimate $\hat{\mu}$, and estimate $\widehat{MSR}_c$ from section 4.2. So, given Theorems 2 and 6, we think that consistency is plausible. However, given the lengthy material in this paper, this is beyond the scope of our theoretical analysis. Hence, similar corollaries for Theorems 6-7 cannot be handled in this paper.
An important fact that applies to all Corollaries here is that we can only have $p<n$ case, as discussed in Remark 3 of Theorem 8.
\section{Simulations}\label{sims}
\subsection{Models and Implementation Details}
In this section, we compare the nodewise regression with several models in a simulation exercise. The two aims of the exercise are to determine whether our method achieves consistency and how our method performs compared to others in the estimation of the constrained maximum Sharpe Ratio, the out-of-sample maximum Sharpe Ratio, and the Sharpe Ratio in global minimum-variance and Markowitz mean-variance portfolios.
The other methods that are used widely in the literature and benefit from high-dimensional techniques are the principal orthogonal complement thresholding (POET) from \cite{fan2013}, the nonlinear shrinkage (NL-LW) and the single factor nonlinear shrinkage (SF-NL-LW) from \cite{lw2017}, and the maximum Sharpe Ratio estimated and sparse regression (MAXSER) from \cite{ao2019}. All models except for the MAXSER are plug-in estimators, where the first step is to estimate the precision/covariance matrix, and the second step is to plug-in the estimate in the desired equation.
The POET uses principal components to estimate the covariance matrix and allows some eigenvalues of $\boldsymbol\Sigma_n$ to be spiked and grow at a rate $O(p)$, which allows common and idiosyncratic components to be identified via principal components analysis and can consistently estimate the space spanned by the eigenvectors of $\boldsymbol\Sigma_n$. However, \citet{fan2013} point out that the absolute convergence rate of the model is not satisfactory for estimating $\boldsymbol\Sigma_n$, and consistency can only be achieved in terms of the relative error matrix.
Nonlinear shrinkage is a method that individually determines the amount of shrinkage of each eigenvalue in the covariance matrix for a particular loss function. The main aim is to increase the value of the lowest eigenvalues and decrease the largest eigenvalues to stabilize the high-dimensional covariance matrix. This nonlinear method is a very novel and excellent idea.
\citet{lw2017} propose a function that captures the objective of an investor using portfolio selection. As a result, they have an optimal estimator of the covariance matrix for portfolio selection for many assets. The SF-NL-LW method extracts a single factor structure from the data before estimating the covariance matrix, which is simply an equal-weighted portfolio with all assets.
Finally, the MAXSER starts with estimating the adjusted squared maximum Sharpe Ratio used in a penalized regression to obtain the portfolio weights. Of all the discussed models, the MAXSER is the only one that does not estimate the precision matrix in a plug-in estimator of the maximum Sharpe Ratio.
Regarding implementation, the POET and both models from \cite{lw2017} are available in the R packages POET \cite{POETR} and nlshrink \cite{nlshrinkR}. The SF-NL-LW needs some minor adjustments following the procedures described in \cite{lw2017}. For the MAXSER, we follow the steps for the non-factor case in \cite{ao2019}, and we use the package lars (\cite{larsR}) for the penalized regression estimation. We estimate the nodewise regression following the steps in Section 3.2 using the glmnet package \cite{glmnet2010} for penalized regressions. We used two alternatives to select the regularization parameter $\lambda$, a $10$-fold cross validation (CV), and the generalized information criterion (GIC) from \cite{zhang2010regularization}.
The GIC procedure starts by fitting $\widehat{\boldsymbol\gamma}_j$ in (\ref{rv15}) for a range of $\lambda_j$ that goes from the intercept-only model to the largest feasible model. This is automatically done by the glmnet package. Then, for the GIC procedure, we calculate the information criterion for a given $\lambda_j$ among the ranges of all possible tuning parameters
\begin{equation}
GIC_j (\lambda_j) = \frac{SSR (\lambda_j)}{n} + q(\lambda_j) \log(p-1) \frac{\ln[\ln(n)]}{n},
\end{equation}
where $SSR(\lambda_j)$ is the sum squared error for a given $\lambda_j$, $q (\lambda_j)$ is the number of variables, given $\lambda_j, $ in the model that is nonzero, and $p$ is the number of assets. The last step is to select the model with the smallest GIC. Once this is done for all assets $j = 1,\dots,p$, we can proceed to obtain $\widehat{\boldsymbol\Gamma}_{GIC}$.
For the CV procedure, we split the sample into $k$ subsamples and fit the model for a range of $\lambda_j$ as in the GIC procedure. However, we will fit models in the subsamples. We always estimate the models in $k-1$ subsamples, leaving one subsample as a test sample, where we compute the mean squared error (MSE). After repeating the procedure using all $k$ subsamples as a test, we finally compute the average MSE across all subsamples and select the $\lambda_j$ for each asset $j$ that yields the smallest average MSE. We can then use the estimated $\widehat{\boldsymbol\gamma}_j$ to obtain $\widehat{\boldsymbol\Gamma}_{CV}$.
\subsection{Data Generation Process and Results}
The DGP is based on a simplified version of the factor DGP in \citet{ao2019}, for $j=1,\cdots, p$:
\begin{equation}
\boldsymbol y_j = \alpha_j + \sum_{k = 1}^K \beta_{j,k} f_k + \boldsymbol e_j,
\end{equation}
where $\boldsymbol y_j$ and $\boldsymbol f_k$ are the monthly asset returns of asset $j$, factor returns of factor $k$ respectively, $\beta_{j,k}$ are the individual stock sensitivities to the factors, and $\alpha_j + e_j$ represent the idiosyncratic component of each stock.
We start with two specifications that correspond to two tables. Table 1 corresponds to 1 factor: excess return of the market portfolio, hence $K=1$, and Table 2 corresponds to 3 factors from the Fama \& French three factors, $K=3$. \footnote{The factors are book-to-market, market capitalization, and the excess return of the market portfolio.} Let $\boldsymbol\mu_{f}$ and $\boldsymbol\Sigma_f$ be the factors' sample mean and covariance matrix. The $\beta$, and $\alpha$ and covariance matrix of residuals: $\widetilde{\boldsymbol\Sigma}_n$ are estimated using a simple least-squares regression using returns from the S\&P500 stocks that were part of the index in the entire period from 2008 to 2017. In each simulation, we randomly select $p$ stocks from the pool with replacement because our simulations require more than the total number of available stocks. We then used the selected stocks to generate individual returns with covariance matrix of errors: $\widehat{\boldsymbol\Sigma}_n = \widetilde{\boldsymbol\Sigma}_n \odot Toeplitz ({\rho})$, where $Toeplitz ({\rho})$ is the $p \times p$ matrix of the form, for (i,j)th element
\[
Toeplitz ({\rho})_{i,j}:= \rho^{|i-j|},
\]
with $\rho=0.25, 0.5, 0.75$. $\boldsymbol A \odot \boldsymbol B $ represents element by element multiplication (Hadamard product) of two square matrices $\boldsymbol A, \boldsymbol B$ of the same dimensions.
Tables 1-2 show the results. The values in each cell show the average absolute estimation error for estimating the square of the Sharpe Ratio. Each eight-column block in the table shows the results for a different sample size. In each of these blocks, the first four columns are for $p = n/2$, and the last four columns are for $p = 3n/2$. MSR, MSR-OOS, GMV-SR, and MKW-SR are the constrained maximum Sharpe Ratio, the out-of-sample maximum Sharpe Ratio, the Sharpe Ratio from the global minimum-variance portfolio, and the Sharpe Ratio from the Markowitz portfolio with target returns set to 1\%, respectively. Therefore, there are four categories to evaluate the different estimates. The MAXSER risk constraint was set to 0.04 following \cite{ao2019}. We ran 100 iterations in each simulation setup.
All bold-face entries in tables show category champions.
Both Tables show that our method achieves consistency, as shown in Theorems. Analyzing $K=3$, Table 2, with $\rho=0.50$ OOS-MSR (the Out Of Sample-Maximum Sharpe Ratio), and Generalized Information Criterion tuning parameter selection, the estimation error at $p=n/2$, with $n=100$ is 1.244, and this error declines to 0.585 at $p=n/2, n=200$, and then declines to 0.321 at $p=n/2, n =400$. So with jointly increasing $n,p$ we show that the error declines, as predicted by our theorems. The main reason is that errors grow with $\sqrt{\ln(p)}$, but decline with $n^{1/2}$ rate. So the number of assets in a large portfolio only affects the error logarithmically. To give another example from Table 2, with $\rho=0.50$, GMV-SR (Global Minimum Variance-Sharpe Ratio) and Cross Validation tuning parameter selection with our method, the estimation error is 0.352 with $p= 3n/2, n=100$, then this error declines to 0.213 with $p=3n/2, n=200$, and further declines to 0.143 with $p= 3n/2, n=400$.
Next, we consider which method achieves the smallest estimation error. Table 1 favors SF-NL-LW (Single Factor Non-Linear Shrinkage of Ledoit-Wolf) since it has a single factor built into this subset of their technique. We get better results in Table 2 ($K=3$) for our methods. We have 4 categories: MSR, OOS-MSR, GMV-SR, MKW-SR corresponding to our Theorems 3-9.
There are nine possibilities in each category (given we are either at $p=n/2$ or $p=3n/2$), representing three choices of sample sizes paired with 3 choices of different Toeplitz structures.
We analyze each category. We start with Table 1. With $p=3n/2$ in OOS-MSR our NW-GIC method has the smallest errors 8 out of 9 categories. When $p=n/2$, MAXSER method dominates all others since it is specifically factor model designed to handle OOS-MSR with $p<n$. In GMV-SR, with $p=n/2$, in 3 out of 9 cases, our NW-GIC dominates. In the other categories in Table 1, non-linear shrinkage method of Ledoit-Wolf (2017) does the best, but our methods come a very close second.
In Table 2, with $K=3$, our methods perform better than in Table 1. In the category of GMV-SR, with $p=3n/2$, out of 9 possible configurations, our methods have the smallest error in 7 cases.
Our methods dominate in the same category, with $p=0.5n$, 5 out of 9 possibilities. In the case of the category of MKW-SR (Markowitz-Sharpe Ratio), our theorems predict that our methods may suffer from a number of factors. We see that non-linear shrinkage methods are the best, and our methods are the second best in this category. In the constrained maximum Sharpe Ratio, (MSR) non-linear shrinkage methods perform the best.
\begin{landscape}
\begin{table}[]\label{tab1}
\caption{Simulation Results -- Single Factor Toeplitz DGP with Real Factors}
\label{tab:sim-toe-sfact-real}
\begin{adjustbox}{max width=1.45\textwidth}
\begin{threeparttable}
\begin{tabular}{lccccccccccccccccccccccccccccc}
\hline
& & & & & & & & & & & \multicolumn{9}{c}{{\ul \textbf{Toeplitz $\rho = 0.25$}}} & & & & & & & & & & \\
& \multicolumn{9}{c}{n = 100} & & \multicolumn{9}{c}{n = 200} & & \multicolumn{9}{c}{n = 400} \\ \cline{2-10} \cline{12-20} \cline{22-30}
& \multicolumn{4}{c}{p = n/2} & & \multicolumn{4}{c}{p = 3n/2} & & \multicolumn{4}{c}{p = n/2} & & \multicolumn{4}{c}{p = 3n/2} & & \multicolumn{4}{c}{p = n/2} & & \multicolumn{4}{c}{p=3n/2} \\ \cline{2-5} \cline{7-10} \cline{12-15} \cline{17-20} \cline{22-25} \cline{27-30}
& MSR & OOS-MSR & GMV-SR & MKW-SR & & MSR & OOS-MSR & GMV-SR & MKW-SR & & MSR & OOS-MSR & GMV-SR & MKW-SR & & MSR & OOS-MSR & GMV-SR & MKW-SR & & MSR & OOS-MSR & GMV-SR & MKW-SR & & MSR & OOS-MSR & GMV-SR & MKW-SR \\ \cline{2-5} \cline{7-10} \cline{12-15} \cline{17-20} \cline{22-25} \cline{27-30}
NW-GIC & 0.517 & 1.062 & 0.680 & 0.114 & & 0.520 & 1.095 & 0.251 & 0.126 & & 0.331 & 0.515 & \textbf{0.211} & 0.073 & & 0.352 & \textbf{0.545} & 0.140 & 0.079 & & 0.208 & 0.262 & \textbf{0.091} & \textbf{0.051} & & 0.219 & \textbf{0.273} & 0.077 & 0.053 \\
NW-CV & 0.517 & 1.061 & 0.634 & 0.112 & & 0.521 & 1.099 & 0.251 & 0.128 & & 0.331 & 0.514 & 0.212 & \textbf{0.072} & & 0.352 & 0.546 & 0.140 & 0.079 & & 0.208 & 0.262 & 0.091 & 0.051 & & 0.219 & 0.273 & 0.078 & 0.053 \\
POET & 0.526 & 1.055 & 0.636 & 0.167 & & 0.522 & 1.095 & 0.260 & 0.144 & & 0.336 & 0.511 & 0.212 & 0.102 & & 0.354 & 0.548 & 0.145 & 0.089 & & 0.212 & 0.263 & 0.096 & 0.066 & & 0.220 & 0.276 & 0.081 & 0.058 \\
NL-LW & \textbf{0.487} & 1.705 & \textbf{0.559} & 0.172 & & \textbf{0.480} & 2.249 & 0.377 & 0.333 & & \textbf{0.301} & 0.961 & 0.322 & 0.216 & & \textbf{0.300} & 1.350 & 0.329 & 0.391 & & \textbf{0.169} & 0.645 & 0.265 & 0.258 & & \textbf{0.163} & 0.931 & 0.333 & 0.416 \\
SF-NL-LW & 0.516 & 1.069 & 0.689 & \textbf{0.110} & & 0.517 & \textbf{1.094} & \textbf{0.249} & \textbf{0.120} & & 0.330 & 0.515 & 0.215 & 0.072 & & 0.350 & 0.545 & \textbf{0.139} & \textbf{0.076} & & 0.207 & 0.263 & 0.091 & 0.051 & & 0.217 & 0.273 & \textbf{0.076} & \textbf{0.051} \\
MAXSER & & \textbf{0.359} & & & & & & & & & & \textbf{0.152} & & & & & & & & & & \textbf{0.098} & & & & & & & \\
& & & & & & & & & & & & & & & & & & & & & & & & & & & & & \\
& & & & & & & & & & & \multicolumn{9}{c}{{\ul \textbf{Toeplitz $\rho = 0.5$}}} & & & & & & & & & & \\
NW-GIC & 0.525 & 1.067 & 0.829 & 0.134 & & 0.529 & \textbf{1.095} & \textbf{0.266} & 0.157 & & 0.342 & 0.521 & \textbf{0.220} & 0.100 & & 0.365 & \textbf{0.552} & 0.161 & 0.113 & & 0.222 & 0.271 & 0.108 & 0.083 & & 0.233 & \textbf{0.283} & 0.108 & 0.089 \\
NW-CV & 0.526 & 1.067 & 0.726 & 0.132 & & 0.531 & 1.099 & 0.267 & 0.159 & & 0.342 & 0.521 & 0.221 & 0.100 & & 0.365 & 0.553 & 0.161 & 0.113 & & 0.222 & 0.271 & 0.108 & 0.084 & & 0.233 & 0.283 & 0.108 & 0.089 \\
POET & 0.535 & 1.061 & 0.721 & 0.190 & & 0.531 & 1.096 & 0.279 & 0.175 & & 0.348 & 0.518 & 0.226 & 0.133 & & 0.367 & 0.556 & 0.167 & 0.124 & & 0.226 & 0.273 & 0.118 & 0.100 & & 0.235 & 0.286 & 0.113 & 0.095 \\
NL-LW & \textbf{0.495} & 1.694 & \textbf{0.558} & 0.151 & & \textbf{0.489} & 2.231 & 0.360 & 0.290 & & \textbf{0.306} & 0.954 & 0.316 & 0.192 & & \textbf{0.313} & 1.340 & 0.289 & 0.342 & & \textbf{0.175} & 0.641 & 0.233 & 0.231 & & \textbf{0.177} & 0.926 & 0.281 & 0.365 \\
SF-NL-LW & 0.523 & 1.076 & 0.819 & \textbf{0.130} & & 0.526 & 1.095 & 0.266 & \textbf{0.151} & & 0.340 & 0.523 & 0.225 & \textbf{0.099} & & 0.363 & 0.553 & \textbf{0.160} & \textbf{0.109} & & 0.220 & 0.273 & \textbf{0.106} & \textbf{0.082} & & 0.232 & 0.284 & \textbf{0.106} & \textbf{0.087} \\
MAXSER & & \textbf{0.363} & & & & & & & & & & \textbf{0.158} & & & & & & & & & & \textbf{0.091} & & & & & & & \\
& & & & & & & & & & & & & & & & & & & & & & & & & & & & & \\
& & & & & & & & & & & \multicolumn{9}{c}{{\ul \textbf{Toeplitz $\rho = 0.75$}}} & & & & & & & & & & \\
NW-GIC & 0.542 & 1.105 & 1.300 & 0.183 & & 0.549 & \textbf{1.131} & 0.318 & 0.227 & & 0.366 & 0.558 & 0.248 & 0.166 & & 0.390 & \textbf{0.593} & 0.224 & 0.189 & & 0.250 & 0.309 & 0.174 & 0.156 & & 0.264 & \textbf{0.326} & 0.193 & 0.170 \\
NW-CV & 0.542 & 1.108 & 1.131 & 0.183 & & 0.551 & 1.144 & 0.319 & 0.231 & & 0.366 & 0.560 & 0.248 & 0.167 & & 0.390 & 0.595 & 0.223 & 0.189 & & 0.250 & 0.311 & 0.174 & 0.157 & & 0.264 & 0.327 & 0.193 & 0.171 \\
POET & 0.553 & 1.104 & 1.086 & 0.248 & & 0.552 & 1.136 & 0.333 & 0.249 & & 0.373 & 0.561 & 0.261 & 0.204 & & 0.393 & 0.601 & 0.235 & 0.203 & & 0.256 & 0.319 & 0.192 & 0.179 & & 0.267 & 0.335 & 0.201 & 0.180 \\
NL-LW & \textbf{0.510} & 1.703 & \textbf{0.548} & \textbf{0.109} & & \textbf{0.510} & 2.215 & 0.334 & \textbf{0.187} & & \textbf{0.324} & 0.956 & 0.298 & \textbf{0.132} & & \textbf{0.337} & 1.339 & 0.226 & 0.232 & & \textbf{0.196} & 0.647 & 0.194 & 0.151 & & \textbf{0.202} & 0.931 & \textbf{0.186} & 0.252 \\
SF-NL-LW & 0.537 & 1.115 & 0.922 & 0.176 & & 0.545 & 1.136 & \textbf{0.315} & 0.220 & & 0.361 & 0.563 & \textbf{0.245} & 0.157 & & 0.387 & 0.597 & \textbf{0.218} & \textbf{0.183} & & 0.244 & 0.313 & \textbf{0.160} & \textbf{0.149} & & 0.261 & 0.330 & 0.187 & \textbf{0.167} \\
MAXSER & & \textbf{0.371} & & & & & & & & & & \textbf{0.169} & & & & & & & & & & \textbf{0.082} & & & & & & & \\ \hline
\end{tabular}
\begin{tablenotes}
\item The table shows the simulation results for the Toeplitz DGP. Each simulation was done with 100 iterations. We used sample sizes $n$ of 100, 200 and 400, and the number of stocks was either $n/2$ or $1.5n$ for the low-dimensional and the high-dimensional case, respectively. Each block of rows shows the results for a different value of $\rho$ in the Toeplitz DGP. The values in each cell show the average absolute estimation error for estimating the square of the Sharpe Ratio.
\end{tablenotes}
\end{threeparttable}
\end{adjustbox}
\end{table}
\end{landscape}
\begin{landscape}
\begin{table}[] \label{tab2}
\caption{Simulation Results -- 3 Factor Toeplitz DGP with Real Factors}
\label{tab:sim-toe-fact-real}
\begin{adjustbox}{max width=1.45\textwidth}
\begin{threeparttable}
\begin{tabular}{lccccccccccccccccccccccccccccc}
\hline
{\ul } & {\ul } & {\ul } & {\ul } & {\ul } & {\ul } & {\ul } & {\ul } & {\ul } & {\ul } & {\ul } & \multicolumn{9}{c}{{\ul Toeplitz $\rho = 0.25$}} & {\ul } & {\ul } & {\ul } & {\ul } & {\ul } & {\ul } & {\ul } & {\ul } & {\ul } & {\ul } \\
& \multicolumn{9}{c}{n = 100} & & \multicolumn{9}{c}{n = 200} & & \multicolumn{9}{c}{n = 400} \\ \cline{2-10} \cline{12-20} \cline{22-30}
& \multicolumn{4}{c}{p = n/2} & & \multicolumn{4}{c}{p = 3n/2} & & \multicolumn{4}{c}{p = n/2} & & \multicolumn{4}{c}{p=3n/2} & & \multicolumn{4}{c}{p = n/2} & & \multicolumn{4}{c}{p = 3n/2} \\ \cline{2-5} \cline{7-10} \cline{12-15} \cline{17-20} \cline{22-25} \cline{27-30}
& MSR & OOS-MSR & GMV-SR & MKW-SR & & MSR & OOS-MSR & GMV-SR & MKW-SR & & MSR & OOS-MSR & GMV-SR & MKW-SR & & MSR & OOS-MSR & GMV-SR & MKW-SR & & MSR & OOS-MSR & GMV-SR & MKW-SR & & MSR & OOS-MSR & GMV-SR & MKW-SR \\
NW-GIC & 0.544 & 1.237 & 0.739 & 0.152 & & 0.579 & \textbf{1.269} & 0.343 & 0.220 & & 0.369 & 0.578 & \textbf{0.235} & 0.117 & & 0.391 & \textbf{0.658} & 0.200 & 0.158 & & 0.242 & 0.311 & \textbf{0.148} & \textbf{0.085} & & 0.254 & \textbf{0.350} & \textbf{0.122} & 0.100 \\
NW-CV & 0.543 & 1.237 & \textbf{0.724} & 0.151 & & 0.580 & 1.309 & \textbf{0.341} & 0.224 & & 0.369 & 0.578 & 0.235 & 0.117 & & 0.391 & 0.658 & \textbf{0.199} & 0.158 & & 0.242 & 0.311 & 0.148 & 0.085 & & 0.254 & 0.350 & 0.122 & 0.100 \\
POET & 0.558 & 1.137 & 1.041 & 0.308 & & 0.590 & 1.641 & 0.463 & 0.391 & & 0.418 & 0.852 & 0.372 & 0.346 & & 0.445 & 2.078 & 0.459 & 0.411 & & 0.338 & 1.279 & 0.436 & 0.371 & & 0.351 & 3.682 & 0.484 & 0.410 \\
NL-LW & \textbf{0.511} & 1.726 & 0.906 & \textbf{0.110} & & \textbf{0.535} & 2.251 & 0.481 & \textbf{0.086} & & \textbf{0.348} & 0.988 & 0.289 & \textbf{0.084} & & \textbf{0.358} & 1.370 & 0.260 & \textbf{0.118} & & \textbf{0.216} & 0.642 & 0.206 & 0.092 & & \textbf{0.223} & 0.953 & 0.206 & 0.169 \\
SF-NL-LW & 0.540 & 1.168 & 0.736 & 0.199 & & 0.567 & 1.292 & 0.349 & 0.238 & & 0.366 & 0.581 & 0.235 & 0.152 & & 0.383 & 0.679 & 0.203 & 0.160 & & 0.240 & 0.321 & 0.157 & 0.090 & & 0.251 & 0.361 & 0.126 & \textbf{0.096} \\
MAXSER & & \textbf{0.375} & & & & & & & & & & \textbf{0.166} & & & & & & & & & & \textbf{0.081} & & & & & & & \\
& & & & & & & & & & & & & & & & & & & & & & & & & & & & & \\
{\ul } & {\ul } & {\ul } & {\ul } & {\ul } & {\ul } & {\ul } & {\ul } & {\ul } & {\ul } & {\ul } & \multicolumn{9}{c}{{\ul Toeplitz $\rho = 0.50$}} & {\ul } & {\ul } & {\ul } & {\ul } & {\ul } & {\ul } & {\ul } & {\ul } & {\ul } & {\ul } \\
NW-GIC & 0.548 & 1.247 & 0.760 & 0.163 & & 0.584 & \textbf{1.273} & 0.352 & 0.232 & & 0.376 & 0.585 & 0.239 & 0.132 & & 0.398 & \textbf{0.666} & 0.213 & 0.172 & & 0.251 & 0.321 & \textbf{0.159} & 0.100 & & 0.263 & \textbf{0.360} & \textbf{0.143} & 0.115 \\
NW-CV & 0.548 & 1.248 & \textbf{0.743} & 0.162 & & 0.585 & 1.314 & \textbf{0.351} & 0.236 & & 0.376 & 0.585 & 0.239 & 0.132 & & 0.398 & 0.666 & \textbf{0.212} & 0.173 & & 0.251 & 0.321 & 0.160 & 0.100 & & 0.263 & 0.360 & 0.143 & 0.115 \\
POET & 0.563 & 1.148 & 1.203 & 0.320 & & 0.595 & 1.642 & 0.471 & 0.402 & & 0.424 & 0.855 & 0.382 & 0.358 & & 0.451 & 2.065 & 0.472 & 0.422 & & 0.346 & 1.276 & 0.451 & 0.383 & & 0.359 & 3.648 & 0.498 & 0.420 \\
NL-LW & \textbf{0.516} & 1.732 & 0.903 & \textbf{0.107} & & \textbf{0.540} & 2.244 & 0.475 & \textbf{0.079} & & \textbf{0.349} & 0.990 & 0.293 & \textbf{0.077} & & \textbf{0.366} & 1.372 & 0.253 & \textbf{0.101} & & \textbf{0.226} & 0.648 & 0.206 & \textbf{0.083} & & \textbf{0.229} & 0.955 & 0.185 & 0.151 \\
SF-NL-LW & 0.543 & 1.181 & 0.753 & 0.206 & & 0.572 & 1.297 & 0.358 & 0.249 & & 0.371 & 0.589 & \textbf{0.237} & 0.160 & & 0.390 & 0.687 & 0.214 & 0.173 & & 0.248 & 0.331 & 0.167 & 0.102 & & 0.260 & 0.371 & 0.146 & \textbf{0.110} \\
MAXSER & & \textbf{0.379} & & & & & & & & & & \textbf{0.168} & & & & & & & & & & \textbf{0.080} & & & & & & & \\
& & & & & & & & & & & & & & & & & & & & & & & & & & & & & \\
{\ul } & {\ul } & {\ul } & {\ul } & {\ul } & {\ul } & {\ul } & {\ul } & {\ul } & {\ul } & {\ul } & \multicolumn{9}{c}{{\ul Toeplitz $\rho = 0.75$}} & {\ul } & {\ul } & {\ul } & {\ul } & {\ul } & {\ul } & {\ul } & {\ul } & {\ul } & {\ul } \\
NW-GIC & 0.556 & 1.295 & 0.812 & 0.189 & & 0.593 & \textbf{1.318} & \textbf{0.377} & 0.263 & & 0.387 & 0.623 & 0.260 & 0.164 & & 0.411 & \textbf{0.708} & 0.246 & 0.206 & & 0.267 & 0.360 & 0.198 & 0.136 & & 0.281 & \textbf{0.404} & 0.192 & 0.150 \\
NW-CV & 0.556 & 1.301 & 0.802 & 0.187 & & 0.595 & 1.340 & 0.383 & 0.268 & & 0.388 & 0.624 & 0.260 & 0.165 & & 0.412 & 0.709 & 0.246 & 0.207 & & 0.268 & 0.362 & 0.198 & 0.137 & & 0.281 & 0.405 & 0.192 & 0.150 \\
POET & 0.571 & 1.200 & 2.071 & 0.347 & & 0.605 & 1.679 & 0.490 & 0.428 & & 0.436 & 0.886 & 0.409 & 0.385 & & 0.464 & 2.068 & 0.501 & 0.448 & & 0.361 & 1.292 & 0.483 & 0.410 & & 0.375 & 3.607 & 0.529 & 0.445 \\
NL-LW & \textbf{0.523} & 1.727 & 0.947 & \textbf{0.106} & & \textbf{0.550} & 2.257 & 0.468 & \textbf{0.071} & & \textbf{0.358} & 1.009 & 0.287 & \textbf{0.066} & & \textbf{0.377} & 1.397 & 0.242 & \textbf{0.068} & & \textbf{0.234} & 0.659 & \textbf{0.195} & \textbf{0.062} & & \textbf{0.242} & 0.979 & \textbf{0.155} & \textbf{0.109} \\
SF-NL-LW & 0.549 & 1.228 & \textbf{0.783} & 0.222 & & 0.581 & 1.346 & 0.382 & 0.272 & & 0.381 & 0.631 & \textbf{0.246} & 0.179 & & 0.402 & 0.731 & \textbf{0.241} & 0.200 & & 0.260 & 0.371 & 0.195 & 0.127 & & 0.275 & 0.417 & 0.188 & 0.141 \\
MAXSER & & \textbf{0.385} & & & & & & & & & & \textbf{0.173} & & & & & & & & & & \textbf{0.079} & & & & & & & \\ \hline
\end{tabular}
\begin{tablenotes}
\item The table shows the simulation results for the Toeplitz DGP. Each simulation was done with 100 iterations. We used sample sizes $n$ of 100, 200 and 400, and the number of stocks was either $n/2$ or $1.5n$ for the low-dimensional and the high-dimensional case, respectively. Each block of rows shows the results for a different value of $\rho$ in the Toeplitz DGP. The values in each cell show the average absolute estimation error for estimating the square of the Sharpe Ratio. \end{tablenotes}
\end{threeparttable}
\end{adjustbox}
\end{table}
\end{landscape}
\section{Empirical Application}\label{emp}
For the empirical application, we use two subsamples. The first subsample uses data from January 1995 to December 2019 with an out-of-sample period from January 2005 to December 2019. We selected all stocks in the S\&P 500 index for at least one month in the out-of-sample period and have data for the entire 1995-2019 period resulting in 382 stocks. The second subsample starts in January 1990 and ends in December 2019 with an out-of-sample period from January 2000 to December 2019. Using the same criterion as the first subsample, the number of stocks was 321, which is around 15\% fewer than the first subsample. The objective is to have an out-of-sample competition between models, and we only estimated GMV and Markowitz portfolios for the plug-in estimators.
The first out-of-sample period includes only the recession of 2008. The second out-of-sample period includes the recessions of 2000 and 2008, and the out-of-sample periods reflect recent history.
The Markowitz return constraint $\rho_1$ is 0.8\% per month, and the MAXSER risk constraint is 4\%. In the low-dimensional experiment, we randomly select 50 stocks from the pool to estimate the models with the same stocks for all windows. We also experimented with 25 stocks but did not report them. That table is available from the authors on demand. In the high-dimensional case, we use all available stocks.
We use a rolling window setup for the out-of-sample estimation of the Sharpe Ratio following \citet{caner2019}. Specifically, samples of size $n$ are divided into in-sample $(1:n_{I})$ and out-of-sample $(n_{I}+1:n)$. We start by estimating the portfolio $\widehat{\boldsymbol w}_{n_I}$ in the in-sample period and the out-of-sample portfolio returns $\widehat{\boldsymbol w}'_{n_I}y_{n_{I+1}}$. Then, we roll the window by one element $(2:n_{I}+1)$ and form a new in-sample portfolio $\widehat{\boldsymbol w}_{n_{I+1}}$ and out-of-sample portfolio returns $\widehat{\boldsymbol w}'_{n_{I+1}}y_{n_{I+2}}$. This procedure is repeated until the end of the sample.
The out-of-sample average return and variance without transaction costs are
\begin{equation*}
\widehat{\boldsymbol \mu}_{os} = \frac{1}{n-n_{I}} \sum_{t=n_I}^{n-1} {\widehat{\boldsymbol w}'_t}\boldsymbol y_{t+1},\quad \widehat{\boldsymbol\Sigma}_{y,os}^2 = \frac{1}{n-n_{I}-1} \sum_{t=n_I}^{n-1} ({\widehat{\boldsymbol w}'_t}\boldsymbol y_{t+1} - \widehat{\boldsymbol\mu}_{os})^2.
\end{equation*}
We estimate the Sharpe Ratios with and without transaction costs. The transaction cost, $c$, is defined as 50 basis points following \citet{demiguel2007optimal}. Let $y_{P,t+1} = {\widehat{\boldsymbol w}'_t}\boldsymbol y_{t+1}$ be the return of the portfolio in period $t+1$; in the presence of transaction costs, the returns will be defined as
\begin{equation*}
y_{P,t+1}^{Net} = y_{P,t+1} - c(1+y_{P,t+1})\sum_{j=1}^p|{\widehat{w}_{t+1,j}}-{\widehat{w}^+_{t,j}}|,
\end{equation*}
where ${\widehat{w}^+_{t,j}} = {\widehat{w}_{t,j}}(1+y_{t+1,j})/(1+y_{t+1,P})$ and $y_{t,j}$ and $y_{t,P}$ are the excess returns of asset $j$ and the portfolio $P$ added to the risk-free rate. The adjustment made in ${\widehat{w}^+_{t,j}}$ is because the portfolio at the end of the period has changed compared to the portfolio at the beginning of the period.
The Sharpe Ratio is calculated from the average return and the variance of the portfolio in the out-of-sample period
\begin{equation*}
SR = \frac{\widehat{\boldsymbol\mu}_{os}}{\widehat{\boldsymbol\Sigma}_{y,os}}.
\end{equation*}
The portfolio returns are replaced by the returns with transaction costs when we calculate the Sharpe Ratio with transaction costs.
We use the same test as \citet{ao2019} to compare the models. Specifically,
\begin{equation}
\label{eq:sharpe}
H_0 ~ : ~ SR_{NW}\leq SR_0 ~~ vs ~~ H_a ~ : ~ SR_{NW}> SR_0,
\end{equation}
where $SR_{NW}$ is the Sharpe Ratio of our feasible nodewise model, which is tested against all remaining models. This is the \citet{jobson1981performance} test with \citet{memmel2003performance} correction.
We also considered the method of \citet{ledoit2008robust} for testing the significance of the winner and using the equally weighted portfolio as a benchmark; the results were very similar and hence are not reported.
We also include an equally weighted portfolio (EW). GMV-NW-GIC and GMV-NW-CV denote the nodewise method with GIC and cross validation tuning parameter choices, respectively, in the global minimum-variance portfolio (GMV).
In each of our feasible nodewise models with GIC, CV, we either use a single-factor model (market as the only factor) or three-factor model. They are denoted GMV-NW-GIC-SF, GMV-NW-GIC-3F for the global minimum variance portfolio analyzed with feasible nodewise method and GIC criterion for tuning parameter choice and single and three-factor models, respectively. In the same way, we define GMV-NW-CV-SF, GMV-NW-CV-3F. We take GMV-NW-GIC-SF as the benchmark to test against all other methods since it generally does well in different preliminary forecasts.
GMV-POET, GMV-NL-LW, and GMV-SF-NL-LW denote the POET, nonlinear shrinkage, and single-factor nonlinear shrinkage methods, respectively, which are described in the simulation section and also used in the global minimum-variance portfolio. The MAXSER is also used and explained in the simulation section. MW denotes the Markowitz mean-variance portfolio, and MW-NW-GIC-SF denotes the feasible nodewise method with GIC tuning parameter selection in the Markowitz portfolio with a single factor. All the other methods with MW headers are analogous and thus self-explanatory.
The results are presented in Tables \ref{tab:emp_sp1} and \ref{tab:emp_sp2}. Table \ref{tab:emp_sp1} shows the results for the 2005-2019 out-of-sample period. Feasible nodewise methods do well in terms of the Sharpe Ratio in Table \ref{tab:emp_sp1}.
For example, with transaction costs in the low-dimensional portfolio category, in terms of Sharpe Ratio (SR) (averaged over the out-of-sample time period), GMV-NW-GIC-SF is the best model. It has an SR of 0.210. In the case of high dimensional case with transaction costs in the same table, GMV-POET and our GMV-NW-GIC-SF virtually tie (difference in favor of POET in fourth decimal) at 0.214 for the Sharpe Ratio.
If we were to analyze only the Markowitz portfolio in Table \ref{tab:emp_sp1}, with transaction costs in high dimensions, MW-NW-GIC-SF has the highest SR of 0.211. Therefore, even in other subcategories of Markowitz portfolio, the feasible nodewise method dominates. Although statistical significance is not established, it is unclear that these significance tests have high power in our high-dimensional cases.
Table \ref{tab:emp_sp2} shows the results for the out-of-sample January 2000-2019 subsample.
We see that feasible nodewise methods dominate all scenarios except for the low-dimensional case with transaction costs. In high dimensionality with transaction costs, GMV-NW-GIC-SF (Markowitz-nodewise-GIC) has an SR of 0.225, and the closest is GMV-POET with 0.204. Also, we experimented with two other out-sample periods of 2005-2017, 2000-2017, and the results are slightly better for our methods, and these can be shared on demand.
\begin{table}[htb]
\caption{Empirical Results -- Out-of-Sample Period from Jan. 2005 to Dec. 2019}
\label{tab:emp_sp1}
\begin{adjustbox}{max width=\textwidth}
\begin{threeparttable}
\begin{tabular}{lccccccccccccccccccc}
\hline
& \multicolumn{9}{c}{{\ul Without TC}} & {\ul } & \multicolumn{9}{c}{{\ul With TC}} \\
& \multicolumn{4}{c}{{\ul Low Dim.}} & {\ul } & \multicolumn{4}{c}{{\ul High Dim.}} & {\ul } & \multicolumn{4}{c}{{\ul Low Dim.}} & {\ul } & \multicolumn{4}{c}{{\ul High Dim.}} \\
& SR & AVG & SD & p-value & & SR & AVG & SD & p-value & & SR & AVG & SD & p-value & & SR & AVG & SD & p-value \\ \cline{1-5} \cline{7-10} \cline{12-15} \cline{17-20}
EW & 0.196 & 0.010 & 0.052 & 0.730 & & 0.197 & 0.010 & 0.048 & 0.644 & & 0.191 & 0.010 & 0.052 & 0.802 & & 0.191 & 0.009 & 0.048 & 0.792 \\
GMV-NW-GIC-SF & \textbf{0.229} & 0.008 & 0.036 & & & 0.236 & 0.008 & 0.032 & & & \textbf{0.210} & 0.008 & 0.036 & & & 0.214 & 0.007 & 0.032 & \\
GMV-NW-CV-SF & 0.226 & 0.008 & 0.036 & 0.590 & & 0.240 & 0.008 & 0.032 & 0.398 & & 0.203 & 0.007 & 0.036 & 0.132 & & 0.192 & 0.006 & 0.032 & 0.002 \\
GMV-NW-GIC-3F & 0.215 & 0.007 & 0.034 & 0.576 & & 0.214 & 0.007 & 0.033 & 0.520 & & 0.191 & 0.007 & 0.034 & 0.424 & & 0.183 & 0.006 & 0.033 & 0.398 \\
GMV-NW-CV-3F & 0.212 & 0.007 & 0.034 & 0.474 & & 0.226 & 0.007 & 0.032 & 0.790 & & 0.183 & 0.006 & 0.034 & 0.278 & & 0.132 & 0.004 & 0.032 & 0.032 \\
GMV-POET & 0.218 & 0.007 & 0.034 & 0.682 & & 0.232 & 0.007 & 0.030 & 0.914 & & 0.203 & 0.007 & 0.034 & 0.822 & & \textbf{0.214} & 0.006 & 0.030 & 0.996 \\
GMV-NL-LW & 0.236 & 0.008 & 0.034 & 0.834 & & 0.236 & 0.007 & 0.030 & 0.998 & & 0.205 & 0.007 & 0.034 & 0.908 & & 0.179 & 0.005 & 0.031 & 0.490 \\
GMV-SF-NL-LW & 0.216 & 0.007 & 0.034 & 0.684 & & \textbf{0.245} & 0.007 & 0.030 & 0.886 & & 0.190 & 0.007 & 0.034 & 0.546 & & 0.184 & 0.006 & 0.030 & 0.600 \\
MW-NW-GIC-SF & 0.229 & 0.008 & 0.034 & 0.970 & & 0.236 & 0.008 & 0.032 & 0.966 & & 0.205 & 0.007 & 0.034 & 0.786 & & 0.211 & 0.007 & 0.032 & 0.706 \\
MW-NW-CV-SF & 0.228 & 0.008 & 0.034 & 0.942 & & 0.242 & 0.008 & 0.032 & 0.620 & & 0.197 & 0.007 & 0.034 & 0.482 & & 0.190 & 0.006 & 0.032 & 0.056 \\
MW-NW-GIC-3F & 0.214 & 0.007 & 0.033 & 0.628 & & 0.217 & 0.007 & 0.033 & 0.606 & & 0.185 & 0.006 & 0.033 & 0.444 & & 0.183 & 0.006 & 0.033 & 0.416 \\
MW-NW-CV-3F & 0.212 & 0.007 & 0.033 & 0.574 & & 0.225 & 0.007 & 0.032 & 0.790 & & 0.177 & 0.006 & 0.033 & 0.302 & & 0.125 & 0.004 & 0.032 & 0.032 \\
MW-POET & 0.223 & 0.007 & 0.032 & 0.880 & & 0.229 & 0.007 & 0.030 & 0.844 & & 0.200 & 0.006 & 0.032 & 0.794 & & 0.207 & 0.006 & 0.030 & 0.840 \\
MW-NL-LW & 0.220 & 0.008 & 0.034 & 0.860 & & 0.235 & 0.007 & 0.030 & 0.980 & & 0.186 & 0.006 & 0.034 & 0.636 & & 0.177 & 0.005 & 0.030 & 0.540 \\
MW-SF-NL-LW & 0.204 & 0.007 & 0.034 & 0.574 & & 0.241 & 0.007 & 0.030 & 0.920 & & 0.175 & 0.006 & 0.034 & 0.482 & & 0.180 & 0.005 & 0.030 & 0.554 \\
MAXSER & 0.161 & 0.010 & 0.065 & 0.510 & & & & & & & 0.024 & 0.002 & 0.066 & 0.116 & & & & & \\ \hline
\end{tabular}
\begin{tablenotes}
\item The table shows the Sharpe Ratio (SR), average returns (Avg), standard deviation (SD) and p-value of the \citet{jobson1981performance} test with \citet{memmel2003performance} correction. We also applied the
\citet{ledoit2008robust} test with circular bootstrap, and the results were very similar; therefore we only report those of the first test in this table. The statistics were calculated from 180 rolling windows covering the period from Jan. 2005 to Dec. 2019, and the size of the estimation window was 120 observations.
\end{tablenotes}
\end{threeparttable}
\end{adjustbox}
\end{table}
\begin{table}[htb]
\caption{Empirical Results -- Out-of-Sample Period from Jan. 2000 to Dec. 2019}
\label{tab:emp_sp2}
\begin{adjustbox}{max width=\textwidth}
\begin{threeparttable}
\begin{tabular}{lccccccccccccccccccc}
\hline
{\ul } & \multicolumn{9}{c}{{\ul Without TC}} & {\ul } & \multicolumn{9}{c}{{\ul With TC}} \\
{\ul } & \multicolumn{4}{c}{{\ul Low Dim.}} & {\ul } & \multicolumn{4}{c}{{\ul High Dim.}} & {\ul } & \multicolumn{4}{c}{{\ul Low Dim.}} & {\ul } & \multicolumn{4}{c}{{\ul High Dim.}} \\ \cline{1-5} \cline{7-10} \cline{12-15} \cline{17-20}
{\ul } & SR & AVG & SD & p-value & & SR & AVG & SD & p-value & & SR & AVG & SD & p-value & & SR & AVG & SD & p-value \\ \cline{1-5} \cline{7-10} \cline{12-15} \cline{17-20}
EW & 0.201 & 0.010 & 0.047 & 0.874 & & 0.210 & 0.010 & 0.047 & 0.546 & & \textbf{0.195} & 0.009 & 0.047 & 0.998 & & 0.203 & 0.010 & 0.047 & 0.758 \\
GMV-NW-GIC-SF & \textbf{0.213} & 0.008 & 0.035 & & & 0.245 & 0.008 & 0.034 & & & 0.195 & 0.007 & 0.035 & & & \textbf{0.225} & 0.008 & 0.034 & \\
GMV-NW-CV-SF & 0.212 & 0.008 & 0.036 & 0.940 & & 0.249 & 0.008 & 0.034 & 0.374 & & 0.191 & 0.007 & 0.036 & 0.454 & & 0.206 & 0.007 & 0.033 & 0.006 \\
GMV-NW-GIC-3F & 0.193 & 0.007 & 0.034 & 0.424 & & 0.224 & 0.007 & 0.031 & 0.498 & & 0.171 & 0.006 & 0.034 & 0.382 & & 0.192 & 0.006 & 0.032 & 0.260 \\
GMV-NW-CV-3F & 0.188 & 0.006 & 0.034 & 0.348 & & 0.231 & 0.007 & 0.031 & 0.700 & & 0.161 & 0.006 & 0.034 & 0.196 & & 0.139 & 0.004 & 0.031 & 0.016 \\
GMV-POET & 0.185 & 0.006 & 0.033 & 0.282 & & 0.222 & 0.007 & 0.032 & 0.416 & & 0.169 & 0.006 & 0.033 & 0.316 & & 0.204 & 0.007 & 0.032 & 0.430 \\
GMV-NL-LW & 0.160 & 0.006 & 0.035 & 0.172 & & 0.232 & 0.007 & 0.029 & 0.838 & & 0.131 & 0.005 & 0.035 & 0.120 & & 0.175 & 0.005 & 0.029 & 0.398 \\
GMV-SF-NL-LW & 0.172 & 0.006 & 0.034 & 0.252 & & 0.242 & 0.007 & 0.028 & 0.934 & & 0.145 & 0.005 & 0.034 & 0.196 & & 0.184 & 0.005 & 0.028 & 0.398 \\
MW-NW-GIC-SF & 0.211 & 0.007 & 0.034 & 0.872 & & 0.243 & 0.008 & 0.032 & 0.868 & & 0.189 & 0.006 & 0.034 & 0.644 & & 0.219 & 0.007 & 0.032 & 0.602 \\
MW-NW-CV-SF & 0.210 & 0.007 & 0.034 & 0.834 & & \textbf{0.249} & 0.008 & 0.032 & 0.656 & & 0.185 & 0.006 & 0.034 & 0.504 & & 0.202 & 0.006 & 0.032 & 0.028 \\
MW-NW-GIC-3F & 0.191 & 0.006 & 0.034 & 0.442 & & 0.226 & 0.007 & 0.031 & 0.584 & & 0.165 & 0.006 & 0.034 & 0.338 & & 0.190 & 0.006 & 0.031 & 0.326 \\
MW-NW-CV-3F & 0.184 & 0.006 & 0.034 & 0.324 & & 0.228 & 0.007 & 0.031 & 0.652 & & 0.153 & 0.005 & 0.034 & 0.162 & & 0.132 & 0.004 & 0.031 & 0.038 \\
MW-POET & 0.181 & 0.006 & 0.032 & 0.282 & & 0.216 & 0.007 & 0.031 & 0.408 & & 0.161 & 0.005 & 0.033 & 0.240 & & 0.195 & 0.006 & 0.031 & 0.402 \\
MW-NL-LW & 0.151 & 0.005 & 0.036 & 0.172 & & 0.229 & 0.007 & 0.029 & 0.782 & & 0.120 & 0.004 & 0.036 & 0.092 & & 0.172 & 0.005 & 0.029 & 0.352 \\
MW-SF-NL-LW & 0.161 & 0.006 & 0.035 & 0.248 & & 0.237 & 0.007 & 0.028 & 0.886 & & 0.131 & 0.005 & 0.035 & 0.152 & & 0.178 & 0.005 & 0.028 & 0.398 \\
MAXSER & 0.040 & 0.004 & 0.088 & 0.294 & & & & & & & -0.039 & -0.004 & 0.099 & 0.364 & & & & & \\ \hline
\end{tabular}
\begin{tablenotes}
\item The table shows the Sharpe Ratio (SR), average returns (Avg), standard deviation (SD) and p-value of the \citet{jobson1981performance} test with \citet{memmel2003performance} correction. We also applied the
\citet{ledoit2008robust} test with circular bootstrap, and the results were very similar; therefore we only report those of the first test in this table. The statistics were calculated from 240 rolling windows covering the period from Jan. 2005 to Dec. 2019, and the size of the estimation window was 120 observations.
\end{tablenotes}
\end{threeparttable}
\end{adjustbox}
\end{table}
In Table \ref{tab:stat}, we analyze turnover, leverage and maximum leverage (equations (\ref{eq:tn}), (\ref{eq:l}) and (\ref{eq:maxl}), respectively) of the portfolios in Tables \ref{tab:emp_sp1}-\ref{tab:emp_sp2}.
The definitions are as follows for turnover:
\begin{equation}
\label{eq:tn}
\mbox{turnover} = \sum_{j=1}^p|{\widehat{w}_{t+1,j}}-{\widehat{w}^+_{t,j}}|,
\end{equation}
\noindent and leverage
\begin{equation}
\label{eq:l}
\mbox{leverage} = \left|\sum_{j = 1}^p \min\{\widehat{w}_{t+1,j},0\}\right|,
\end{equation}
\noindent and maximum leverage
\begin{equation}
\label{eq:maxl}
\mbox{max leverage} = \max_j\{\left| \min\{\widehat{w}_{t+1,j},0\} \right|\}.
\end{equation}
It is clear that in Table \ref{tab:stat} in terms of turnover, leverage, maximum leverage, GMV-POET and GMV-NW-GIC-SF do well, with the best and close to best respectively if we discount EW portfolios.
\subsection{Time Series of Sharpe Ratios and Turnover}
Figures \ref{fig:TS_LD} and \ref{fig:TS_HD} shows Global Minimum Variance results of the NW-GIF-SF, the POET and the SF-NL-LW models with transaction costs. The results were obtained through a 24 months rolling window with the out-of-sample returns from the 2000-2019 experiment, which yields time-series that start in 2002 and end in 2019 for the Sharpe Ratio and the turnover. The main conclusion from the figures is that Nodewise works better in terms of the Sharpe Ratio in deep recessions like the 2008 crisis, but Nonlinear Shrinkage and POET are \textnormal{sup}erior when we have long periods of normality in the markets. Nodewise also delivers better Sharpe Ratios during the recovery of the crisis. On the turnover side, Nodewise and POET consistently have lower turnover than Nonlinear Shrinkage with POET being the overall lowest. However, during the 2008 crisis, especially in the high dimension setup, POET had a higher turnover than Nodewise.
\newpage
\begin{table}[htb]
\centering
\caption{Turnover and Leverage}
\label{tab:stat}
\begin{adjustbox}{max width=0.65\textwidth}
\begin{threeparttable}
\begin{tabular}{lccccccc}
\hline
& \multicolumn{7}{c}{{\ul \textbf{2005-2019 Subsample}}} \\
& \multicolumn{3}{c}{{\ul Low Dimension}} & & \multicolumn{3}{c}{{\ul High Dimension}} \\
& Turnover & Leverage & Max Leverage & & Turnover & Leverage & Max Leverage \\ \cline{2-4} \cline{6-8}
EW & 0.053 & 0.000 & 0.000 & & 0.054 & 0.000 & 0.000 \\
GMV-NW-GIC-SF & 0.125 & 0.312 & 0.042 & & 0.130 & 0.376 & 0.009 \\
GMV-NW-CV-SF & 0.160 & 0.311 & 0.040 & & 0.302 & 0.395 & 0.014 \\
GMV-NW-GIC-3F & 0.148 & 0.380 & 0.048 & & 0.186 & 0.528 & 0.013 \\
GMV-NW-CV-3F & 0.190 & 0.382 & 0.049 & & 0.593 & 0.567 & 0.030 \\
GMV-POET & 0.096 & 0.288 & 0.043 & & 0.096 & 0.299 & 0.007 \\
GMV-NL-LW & 0.198 & 0.420 & 0.057 & & 0.325 & 0.807 & 0.024 \\
GMV-SF-NL-LW & 0.163 & 0.383 & 0.050 & & 0.341 & 0.904 & 0.025 \\
MW-NW-GIC-SF & 0.154 & 0.331 & 0.046 & & 0.150 & 0.382 & 0.009 \\
MW-NW-CV-SF & 0.191 & 0.329 & 0.044 & & 0.322 & 0.402 & 0.014 \\
MW-NW-GIC-3F & 0.179 & 0.401 & 0.052 & & 0.207 & 0.539 & 0.013 \\
MW-NW-CV-3F & 0.220 & 0.401 & 0.051 & & 0.626 & 0.582 & 0.030 \\
MW-POET & 0.128 & 0.306 & 0.046 & & 0.117 & 0.307 & 0.008 \\
MW-NL-LW & 0.220 & 0.440 & 0.064 & & 0.327 & 0.814 & 0.024 \\
MW-SF-NL-LW & 0.184 & 0.400 & 0.052 & & 0.344 & 0.912 & 0.025 \\
MAXSER & 1.766 & 0.421 & 0.200 & & & & \\
& & & & & & & \\
& \multicolumn{7}{c}{{\ul \textbf{2000-2019 Sub Sample}}} \\ \cline{2-4} \cline{6-8}
EW & 0.056 & 0.000 & 0.000 & & 0.056 & 0.000 & 0.000 \\
GMV-NW-GIC-SF & 0.120 & 0.283 & 0.049 & & 0.127 & 0.342 & 0.011 \\
GMV-NW-CV-SF & 0.142 & 0.279 & 0.048 & & 0.278 & 0.361 & 0.014 \\
GMV-NW-GIC-3F & 0.144 & 0.355 & 0.053 & & 0.192 & 0.541 & 0.016 \\
GMV-NW-CV-3F & 0.181 & 0.353 & 0.053 & & 0.557 & 0.572 & 0.030 \\
GMV-POET & 0.097 & 0.290 & 0.038 & & 0.107 & 0.322 & 0.009 \\
GMV-NL-LW & 0.196 & 0.396 & 0.068 & & 0.311 & 0.782 & 0.027 \\
GMV-SF-NL-LW & 0.173 & 0.383 & 0.062 & & 0.310 & 0.849 & 0.026 \\
MW-NW-GIC-SF & 0.142 & 0.296 & 0.050 & & 0.148 & 0.351 & 0.011 \\
MW-NW-CV-SF & 0.165 & 0.292 & 0.048 & & 0.299 & 0.369 & 0.014 \\
MW-NW-GIC-3F & 0.165 & 0.368 & 0.054 & & 0.209 & 0.548 & 0.016 \\
MW-NW-CV-3F & 0.203 & 0.364 & 0.054 & & 0.582 & 0.581 & 0.030 \\
MW-POET & 0.121 & 0.301 & 0.041 & & 0.126 & 0.333 & 0.009 \\
MW-NL-LW & 0.214 & 0.409 & 0.071 & & 0.313 & 0.787 & 0.027 \\
MW-SF-NL-LW & 0.197 & 0.395 & 0.067 & & 0.314 & 0.855 & 0.025 \\
MAXSER & 1.860 & 0.371 & 0.201 & & & & \\ \hline
\end{tabular}
\begin{tablenotes}
\item The table shows the average turnover, average leverage and average max leverage for all portfolios across all out-of-sample windows. The top panel shows the results for the 2000-2019 out-of-sample period, and the second panel shows the results for the 2005-2019 out-of-sample period.
\end{tablenotes}
\end{threeparttable}
\end{adjustbox}
\end{table}
\newpage
\begin{figure}
\centering
\includegraphics[width=0.8\textwidth]{TS_LD.eps}
\caption{24 months rolling Sharpe Ratio and turnover - Low Dimension with transaction costs}
\label{fig:TS_LD}
\end{figure}
\begin{figure}[ht]
\centering
\includegraphics[width=0.8\textwidth]{TS_HD.eps}
\caption{24 months rolling Sharpe Ratio and turnover - High Dimension with transaction costs}
\label{fig:TS_HD}
\end{figure}
\section{Conclusion}
We provide a hybrid factor model combined with nodewise regression method that can control for risk and obtain the maximum expected return of a large portfolio. Our result is novel and holds even when $p>n$. We allow for an increasing number of factors, with possible unbounded largest eigenvalue of the covariance matrix of errors. Sparsity is assumed on the precision matrix of errors rather than the covariance matrix of errors. We also show that the maximum out-of-sample Sharpe Ratio can be estimated consistently. Furthermore, we also develop a formula for the maximum Sharpe Ratio when the sum of the weights of the portfolio is one. A consistent estimate for the constrained case is also shown. Then, we extended our results to the consistent estimation of the Sharpe Ratios in two widely used portfolios in the literature. It will be essential to extend our results to more restrictions on portfolios.
\bibliographystyle{chicagoa}
\bibliography{maxsrsep14-mc}
\newpage