EconBase
← Back to paper

L2-Relaxation: With Applications to Forecast Combination and Portfolio Analysis

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

106,580 characters · 15 sections · 99 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

$_2$-Relaxation: With Applications to Forecast Combination and Portfolio Analysis

abstractThis paper tackles forecast combination with many forecasts or minimum variance portfolio selection with many assets. A novel convex problem called $\ell_{2}$-relaxation is proposed. In contrast to standard formulations, $\ell _{2}$-relaxation minimizes the squared Euclidean norm of the weight vector subject to a set of relaxed linear inequality constraints. The magnitude of relaxation, controlled by a tuning parameter, balances the bias and variance. When the variance-covariance (VC) matrix of the individual forecast errors or financial assets exhibits latent group structures --- a block equicorrelation matrix plus a VC for idiosyncratic noises, the solution to $\ell_{2}$-relaxation delivers roughly equal within-group weights. Optimality of the new method is established under the asymptotic framework when the number of the cross-sectional units $N$ potentially grows much faster than the time dimension $T$. Excellent finite sample performance of our method is demonstrated in Monte Carlo simulations. Its wide applicability is highlighted in three real data examples concerning empirical applications of microeconomics, macroeconomics, and finance. { } { Key Words: Forecast combination puzzle; high dimension; latent group; machine learning; portfolio analysis; optimization.} { JEL Classification: C22, C53, C55 }

\singlespacing

\singlespacing

{ We thank the editor and two referees for their valuable suggestions. We are also grateful to Bruce Hansen, Yingying Li, Esfandiar Maasoumi, Jack Porter, Yuying Sun, Aman Ullah, Xia Wang, Xinyu Zhang, and Xinghua Zheng for their helpful comments. Shi acknowledges financial support from Hong Kong Research Grants Council (RGC) No.14500118. Su gratefully acknowledges the support from Natural Science Foundation of China (No.72133002). Xie's research is supported by the Natural Science Foundation of China (No.72173075), the Shanghai Research Center for Data Science and Decision Technology, and the Fundamental Research Funds for the Central Universities. Address correspondence: Zhentao Shi: [email removed], School of Economics, Georgia Institute of Technology, 205 Old C.E. Building, 221 Bobby Dodd Way, Atlanta, GA 30332, U.S.A., and Department of Economics, 928 Esther Lee Building, the Chinese University of Hong Kong, Shatin, New Territories, Hong Kong SAR, China. Liangjun Su: [email removed], School of Economics and Management, Tsinghua University, Beijing, China. Tian Xie: [email removed], College of Business, Shanghai University of Finance and Economics, Shanghai, China.}

\singlespacing

Introduction

Forecast combination assigns weights to individual experts to reduce forecast errors, and portfolio management assigns weights to financial assets to reduce risk exposure. The classical approach to both problems, to be formally laid out in Section (ref), solves a quadratic optimization

equation[equation omitted — 217 chars of source]

where $N$ is the number of forecasts or assets, $\mathbf{1}_{N}$ is a column of $N$ ones, and $\widehat{\boldsymbol{\Sigma }}$ is a variance-covariance (VC) matrix estimate computed from $T$ time series observations. When $\widehat{\boldsymbol{\Sigma }}$ is invertible, the explicit solution to (ref) is

equation[equation omitted — 237 chars of source]

where \textquotedblleft C\textquotedblright\ in the superscript refers to the classical approach.

Consider, for simplicity, the case when $\widehat{\boldsymbol{\Sigma }}$ is estimated as the plain sample VC matrix computed from $T$ observations of $ N\times 1$ vectors, where each entry represents a forecast error in forecast combination bates1969combination, or an asset's excess return in a portfolio markowitz1952portfolio. When $T\gg N$, the classical solution (ref) is valid as $\widehat{\boldsymbol{ \Sigma }}$ is generally invertible. However, $\widehat{\boldsymbol{\Sigma }}$ may suffer ill-posedness when $N$ is comparable to $T$, and it is singular when $N$ is larger than $T$, which invalidates (ref). Such defects are well recognized in the empirical literature. When $N$ is large, $\widehat{\mathbf{w}}^{\mathrm{C}}$ often performs poorly because of the difficulty in estimating the large population VC matrix with precision. Rather, the simple average (namely, equal-weight) routinely outperforms. In forecast combination, this empirical fact is known as the forecast combination puzzle Clemen1989, Stock_Watson2004. Parallelly, in portfolio management the so-called naive $1/N$ diversification strategy is found to achieve robust out-of-sample gains compared to many sophisticated alternatives demiguel2009optimal, demiguel2009generalized.

In this paper, we propose a new estimation technique on the weights, to be presented in Section (ref) below. This method minimizes the squared Euclidean norm ($\ell _{2}$-norm) of the weight vector subject to a relaxed version of the first order conditions from the minimization problem in ((ref)), yielding the $\ell _{2}$- relaxation problem. The strategy is similar in spirit to the $\ell _{1}$ -relaxation in Dantzig selector a la candes2007dantzig. Interestingly, \ $\ell _{2}$-relaxation incorporates as special cases the simple average (equal-weight) strategy by setting the tuning parameter to be sufficiently large, and the classical optimal weighting scheme when the tuning parameter is zero. The tuning parameter helps to balance the bias and variance and to deliver roughly equal groupwise weights when the VC matrix exhibits a certain latent group structure. This is consistent with the intuition that when the VC matrix displays an exactly block equicorrelation structure, one should assign the same weight to all individual units within the same group whereas potentially distinct weights to the units in different groups. When the VC matrix is contaminated with a noise component, we show that the resultant $\ell _{2}$-relaxed weights are close to the infeasible groupwise equal weights.

The approximate latent group structure is not a man-made artifact. It is inherent in many models and applications. For example, it emerges when a factor structure dwells in the forecasting errors and the factor loadings are directly governed by certain latent group structures or are approximable by a few values (see Examples (ref) and (ref) in Section 3.1). It emerges when forecast combinations are based on a large number of forecast models with a fixed number of predictive regressors (see Example (ref) in Section 3.1). It also emerges when industrial classification serves as a proxy of the clustering pattern in the VC matrix of returns of stocks (see Example (ref) in Section 3.1).

Given the latent group structure, we establish two main theoretical results under a high dimensional asymptotic framework in which the number of individual units can be much larger than the time series dimension. (i) The estimated weights of $\ell _{2}$-relaxation converges to the within-group equal-weight solution (see Theorem (ref) in Section 3.2), and (ii) The empirical risk based on $\ell _{2}$-relaxation approaches the risk given by the oracle group information (see Theorem (ref) in Section 3.2). We assess the finite sample behavior of $\ell _{2}$-relaxation in Monte Carlo simulations. Compared with the oracle estimator and some popular off-the-shelf machine learning estimators, $\ell _{2}$-relaxation performs well under various data generating processes (DGPs). We further evaluate its empirical accuracy in three real data examples covering box office prediction, inflation forecast by surveyed professionals, and financial market portfolios. These examples showcase the wide applicability of $\ell _{2}$-relaxation.

Literature Review. This paper stands on several strands of vast literature. Forecast combination is reviewed by Clemen1989 and elliott2016 up to the points of their writing. Averaging forecasts appear to be a more robust procedure than the so-called optimal combination bates1969combination, and a reasonable explanation suggests that the errors on the estimation of the weights can be large and thus dominate the gains from the use of optimal combination (see, e.g., Smith_Wallis2009, Smith_Wallis2009, CMVW2016, CMVW2016).

A lesson learned from this literature is that it is unwise to include all possible variables; limiting the number of unknown parameters can help reduce estimation errors. This stylized fact has catalyzes the adoption of various shrinkage, regularization, and machine learning techniques. See, e.g., hansen2007least, Conflitti2015optimal, Bayer2018combining, Wilms2018, Kotchoni2019macroeconomic, Coulombe2020machine, diebold2021aggregation, and Roccazzella2020optimal. In particular, elliott2013complete propose a complete subset regression (CSR) approach to forecast combinations by using equal weights to combine forecasts based on the same number of predictive variables. Diebold_Shin2019 bring forth the partially egalitarian Lasso (peLASSO) procedures that discard some forecasts and then shrink the remaining forecasts toward the equal weights. Our $\ell _{2}$-relaxation adds a new way to regularize the weight estimation and includes the strategy of Diebold_Shin2019 as a special case.

Our paper is related to the burgeoning literature on latent group structures in panel data analysis; see, e.g., bonhomme2015grouped, su2016identifying, Su_Ju2018, Su_Wang_Jin2019, vogt2017classification, vogt2020multiscale, and bonhomme2022discretizing. While most of these previous studies focus on the recovery of the latent group structures in the conditional mean model, in our paper the group pattern is a latent structure in the VC matrix that encourages parameter parsimony and facilitates estimation accuracy. We do not attempt to recover the group identities.

Lastly, there is statistical and financial econometric literature on the estimation of large VC matrix. See Disatnik_Katz2012, fanli2012vast, Fan_Liao_Mincheva2013, fan2016incorporating, Ledoit_Wolf2017, and ao2019approaching, among many others. In particular, Ledoit_Wolf2004 use Bayesian methods for shrinking the sample correlation matrix to an equicorrelated target and show that this helps select portfolios with low volatility compared to those based on the sample correlation; Ledoit_Wolf2017 promote a nonlinear shrinkage estimator that is more flexible than the previous linear shrinkage estimators. Instead of regularizing the VC matrix, $\ell _{2}$-relaxation shrinks the weights and it can be used in conjunction with a high dimensional VC estimator.

Our paper complements the literature from the following aspects. Firstly, in terms of combination techniques, we corroborate in theory and in numerical experiments that $\ell_2$-relaxation is a competitive and easy-to-implement procedure. Secondly, unlike most panel data group structure papers, we focus on improvement of the out-of-sample performance and do not attempt to recover the membership for each individual, and a latent community or group structure is assumed for statistical optimality. Finally, while the dominating method in the large scale portfolio analysis shrinks the entries of the VC matrix, we take a viable alternative to directly discipline the weights. In summary, within the unified framework of forecast combination and portfolio optimization, $\ell _{2}$-relaxation is an innovative method with asymptotic guarantee under latent group structures.

Organization. The rest of the paper is organized as follows. Section (ref) motivates and introduces the $\ell _{2}$ -relaxation problem. Section (ref) studies the statistical properties of the estimator and establishes its asymptotic optimality under the latent group structures. Section (ref) reports Monte Carlo simulation results. The new method is applied to three datasets in Section (ref). All theoretical results are proved in Appendix A, and additional numerical results are contained in Appendix B.

Notation. Let \textquotedblleft $:=$\textquotedblright\ signify a definition, \textquotedblleft $\otimes $\textquotedblright\ be the Kronecker product, and $a\wedge b=\min \left\{ a,b\right\} $. We write $a\asymp b$ when both $a/b$ and $b/a$ are stochastically bounded. For a random variable $ x$, we write its population mean as $E\left[ x\right] $; for a sample $ \left( x_{1},\ldots ,x_{T}\right) $, we write its sample mean as $\mathbb{E} _{T}\left[ x_{t}\right] :=T^{-1}\sum_{t=1}^{T}x_{t}$. A plain $b$ denotes a scalar, a boldface lowercase $\mathbf{b}$ denotes a vector, and a boldface uppercase $\mathbf{B}$ denotes a matrix. The $\ell _{1}$-norm and $\ell _{2}$ -norm of $\mathbf{b}=(b_{1},...,b_{n})^{\prime }$ are written as $\left\Vert \mathbf{b}\right\Vert _{1}:=\sum_{i=1}^{n}\left\vert b_{i}\right\vert $ and $ \left\Vert \mathbf{b}\right\Vert _{2}:=\left( \sum_{i=1}^{n}b_{i}^{2}\right) ^{1/2},$ respectively. For a generic index set $\mathcal{G}\subset \left[ N \right] :=\left\{ 1,\ldots ,N\right\} $, we denote $\left\vert \mathcal{G} \right\vert $ as the cardinality of $\mathcal{G}$, and $\mathbf{b}_{\mathcal{ G}}=\left( b_{i}\right) _{i\in \mathcal{G}_{k}}$ as the $\left\vert \mathcal{ G}\right\vert $-dimensional subvector. $\phi _{\max }\left( \mathbf{\cdot } \right) $ and $\phi _{\min }\left( \cdot \right) $ represent the maximum and minimum eigenvalues of a real symmetric matrix, respectively. For a generic $ n\times m$ matrix $\mathbf{B}$, define the sup-norm as $\left\Vert \mathbf{B} \right\Vert _{\infty }:=\max_{i\leq n,j\leq m}\left\vert b_{ij}\right\vert $ , and the maximum column $\ell _{2}$ matrix norm as $\left\Vert \mathbf{B} \right\Vert _{c2}:=\max\nolimits_{j\leq m}\Vert \mathbf{B}_{\cdot j}\Vert _{2}$, where $\mathbf{B}_{\cdot j}$ is the $j$-th column. $\boldsymbol{0} _{n} $ and $\boldsymbol{1}_{n}$ are $n\times 1$ vectors of zeros and ones, respectively, and $\mathbf{\mathbf{I}}_{n}$ is the $n\times n$ identity matrix.

Formulation

Classical Approaches

In this section, we fix the ideas by characterizing the similarities between the classical forecast combination problem and portfolio optimization. We start with the former. Suppose that $y_{t+1}$ is an outcome variable, and there are $N$ forecasts for $y_{t+1}$, stacked as $\mathbf{f} _{t}=(f_{1t},...,f_{Nt})^{\prime }$, available at time $t$. The time dimension is indexed by $t\in \left[ T\right] :=\{1,2,...,T\}$, and the cross-sectional units are indexed by $i\in \left[ N\right] $. An $N\times 1$ weight vector $\mathbf{w}=(w_{i})_{i\in \lbrack N]}$ will linearly combine the forecasts into $\mathbf{w}^{\prime }\mathbf{f}_{t}$. We are interested in finding the weight $\mathbf{w}$ to minimize the mean squared forecast error (MSFE) of the combined forecast error $(y_{t+1}-\mathbf{w}^{\prime } \mathbf{f}_{t})$.

We collect the individual forecast error $e_{it}:=y_{t+1}-f_{it}$, and denote $\mathbf{e}_{t}:=(e_{it})_{i\in \lbrack N]}$ and $\bar{\mathbf{e}} :=T^{-1}\sum_{t=1}^{T}\mathbf{e}_{t}$.\footnote{bates1969combination assume unbiased forecasts and thus no demeaning is necessary in the construction of $\widehat{\boldsymbol{\Sigma }}$ in ((ref)) below. Here we accommodate potential biases of individual forecasts by the centered sample variance in order to present a unified framework for both forecast combination and portfolio optimization (see (ref)). Section (ref) in the Online Appendix shows that $\widehat{ \boldsymbol{\Sigma }}$ in ((ref)) copes with biased forecasts.} We compute its plain sample VC matrix\footnote{ We use the plain sample VC here to simplify the presentation. Alternative VC estimators tailored for high dimensional contexts can also be employed. In Section (ref), we report simulation results from the shrinkage VC estimators Ledoit_Wolf2004, Ledoit_Wolf2020 along with those from the plain sample VC estimator.}

equation[equation omitted — 168 chars of source]

Traditionally, the weights are determined by solving (ref) .

Forecast combination is intrinsically related to the mean-variance analysis of portfolio selection markowitz1952portfolio. Given $N$ financial assets of excess (relative to a risk-free asset) return $\mathbf{r} _{t}=(r_{it})_{i\in \lbrack N]}$, write the sample average return $\bar{ \mathbf{r}}:=T^{-1}\sum_{t=1}^{T}\mathbf{r}_{t}$ and the plain sample VC matrix

equation[equation omitted — 169 chars of source]

The weight vector $\mathbf{w}$ can be solved from

equation[equation omitted — 275 chars of source]

where $r^{\ast }$ is a user-specified target return. It is recognized that when there are many assets, precise estimation of the mean returns is a challenging task merton1980estimating, and the recent literature shifts to the minimum variance portfolio (MVP). MVP drops the linear restriction on returns (i.e., $\bar{\mathbf{r}}^{\prime }\mathbf{w}\geq r^{\ast }$), which leads to an optimization problem identical to (ref).\footnote{ See linton2019financial for a textbook treatment, demiguel2009generalized and demiguel2009generalized for extensive empirical comparisons, and fan2012vast, cai2020high and ding2021high for latest advancements of MVP.}

As forecast combination and MVP share the same form, we focus on the optimization problem in (ref). It can be rewritten as an unconstrained Lagrangian problem $\mathbf{w}^{\prime }\widehat{\boldsymbol{ \Sigma }}\mathbf{w}/2 +\gamma \left( \mathbf{w}^{\prime }\boldsymbol{1} _{N}-1\right) ,$ where $\gamma $ is the Lagrangian multiplier. The corresponding Kuhn-Karush-Tucker (KKT) conditions are:

equation[equation omitted — 207 chars of source]

The invertibility of the estimated VC matrix is not innocuous in high dimensional settings.\footnote{ The term \textquotedblleft high dimensional\textquotedblright\ means that the number of unknown parameters (in our context, $N$) is comparable to or larger than the sample size $T$. We will allow for $N/T\rightarrow c\in (0,\infty )$ as $\left( N,T\right) \rightarrow \infty $.} For example, when $ \widehat{\boldsymbol{\Sigma }}$ is the plain sample VC matrix, it must be singular when $N>T.$ Consider the case where $N$ is of similar magnitude to $T$ but $N<T$. Even if $\widehat{\boldsymbol{\Sigma }}$ is non-singular, a few sample eigenvalues of $\widehat{\boldsymbol{\Sigma }}$ are likely to be close to zero, leading to a numerically unstable solution when taking the matrix inverse in ((ref)).

Relaxation

To stabilize the numerical solution, we are inspired by the Dantzig selector candes2007dantzig and the relaxed empirical likelihood shi2016econometric to consider relaxing the sup-norm of the KKT condition as follows:

equation[equation omitted — 327 chars of source]

where $\tau $ is a tuning parameter to be specified by the user. We call the programming in ((ref)) the $\ell _{2}$-relaxation problem, and denote its solution as $\widehat{\mathbf{w}}=\widehat{\mathbf{w}}_{\tau },$ where the dependence of $\widehat{\mathbf{w}}$ on $\tau $ is often suppressed for notational conciseness.

When $\tau =0$, the solution $\widehat{\mathbf{w}}$ is characterized by the KKT conditions in ((ref)), and $\widehat{\mathbf{w}}^{\mathrm{C}}$ in ((ref)) is the unique solution when $\widehat{ \boldsymbol{\Sigma }}$ is invertible. Thus $\ell _{2}$-relaxation keeps the classical approach as a special case. Constraints in ((ref)) are feasible for any $\tau \geq 0$. The solution to ((ref)) is always unique because the objective is a strictly convex function and the feasible set is a closed convex set. The tuning parameter $\tau $ plays a crucial role in balancing the bias and variance: the bias is small when $\tau $ is small, whereas the variance is small when $\tau $ is large.\footnote{ A numerical illustration is provided in Appendix (ref).} If $ \tau $ is sufficiently large, say $\tau \geq \max_{i\in \left[ N\right] }| \widehat{\boldsymbol{\Sigma }}_{i\cdot }\boldsymbol{1}_{N}|/N$, the second constraint in ((ref)) is slack and thus irrelevant to the minimization. As a result, the simple average weight $N^{-1}\boldsymbol{1} _{N}$ solves ((ref)). In addition, relaxing $\tau $ from 0 reduces the sensitivity of the weights to the noise in the estimated VC matrix to prevent in-sample over-fitting.

On the other hand, $\ell _{2}$-relaxation can be motivated from the information theory as in diebold2021aggregation. The choice of $\ell _{2}$-norm can be viewed as a special case of R{\'{e}}nyi's cross-entropy renyi1961measures. This particular choice is for convenience because: (i) the dual of the Euclidean norm with respect to the inner product is the Euclidean norm itself and (ii) it accommodates $w_{i}<0$, which should not be ruled out in applications of forecast combination and portfolio analysis. By choosing a positive value of $\tau ,$ the constraints in ((ref)) yield a feasible set that is larger than that associated with $\tau =0.$ Let $\overline{w}:=\frac{1}{N} \sum_{i=1}^{N}w_{i}.$ Then

equation*[equation* omitted — 209 chars of source]

where the last equality holds under the constraint $\mathbf{w}^{\prime } \boldsymbol{1}_{N}=1.$ Clearly, the $\ell _{2}$-relaxation aims to minimize the sample variance of the weights over the feasible set and it effectively eliminates unnecessary variations across the individual weights. As we shall see, in the presence of a latent group structure in the dominant component of the VC matrix, the $\ell _{2}$-relaxation shrinks the individual weights to the group mean. This allows our estimator to include the widely used simple average (SA) estimator as a special case.\footnote{ There are other possibilities for new estimators that combine a particular entropy and a feasible set defined by a geometric structure tailored for a high-dimensional economic or financial problem of interest. We will need further exploration to see whether they can include some popular estimators as special cases.}

Theoretical Analysis

Latent Group Structures

In this section, we impose latent group structures on $\widehat{\boldsymbol{ \Sigma}}$ or its population expectation $E[\widehat{\boldsymbol{\Sigma}}]$ and then study the implications on the $\ell_{2}$-relaxed estimates of the weights.

Statistical analysis of high dimensional problems typically postulates certain structures on the data generating process for dimension reduction. For example, variable selection methods such as Lasso tibshirani1996regression and SCAD fan2001variable are motivated from regressions with sparsity, meaning most of the regression coefficients are either exactly zero or approximately zero. Similarly, in large VC estimation, various structures have been considered in the literature. bickel2008regularized impose many off-diagonal elements to be zero, engle2012dynamic assume a block equicorrelation structure, and Ledoit_Wolf2004 use Bayesian methods for shrinking the sample correlation matrix to an equicorrelated target, to name just a few.

Imposing latent group structures is an alternative way to reduce dimensions, which now has grown into a burgeoning literature. To analyze (ref) in depth in the high dimensional framework, we assume $\widehat{\boldsymbol{ \Sigma }}=\{\widehat{\Sigma }_{ij}\}_{i,j\in \left[ N\right] }$ can be approximated by a block equicorrelation matrix:

equation[equation omitted — 148 chars of source]

where $\widehat{\boldsymbol{\Sigma }}^{\ast }=\{\widehat{\Sigma }_{ij}^{\ast }\}_{i,j\in \left[ N\right] }$ is a block equicorrelation matrix and $ \widehat{\boldsymbol{\Sigma }}^{e}=\{\widehat{\Sigma }_{ij}^{e}\}_{i,j\in \left[ N\right] }$ denotes the deviation of $\widehat{\boldsymbol{\Sigma }}$ from the block equicorrelation matrix. We write

equation[equation omitted — 253 chars of source]

where $\mathbf{Z}=\{Z_{ik}\}$ denotes an $N\times K$ binary matrix providing the cluster membership of each individual forecast, i.e., $Z_{ik}=1$ if forecast $i$ belongs to group $\mathcal{G}_{k}\subset \left[ N\right] $ and $ Z_{ik}=0$ otherwise, and $\widehat{\boldsymbol{\Sigma }}^{\mathrm{co}}=\{ \widehat{\Sigma }_{kl}^{\mathrm{co}}\}_{k,l\in \left[ K\right] }$ is a $ K\times K$ symmetric positive definite matrix. Here, the superscript \textquotedblleft co\textquotedblright\ stands for \textquotedblleft core\textquotedblright . Note that $\widehat{\Sigma }_{ij}^{\ast }=\widehat{ \Sigma }_{kl}^{\mathrm{co}}$ if $i\in \mathcal{G}_{k}\text{ and }j\in \mathcal{G}_{l}.$

One can observe $\widehat{\boldsymbol{\Sigma }}$ from the data but not $ \widehat{\boldsymbol{\Sigma }}^{\ast }$. We will be precise about the definition of \textquotedblleft approximation\textquotedblright\ for $ \widehat{\boldsymbol{\Sigma }}^{e}$ in Assumption (ref) later. Let $N_{k}:=\left\vert \mathcal{G}_{k}\right\vert $ be the number of individuals in the $k$th group, and thus $N = \sum_{k=1}^{K}N_{k}$. For ease of notation and after necessary re-ordering the $N$ forecast units, we write

equation[equation omitted — 192 chars of source]

in which the units in the same group cluster together in a block. The re-ordering is for the convenience of notation only. The theory to be developed is irrelevant to the ordering of individuals, and does not require the knowledge about the membership matrix $\mathbf{Z}$.

We now motivate the decomposition ((ref)) using five examples.

exampleChan_Pauwels2018 assume the existence of a \textquotedblleft best\textquotedblright\ unbiased forecast $f_{0t}$ of variable $y_{t+1}$ with an associated forecast error $e_{0t}$, and the forecast error $e_{it}$ of model $i$ can be decomposed as \begin{equation*} e_{it}=e_{0t}+u_{it}, \end{equation*} where $e_{0t}$ represents the forecast error from the best forecasting model, and $u_{it}$ is the deviation of $e_{it}$ from the best forecasting model. Assuming $E\left[ u_{it}\right] =0 $ and $E\left[ e_{0t}u_{it}\right] =0$ for each $i,$ the VC of $\mathbf{e}_{t}$ can be written as $\mathbf{ \Sigma }_0 =E\left[ \mathbf{e}_{t}\mathbf{e}_{t}^{\prime }\right] =E\left[ e_{0t}^{2}\right] \mathbf{1}_{N} \mathbf{1}_{N}^{\prime }+E[\mathbf{u}_{t} \mathbf{u}_{t}^{\prime }],$ where $\mathbf{u}_{t}=(u_{1t},...,u_{Nt})^{ \prime }$. At the sample level, we have $\widehat{\boldsymbol{\Sigma }}= \widehat{\boldsymbol{\Sigma }}^{\ast }+\widehat{\boldsymbol{\Sigma }}^{e},$ where \begin{equation*} \widehat{\boldsymbol{\Sigma }}=\mathbb{E}_{T}\left[ \mathbf{e}_{t}\mathbf{e} _{t}^{\prime }\right] , \widehat{\boldsymbol{\Sigma }}^{\ast }= \mathbb{E}_{T}\left[ e_{0t}^{2}\right] \mathbf{1}_{N}\mathbf{1}_{N}^{\prime }, and \widehat{\boldsymbol{\Sigma }}^{e}=\mathbb{E}_{T}[\mathbf{u} _{t}\mathbf{u}_{t}^{\prime }]+\mathbf{1}_{N}\mathbb{E}_{T}[e_{0t}\mathbf{u} _{t}^{\prime }]+\mathbb{E}_{T}[e_{0t}\mathbf{u}_{t}]\mathbf{1}_{N}^{\prime }. \end{equation*} In this case, all the $N$ forecast units belong to the same group $\mathcal{G }_{1}$ as rank$(\widehat{\boldsymbol{\Sigma }}^{\ast })=1$.
exampleConsider that each individual forecast $f_{it}$ is generated from a factor model \begin{equation} f_{it}=\boldsymbol{\lambda }_{g_{i}}^{\prime }\boldsymbol{\eta }_{t}+u_{it}, \end{equation} where $\boldsymbol{\lambda }_{g_{i}}$ is a $q\times 1$ vector of factor loadings, $\boldsymbol{\eta }_{t}$ is a $q\times 1$ vector of latent factors, and $u_{it}$ is an idiosyncratic shock. Here $g_{i}$ denotes individual $i$'s membership, i.e., it takes value $k$ if individual $i$ belongs to group $\mathcal{G}_{k}$ for $k\in \lbrack K]$ and $i\in \lbrack N] $. Similarly, assume $y_{t+1}=\boldsymbol{\lambda }_{y}^{\prime } \boldsymbol{\eta }_{t}+u_{y,t+1}$, with $E\left[ u_{it}|\boldsymbol{\eta } _{t}\right] =0$ and $E\left[ u_{y,t+1}|\boldsymbol{\eta }_{t}\right] =0$ and $E\left[ u_{it}u_{y,t+1}|\boldsymbol{\eta }_{t}\right] =0.$\footnote{ Other than those $q$ factors in $\boldsymbol{\eta }_{t}$, the additional latent factor $u_{y,t+1}$ in $y_{t+1}$ is unforeseeable at time $t$. In other words, given the information set $\mathcal{I}_{t}$ that contains $ \left( \{f_{it}\}_{i\in \left[ N\right] },\boldsymbol{\eta }_{t}\right) $ and $\mathbf{u}_{t}$ at time $t$, the error $u_{y,t+1}=y_{t+1}-\boldsymbol{ \lambda }_{y}^{\prime }\boldsymbol{\eta }_{t}=y_{t+1}-E\left( y_{t+1}| \mathcal{I}_{t}\right) $ must be orthogonal to $\mathcal{I}_{t}.$ Then $E \left[ u_{it}u_{y,t+1}|\mathcal{I}_{t}\right] =0$ implies $E\left[ u_{it}u_{y,t+1}|\boldsymbol{\eta }_{t}\right] =0$ by the law of iterated expectations.} For simplicity, we also assume conditional homoskedasticity var$\left( \mathbf{u}_{t}|\boldsymbol{\eta }_{t}\right) =\mathbf{\mathbf{ \Omega }}_{u}$ and the factor loadings are nonstochastic. Then individual $i$ 's forecast error is \begin{equation*} e_{it}=y_{t+1}-f_{it}=[\left( \boldsymbol{\lambda }_{y}-\boldsymbol{\lambda } _{g_{i}}\right) ^{\prime }\boldsymbol{\eta }_{t}+u_{y,t+1}]-u_{it}=\lambda _{g_{i}}^{\dag \prime }\boldsymbol{\eta }_{t}^{\dag }-u_{it}, \end{equation*} where $\boldsymbol{\eta }_{t}^{\dag }:=(\boldsymbol{\eta }_{t}^{\prime },u_{y,t+1})^{\prime }$ and $\boldsymbol{\lambda }_{g_{i}}^{\dag }:=(( \boldsymbol{\lambda }_{y}-\boldsymbol{\lambda }_{g_{i}})^{\prime },1)^{\prime }$, or equivalently $\mathbf{e}_{t}=\mathbf{\Lambda }^{\dag } \boldsymbol{\eta }_{t}^{\dag }-\mathbf{u}_{t}$ in a vector form, where $ \mathbf{\Lambda }^{\dag }\mathbf{:}=\left( \lambda _{g_{1}}^{\dag },\ldots ,\lambda _{g_{N}}^{\dag }\right) ^{\prime }$. The population VC of $\mathbf{e }_{t}$ is given by $\mathbf{\Sigma }_{0}=E\left[ \mathbf{e}_{t}\mathbf{e} _{t}^{\prime }\right] =\mathbf{\Lambda }^{\dag }E[\boldsymbol{\eta } _{t}^{\dag }\boldsymbol{\eta }_{t}^{\dag \prime }]\mathbf{\Lambda }^{\dag \prime }+\mathbf{\Omega }_{u}.$ Decompose the sample VC as $\widehat{ \boldsymbol{\Sigma }}=\widehat{\boldsymbol{\Sigma }}^{\ast }+\widehat{ \boldsymbol{\Sigma }}^{e},$ where\footnote{ The conclusion here also holds for the centered version of the VC matrix: $ \widehat{\boldsymbol{\Sigma }}=\mathbb{E}_{T}\left[ \left( \mathbf{e}_{t}- \mathbf{\bar{e}}\right) (\mathbf{e}_{t}-\mathbf{\bar{e}})^{\prime }\right] $ with more complicated notation.} \begin{equation*} \widehat{\boldsymbol{\Sigma }}=\mathbb{E}_{T}\left[ \mathbf{e}_{t}\mathbf{e} _{t}^{\prime }\right] , \widehat{\boldsymbol{\Sigma }}^{\ast }= \mathbf{\Lambda }^{\dag }\mathbb{E}_{T}[\boldsymbol{\eta }_{t}^{\dag } \boldsymbol{\eta }_{t}^{\dag \prime }]\mathbf{\Lambda }^{\dag \prime }, and \widehat{\boldsymbol{\Sigma }}^{e}=\mathbb{E}_{T}[\mathbf{u}_{t}\mathbf{ u}_{t}^{\prime }]-\mathbf{\Lambda }^{\dag }\mathbb{E}_{T}[\boldsymbol{\eta } _{t}^{\dag }\mathbf{u}_{t}^{\prime }]-\mathbb{E}_{T}[\mathbf{u}_{t} \boldsymbol{\eta }_{t}^{\dag \prime }]\mathbf{\Lambda }^{\dag \prime }. \end{equation*} By construction, the core matrix has element $\widehat{\Sigma }_{kl}^{ \mathrm{co}}=\lambda _{k}^{\dag \prime }\mathbb{E}_{T}[\boldsymbol{\eta } _{t}^{\dag }\boldsymbol{\eta }_{t}^{\dag \prime }]\lambda _{l}^{\dag }$ for $ k,l\in \left[ K\right] $, the equicorrelation matrix has element $\widehat{ \Sigma }_{ij}^{\ast }=\widehat{\Sigma }_{kl}^{\mathrm{co}}$ if $i\in \mathcal{G}_{k}$ and $j\in \mathcal{G}_{l}$, and rank$(\widehat{\boldsymbol{ \Sigma }}^{\ast })\leq \left( q+1\right) \wedge K$.
remarkWe emphasize that our theory below does not require the knowledge about the group membership for individual forecasts. Alternatively, one can estimate the multi-factor structure in ((ref)) by the principal component analysis (PCA) and then apply either the $K$-means algorithm or the sequential binary segmentation algorithm Wang_Su2021 to the estimated factor loadings to identity the true group membership. Then one can impose the recovered group structure before computing classical weights. This method is computationally involved and is subject to the usual classification error issue: the presence of classification error in finite samples is carried upon and thus adversely affects the subsequent estimation of the weights. In contrast, the advantage of $\ell _{2}$-relaxation is that it is computationally simple as it directly works with the sample moments and hence bypasses the factor structure and the group membership.\footnote{ It is worth mentioning that Hsiao_Wan2014 assume that the forecast errors exhibit a multi-factor structure, but they do not assume the presence of $K$ latent groups in the $N$ factor loadings and write $\boldsymbol{ \lambda }_{i}$ in place of $\boldsymbol{\lambda }_{g_{i}}$. In the absence of the latent group structures among the factor loadings $\{\boldsymbol{ \lambda }_{i}\}_{i\in \lbrack N]}$, the dominant component in $\widehat{ \boldsymbol{\Sigma }}$ will have a low-rank structure but not a latent group structure. Analyses of this case will be different from the current paper, which we leave for future research.}

Latent groups may be present not only in approximate factor models, as in the above two motivating examples, but also in some forecast problems in which multi-factor structures are implicit. Here follows such an example.

exampleSuppose that the outcome variable $y_{t+1}$ is generated via the process \begin{equation} y_{t+1}=\mathbf{x}_{t}^{\prime }\boldsymbol{\theta }^{0}+u_{t+1}\,\,\, for t=-T_{0},...,-1,0,1,... \end{equation} where $\mathbf{x}_{t}=(x_{j,t})_{j=1}^{p}$ is a $p\times 1$ vector of potential predictive variables, $\boldsymbol{\theta }^{0}=(\theta _{j}^{0})_{j=1}^{p}$ is a $p\times 1$ vector of regression coefficients, and $u_{t+1}$ is the error term such that $E\left[ u_{t+1}|\mathbf{x}_{t}\right] =0$ and $E\left[ u_{t+1}^{2}|\mathbf{x}_{t}\right] =\sigma _{u}^{2}$. Due to costly data collection or ignorance, the forecaster $i$ utilizes only a subset $\mathbf{x}_{S_{i},t}$ of $\mathbf{x}_{t}$, where $S_{i}\subset \lbrack p]$, to exercise prediction with the OLS estimate. Let $\widehat{ \boldsymbol{\theta }}_{S_{i},t}=(\sum_{l=-T_{0}+1}^{t}\mathbf{x}_{S_{i},l-1} \mathbf{x}_{S_{i},l-1}^{\prime })^{-1}\sum_{l=-T_{0}+1}^{t}\mathbf{x} _{S_{i},l-1}\mathbf{y}_{l},$ and $\widehat{\boldsymbol{\theta }}_{i,t}$ be the sparse $p\times 1$ vector that embeds the corresponding $\widehat{ \boldsymbol{\theta }}_{S_{i},t}$ so that $(\widehat{\boldsymbol{\theta }} _{i,t})_{S_{i}}=\widehat{\boldsymbol{\theta }}_{S_{i},t}$ and $(\widehat{ \boldsymbol{\theta }}_{i,t})_{[p]\backslash S_{i}}=\mathbf{0}$. We consider two forecasting schemes: the fixed window and the rolling window. (i) In the case of a fixed estimation window, the $i$th forecast of $y_{t+1}$ is given by $f_{it}:=\mathbf{x}_{S_{i},t}^{\prime }\widehat{\boldsymbol{ \theta }}_{S_{i},0}$ for $t\geq 1$. The associated forecast error\ is \begin{equation*} e_{it}=y_{t+1}-f_{it}=y_{t+1}-\mathbf{x}_{t}^{\prime }\widehat{\boldsymbol{ \theta }}_{i,0}=u_{t+1}+\mathbf{x}_{t}^{\prime }(\boldsymbol{\theta }^{0}- \boldsymbol{\theta }_{i}^{0})+\epsilon _{i,t}, \end{equation*} where $\boldsymbol{\theta }_{i}^{0}:=$ plim$_{T_{0}}\widehat{\boldsymbol{ \theta }}_{i,0}$ and $\epsilon _{i,t}:=\mathbf{x}_{t}^{\prime }(\widehat{ \boldsymbol{\theta }}_{i}-\boldsymbol{\theta }_{i}^{0}).$ This is a $\left( p+1\right) $-factor model with factors $(\mathbf{x}_{t}^{\prime },u_{t+1})^{\prime }$ and factor loadings $((\boldsymbol{\theta }^{0}- \boldsymbol{\theta }_{i}^{0})^{\prime },1)^{\prime }.$ (ii) In the case of a rolling window, the forecast error is \begin{equation*} e_{it}=y_{t+1}-\mathbf{x}_{S_{i},t}^{\prime }\widehat{\boldsymbol{\theta }} _{S_{i},t}=u_{t+1}+\mathbf{x}_{t}^{\prime }(\boldsymbol{\theta }^{0}- \widehat{\boldsymbol{\theta }}_{i,t})=u_{t+1}+\mathbf{x}_{t}^{\prime }( \boldsymbol{\theta }^{0}-\boldsymbol{\theta }_{i}^{0})+\epsilon _{i,t}, \end{equation*} where $\widehat{\boldsymbol{\theta }}_{i,t}\overset{p}{\rightarrow } \boldsymbol{\theta }_{i}^{0}$ as $T_{0}\rightarrow \infty $ is assumed to hold uniformly in $\left( i,t\right) $ under some regularity conditions that include the covariance stationarity, and $\epsilon _{i,t}:=\mathbf{x} _{t}^{\prime }(\widehat{\boldsymbol{\theta }}_{i,t}-\boldsymbol{\theta } _{i}^{0})$. Therefore, we have an approximate $\left( p+1\right) $-factor model with factors $(\mathbf{x}_{t}^{\prime },u_{t+1})^{\prime }$ and factor loadings $((\boldsymbol{\theta }^{0}-\boldsymbol{\theta }_{i}^{0})^{\prime },1)^{\prime }$. Similar analysis applies to the rolling window of fixed length $L$, in which the forecaster $i$ estimates the coefficient by $ \widehat{\boldsymbol{\theta }}_{S_{i},t}^{L}=(\sum_{l=t-L+1}^{t}\mathbf{x} _{S_{i},t-1}\mathbf{x}_{S_{i},t-1}^{\prime })^{-1}\sum_{l=t-L+1}^{t}\mathbf{x }_{S_{i},t-1}\mathbf{y}_{t}$. In either case, $e_{it}$ exhibits a factor structure where the factor loadings are given by $((\boldsymbol{\theta }^{0}-\boldsymbol{\theta } _{i}^{0})^{\prime },1)^{\prime }.$ When $\boldsymbol{\theta }_{i}^{0}$ exhibits a latent group structure (see the next example), say, $\boldsymbol{ \theta }_{i}^{0}=\boldsymbol{\theta }_{g_{i}}^{0}$ with $g_{i}$ being as defined in the last example, the forecast error reduces to that in Example (ref) with \begin{equation*} e_{it}=\lambda _{g_{i}}^{\dag \prime }\boldsymbol{\eta }_{t}^{\dag }-u_{i,t}, \end{equation*} where $\lambda _{g_{i}}^{\dag }:=((\boldsymbol{\theta }^{0}-\boldsymbol{ \theta }_{g_{i}}^{0})^{\prime },1)^{\prime },$ $\boldsymbol{\eta }_{t}^{\dag }=(\mathbf{x}_{t}^{\prime },u_{t+1})^{\prime },$ and $u_{i,t}=-\epsilon _{i,t}.$ Then $\widehat{\boldsymbol{\Sigma }},$ $\widehat{\boldsymbol{\Sigma }}^{\ast },$ and $\widehat{\boldsymbol{\Sigma }}^{e}$ can be defined as in Example (ref).

The next example is a simple linear regression that yields a two-group structure in the VC matrix, and it can be easily extended to the multiple group structure by allowing for multiple regressors to have predictive power.

exampleWe reuse the notation in Example (ref) while focus on a special case where only one regressor inside $\mathbf{x} _{t},$ say, $x_{1,t},$ has predictive power and we employ the fixed window scheme to forecast. Then $\theta _{1}^{0}\neq 0$ and $\theta _{j}^{0}=0$ for all $j=2,...,p,$ where we allow $p$ to diverge to infinity slowly. We can divide the $N$ forecasts into two groups according to whether $1\in S_{i}$, i.e., whether the only predictive regressor $x_{1,t}$ is included in the $i$ th forecasting model. Without loss of generality, we assume $E\left[ \mathbf{ x}_{t}\right] =\mathbf{0}$ and $E\left[ \mathbf{x}_{t}\mathbf{x}_{t}^{\prime }\right] =\mathbf{I}_{p}$. Furthermore, we assume that $x_{1,t}$ is included in forecasting model $i$ as the first element in $\mathbf{x}_{S_{i},t}$ for\ $i\in \mathcal{G}_{1}=[N_{1}]$ while it is excluded for $i\in \mathcal{G} _{2}=\{N_{1}+1,....,N\}$. Intuitively, the first $N_{1}$ forecasting models are correctly specified for the conditional mean of $y_{t+1}$ while the other $N_{2}:=N-N_{1}$ models are misspecified. Note that \begin{equation*} e_{it}=y_{t+1}-\mathbf{x}_{S_{i},t}^{\prime }\widehat{\boldsymbol{\theta }} _{S_{i},0}=\left[ u_{t+1}+(x_{1,t}\theta _{1}^{0}-\mathbf{x} _{S_{i},t}^{\prime }\boldsymbol{\theta }_{S_{i},0}^{0})\right] +\mathbf{x} _{S_{i},t}^{\prime }(\boldsymbol{\theta }_{S_{i},0}^{0}-\widehat{\boldsymbol{ \theta }}_{S_{i},0})=v_{it}+s_{it}, \end{equation*} where $v_{it}:=u_{t+1}+(x_{1,t}\theta _{1}^{0}-\mathbf{x}_{S_{i},t}^{\prime } \boldsymbol{\theta }_{S_{i},0}^{0}),\ s_{it}:=\mathbf{x}_{S_{i},t}^{\prime }( \boldsymbol{\theta }_{S_{i},0}^{0}-\widehat{\boldsymbol{\theta }}_{S_{i},0})$ and $\boldsymbol{\theta }_{S_{i},0}^{0}$ is the probability limit of $ \widehat{\boldsymbol{\theta }}_{S_{i},0}.$ Under some regularity conditions, the effect of the parameter estimation error $s_{it}$ can be made as small as possible for a sufficiently large $T_{0}$ as $\Vert \boldsymbol{\theta } _{S_{i},0}^{0}-\widehat{\boldsymbol{\theta }}_{S_{i},0}\Vert _{2}=O_{p}((T_{0}/p_{i})^{-1/2})$, where $p_{i}$ is the number of regressors in the $i$th model that can be divergent to infinity too. The orthonormal regressors imply $\boldsymbol{\theta }_{S_{i},0}^{0}=\left( \theta _{1}^{0}, \mathbf{0}_{p_{i}-1}^{\prime }\right) ^{\prime }$ for $i\in \mathcal{G}_{1}$ and $\boldsymbol{\theta }_{S_{i},0}^{0}=\mathbf{0}_{p_{i}}$ for $i\in \mathcal{G}_{2}$. Let $\mathbf{v}_{t}:=\left( v_{1t},...,v_{Nt}\right) ^{\prime }$, $\mathbf{\bar{v}}:=\mathbb{E}_{T}(\mathbf{v}_{t})$, and $ \widehat{\mathbf{V}}:=\mathbb{E}_{T}\left[ \left( \mathbf{v}_{t}-\mathbf{ \bar{v}}\right) (\mathbf{v}_{t}-\mathbf{\bar{v}})^{\prime }\right] $. Define $\mathbf{s}_{t}$ and $\mathbf{\bar{s}}$ analogously. Then $\widehat{ \boldsymbol{\Sigma }}=\mathbb{E}_{T}\left[ \left( \mathbf{e}_{t}-\mathbf{ \bar{e}}\right) (\mathbf{e}_{t}-\mathbf{\bar{e}})^{\prime }\right] =\widehat{ \boldsymbol{\Sigma }}^{\ast }+\widehat{\boldsymbol{\Sigma }}^{e},$ where \begin{eqnarray*} \widehat{\boldsymbol{\Sigma }}^{\ast }&:= &plim_{T\rightarrow \infty } \widehat{\mathbf{V}}=\left( \begin{array}{cc} \sigma _{u}^{2}\boldsymbol{1}_{N_{1}}\boldsymbol{1}_{N_{1}}^{\prime } & \sigma _{u}^{2}\boldsymbol{1}_{N_{1}}\boldsymbol{1}_{N_{2}}^{\prime } \\ \sigma _{u}^{2}\boldsymbol{1}_{N_{2}}\boldsymbol{1}_{N_{1}}^{\prime } & \left[ \sigma _{u}^{2}+\left( \theta _{1}^{0}\right) ^{2}\right] \boldsymbol{ 1}_{N_{2}}\boldsymbol{1}_{N_{2}}^{\prime } \end{array} \right) , \\ \widehat{\boldsymbol{\Sigma }}^{e}&:= &(\widehat{\mathbf{V}}-plim _{T\rightarrow \infty }\widehat{\mathbf{V}})+\mathbb{E}_{T}\left[ \left( \mathbf{s}_{t}-\mathbf{\bar{s}}\right) (\mathbf{s}_{t}-\mathbf{\bar{s}} )^{\prime }+\left( \mathbf{v}_{t}-\mathbf{\bar{v}}\right) (\mathbf{s}_{t}- \mathbf{\bar{s}})^{\prime }+\left( \mathbf{s}_{t}-\mathbf{\bar{s}}\right) ( \mathbf{v}_{t}-\mathbf{\bar{v}})^{\prime }\right] . \end{eqnarray*} can be easily verified.
remarkExample (ref) offers a setting in which the strategy of Diebold_Shin2019 is optimal. Intuitively, in the presence of two groups of forecasts (say $\mathcal{G}_{1}$ and $\mathcal{G}_{2}$) with the same forecast variance among each group, if the covariance between the good (those in $\mathcal{G}_{1},$ say) and bad (those in $\mathcal{G}_{2},$ say) forecasts is the same as the variance of the good forecasts, an optimal forecast combination should assign zero weight to the group of bad forecasts and equal nonzero weight to the group of good forecasts. Lemma (ref) below suggests that if $\widehat{\boldsymbol{\Sigma }}^{\ast }$ was observed and used, the optimal strategy would assign $1/N_{1}$ weight to each of the first $N_{1}$ forecasts and 0 weight to each of the last $N_{2}$ forecasts. When $\widehat{\boldsymbol{\Sigma }}^{\ast }$ is replaced by the feasible version $\widehat{\boldsymbol{\Sigma }},$ our theory below ensures that the $ \ell _{2}$-relaxation assigns approximately $1/N_{1}$ weight to each of the first $N_{1}$ forecasts and approximately 0 weight to each of the last $N_{2} $ forecasts.

Lastly, we give an example that illustrates the use of group structure in portfolio analysis.

exampleVolatility matrix is a fundamental component for portfolio analysis. To reduce the complexity in estimating a vast VC matrix, engle2012dynamic employ the Standard Industrial Classification (SIC) to assign the equicorrelated blocks. Using MVP, clements2015benefits find evidence in favor of equicorrelation across portfolio sizes. Each of these papers explicitly specifies a criterion, either SIC or portfolio size, to allocate an individual's group identity. In contrast, no knowledge about the membership is required to implement $\ell _{2}$-relaxation; block equicorrelation is taken as a latent structure.

Next, we specify an asymptotic target for the $\ell _{2}$-relaxation estimator $\widehat{\mathbf{w}}$. Consider the oracle problem of $\ell _{2}$ -relaxation with an infeasible $\widehat{\boldsymbol{\Sigma }}^{\ast }$:

equation[equation omitted — 308 chars of source]

Denote the solution to the above problem as $\mathbf{w}^{\ast }$. Lemma (ref) below shows that the squared $\ell _{2}$-norm objective function produces the within-group equally weighted solution $\mathbf{w} ^{\ast }$. The problem ((ref)) with $\left( N+1\right) $ free parameters is effectively reduced to merely $\left( K+1\right) $ free parameters in the oracle problem ((ref)).

lemmaThe solution to ((ref)) takes within-group equal values in the form \begin{equation*} \mathbf{w}^{\ast }=\left( N^{-1}b_{01}^{\ast }\boldsymbol{1}_{N_{1}}^{\prime },\cdots ,N^{-1}b_{0K}^{\ast }\boldsymbol{1}_{N_{K}}^{\prime }\right) ^{\prime }, \end{equation*} where the expression of $\left( b_{0k}^{\ast }\right) _{k\in \lbrack K]}$ is given in Equation ((ref)) in the Online Appendix.

The use of squared $\ell _{2}$-norm in ((ref)) yields the same weight across units in each group for the oracle problem. When $ \widehat{\boldsymbol{\Sigma }}^{\ast }$ is replaced by its feasible version $ \widehat{\boldsymbol{\Sigma }},$ we will show that $\ell _{2}$-relaxation guarantees that the weights are approximately equal within each group so that $\widehat{\mathbf{w}}$ and $\mathbf{w}^{\ast }$ are sufficiently close.

Asymptotic Theory

We study the asymptotic properties of the $\ell _{2}$-relaxed estimator in this section. We consider a triangular array of models indexed by $T$ and $N$ , both passing to infinity. Let $\phi _{NT}:=\sqrt{ (\log N) / (T\wedge N)} \rightarrow 0$. Note that we allow both $N\gg T $ (as in standard high dimensional problems) and $T\gg N$ or $T\asymp N$ in view of $\phi _{NT}$. But we rule out the traditional case of \textquotedblleft fixed $N$ and large $T$\textquotedblright , which has been covered by the classical approach (ref) and (ref).

In (ref) $\widehat{\mathbf{\Sigma }}$ is decomposed into $ \widehat{\mathbf{\Sigma }}^{\ast }$ and $\widehat{\mathbf{\Sigma }}^{e}$, and in (ref) $\widehat{\mathbf{\Sigma }}^{\ast }$ is characterized by $\widehat{\mathbf{\Sigma }}^{\mathrm{co}}$. Let $ \boldsymbol{\Sigma }_{0}^{e}:=E[\widehat{\boldsymbol{\Sigma }}^{e}]$, $ \boldsymbol{\Delta }^{e}:=\widehat{\boldsymbol{\Sigma }}^{e}-\boldsymbol{ \Sigma }_{0}^{e}$, $\mathbf{\Sigma }_{0}^{\ast }=E[\widehat{\mathbf{\Sigma }} ^{\ast }],$ $\boldsymbol{\Sigma }_{0}^{\mathrm{co}}:=E[\widehat{\boldsymbol{ \Sigma }}^{\mathrm{co}}]$, and $\boldsymbol{\Delta }^{\mathrm{co}}:=\widehat{ \boldsymbol{\Sigma }}^{\mathrm{co}}-\boldsymbol{\Sigma }_{0}^{\mathrm{co}}$. We impose regularity conditions on the population matrices and the sampling errors.

assumptionThere are positive finite constants $C_{e0}$, $ \underline{c}$, and $\overline{c}$ such that: \begin{enumerate} • $\phi_{\max}\left(\boldsymbol{\Sigma}_{0}^{e}\right)=O(\sqrt{N} \phi_{NT})$, $\Vert\boldsymbol{\Sigma}_{0}^{e}\Vert_{c2}\leq C_{e0} \cdot \phi_{\max}\left(\boldsymbol{\Sigma}_{0}^{e}\right)$, and $\left\Vert \boldsymbol{\Delta}^{e}\right\Vert _{\infty}=O_{p}((T/\log N)^{-1/2})$; • $\underline{c}\leq\phi_{\min}\left(\boldsymbol{\Sigma}_{0}^{\mathrm{co} }\right)\leq\phi_{\max}\left(\boldsymbol{\Sigma}_{0}^{\mathrm{co}}\right)\leq \overline{c}$, and $\left\Vert \boldsymbol{\Delta}^{\mathrm{co}}\right\Vert _{\infty}=O_{p}((T/\log N)^{-1/2})$. \end{enumerate}

The first condition in Assumption (ref)(a) allows the maximum eigenvalue of the $N\times N$ matrix $\boldsymbol{\Sigma }_{0}^{e}$ to diverge to infinity, but at a limited rate $\sqrt{N}\phi _{NT}$. The second condition in (a) is similar to but weaker than the absolute row-sum condition that is frequently used to model weak cross-sectional dependence; see, e.g., Fan_Liao_Mincheva2013. The third condition in (a) requires that the sampling error of $\Delta _{ij}^{e}$ be controlled by $ (T/\log N)^{-1/2}$ uniformly over $i$ and $j$ so that each element of $ \widehat{\boldsymbol{\Sigma }}^{e}$ should not deviate too much from its population mean $\boldsymbol{\Sigma }_{0}^{e}$. This condition can be established under some low-level assumptions; see, e.g., Chapter 6 in Wainwright2019high. Assumption (ref)(b) bounds all eigenvalues of the population core away from 0 and infinity, and impose similar stochastic order on $\boldsymbol{\boldsymbol{\Delta }}^{\mathrm{co}}$ as that on $\boldsymbol{\Delta }^{e}$ in Assumption (ref)(a). Because ${\widehat{\mathbf{\Sigma }}}^{\mathrm{co}}$ is a low-rank matrix, the restriction on $\mathbf{\Delta }^{\mathrm{co}}$ is very mild and the sample error of the feasible VC $\widehat{\mathbf{\Sigma }}-E[\widehat{ \mathbf{\Sigma }}]=\mathbf{\Delta }^{e}+\mathbf{\Delta }^{\mathrm{co}}$ is primarily determined by $\mathbf{\Delta }^{e}$.

example(Example (ref), cont.) Following the notation of Example (ref), we can decompose the population variance-covariance matrix $ \mathbf{\Sigma }:=E\left[ \mathbf{e}_{t}\mathbf{e}_{t}^{\prime }\right] $ as $\mathbf{\Sigma =\Sigma }_{0}^{\ast }+\mathbf{\Sigma }_{0}^{e},$ where $ \mathbf{\Sigma }_{0}^{\ast }=\mathbf{\Lambda }^{\dag }E[\boldsymbol{\eta } _{t}^{\dag }\boldsymbol{\eta }_{t}^{\dag \prime }]\mathbf{\Lambda }^{\dag \prime },$ and $\mathbf{\Sigma }_{0}^{e}=\mathbf{\Omega }_{x}=\left\{ \Omega _{x,ij}\right\} .$ The corresponding sampling error is \begin{equation*} \Delta _{ij}^{e}=\left\{ \mathbb{E}_{T}\left[ \epsilon _{i,t}\epsilon _{j,t} \right] -\Omega _{x,ij}\right\} -\sum_{l\in \{i,j\}}\left\{ (\boldsymbol{ \lambda }_{y}-\boldsymbol{\lambda }_{g_{l}})^{\prime }\mathbb{E}_{T}\left[ \boldsymbol{\eta }_{t}(u_{y,t+1}-u_{l,t})\right] +\mathbb{E}_{T}\left[ u_{y,t+1}u_{l,t}\right] \right\} . \end{equation*} Then the first part of Assumption (ref)(a) is satisfied as long as $\phi _{\max }\left( \boldsymbol{\Omega }_{x}\right) _{\text{{}}}=O(\sqrt{ N}\phi _{NT})$. For the sampling error matrix, if \begin{equation*} \max_{i,j\in \lbrack N]}\left\{ \left\vert \mathbb{E}_{T}\left[ \boldsymbol{ \eta }_{t}(u_{y,t+1}-u_{i,t})\right] \right\vert +\left\vert \mathbb{E}_{T} \left[ u_{y,t}u_{i,t}\right] \right\vert +\left\vert \mathbb{E}_{T}\left[ u_{i,t}u_{i,j}\right] -\Omega _{ij,x}\right\vert \right\} =O_{p}((T/\log N)^{-1/2}), \end{equation*} then $\left\Vert \boldsymbol{\boldsymbol{\Delta }}^{e}\right\Vert _{\infty }=O_{p}((T/\log N)^{-1/2})$ is satisfied as well.

The extent of relaxation in ((ref)) is controlled by the tuning parameter $\tau $, which is to be chosen by cross validations (CV) in simulations and applications. We spell out admissible range of $\tau $ in Assumption (ref)(a) below. Assumption (ref)(b) restricts $\underline{r} := \min_{ k\in [K] } N_k / N$ relative to $K$.

assumptionAs $\left(N,T\right)\rightarrow\infty,$ \begin{enumerate} • $\sqrt{K}\phi_{NT}/\tau+K^{5/2}\tau\to0$; • $\underline{r}\asymp K^{-1}$. \end{enumerate}

In order to meet the condition $\sqrt{K}\phi _{NT}/\tau \rightarrow 0$ in Assumption (ref)(a), it suffices to specify

equation*[equation* omitted — 50 chars of source]

for some slowly diverging sequence $D_{\tau }$ as $\left( N,T\right) \rightarrow \infty $, for example, $\log \log \left( N\wedge T\right)$. If $ K $ is finite, this specification implies that $\tau $ should shrink to zero at a rate slightly slower than $\phi _{NT}$. We allow $K\rightarrow \infty $ , provided $K^{5/2}\tau \rightarrow 0$ so that the sampling error in $ \widehat{\boldsymbol{\Sigma }}^{e}$ would not offset the dominant grouping effect of the $\ell _{2}$-relaxation in the presence of latent groups in $ \widehat{\boldsymbol{\Sigma }}^{\mathrm{\ast }}$. The particular rate $ K^{5/2}\tau $ will appear as the order of convergence in Theorem (ref) below. Though the exact number of groups $K$ is usually unknown in reality, if the researcher believes that $K$ is asymptotically dominated by some explicit rate function $\bar{K}_{N,T}$ of $N$ and $T$ in that $ \limsup K/\bar{K}_{N,T}<1$, say $\bar{K}_{N,T}=\left( N\wedge T\right) ^{1/7} $, then all the following theoretical results still hold if $K$ is replaced by $\bar{K}_{N,T}$ and $\tau $ is replaced by $\tau _{\bar{K} }=C_{\tau }\bar{K}_{N,T}^{1/2}\phi _{NT}$ for some positive constant $ C_{\tau }$, and Assumption (ref)(a) is replaced by $\bar{K} _{N,T}^{1/2}\phi _{NT}/\tau +\bar{K}_{N,T}^{5/2}\tau \rightarrow 0$ accordingly.

Assumption (ref)(b) requires that the smallest relative group size $\underline{r}$ be proportional to the reciprocal of $K$. If a group included too few members, the weight of the group would be too small to matter and thus the associated coefficients too difficult to estimate. Assumption (ref) (b) is, indeed, a simplifying condition for notational conciseness. If we drop it, $\underline{r}$ will appear in the rates of convergence in all the following results, which complicates the expressions but adds no new insight.

Recall that the oracle weight vector $\mathbf{w}^{*}$ shares equal weights within each group. Theorem (ref) establishes meaningful convergence for $\widehat{\mathbf{w}}$ to $\mathbf{w}^{*}$.

theoremUnder Assumptions (ref) and (ref), we have \begin{equation*} \left\Vert \widehat{\mathbf{w}}-\mathbf{w}^{*}\right\Vert _{2}=O_{p}\left(N^{-1/2}K^{2}\tau\right)=o_{p}\left(N^{-1/2}\right) and \left\Vert \widehat{\mathbf{w}}-\mathbf{w}^{*}\right\Vert _{1}=O_{p}\left(K^{2}\tau\right)=o_{p}\left(1\right). \end{equation*}
remarkWhen we work with weight vectors of growing dimension, we must be cautious about the rate of convergence. As a trivial example, consider the simple average weight $\widehat{\mathbf{w}}^{\mathrm{SA}}=\boldsymbol{1}_{N}/N$ and an ad hoc oracle weight of two groups $\mathbf{w}^{\ast }=(0.5\cdot \mathbf{1}_{0.5N}^{\prime },1.5\cdot \mathbf{1}_{0.5N}^{\prime })^{\prime }/N $.\footnote{ Without loss of generality, we assume $N$ is an even number here.} In this case, the $\ell _{2}$-distance \begin{equation*} \left\Vert \widehat{\mathbf{w}}^{\mathrm{SA}}-\mathbf{w}^{\ast }\right\Vert _{2}=\Vert 0.5\cdot \boldsymbol{1}_{N}/N\Vert _{2}=0.5/\sqrt{N}\rightarrow 0 \end{equation*} while $\left\Vert \widehat{\mathbf{w}}^{\mathrm{SA}}-\mathbf{w}^{\ast }\right\Vert _{1}=0.5$. It is thus only non-trivial if we manage to show $ \left\Vert \widehat{\mathbf{w}}-\mathbf{w}^{\ast }\right\Vert _{1}=o_{p}\left( 1\right) $ and $\left\Vert \widehat{\mathbf{w}}-\mathbf{w} _{0}^{\ast }\right\Vert _{2}=o_{p}\left( N^{-1/2}\right) $, which is achieved by Theorem (ref).

The convergence further implies desirable oracle (in)equalities in Theorem (ref) below. It shows that the empirical risk under $\widehat{ \mathbf{w}}$ would be asymptotically as small as if we knew the oracle object $\widehat{\boldsymbol{\Sigma }}^{\ast }$.

theorem[Oracle (in)equalities] Under the assumptions in Theorem (ref), we have \begin{enumerate} • $\widehat{\mathbf{w}}^{\prime }\widehat{\boldsymbol{\Sigma }}\widehat{ \mathbf{w}}=\mathbf{w}^{\ast \prime }\widehat{\boldsymbol{\Sigma }}^{\ast } \mathbf{w}^{\ast }+O_{p}\left( \tau K^{5/2}\right) .$ \end{enumerate} Furthermore, let $\widehat{\boldsymbol{\Sigma }}^{\mathrm{new}}$ and $\widehat{\boldsymbol{\Sigma }}^{\ast \mathrm{new}}$ be the counterparts of $\widehat{\boldsymbol{\Sigma }}$ and $\widehat{\boldsymbol{\Sigma }} ^{\ast }$ from a new testing sample of $T^{\mathrm{new}}$ observations where $T^{\mathrm{new}}\asymp T$. The testing sample can be either dependent or independent of the training dataset used to estimate $\widehat{\mathbf{w}}$ and $\mathbf{w}^{\ast }$. If the testing dataset is generated by the same DGP as that of the training dataset, then \begin{enumerate} \setcounter{enumi}{1} • $\widehat{\mathbf{w}}^{\prime }\widehat{\boldsymbol{\Sigma }}^{\mathrm{ new}}\widehat{\mathbf{w}}=\mathbf{w}^{\ast \prime }\widehat{\boldsymbol{ \Sigma }}^{\ast \mathrm{new}}\mathbf{w}^{\ast }+O_{p}\left( \tau K^{5/2}\right) $; • $\widehat{\mathbf{w}}^{\prime }\widehat{\boldsymbol{\Sigma }}\widehat{ \mathbf{w}}\leq \widehat{\mathbf{w}}^{\prime }\widehat{\boldsymbol{\Sigma }} ^{\mathrm{new}}\widehat{\mathbf{w}}\leq Q(\boldsymbol{\Sigma } _{0})+O_{p}(\tau K^{5/2}),$ where $\boldsymbol{\Sigma }_{0}=\boldsymbol{ \Sigma }_{0}^{\ast }+\boldsymbol{\Sigma }_{0}^{e}$ and $Q(\boldsymbol{\Sigma }_{0}):=\min_{\mathbf{w}^{\prime }\boldsymbol{1}_{N}=1}\,\mathbf{w}^{\prime } \boldsymbol{\Sigma }_{0}\mathbf{w} $. \end{enumerate}

Theorem (ref)(a) is an in-sample oracle equality, and (b) is an out-of-sample oracle equality. Because the magnitude of the idiosyncratic shock is controlled by Assumption (ref), the convergence of the weight estimator in Theorem (ref) allows the sample risk $\widehat{\mathbf{w}}^{\prime } \widehat{\boldsymbol{\Sigma }}\widehat{\mathbf{w}}$ to approximate the oracle risk $\mathbf{w}^{\ast \prime }\widehat{\boldsymbol{\Sigma }}^{\ast } \mathbf{w}^{\ast }$. The approximation is nontrivial by noting that $\mathbf{ w}^{\ast \prime }\widehat{\boldsymbol{\Sigma }}^{\ast }\mathbf{w}^{\ast }$ and $\mathbf{w}^{\ast \prime }\widehat{\boldsymbol{\Sigma }}^{\ast \mathrm{ new}}\mathbf{w}^{\ast }$ are bounded away from 0 given the low rank structure of $\widehat{\boldsymbol{\Sigma }}^{\ast \mathrm{new}}$ and that $ \tau K^{5/2}\rightarrow 0$ under Assumption (ref)(a). In other words, the risk of our sample estimator would be as low as if we were informed of the infeasible oracle group membership, up to an asymptotically negligible term $O_{p}\left( \tau K^{5/2}\right)$.

While our $\ell _{2}$-relaxation regularizes the combination weights, there is another line of literature of regularizing the high dimensional VC estimation or its inverse (the precision matrix); see bickel2008regularized, Fan_Liao_Mincheva2013, and the overview by fan2016overview. Theorem (ref)(c) implies that the in-sample and out-of-sample risks coming out of $\ell _{2}$ -relaxation are comparable with the resultant risk from estimating the high dimensional VC matrix. Given that the population VC $\boldsymbol{\Sigma } _{0} $ is the target of high dimensional VC estimation, in forecast combination $Q(\boldsymbol{\Sigma }_{0})$ is bates1969combination's optimal risk, and in portfolio analysis $Q(\boldsymbol{\Sigma }_{0})$ is the global minimum risk. VC $\boldsymbol{\Sigma }_{0}$ takes into account both the low rank component $\boldsymbol{\Sigma }_{0}^{\ast }$ and the high rank component $\boldsymbol{\Sigma }_{0}^{e}$. Even if $\boldsymbol{\Sigma }_{0}$ can be estimated so well that the estimation error is completely eliminated, our out-of-sample risk $\widehat{\mathbf{w}}^{\prime }\widehat{\boldsymbol{ \Sigma }}^{\mathrm{new}}\widehat{\mathbf{w}}$ is within an $O_{p}\left( \tau K^{5/2}\right) $ tolerance level of $Q(\boldsymbol{\Sigma }_{0})$.

remarkWe establish original asymptotic results to support this new $\ell _{2}$ -relaxation method, although they are relegated to the Online Appendix due to their technical nature and the limitations of space. Here we give a roadmap of the theoretical construction. The duality between the sup-norm constraint and the $\ell _{1}$-norm leads to the dual problem ((ref)), which is a linearly constrained quadratic optimization. This dual resembles Lasso tibshirani1996regression in view of its $\ell _{1}$-penalty. Rather than directly working with the primal problem ((ref)), we first develop the asymptotic convergence in the dual. Studies of high dimensional regressions have offered a few inequalities for Lasso to handle sparse regressions. We sharpen these techniques in our context to cope with groupwise sparsity in an innovative way. Once the convergence of the high dimensional parameters in the dual problem is established (See Theorem (ref)), the convergence of the combination weights follows in Theorem (ref), and then the asymptotic optimality in Theorem (ref) proceeds.

Monte Carlo Simulations

In this section, we illustrate the performance of the proposed $\ell _{2}$ -relaxation method via Monte Carlo simulations. We consider two different simulation settings corresponding to the forecasting combination and portfolio optimization in Section (ref).

This paper's numerical works are implemented in MATLAB, with the VC estimates described in Box 1. With modern convex optimization modeling languages and open-source convex solvers, the quadratic optimization with constraints such as ((ref)) can be handled with ease even when $N$ is in hundreds or thousands. Proprietary convex solvers can also be called upon for further speed gain in numerical operations; see gaoimplementing.

center[center omitted — 1,218 chars of source]

Forecast Combination

We assume the simulated data follow a group pattern with the same number of members in each group, i.e., $N_{k}=N/K$ for each $k\in \left[ K\right] $. Let $\mathbf{\Psi }^{\text{co}}$ be a $K\times K$ symmetric positive definite matrix, and $\mathbf{\Psi }=\mathbf{\Psi }^{\text{co}}\otimes ( \mathbf{1}_{N_{1}}\mathbf{1}_{N_{1}}^{\prime })$ be its $N\times N$ block equicorrelation matrix. We consider three DGPs, in which we start with a baseline model of independent factors, and then allow dynamic factors and approximate factors.

description• DGP 1. The baseline model generates $N$ forecasters from \begin{equation} \mathbf{f}_{t}=\mathbf{\Psi }^{1/2}\boldsymbol{\eta }_{t}+\mathbf{u}_{t}, \end{equation} where $\mathbf{\Psi }^{1/2}=N_{1}^{-1/2}(\mathbf{\Psi }^{\text{co} })^{1/2}\otimes (\mathbf{1}_{N_{1}}\mathbf{1}_{N_{1}}^{\prime })$, $ \boldsymbol{\eta }_{t}\sim N(\boldsymbol{0},\mathbf{I}_{N})$ is independent of the idiosyncratic noise $\mathbf{u}_{t}\sim N(\boldsymbol{0},\mathbf{ \Omega } _{u})$, and the latter is independent across $t$. • DGP 2. We extend the baseline model by allowing for temporal serial dependence in $\left\{ \boldsymbol{\eta }_{t}\right\} $. Specifically, for each $i$, we generate $\eta _{it}$ from an AR(1) model \begin{equation*} \eta _{it}=\rho _{i}\eta _{i,t-1}+\epsilon _{it}^{\eta }, \end{equation*} where $\rho _{i}\sim \mathrm{Uniform}(0,0.9)$ is a random autoregressive coefficient, the noise $\epsilon _{it}^{\eta }\sim \mathrm{i.i.d.\;} N(0,1-\rho _{i}^{2})$, and the initial values $\eta_{i0}\sim \mathrm{i.i.d.\; }N(0,1)$. • DGP 3. The equal factor loadings within a group can be an approximation of more general factor loading configurations. This DGP experiments with another extension to the baseline model by defining $\tilde{ \mathbf{\Psi }}^{1/2}:={\mathbf{\Psi }}^{1/2}+\mathrm{i.i.d.\;} N(0,N_{1}^{-1/2})$ as a perturbed factor loading matrix to replace ${\mathbf{ \Psi }}^{1/2}$ in (ref).

The target variable is generated as $y_{t+1}={\mathbf{w}_{\psi }^{\ast }} ^{\prime }\mathbf{\Psi }^{1/2}\boldsymbol{\eta }_{t}+u_{y,t+1},$ where $ u_{y,t+1}\sim N(0,\sigma _{y}^{2})$ is independent of $\boldsymbol{\eta } _{t} $ and $\mathbf{u}_{t},$ and $\mathbf{w}_{\psi }^{\ast }=[(\mathbf{\Psi } ^{\text{co}})^{-1}\mathbf{1}_{K}]\otimes \mathbf{1}_{N_{1}}/[N_{1}\mathbf{1} _{K}^{\prime }(\mathbf{\Psi }^{\text{co}})^{-1}\mathbf{1}_{K}].$ The forecast error vector is

equation*[equation* omitted — 225 chars of source]

and its population VC can be written as

equation*[equation* omitted — 394 chars of source]

By construction, $\mathbf{w}_{\psi }^{\ast } = \arg \underset{\mathbf{w} ^{\prime }\mathbf{1}_{N}=1}{\min }\,\mathbf{w}^{\prime }\mathbf{\Sigma }_{0} \mathbf{w}$.

We compare the following estimators of $\mathbf{w}$, all subject to the restriction $\mathbf{w}^{\prime }\mathbf{1}_{N}=1$: (i) the oracle estimator with known group membership; (ii) simple averaging (SA); (iii) the $\ell _{2} $-relaxation estimator with $\tau = 0$ ($\ell_2$-relax$^0$); (iv) Lasso; (v) Ridge; (vi) the principle component (PC) grouping estimator; and (vii) $\ell _{2}$-relaxation with three different $\widehat{\boldsymbol{ \Sigma }}$ estimators in Box 1.

remarkWe elaborate the rivalries. The oracle estimator takes advantage of the true group membership in the DGP. Given information about the group membership, we reduce the $N$ forecasters to $K$ forecasters $f_{(g_{k}),t}=N_{k}^{-1} \sum_{i\in \mathcal{G}_{k}}f_{it}$ for $k\in \lbrack K]$, and use the low-dimensional ((ref)) to find the optimal weights. For Lasso and Ridge, we recenter the weights toward the SA weights for a fair comparison: \begin{eqnarray*} \widehat{\mathbf{w}}_{Lasso}\ &=&\arg \underset{\mathbf{w}^{\prime } \mathbf{1}_{N}=1}{\min }\,\frac{1}{2}\mathbf{w}^{\prime }\widehat{ \boldsymbol{\Sigma }}\mathbf{w}+\tau \Vert \mathbf{w}-\mathbf{1}_{N}/N\Vert _{1} \\ \widehat{\mathbf{w}}_{Ridge}\ &=&\arg \underset{\mathbf{w}^{\prime } \mathbf{1}_{N}=1}{\min }\,\frac{1}{2}\mathbf{w}^{\prime }\widehat{ \boldsymbol{\Sigma }}\mathbf{w}+\tau \Vert \mathbf{w}-\mathbf{1}_{N}/N\Vert _{2}^{2} \end{eqnarray*} where $\tau $ is the tuning parameter. Furthermore, we estimate the group membership in PC as follows. We compute the $T\times N$ in-sample forecasters' error matrix $\widehat{\mathbf{E}}=(\widehat{\mathbf{e}} _{1},...,\widehat{\mathbf{e}}_{T})^{\prime }$, save the associated $N\times N $ factor loading matrix $\widehat{\mathbf{\Gamma }}$ of the singular decomposition $\widehat{\mathbf{E}}=\widehat{\mathbf{U}}\widehat{\mathbf{D}} \widehat{\mathbf{\Gamma }}^{\prime }$, where $\widehat{\mathbf{D}}$ is the \textquotedblleft diagonal\textquotedblright\ matrix of the singular values in descending order. We extract the the first $q$ columns of $\widehat{ \mathbf{\Gamma }}$, and perform the standard $K$-means clustering algorithm to partition the factor loading vectors into $K$ estimated groups $\widehat{ \mathcal{G}}_{k}$, $k\in \lbrack K]$. We use the true $K$ and try $q=5,10,20$ to avoid tuning on these hyperparameters in this PC grouping procedure.

We estimate the weights $\widehat{\mathbf{w}}$ with the training sample $ \{(y_{t+1},\mathbf{f}_{t}),$ $t\in \lbrack T])$, and then cast the one-step-ahead prediction $\widehat{\mathbf{w}}^{\prime }\mathbf{f}_{T+1}$ for $y_{T+2}$. The above exercise is repeated to evaluate the MSFE $E\left[ (y_{T+2}-\widehat{\mathbf{w}}^{\prime }\mathbf{f}_{T+1})^{2}\right] -\sigma _{y}^{2}$ of each estimator, where the unpredictable components $\sigma _{y}^{2}$ in the MSFE is subtracted and the mathematical expectations are approximated by empirical averages of 1000 simulation replications.\footnote{ We also report the mean absolute forecast error (MAFE), which is not covered by our theory; see Online Appendix (ref) for reference.}

We experiment with three training sample sizes $T=50,100,200$ with the corresponding $K=2,4,6$ and $N=100,200,300$, respectively. We specify

equation*[equation* omitted — 267 chars of source]

and $\mathbf{\Omega }_{u}=\sigma _{u}^{2}\mathbf{I}_{N}$ with $\sigma _{u}=5$ . To highlight the effect of the signal-to-noise ratio (SNR) on the forecast accuracy, we specify $\sigma _{y}=1$ as the low-signal design (with SNR around $3:7$) and $\sigma _{y}=0.1$ as the high-signal design (with SNR around $7:3$). Online Appendix (ref) details the formula of the SNR for our setting.

To implement $\ell _{2}$-relaxation, one needs to choose the tuning parameter $\tau .$ Even though it is beyond the scope of the current paper to provide a formal theoretical analysis on the choice of $\tau $, according to our experience gained from extensive experiments, the commonly used cross-validation (CV) method or its time-series-adjusted version works fairly well in simulations and applications. Here, the tuning parameters for DGPs 1 and 3 are obtained by the conventional 5-fold CV through a grid search, detailed in Box 2; we have also tried the 3-fold CV and the 10-fold CV, and the results are qualitatively intact. This 5-fold CV that randomly permutes the data accounts for neither the chronological order nor the serial correlation of the time series data, however. Practitioners usually resort to the out-of-sample (OOS) evaluation instead.\footnote{ See bergmeir2012 and mirakyan2017, among others. See also arlot2010 for a survey of cross-validation procedures for model selection.} Algorithm for the OOS evaluation, applied to DGP 2, is provided in Box 3.

center[center omitted — 1,737 chars of source]

Figure (ref) illustrates the estimated weights of a typical replication under DGP 1. The four rows of sub-figures correspond to the oracle, the $ \ell _{2}$-relaxation with $\widehat{\boldsymbol{\Sigma }}_{2}$, Lasso, and ridge, respectively; the three columns represent the results under $K=2,4,6$ , respectively. We unify the scale of axes for the four subplots in each column to facilitate comparison. For each sub-figure, the estimated weights are plotted against $\left[ N\right] $. Although individuals are not explicitly classified into groups, $\ell _{2}$-relaxation estimates exhibits grouping patterns that mimic the oracle weights. Such patterns are observed in neither Lasso nor Ridge.

figure[figure omitted — 139 chars of source]

The six panels in Table (ref) report the out-of-sample prediction accuracy by MSFE for all three DGPs with low and high SNRs, respectively. \footnote{ Besides $\ell_2$-relaxation, Ridge and Lasso also require tuning parameters, which are selected in the same fashion. For DGPs 1 and 3, we use the 5-fold cross-validation (Box 2); and for DGP 2, we use the out-of-sample evaluation approach (Box 3). For the best empirical performance, we use the nonlinear shrinkage VC estimator $\widehat{ \boldsymbol{\Sigma }}_{2}$ for $\widehat{\boldsymbol{\Sigma }}$. In addition, for the Ridge estimation, it is easy to verify that centering the weights around $1/N$ or any other constant yields the same estimator $ \widehat{\mathbf{w}}_{\text{Ridge}}$ when the constraint $\mathbf{w}^{\prime }\mathbf{1}_{N}=1$ is imposed.} The first three columns show the settings of $T$, $N$ and $K$, and the following columns show the MSFEs of the labeled estimators. All estimators have stronger performance under a high SNR than that under a low SNR. The rankings of relative performance among the six estimators are similar across different DGPs despite that the additional factor loading noises enlarge the MSFEs of all estimators from DGP 3 relative to those from DGP 1.

table[table omitted — 3,119 chars of source]

The infeasible grouping information helps the oracle estimator to prevail in all cases. Regardless which $\widehat{\boldsymbol{\Sigma }}$ estimator is employed, $\ell _{2}$-relaxation outperforms feasible competitors and its MSFE approaches that of the oracle estimator. The $\ell _{2}$-relaxation with $\widehat{\boldsymbol{\Sigma }}_{2}$ generally achieves the best performance among all feasible estimators. Lasso and Ridge are in general better than the PC estimator. Given the group structures in the DGP, SA in general lags far behind the other feasible estimators that learn the combination weights from the data. Notice that the MSFEs by Oracle, $\ell_2$ -relax$^0$, Lasso, Ridge, and the $\ell _{2}$-relaxation decrease as $T$ grows along with $N$. However, the results by SA and PC under all values of $ q$ may diverge as $(T,N)$ increases.

Our results are not sensitive to different evaluation methods. bergmeir2018 argue that the standard 5-fold CV is valid in purely autoregressive models with uncorrelated errors. Simulation results for DGP 2 by the conventional 5-fold CV are reported in Appendix (ref). In summary, we observe robust performance of $\ell _{2}$-relaxation superior to the other feasible estimators across the DGP designs, signal strength, and CV methods.

Portfolio Analysis

We extend fan2012vast's design to a simulated Fama-French five-factor model. Besides the market factor, fama2015 identify four additional factors capturing the size, value, profitability, and investment patterns in average stock returns. Let $R_{i}$ be the excessive return of the $i$th stock. The five-factor model is similar to (ref):

equation[equation omitted — 103 chars of source]

where $\boldsymbol{\lambda }_{i}=\{\lambda _{ij}\}_{j=1}^{5}$ is the vector of 5 factor loadings.

table[table omitted — 1,844 chars of source]

We simulate the returns for $N=100$ assets and $T=240$ months. The factors and factor loadings are generated from the multivariate normal distributions $N(\mathbf{\mu }_{f},\mathbf{cov}_{f})$ and $N(\mathbf{\mu }_{\lambda }, \mathbf{cov}_{\lambda })$, respectively. The values of the parameters $( \mathbf{\mu }_{b},\mathbf{cov}_{b},\mathbf{\mu }_{f},\mathbf{cov}_{f})$ are displayed in Table (ref), which are calibrated to the 2001--2020 real market data of Fama and French 100 portfolios on the size and book-to-market (Size-BM), size and investment (Size-INV), and size and operating profitability (Size-OP).\footnote{ The factor and portfolio data are available at \url{https://mba.tuck.dartmouth.edu/pages/faculty/ken.french/data _library.html}.} The idiosyncratic noises are generated from $N(\mathbf{0}, \mathbf{cov}_{u})$, where $\mathbf{cov}_{u}$ is the sample VC matrix of the residuals from the OLS estimation of (ref).

Instead of MFSE, we use the Sharpe ratio as the criterion for MVP. The $\ell _{2} $-relaxation estimator allows negative weights, which correspond to short positions of financial assets. For each repetition, the following rolling window estimation is considered with window length $L$. We avoid recursively training models for each month. Following a similar strategy in xiu2020, we train and roll forward once every year, as elaborated in Box 4.

center[center omitted — 792 chars of source]
table[table omitted — 1,003 chars of source]

We consider training data of lengths $L=60$ (5 years) and 120 (10 years). The $\ell _{2}$-relaxation estimator is compared with SA and the gross exposure constraints (GEC) methods fan2012vast:

equation*[equation* omitted — 236 chars of source]

where the exposure constraint is set as $c=1$ (no short exposure) or $c=2$ (allowing 50% short exposure).\footnote{ Similar to the numerical implementation of Lasso and Ridge, throughout this paper GEC is estimated with the VC matrix $\widehat{\boldsymbol{\Sigma }} _{2} $ for a fair comparison with the best $\ell _{2}$-relaxation outcomes in most cases.} Table (ref) reports the Sharpe ratios averaged over 1000 replications. The three panels correspond to portfolios sorted by Size-BM, Size-INV, and Size-OP, respectively. SA performs poorly in terms of yielding the lowest Sharpe ratios in all cases. GEC with a short exposure of 50% is better than those without. For the case of $\ell _{2}$-relaxation, the three choices of $\widehat{\boldsymbol{\Sigma }}$ lead to similar Sharpe ratios, which all outperform GEC and SA.

Empirical Applications

In this section, we explore three empirical examples. In the first two applications, we assess the MSFE\footnote{ The MAFE results are available in Appendix (ref).} of a microeconomic study of forecasting box office and a macroeconomic exercise for the survey of professional forecasters (SPF). The last one is a financial application of MVP evaluated by the Sharpe ratio.

Box Office

The motion picture industry devotes enormous resources to marketing in order to influence consumer sentiment toward their products. These resources are intended to reduce the supply-demand friction on the market. On the supply side, movie making is an expensive business; on the demand side, however, the audience's taste is notoriously fickle. Accurate prediction of box office is financially crucial for motion picture investors.

Based on the data of Hollywood movies released in North America between October 1, 2010 and June 30, 2012, lehrer2017 demonstrate the sound out-of-sample performance of the prediction model averaging (PMA). We revisit their dataset of 94 cross-sectional observations (movies), 28 non-constant explanatory variables and 95 candidate forecasters according to a multitude of model specifications. Guided by the intuition that the input variables capturing similar characteristics are \textquotedblleft closer\textquotedblright\ to one another, lehrer2017 cluster input variables into six groups in their Appendix D.1:

align*[align* omitted — 505 chars of source]

Since the 95 forecasters are generated based on these input variables, the potential group patterns may help $\ell _{2}$-relaxation achieve more accurate forecasts than other off-the-shelf machine learning shrinkage methods in this setting.

Following lehrer2017, we randomly rearrange the full sample with $ n=94 $ movies into a training set of the size $n_{\mathrm{tr}}$ and an evaluation set of the size $n_{\mathrm{ev}}=n-n_{\mathrm{tr}}$, which we experiment with $n_{\mathrm{ev}}=10,$ $20,$ $30$, and 40. We repeat this procedure for 1,000 times and evaluate the MSFE of PMA, $\ell_2$-relax$^0$, the CSR elliott2013complete, the peLASSO Diebold_Shin2019, Lasso, Ridge, and $\ell _{2}$-relaxation. We choose the number of subset regressors to be 10 and 15 for the CSR approach, denoted as CSR$_{10}$ and CSR$_{15}$, respectively. Since the total number of candidate models is too large to handle, we follow genre2013 and randomly pick 10000 candidate models instead. For peLASSO, we follow Diebold_Shin2019 and conduct a two-step estimation (first Lasso, then Ridge). Movies are viewed as independent observations and thus the tuning parameters are chosen by the conventional 5-fold CV. We conduct a grid search from 0 to 5 with increment 0.1.

table[table omitted — 853 chars of source]

Since the magnitude of MSFEs varies on the evaluation sizes $n_{\mathrm{ev}}$ , we report in Table (ref) the mean risk relative to that of PMA for convenience of comparison. Entries smaller than 1 indicate better performance relative to that of PMA. While PMA is known to outperform Lasso and Ridge in lehrer2017, in this exercise it also outperforms $\ell_2$ -relax$^0$, CSR, and peLasso. Shrinkage toward the global equal weight $1/N$ or toward 0 is not favored in this experiment. $\ell _{2}$-relaxation, on the other hand, yields lower risk than PMA under any $\widehat{\boldsymbol{ \Sigma }}$, and the edge generally increases with the value of $n_{\mathrm{ev }}$.

To demonstrate the potential grouping pattern in the data, we show the estimated weights of a typical replication on $n_{\mathrm{ev}}=10$ and $\tau =1$ in Figure (ref). The pattern is similar for other values of $n_{ \mathrm{ev}}$. The vertical axis of Figure (ref) represents the estimated weights, and the horizontal axis shows all the 95 forecasters order by the weights from high to low. In addition, we divide the weights into five groups according to the following manually selected intervals: $ (-\infty ,-0.05)$, $(0.05,0)$, $(0,0.05)$, $(0.05,0.2)$, and $(0.2,+\infty )$ . Circles and solid-lines represent the weights and group means, respectively, in Figure (ref). The results demonstrate potential latent grouping pattern. Interesting, more than 40% of the individual models receive negative weights.

figure[figure omitted — 175 chars of source]

Inflation

Firms, consumers, as well as monetary policy authorities count on the outlook of inflation to make rational economic decisions. Besides model-based inflation forecasts published by government and research institutes, SPF reports experts' perceptions about the price level movement in the future. A long-standing myth of forecast combination lies in the robustness of the simple average which extract the mean or median as a predictor in a simple linear regression, as documented by ang2007macro. Recent research shows modern machine learning methods can assist by assigning data-driven weights to individual forecasters to gather disaggregate information; see, e.g., Diebold_Shin2019.

The European Central Bank's SPF inquires many professional institutions for their expectations of the euro-zone macroeconomic outlook. We revisit genre2013's harmonized index of consumer prices (HICP) dataset, which covers 1999Q1--2018Q4. The experts were asked about their one-year- and two-year-ahead predictions. The raw data record 119 forecasters in total, but are highly unbalanced with plenty of missing values, mainly due to entry and exit in the long time span. We follow genre2013 to obtain 30 qualified forecasters by first filtering out irregular respondents if he or she missed more than 50% of the observations, and then using a simple AR(1) regression to interpolate the missing values in the middle.

table[table omitted — 590 chars of source]

Our benchmark is the simple average (SA) on all 30 forecasters. We compare the forecast errors of SA, $\ell_2$-relax$^0$, Lasso, Ridge, and $\ell _{2}$ -relaxation. We use a rolling window of 40 quarters for estimation. The tuning parameters are selected by the OOS approach described in Box 2 with grid search from 0 to 5 with increment 0.01. The results of relative risks are presented in Table (ref), with the MSFEs of SA standardized as 1. $\ell_2$-relax$^0$ performs worse than SA. Lasso and Ridge yield roughly 15% improvement relative to SA. $\ell _{2}$-relaxation exhibits robust performance under all choices of $\widehat{\boldsymbol{ \Sigma }}$.

figure[figure omitted — 156 chars of source]

Since we do not directly observe the underlying factors based upon which the forecasters make decisions, we illustrate in Figure (ref) the estimated weights associated with the 30 forecasters of a typical roll from 1999Q1 to 2008Q4 and $\tau =0.02$. Sub-figures (a) and (b) are associated with results of one-year-ahead and two-year-ahead forecasting, respectively. The horizontal axis shows the forecasters and the vertical axis represents the estimated weights. The weights can be roughly categorized into five groups according to the following manually selected intervals: $(-\infty ,-0.3)$, $(-0.3,-0.1)$, $(-0.1,0.2)$, $(0.2,0.4)$, and $(0.4,+\infty )$. In both sub-figures, the spread of the weights deviates the equal-weight SA strategy. In sub-figure (b) the weights are more concentrated around 0 than sub-figure (a), reflecting the challenges to forecast over a longer horizon.

Fama and French 100 Portfolios

Here we mimic the simulation design in Section (ref) but feed the algorithms with the real 2001--2020 Fama and French 100 monthly portfolios on Size-BM, Size-INV, and Size-OP. The empirical results are presented in Table (ref).

table[table omitted — 1,038 chars of source]

The benchmark SA suggested by demiguel2009generalized delivers higher Sharpe ratios under the longer rolling window, indicating substantial noise in the simple aggregation over the cross section when $L$ is small. The performance of SA is eclipsed by GEC and $\ell _{2}$-relaxation in all cases. GEC without short exposure ($c=1$) wins the Size-INV with $L=60$, and that with 50% short exposure ($c=2$) wins the Size-OP with $L=60$. In four out of six cases, nevertheless, $\ell _{2}$-relaxation delivers the highest Sharpe ratios.

figure[figure omitted — 410 chars of source]

To better understand the behavior of $\ell _{2}$-relaxation, we plot in Figure (ref) its estimated weights of a typical estimation window with $ L=60$ and 120, and the value of tuning parameter set to $\tau =1$. We assign the 100 portfolios into 5 groups in each case and plot the group means by the red horizontal lines. The weights are manually categorized into five intervals: $(-\infty ,-0.05)$, $(-0.05,0)$, $(0,0.05)$, $(0.05,0.15)$, and $ (0.15,+\infty )$. The distribution of the weights are similar across $L$ in the sub-figure for the Size-BM sorted portfolios. Under the Size-INV sorting, however, the weights for $L=60$ are less spread and many are close to zero, and the weights under the Size-OP portfolios share similar patterns. This phenomenon helps to explain the high Sharpe ratio of GEC under $L=60$: intuitively, its $\ell _{1}$-norm restriction $\Vert \mathbf{w} \Vert _{1}\leq c$ would shrink all weights toward zero and moreover push many small weights to be exactly zero.

Acknowledging that the theoretical setup of parsimonious factors and group structure is an approximation of the real financial world at best, in future studies we are interested in investigating an enhanced $\ell _{2}$ -relaxation with the exposure constraint $\Vert \mathbf{w}\Vert _{1}\leq c$.

Conclusion

This paper presents a new machine learning algorithm, namely, $\ell _{2}$ -relaxation. When the forecast error VC or the portfolio VC can be approximated by a block equicorrelation structure, we establish its consistency and asymptotic optimality in the high dimensional context. Simulations and real data applications demonstrate excellent performance of the $\ell _{2}$-relaxation method.

Our work raises several interesting issues for further research. First, we have not studied the optimal choice of the tuning parameter $\tau $ or provided a formal justification for the use of the cross-validated $\tau . $ Recently, wu2020survey have reviewed the literature of tuning parameter selection for high dimensional regressions and discussed various strategies to choose the tuning parameter to achieve either prediction accuracy or support recovery such as the $L$-fold cross-validation, $m$ -out-of-$n$ bootstrap and extended Bayesian information criterion (BIC). chetverikov2021cross have studied the theoretical properties of Lasso based on cross-validated choice of the tuning parameter. It will be interesting to study whether we can draw support from these papers to provide theoretical guidance concerning the choice of $\tau $ in our context.

Second, additional restrictions can be imposed to accompany the $\ell _{2}$ -relaxation problem. For example, if sparsity is desirable, we may consider adding the exposure constraint $\left\Vert \mathbf{w}\right\Vert _{1}\leq c$ for some tuning parameter $c$, which echoes the idea of mixed $\ell _{1}$- and $\ell _{2}$-penalty of the elastic net method by Zou_Hastie2005. Another example is to incorporate the constraints $w_{i}\geq 0$ for all $i$ if non-negative weights are desirable jagannathan2003risk. Third, our $\ell _{2}$-relaxation is motivated from the MSFE loss function, it is possible to consider other forms of relaxation if the other forms of loss functions (e.g., MAFE) are under investigation. Last but not least, the $ \ell _{2}$-relaxation theory in this paper requires the dominant component $ \mathbf{\hat{\Sigma}}^{\ast }$ of the VC matrix to have a latent group structure. It is desirable to extend the theory to the case where $\mathbf{ \hat{\Sigma}}^{\ast }$ has only a low-rank structure instead of a latent group structure. We shall explore some of these topics in future works.

{{