Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
78,855 characters · 14 sections · 52 citation commands
Large Volatility Matrix Analysis Using Global and National Factor Models
\pagenumbering{arabic}
\doublespacing
High dimensional factor analysis and principal component analysis (PCA), which are powerful tools for dimension reduction, have been extensively studied bai2003inferential, bernanke2005measuring, stock2002forecasting. They have various applications in economics and finance, such as forecasting macroeconomic variables and optimal portfolio allocation. Recently, a multi-level factor structure with global and local factors has received increasing attention ando2016panel, bai2015identification, choi2018multilevel, han2021shrinkage. The global factors impact on all individuals, while the local factors only impact on those in the specific group. The local group can be defined by regional, country, or industry level. In many economic and financial applications, the local or country factors naturally exist. For example, kose2003international characterized the comovement of international business cycles on the global, regional, and country levels by imposing a multi-level factor structure on a dynamic factor model. moench2013dynamic showed that the local factors play an important role in explaining the U.S. real activities. See also ando2015asset, ando2017clustering, aruoba2011globalization, gregory1999common, hallin2011dynamic for related articles. Hence, when the local factor structure is ignored, the conventional factor analysis might yield misleading results.
Based on the factor models, several large volatility matrix estimation procedures have been developed to account for the strong cross-sectional correlation in the stock market. For example, fan2008high studied the impacts of covariance matrix estimation on optimal portfolio allocation and portfolio risk assessment when the factors are observable. In contrast, fan2013large considered latent factor models and developed the covariance matrix estimation procedure by PCA and thresholding, which is called the principal orthogonal component thresholding (POET). In this procedure, the unobservable factors can be consistently estimated with a large number of assets. Recent studies, such as ait2017using, fan2016incorporating, fan2018large, fan2019structured, jung2022next, wang2017asymptotics, have also been conducted along this direction. However, when considering the global market, we often observe not only the global common factors but also the nation-specific risk factors kose2003international. That is, the single-level factor model cannot sufficiently explain volatility dynamics.
This paper proposes a novel large volatility matrix estimation procedure that incorporates a global and national factor structure as well as a sparse idiosyncratic volatility matrix. Specifically, we consider the global asset market, and, to account for the nation-specific risk factors, we assume that there are latent multi-level factors, such as the global common factors and national risk factors. Since the national risk is a regional risk, we further assume that the local factor membership is known. Under this latent multi-level factor models, we first use the PCA procedure to capture the latent global factors. Then, after removing the latent global factors, we apply the PCA procedure in each national group to accommodate the latent national factors. Finally, to account for the sparse idiosyncratic volatility matrix, we adopt an adaptive thresholding scheme, which we call this the Double Principal Orthogonal complEment Thresholding (Double-POET). We then derive convergence rates for Double-POET and its inverse under different matrix norms. When the local factor membership is unknown, we suggest the regularized spectral clustering method amini2013pseudo to detect the latent local structure. We further discuss the benefit of the proposed Double-POET compared to the regular POET procedure. For example, if we ignore the local factors and treat them as idiosyncratic and employ the POET procedure, the POET estimator can be inconsistent. When estimating the local volatility matrix for each country, Double-POET can estimate global factors better than the regular POET estimator. That is, we can find the blessing of dimensionality. The empirical study supports the theoretical findings.
The rest of the paper is organized as follows. Section (ref) sets up the model and proposes the Double-POET estimation procedure. Section (ref) presents an asymptotic analysis of the Double-POET estimators. The merits of the proposed method are illustrated by a simulation study in Section (ref) and by real data application on portfolio allocation in Section (ref). In Section (ref), we conclude the study. All proofs are presented in Appendix (ref).
Throughout this paper, let $\lambda_{\min}(\mbox{\bf A})$ and $\lambda_{\max}(\mbox{\bf A})$ denote the minimum and maximum eigenvalues of a matrix $\mbox{\bf A}$, respectively. In addition, we denote by $\|\mbox{\bf A}\|_{F}$, $\|\mbox{\bf A}\|_{2}$ (or $\|\mbox{\bf A}\|$ for short), $\|\mbox{\bf A}\|_{\infty}$, and $\|\mbox{\bf A}\|_{\max}$ the Frobenius norm, operator norm, $l_{\infty}$-norm and elementwise norm, which are defined respectively as $\|\mbox{\bf A}\|_{F} = \mathrm{tr}^{1/2}(\mbox{\bf A}'\mbox{\bf A})$, $\|\mbox{\bf A}\|_{2} = \lambda_{\max}^{1/2}(\mbox{\bf A}'\mbox{\bf A})$, $\|\mbox{\bf A}\|_{\infty} = \max_{i}\sum_{j}|a_{ij}|$, and $\|\mbox{\bf A}\|_{\max} = \max_{i,j}|a_{ij}|$. When $\mbox{\bf A}$ is a vector, the maximum norm is denoted as $\|\mbox{\bf A}\|_{\infty}=\max_{i}|a_{i}|$, and both $\|\mbox{\bf A}\|$ and $\|\mbox{\bf A}\|_{F}$ are equal to the Euclidean norm. We denote $\mathrm{diag}(\mbox{\bf A}_{1},\ldots, \mbox{\bf A}_{n})$ with the diagonal block entries as $\mbox{\bf A}_{1},\ldots, \mbox{\bf A}_{n}$.
We consider the following multi-level factor model:
where $y_{it}$ is the observed data for the $i$th individual at time $t$, for $i = 1,\ldots, p$, and $t = 1,\ldots,T$; $G_{t}$ is a $k \times 1$ vector of unobservable “global" common factors, $b_{i}$ indicates the corresponding factor loadings; $f_{t}^{g_{i}}$ is an $r_{g_i} \times 1$ vector of unobservable “local" factors that only affect the group $g_{i}$, $\lambda_{i}^{g_{i}}$ indicates the corresponding factor loadings in each group; and $u_{it}$ is an idiosyncratic error component. We denote the cluster or group membership as $g_{i} \in \{1,\dots, J\}$. Throughout the paper, we assume that the group membership $\{g_{i}\}_{i=1}^{p}$ is known and the global and local factors are latent. In this paper, we consider the nation-specific local factors; thus, the group membership is the country. In addition, the numbers of global factors and the number of local factors in each group are fixed.
Given the group membership, for $g_i = j$, we can stack the observations and denote them as $y_{t} \equiv ({y_{t}^{1}}',\dots, {y_{t}^{J}}')'$, where $y_{t}^j = (y_{1t}, \dots, y_{p_jt})'$ and the number of individuals $p_j$ within group $j$ such that $p = \sum_{j=1}^{J}p_j$. Let $F_{t} = {(f_{t}^{1}}', \dots, {f_{t}^{J}}')'$, where $f_{t}^{j}$ is an $r_j \times 1$ vector of local factors and the number of factors $r_j$ for each group $j$ such that $r = \sum_{j=1}^{J} r_j$. We define the $p \times r$ block diagonal matrix as $$\mbox{\boldmath $\Lambda$} = \mathrm{diag}(\Lambda^1, \Lambda^2, \dots, \Lambda^J),$$ where $\Lambda^{j} = (\lambda_{1}^{j}, \dots, \lambda_{p_{j}}^{j})'$ is a $p_j \times r_j$ matrix of local factor loadings for each $j$. Then, the model ((ref)) can be written in vector form as follows:
where $\mbox{\bf B} = (b_1, \dots, b_p)'$ and $u_{t} = (u_{1t}, \dots, u_{pt})'$. In matrix notation, we have
where $\mbox{\bf Y} = (y_{1},\dots,y_{T})'$ is a $T \times p$ matrix of observed data, $\mbox{\bf G} = (G_{1},\dots, G_{T})'$ is a $T \times k$ matrix of global factors, $\mbox{\bf F} = (F_{1}, \dots, F_{T})'$ is a $T \times r$ matrix of local factors, and $\mbox{\bf U} = (u_{1}, \dots, u_{T})'$ is a $T\times p$ matrix of idiosyncratic errors that are uncorrelated with global and local factors. Throughout the paper, we further assume that global and local factors are orthogonal each other.
In this paper, the parameter of interest is the $p \times p$ covariance matrix, $\mbox{\boldmath $\Sigma$}$, of $y_{t}$ as well as its inverse. Given model ((ref)), the covariance matrix can be written as
where $\mbox{\boldmath $\Sigma$}_{u}$ is a sparse idiosyncratic covariance matrix. We can demonstrate the multi-level factor analysis in the presence of spiked eigenvalues at different levels. Specifically, decomposition ((ref)) can be written as
where $\mbox{\boldmath $\Sigma$}_{E} = \mbox{\boldmath $\Lambda$} \mathrm{cov}(F_{t})\mbox{\boldmath $\Lambda$}' + \mbox{\boldmath $\Sigma$}_{u}.$ Decomposition (ref) is a usual single-level factor-based covariance matrix. Thus, based on (ref), we can apply the regular POET procedure. However, the eigenvalue of the covariance matrix $\mbox{\boldmath $\Sigma$}_{E}$ diverges, which makes the POET procedure inefficient. In Section (ref), we discuss this inefficiency of the regular POET procedure. To further accommodate the local factor structure, we decompose the covariance matrix $\mbox{\boldmath $\Sigma$}_{E}$ as follows:
where the $j$th group covariance matrix $\mbox{\boldmath $\Sigma$}_{E}^{j}$ is a $p_j \times p_j$ diagonal block of $\mbox{\boldmath $\Sigma$}_{E}$. We assume that the idiosyncratic covariance matrix $\mbox{\boldmath $\Sigma$}_{u}=(\sigma_{u,ij})_{p\times p}$ is sparse as follows:
for some $q \in [0,1)$, where $m_p$ diverges slowly, such as $\log p$. Intuitively, after removing the global and local factor components, most of pairs are weakly correlated bickel2008covariance, cai2011adaptive. Theoretically, since $\|\mbox{\boldmath $\Sigma$}_{u}\| \leq \|\mbox{\boldmath $\Sigma$}_{u}\|_{1} = O(m_{p})$, when $m_{p} = o(p_{j})$ for all $j \leq J$, there are distinguished eigenvalues between the local factor components and the idiosyncratic error components. In light of this, in many applications, the sparsity condition on the factor model residuals has been considered boivin2006more, fan2016incorporating. Thus, $\mbox{\boldmath $\Sigma$}_{E}^{j}$ is the low-rank plus sparse structure, which has been widely used ait2017using, cai2013sparse, candes2009exact, fan2019structured, fan2008high, fan2013large, johnstone2009consistency, ma2013sparse, negahban2011estimation. When considering the U.S. market only, we have the usual single-level factor-based form. Then, based on (ref), we can apply the regular POET procedure to estimate the local covariance matrix. However, this approach does not use assets outside of the local group, which causes inefficiency. We also discuss this inefficiency in Section (ref). To accommodate the multi-level factor structure, we propose a novel large covariance matrix estimation procedure in the following section.
To make model ((ref)) identifiable, without loss of generality, we impose normalization conditions: $\mathrm{cov}(G_{t}) = \mbox{\bf I}_{k}$ and $\mathrm{cov}(f_{t}^{j}) = \mbox{\bf I}_{r_j}$, where $G_{t}$ and $f_{t}^{j}$ are uncorrelated with each other; both $\mbox{\bf B}'\mbox{\bf B}$ and ${\Lambda^{j}}'\Lambda^{j}$ for $j \in \{1, \dots, J\}$ are diagonal matrices. We also assume the pervasiveness conditions, such that (i) the eigenvalues of $p^{-a_{1}}\mbox{\bf B}'\mbox{\bf B}$ are strictly greater than zero, and (ii) the eigenvalues of $p_{j}^{-a_{2}}{\Lambda^{j}}'\Lambda^{j}$ are strictly greater than zero for each $j$, where $a_{1}, a_{2} \in (0,1]$ are the strengths of global and local factors, respectively. Then, the first $k$ eigenvalues of $\mbox{\bf B} \mathrm{cov}(G_{t})\mbox{\bf B}'$ diverge at rate $O(p^{a_{1}})$, while the first $r$ eigenvalues of $ \mbox{\boldmath $\Lambda$} \mathrm{cov}(F_{t})\mbox{\boldmath $\Lambda$}'$ diverge at rate $O(p^{ca_{2}})$, which does not grow too fast as $p^{a_1} \rightarrow \infty$. We note that $p^c$ is related to the number of stocks in each country. In addition, all eigenvalues of $\mbox{\boldmath $\Sigma$}_u$ are bounded. Let $\mbox{\boldmath $\Gamma$} = \mathrm{diag}(\delta_1,\dots, \delta_k)$ be the leading eigenvalues of $\mbox{\boldmath $\Sigma$}$ and $\mbox{\bf V} = (v_1,\dots, v_k)$ be their corresponding leading eigenvectors.
To accommodate the multi-level factor model (ref), we propose a non-parametric estimator of $\mbox{\boldmath $\Sigma$}$ as follows:
The above procedure can be equivalently represented by a least squares method. In particular, the global and local factor matrices and corresponding loading matrices are estimated as follows. To obtain the global factor part, we first solve the following optimization problem:
subject to the normalization constraints such that
The columns of $\widehat{\mbox{\bf G}}/\sqrt{T}$ are the eigenvectors corresponding to the $k$ largest eigenvalues of the $T\times T$ matrix $T^{-1}\mbox{\bf Y}\mbox{\bf Y}'$ and $\widehat{\mbox{\bf B}} = (\hat{b}_{1}, \dots, \hat{b}_{p})' = T^{-1}\mbox{\bf Y}'\widehat{\mbox{\bf G}}$. Given $\widehat{\mbox{\bf G}}$ and $\widehat{\mbox{\bf B}}$, let a $T \times p$ matrix $\widehat{\mbox{\bf E}} = \mbox{\bf Y}-\widehat{\mbox{\bf G}}\widehat{\mbox{\bf B}}' \equiv (\widehat{E}^{1},\dots, \widehat{E}^{J}) $. Then, with $\widehat{\mbox{\bf E}}$, we estimate the local factor part as follows:
subject to the normalization constraints such that
Then, we obtain $\widehat{\mbox{\bf F}} = (\widehat{F}^{1}, \dots, \widehat{F}^{J})$ and $\widehat{\mbox{\boldmath $\Lambda$}} = \mathrm{diag}(\widehat{\Lambda}^{1}, \dots, \widehat{\Lambda}^{J})$, where the columns of $\widehat{F}^{j}/\sqrt{T}$ are the eigenvectors corresponding to the $r_{j}$ largest eigenvalues of the $T\times T$ matrix $T^{-1}\widehat{E}^{j}\widehat{E}^{j\prime}$ and $\widehat{\Lambda}^{j} = (\hat{\lambda}_{1}^{j}, \dots, \hat{\lambda}_{p_{j}}^{j}) = T^{-1} \widehat{E}^{j\prime}\widehat{F}^{j}$ for $j=1,2,\dots,J$. Finally, we apply the above adaptive thresholding method to $\widehat{\mbox{\boldmath $\Sigma$}}_{u} = T^{-1}\widehat{\mbox{\bf U}}'\widehat{\mbox{\bf U}}$, where $\widehat{\mbox{\bf U}} = \mbox{\bf Y} -\widehat{\mbox{\bf G}}\widehat{\mbox{\bf B}}' - \widehat{\mbox{\bf F}}\widehat{\mbox{\boldmath $\Lambda$}}'$. Similar to the decomposition ((ref)), we have the following substitution estimators:
Similar to the proofs of Theorem 1 of fan2013large, we can show that the estimators ((ref)) based on PCA and ((ref)) based on least squares approaches are equivalent. In particular, by the Eckart-Young theorem, the estimators $\widehat{\mbox{\bf B}}$ and $\widehat{\Lambda}^j$, after normalization, are the first $k$ and $r_{j}$ empirical eigenvectors of the sample covariance matrix $\widehat{\mbox{\boldmath $\Sigma$}}$ and $\widehat{\mbox{\boldmath $\Sigma$}}_{E}^j$, respectively. Then, there exist orthogonal matrices $\mbox{\bf H}$ and $\mbox{\bf J}$ such that
where $\widetilde{\omega}_{T}$ and $\omega_T$ are defined in Section (ref) and Theorem (ref), respectively. The above rates can be easily obtained using the preliminary results in Appendix (ref).
In summary, given the knowledge of group membership, we employ PCA on each diagonal block of the remaining components of the sample covariance matrix, after removing the first $k$ principal components. Then, we apply thresholding on the remaining residual components. We call this procedure the Double Principal Orthogonal complEment Thresholding (Double-POET). Double-POET is the two-step estimation procedure, which makes it possible to estimate global and local factors separately by considering the block structure on the local factors. In contrast, if we employ the single step estimation procedure, such as POET, the local factor structure cannot be explained well. We discuss the theoretical inefficiency of POET in Sections (ref) and (ref), and the numerical study illustrated in Section (ref) shows that Double-POET outperforms POET.
In this section, we establish the asymptotic properties of the Double-POET estimator. To do this, we impose the following technical conditions.
To measure large matrix estimation errors, we consider the relative Frobenius norm, introduced by stein1961estimation: $$ \|\widehat{\mbox{\boldmath $\Sigma$}}-\mbox{\boldmath $\Sigma$}\|_{\Sigma} = p^{-1/2}\|\mbox{\boldmath $\Sigma$}^{-1/2}\widehat{\mbox{\boldmath $\Sigma$}}\mbox{\boldmath $\Sigma$}^{-1/2}-\mbox{\bf I}_{p}\|_{F}. $$ Note that the factor $p^{-1/2}$ performs normalization, such that $\|\mbox{\boldmath $\Sigma$}\|_{\Sigma} = 1$. fan2013large showed that, under this relative Frobenius norm, the POET estimator is consistent as long as $p=o(T^{2})$, while the sample covariance is consistent only if $p=o(T)$ in the approximate single-level factor model. In Section (ref), we will compare the convergence rates of POET and Double-POET in a multi-level factor model.
The following theorem provides the convergence rates under various norms.
In this subsection, we analyze the regular POET of fan2013large in the multi-level factor models and compare it with the proposed Double-POET.
The regular POET method only captures the global factors and regards both the local factor structure and idiosyncratic terms as the idiosyncratic part. That is, the sparsity level of $\mbox{\boldmath $\Sigma$}_{E} = (\varepsilon_{ij})_{p\times p}$ is $$ \mu_{p} = \max_{i\leq p}\sum_{j\leq p} |\varepsilon_{ij}|^{q}, $$ for some $q \in [0,1)$. We note that when $q=0$, $\mu_{p} \asymp p^{c}$, which corresponds to the maximum number of nonzero elements in each row of $\mbox{\boldmath $\Sigma$}_{E}$. Then, the thresholding approach is applied to $\widehat{\mbox{\boldmath $\Sigma$}}_{E}$ instead of $\widehat{\mbox{\boldmath $\Sigma$}}_{u}$ to obtain $\widehat{\mbox{\boldmath $\Sigma$}}_{E}^{\mathcal{T}}$. Therefore, the regular POET estimator is defined as $$ \widehat{\mbox{\boldmath $\Sigma$}}^{\mathcal{T}} = \widehat{\mbox{\bf V}}\widehat{\mbox{\boldmath $\Gamma$}}\widehat{\mbox{\bf V}}' + \widehat{\mbox{\boldmath $\Sigma$}}_{E}^{\mathcal{T}}. $$ Similar to the proofs of fan2013large, we can show that the regular POET yields
where $\widetilde{\omega}_{T} = p^{\frac{5}{2}(1-a_{1})}\sqrt{\log p/T} + 1/p^{\frac{5}{2}a_{1}-\frac{3}{2}-c}$. We compare the convergence rates of the regular POET and proposed Double-POET. For example, under the relative Frobenius norm, when $q = 0$, $m_{p} = O(1)$, $a_{1} =1$, and $a_{2}=1$, we have
The number of assets within a group, $p_j \asymp p^{c}$, for some constant $c>0$ plays a crucial role in a convergence rate. Theorem (ref) shows that $\|\widehat{\mbox{\boldmath $\Sigma$}}^{\mathcal{D}} - \mbox{\boldmath $\Sigma$}\|_{\Sigma}$ can be convergent as long as $p=o(T^{2})$ and $\frac{1}{4}<c<\frac{3}{4}$. In contrast, the rate of the regular POET estimator, $\|\widehat{\mbox{\boldmath $\Sigma$}}^{\mathcal{T}}-\mbox{\boldmath $\Sigma$}\|_{\Sigma}$, does not converge if $c>\frac{1}{2}$ or $p^{\alpha}>T$ with $\alpha = \min\{\frac{1}{2}, 2c\}$. In other words, the regular POET requires the number of assets in each country to be small enough to avoid the curse of dimensionality. However, the number of assets in each country is larger than the number of countries; thus, it is more reasonable to assume $c>\frac{1}{2}$. Therefore, under the global and national factor models, the regular POET does not provide a consistent estimator in terms of the relative Frobenius norm. Furthermore, when the global factors are weak (i.e., $a_{1}<1$) and the local factors are strong (i.e., $a_{2}=1$), Double-POET is equal to or better than POET in terms of the convergence rate under the relative Frobenius norm.
When the signal of local factors is strong, we often consider the local factors as the global weak factors. Then, we can apply the regular POET method with the global and local factors. Theoretically, to identify the latent factors, we additionally need to impose an orthogonality condition between the global and local factor loadings, $\mbox{\bf B}$ and $\mbox{\boldmath $\Lambda$}$. In this section, under this orthogonality condition, we investigate asymptotic behaviors of the Double-POET procedure and compare POET with Double-POET.
Under the orthogonality condition, we first obtain the following modified theoretical results for Double-POET.
We now consider the POET estimator that regards both the global and local factors as the global factors under the orthogonality condition. We call this the POET2 estimator. The estimator is then defined as follows: $$ \widehat{\mbox{\boldmath $\Sigma$}}_{2}^{\mathcal{T}} = \widehat{\mbox{\bf V}}_{2}\widehat{\mbox{\boldmath $\Gamma$}}_{2}\widehat{\mbox{\bf V}}_{2}' + \widehat{\mbox{\boldmath $\Sigma$}}_{u,2}^{\mathcal{T}}, $$ where $\widehat{\mbox{\boldmath $\Gamma$}}_2 = \mathrm{diag}(\widehat{\delta}_{1}, \dots, \widehat{\delta}_{k+r})$ and $\widehat{\mbox{\bf V}}_{2}=(\widehat{v}_{1},\dots, \widehat{v}_{k+r})$. Then, when $a_{1}= a_2 =1$, the POET2 estimator yields
where $\widetilde{\omega}_{T,2} = p^{7(1-c)}\sqrt{\log p/T} + m_{p}/p^{7c-6}$. We note that the additional term $m_{p}/p^{7c-6}$ in $\widetilde{\omega}_{T,2}$ can be negligible only when $c > \frac{6}{7}$ as $p$ increases. This is because the POET2 estimator includes noises on the off-diagonal blocks of the local factor component. Therefore, it requires the number of groups to be sufficiently small. Otherwise, the POET2 estimator does not perform well. That is, since the local factors are considered as the weak factors of the global factor, the local factor should have enough signals.
We compare the convergence rates of the POET2 and Double-POET estimators under the orthogonality condition. When $q = 0$, $m_{p} = O(1)$, $a_{1} =1$, and $a_{2}=1$, we have
When $c<1$, the convergece rate of Double-POET is faster than that of POET2. Furthermore, $\|\widehat{\mbox{\boldmath $\Sigma$}}^{\mathcal{D}} - \mbox{\boldmath $\Sigma$}\|_{\Sigma}$ can be consistent as long as $p=o(T^{2})$ and $c>\frac{1}{4}$, while $\|\widehat{\mbox{\boldmath $\Sigma$}}_{2}^{\mathcal{T}}-\mbox{\boldmath $\Sigma$}\|_{\Sigma}$ can be consistent when $p^{15(1-c)+\frac{1}{2}}=o(T)$ and $c > \frac{9}{10}$. This implies that POET2 performs well only for the multi-level factor model with a small number of groups (i.e., weak factors with enough signals). We also note that when the local factor has the same signal as the global factor ($c=1$), the POET2 and Double-POET procedures have the same convergence rate.
In this section, we demonstrate the blessing of dimensionality for country-wise covariance matrix estimation by using the proposed Double-POET procedure. For the country $j$, we have
where $B^{j} = (b_{1}^{j},\dots, b_{p_j}^{j})'$, $\Lambda^{j} = (\lambda_{1}^{j},\dots, \lambda_{p_j}^{j})'$, and $u_{t}^{j} = (u_{1t}, \dots, u_{p_{j}t})'$. Then, the covariance matrix of the $j$th country can be written as
Then, the sparsity level of $\mbox{\boldmath $\Sigma$}_{u}^{j}$ is $$ m_{p_j} := \max_{i\leq p_j}\sum_{j\leq p_j}|\sigma_{u,ij}|^{q}, \text{ for some } q \in [0,1). $$ To estimate the $j$th country's covariance matrix $\mbox{\boldmath $\Sigma$}^{j}$, assuming that the common factors include both global and national factors with the orthogonality condition between their factor loadings, the regular POET method can be applied with $k+r_j$ principal components using the $j$th country's stock market data (i.e., $T \times p_j$ observation matrix), denoted by $\widehat{\mbox{\boldmath $\Sigma$}}^{j,\mathcal{T}}$. As long as $p_j \rightarrow \infty$, the asymptotic results of fan2013large can be directly applied by replacing $p$ by $p_j \asymp p^{c}$ for $c \in (0,1]$. In contrast, we suggest the Double-POET method by extracting the $j$th diagonal block of $\widehat{\mbox{\boldmath $\Sigma$}}^{\mathcal{D}}$, denoted by $(\widehat{\mbox{\boldmath $\Sigma$}}^{\mathcal{D}})^{j}$. Then, when estimating the global factor component, Double-POET uses more data, which might result in a more accurate global factor estimator. Specifically, we have the following convergence rates of Double-POET estimator for the local covariance matrix.
To implement Double-POET, we need to determine the number of factors. In this section, we describe a data-driven method for determining the number of global and national factors.
We suggest a modified version of the eigenvalue ratio method proposed by ahn2013eigenvalue as follows. We first consider a model selection problem between the multi-level factor model and the single-level factor model, and then propose estimators for the number of factors. Let $\widehat{\delta}_{m}$ be the $m$th largest eigenvalue of the sample covariance matrix and $\mathrm{ER}(m) = \widehat{\delta}_{m}/\widehat{\delta}_{m+1}$ be the $m$th eigenvalue ratio. Under the multi-level factor model (i.e., $k>0$ and $r>0$), the first $k$ eigenvalues of the sample covariance matrix are asymptotically determined by the eigenvalues of $\mbox{\bf B}\mathrm{cov}(G_{t})\mbox{\bf B}'$, the next $r = \sum_{j=1}^{J}r_{j}$ eigenvalues by the eigenvalues of $\mbox{\boldmath $\Lambda$}\mathrm{cov}(F_t)\mbox{\boldmath $\Lambda$}'$, and the other eigenvalues by those of the idiosyncratic covariance matrix. Accordingly, when $a_{1}=1$ and $a_{2}=1$, $\mathrm{ER}(m) = O_{P}(1)$ for $m\neq k$ and $m\neq k+r$, $\mathrm{ER}(k) = O_{P}(p^{1-c})$, and $\mathrm{ER}(k+r) = O_{P}(p^{c})$. This implies that there are two diverging eigenvalue ratios when $0<c<1$, while all other ratios of two adjacent eigenvalues are asymptotically bounded. In contrast, under the single-level factor model, there exists only one diverging eigenvalue ratio. We define $$ \hat{k}_{1} = \max_{1\leq m \leq k_{\max}}\mathrm{ER}(m) \;\;\; \text{ and } \;\;\; \hat{k}_{2} = \max_{1\leq m \neq \hat{k}_{1} \leq k_{\max}}\mathrm{ER}(m), $$ for a prespecified $k_{\max} < \min(p,T)$. Let $\varphi_{p}$ be the tuning parameter, which grows slowly. In practice, we set $\varphi_{p} = d\log p$ for a positive constant $d$. We then select the single-level factor model when $\mathrm{ER}(\hat{k}_{1}) > \varphi_{p}$ and $\mathrm{ER}(\hat{k}_{2}) \leq \varphi_{p}$, and one can apply the regular POET with $\hat{k}_{1}$ factors. When $\mathrm{ER}(\hat{k}_{2}) > \varphi_{p}$, we select the multi-level factor model and estimate the number of global factors $k$ by $\hat{k} = \min\{\hat{k}_{1}, \hat{k}_{2}\}$. Under the following technical conditions, we can show the consistency of $\hat{k}$.
Then, we have the following theorem.
Given the consistently estimated number of global factors, we can apply the existing methods to consistently estimate the number of local factors $r_j$ alessi2010improved, bai2002determining, choi2018multilevel, giglio2021asset, onatski2010determining, trapani2018randomized. Throughout the paper, we use the eigenvalue ratio method of ahn2013eigenvalue as follows: for each group $j$, let $\widehat{\kappa}_{r}^{j}$ be the $r$th largest eigenvalues of the $p_{j} \times p_{j}$ matrix $\widehat{\mbox{\boldmath $\Sigma$}}_{E}^{j}$ in the Section (ref). Then, $r_j$ can be estimated by $$ \hat{r}_{j} = \max_{1 \leq r \leq r_{j,\max}}\frac{\widehat{\kappa}_{r}^{j}}{\widehat{\kappa}_{r+1}^{j}}, $$ for a predetermined $r_{j,\max}$.
In this paper, we assume that local factors are governed by the national regional risk factors, which provides the membership of the local factors. However, in practice, the membership of the local factors is unknown. In this section, we discuss how to use the proposed Double-POET procedure for the unknown local factor membership.
The Double-POET procedure is working as long as the membership for the local factor is known. Thus, to harness the Double-POET, we need to classify assets and find the latent local structure. As discussed in the previous sections, we can estimate the global factor even if the local factor structure is unknown. Thus, we can consistently estimate $\mbox{\boldmath $\Sigma$}_{E}$ defined in (ref), which is the remaining covariance matrix after subtracting the global factor part. Under some local factor structure, $\mbox{\boldmath $\Sigma$}_{E}$ is a form of a block diagonal matrix and its block membership represents the local factor membership. Thus, we can detect the group membership, based on $\mbox{\boldmath $\Sigma$}_{E}$. Specifically, to adjust the scale problem of the convariance matrix, we calculate the correlation matrix of $\mbox{\boldmath $\Sigma$}_{E}$. Since the sign of the correlations does not contain the membership information, we use the absolute values of the correlations. We denote this matrix by $\mathbf{L}$. We consider $\mathbf{L}$ as the adjacency matrix of the local factor network. Based on the adjacency matrix, many models and methodologies have been developed to detect and identify group memberships. Examples include RatioCut hagen1992new, Ncut shi2000normalized, spectral clustering method lei2013consistency, mcsherry2001spectral, rohe2011spectral, regularized spectral clustering amini2013pseudo, semi-definite programming cai2015robust, hajek2015achieving, Newman-Girvan Modularity girvan2002community, and maximum likelihood estimation amini2013pseudo, bickel2009nonparametric. To detect the membership matrix, we employ the regularized spectral clustering (RSC) amini2013pseudo.
The regularized spectral clustering (RSC) is based on the regularized row and column normalized adjacency matrix (or the regularized graph Laplacian) chaudhuri2012spectral, qin2013regularized,
where the degree matrix is denoted by $\mbox{\bf D}= {\rm diag}( \hat{d}_1,\ldots, \hat{d}_p)$ with $\hat{d}_i = \sum_{j=1}^p L_{ij}$, $\mbox{\bf D} _a= \mbox{\bf D} +a \mbox{\bf I}$, $\mbox{\bf I}$ denotes the identity matrix, and $a\geq 0$ is the regularization parameter. For the numerical study, we use the average node degree as the regularization parameter $a$. Then, calculate the eigenvector matrix corresponding to the $K$ largest eigenvalue of $\tilde{\mbox{\bf L}}_{deg, a}$. Using the eigenvector matrix, we apply the k-means clustering procedure and identify the local factor groups. Unfortunately, we cannot observe $\mbox{\boldmath $\Sigma$}_{E}$, and so we estimate it using the POET procedure. Then, using the plug-in method, we estimate the regularized row and column normalized adjacency matrix. Under some regularity condition, we can show that the ratio of the misclassification goes to zero joseph2016impact, qin2013regularized. With this estimated membership, we can apply the proposed Double-POET.
In this section, we conducted simulations to examine the finite sample performance of the proposed Double-POET method. We considered the following multi-level factor model: $$ y_{it} = \sum_{l=1}^{k}b_{il}G_{tl} + \sum_{s=1}^{r_{j}}\lambda_{is}^{j}f_{ts}^{j} + u_{it}, $$ where the global factors and national factors, $G_{tl}$ and $f_{ts}^{j}$, respectively, were all drawn from $\mathcal{N}(0,1)$. The global factor loadings $\{b_{i1}, \dots, b_{ik}\}_{i\leq p}$ were drawn from $\mathcal{N}(\mu_{B}, I_{k})$, where each element of $\mu_{B}$ is i.i.d. Uniform$(-0.5,0.5)$; for each $j$, the local factor loadings $\{\lambda_{i1}^{j}, \dots, \lambda_{ir_{j}}^{j}\}_{i\leq p_j}$ were drawn from $\mathcal{N}(\mu_{\lambda^{j}}, I_{r_{j}})$, where each element of $\mu_{\lambda^j}$ is i.i.d. Uniform$(-0.3,0.3)$.
We generated the idiosyncratic errors as follows. Let $D = \mathrm{diag}(d_{1}^{2},\dots,d_{p}^{2})$, where each $\{d_{i}\}$ was generated independently from Gamma $(\alpha, \beta)$ with $\alpha = \beta = 100$. We set $s = (s_{1},\dots, s_{p})'$ to be a sparse vector, where each $s_{i}$ was drawn from $\mathcal{N}(0,1)$ with probability $\frac{m}{\sqrt{p}\log{p}}$, and $s_{i} = 0$ otherwise. Then, we set a sparse error covariance matrix as $\Sigma_u = D + ss' - \mathrm{diag}\{s_{1}^{2},\dots,s_{p}^{2}\}$. In the simulation, we generated $\Sigma_u$ until it is positive definite. Note that varying $m>0$ enables us to control the sparsity level, and we chose $m = 0.3$. Finally, we generated $\{u_{t}\}_{t\leq T}$ from i.i.d. $\mathcal{N}(0, \Sigma_{u})$.
In this simulation study, we chose the number of periods $T = 300$ and the numbers of factors as $k=3$ and $r_{j}=2$ for each $j$. Then, we considered two cases: (i) increasing $p$ from 60 to 600 in increments of 30 with a fixed $J=10$ (i.e., each $p_j = p/10$), and (ii) increasing $J$ from 2 to 20 with a fixed $p_j=30$. Each simulation is replicated 200 times.
For comparison, we employed Double-POET, POET, and the sample covariance matrix (SamCov) to estimate the true covariance matrix of $y$, $\mbox{\boldmath $\Sigma$}$. The estimation errors are measured in the following norms: $\|\widehat{\mbox{\boldmath $\Sigma$}} -\mbox{\boldmath $\Sigma$}\|_{\Sigma}$, $\|\widehat{\mbox{\boldmath $\Sigma$}} -\mbox{\boldmath $\Sigma$}\|_{\max}$, and $\|(\widehat{\mbox{\boldmath $\Sigma$}})^{-1} - \mbox{\boldmath $\Sigma$}^{-1}\|$, where $\widehat{\mbox{\boldmath $\Sigma$}}$ is one of the covariance matrix estimators. For the POET estimation, we estimated the covariance matrix with two different numbers of factors: (i) POET uses the $k$ number of factors, and (ii) POET2 uses $k + r$ factors, where $r = J\times r_{j}$. For Double-POET, we considered two different cases: (i) D-POET with the known group membership, and (ii) D-POET(RSC) suggested in Section (ref) with the unknown group membership. The proposed numbers of global and local factors estimation method in Section (ref) is applied with $k_{\max}=10+r$ and $r_{j,\max}=10$ for each estimation. In addition, we employed the soft thresholding scheme for both POET and Double-POET.
Figures (ref) and (ref) plot the averages of $\|\widehat{\mbox{\boldmath $\Sigma$}} -\mbox{\boldmath $\Sigma$}\|_{\Sigma}$, $\|\widehat{\mbox{\boldmath $\Sigma$}} -\mbox{\boldmath $\Sigma$}\|_{\max}$ and $\|(\widehat{\mbox{\boldmath $\Sigma$}})^{-1} - \mbox{\boldmath $\Sigma$}^{-1}\|$ from different methods against $p$ and $J$, respectively. From Figure (ref), we find that D-POET performs the best. When comparing the POET-type procedures, POET2 performs better than POET. This may be because POET ignores important local factors. D-POET(RSC) with the unknown group membership performs better than POET2 under different norms. This confirms that the RSC method in Section (ref) can detect the membership well. Under the max norm, all estimators except POET perform roughly the same. This is because the thresholding or imposing the local factor structure affects mainly the elements of the covariance matrix that are nearly zero, and the elementwise norm depicts the magnitude of the largest elementwise absolute error. Figure (ref) shows the similar results except when $J$ is extremely small. This may be because for the small $J$ ($J=2$), we have weak global factors rahter than local factors. We note that the estimation errors of POET2 increase as $J$ grows because the estimator includes more noises on the off-diagonal blocks of the local factor component. The above results support the theoretical findings presented in Section (ref).
We further explored the performance of Double-POET when the group membership is misclassified. In Figure (ref), we report the average errors under different norms against the misclassified rate of the group membership for Double-POET(Mix) with $J = 20$ and $p_j= 20$. Figure (ref) implies that Double-POET method performs better than POET2 unless the misclassification rate is greater than 3%.
We now demonstrate the blessing of dimensionality using Double-POET when estimating the local covariance matrix $\mbox{\boldmath $\Sigma$}^{j}$ and its inverse. We generated the data as above. The SamCov and POET estimators are obtained using the group sample (i.e., $T \times p_j$ observation matrix). For the POET method, we used the number of factors as $k + r_j = 5$ in each estimation. The Double-POET and POET2 estimators for $\mbox{\boldmath $\Sigma$}^{j}$ are the $j$th diagonal block of $\widehat{\mbox{\boldmath $\Sigma$}}^{\mathcal{D}}$ and $\widehat{\mbox{\boldmath $\Sigma$}}_{2}^{\mathcal{T}}$, respectively. We calculated $\|\widehat{\mbox{\boldmath $\Sigma$}}^{j} - \mbox{\boldmath $\Sigma$}^{j}\|_{ \Sigma^{j}}$ and $\|(\widehat{\mbox{\boldmath $\Sigma$}}^{j})^{-1} - (\mbox{\boldmath $\Sigma$}^{j})^{-1}\|$ for $j=1$. Note that we do not present the results of max norm, $\|\widehat{\mbox{\boldmath $\Sigma$}}^{j} - \mbox{\boldmath $\Sigma$}^{j}\|_{ \max}$, as all estimators perform very similar to that shown in Figure (ref).
Figure (ref) depicts the averages of $\|\widehat{\mbox{\boldmath $\Sigma$}}^{j} - \mbox{\boldmath $\Sigma$}^{j}\|_{ \Sigma^{j}}$ and $\|(\widehat{\mbox{\boldmath $\Sigma$}}^{j})^{-1} - (\mbox{\boldmath $\Sigma$}^{j})^{-1}\|$ for Double-POET, POET, and the sample covariance matrix against $p_j$ with a fixed $J=10$, while Figure (ref) plots their average errors against $J$ with fixed $p_j=30$. Figures (ref) and (ref) show that Double-POET has smaller estimation errors than other methods under different norms. In Figure (ref), Double-POET significantly outperforms POET when $p_j$ is small, while the estimation error gap between Double-POET and POET decreases as $p_j$ grows. The estimation error of POET2 under the relative Frobenius norm tends to increase as $p_j$ grows. This is because, even though the global factor component can be estimated more accurately, the estimator includes redundant information on the local factor part. In contrast, Figure (ref) shows that, except when $J$ is very small, the Double-POET method constantly dominates the other methods. This might be because by using other local groups' information, Double-POET can estimate global factors better than POET, especially when $p_j$ is small. That is, the proposed Double-POET enjoys the blessing of dimensionality, which corresponds to the theoretical analysis discussed in Section (ref).
We also examined the performance of the modified eigenvalue ratio (MER) estimator introduced in Section (ref) for the number of global factors. For comparison, we employed alternative estimators: the BIC$_3$ estimator of bai2002determining, the ED estimator of onatski2010determining, the ER estimator of ahn2013eigenvalue, and the estimator of alessi2010improved (ABC). We used the same data generating process as before but considered different number of global factors, $k \in \{0, 3, 6\}$, and sample sizes, $p \in \{100, 200, 300\}$ and $T \in \{150, 300\}$. We set $k_{\max} = 10 + r$ and $\varphi_{p} = 0.3\log p$ for MER estimator, and $k_{\max} = 10$ for the other estimators. We calculated the average of $\hat{k}$ and the percentage correctly determining $k$ over 500 replications.
The results are reported in Table (ref). We find that the proposed MER method performs the best across the different number of global factors. In particular, when $k=0$, MER and ED tend to estimate correctly, while other estimators tend to overestimate $k$. We note that ER does not include the case of $k=0$. When $k>0$, all estimators except BIC$_3$ perform well, but MER and ER slightly outperform ED and ABC. MER often underestimates $k$ if both $p$ and $T$ are small, but it selects $k$ correctly as $p$ grows. The results demonstrate that the proposed MER method is appropriate for the model selection problem for both the multi-level factor model and the single-level factor model.
In this section, we applied the proposed Double-POET method to a minimum variance portfolio allocation study using global stock data. We collected daily transaction prices of international stock markets over 20 countries by the total market capitalization from January 2, 2016, to December 31, 2021. We selected top 100 firms at most for each country based on the market cap and used weekly log-returns to mitigate the effect of different trading hours. We excluded stocks with missing returns and no variation in this period. This operation leads to total $1892$ stocks for this period. The distribution of our sample is presented in Table (ref).
We calculated the Double-POET, POET, and SamCov estimators for each month. In both Double-POET and POET procedures, we estimated the idiosyncratic volatility matrix based on 11 Global Industrial Classification Standard (GICS) sectors ait2017using, fan2016incorporating. For example, the idiosyncratic components for the different sectors were set to zero, and we maintained these for the same sector. This location-based thresholding preserves positive definiteness and corresponds to the hard-thresholding scheme with sector information. We varied the number of (global) factors $k$ from 1 to 5 for both Double-POET and POET. To determine the number of local factors for Double-POET, we used the eigenvalue ratio method suggested by ahn2013eigenvalue with $r_{j, \max} = 10$. We also considered POET2 using $k + \sum_{j=1}^{20}\hat{r}_j$ number of factors from the best performing Double-POET estimator.
To analyze the out-of-sample portfolio allocation performance, we considered the following constrained minimum variance portfolio allocation problem fan2012vast, jagannathan2003risk:
where $\mathbf{1} = (1,\dots,1)^{\top} \in \mathbb{R}^{p}$, the gross exposure constraint $c$ was varied from 1 to 4, and $\widehat{\mbox{\boldmath $\Sigma$}}$ is one of the volatility matrix estimators from Double-POET, POET, and SamCov. We constructed the portfolio at the beginning of each month, based on the stock weights calculated using the data from the past 24 months ($T=104$). We then held the portfolio for one month and calculated the square root of the realized volatility using the weekly portfolio log-returns. Their average was used for the out-of-sample risk. We considered five out-of-sample periods: 2018, 2019, 2020, 2021 and the whole period (2018-2021).
Figure (ref) depicts the out-of-sample risks of the portfolios constructed by SamCov, POET, POET2, Double-POET, and Double-POET(RSC) with $k=5$ against the exposure constraint. From Figure (ref), we find that Double-POET outperforms POET for the same number of factors $k$, while $k=3$ and $k=5$ yields the best performances for Double-POET and POET, respectively. Double-POET($k=3$) reduces the minimum risks by 4.3%--6.7% compared to POET($k=5$). We confirmed that for the purpose of portfolio allocation based on the global stock market, the Double-POET based on the global and national factor model outperforms POET based on the single-level factor model. Thus, we can conjecture that incorporating the latent global and national factor models helps account for the global market dynamics. We note that Double-POET(RSC) does not perform well due to misclassified group membership. One possible explanation is that there might be other type of local factors in addition to the nation-specific factors. For example, if the industrial risk and the national risk are nested on the idiosyncratic part after removing the global factors, the suggested RSC method cannot properly detect only the national group membership. This is an interesting research topic to control the unknown nested local groups, so we leave it for a future study.
We also conducted the country-wise volatility matrix estimation based on SamCov, POET, POET2, and Double-POET and applied them to the same portfolio allocation problem as before. Specifically, we used the sample of each country for POET and SamCov procedures. The Double-POET and POET2 estimators for each local covariance matrix are obtained by extracting each diagonal block of the large Double-POET and POET2 estimators, respectively. Figure (ref) presents the portfolio behavior of the top 100 stocks in the US. Double-POET shows stable results and reduces the minimum risks by 1.8%--4.0% compared to POET except the period 2021. We note that a powerful bull market lasted in 2021, and the risks could be sufficiently explained by only the first principal component (i.e., the market factor), so that, POET($k = 1$) performs the best in this period. Nevertheless, the overall results indicate that the proposed Double-POET method enjoys the blessing of dimensionality.
Figure (ref) shows the results of other 19 countries for the whole period. Except for five countries (GB, IN, IT, TH, and ZA), Double-POET outperforms POET. We note that the SamCov estimator sometimes outperforms others for a few countries. Overall, these results indicate that Double-POET can accurately estimate the global factors by harnessing other countries' observations.
This paper proposes a novel large volatility matrix inference procedure based on the latent global and national factor models. We show the asymptotic behaviors of the proposed Double-POET method and discuss its blessing of dimensionality and efficiency of estimating a large volatility matrix compared to the regular POET procedure. To determine the number of global factors, we extend the eigenvalue ratio procedure ahn2013eigenvalue. In addition, when the membership of the local factors is unknown, we suggest the regularized spectral clustering method to find the latent local structure.
In the empirical study, in terms of portfolio allocation, the proposed estimator shows the best performance. It confirms the presence of the national factor structure in global financial markets, which provides the theoretical basis for employing the Double-POET method. In addition, for the country-wise covariance matrix estimation, the Double-POET procedure yields the blessing of dimensionality by accurately estimating the latent global factors using the information outside the local group.
In this paper, we focus on the national risk factor as only the local factor. However, in practice, there could be other types of risk factors that are nested in the local level. Thus, it is interesting and important to develop a large volatility matrix estimation procedure based on unknown-membership local factor models with the nested local-level group factors. We leave this for a future study.
The authors thank the Editor Professor Torben Andersen, the Associate Editor, and two referees for their careful reading of this paper and valuable comments. The research of Donggyu Kim was supported in part by the National Research Foundation of Korea (NRF) (2021R1C1C1003216).