Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
71,324 characters · 16 sections · 51 citation commands
Next Generation Models for Portfolio Risk Management: An Approach Using Financial Big Data
The widespread failure of portfolio risk management models observed in the 2008 financial crisis raises the question of how risk managers can analyze the risks of investment portfolios while taking into account systematic risks caused by a systemic crisis. Extant literature has developed a range of risk measurement methods, for example, Value at Risk (VaR), expected shortfall (ES), entropic Value at Risk (EVaR), and superhedging price acerbi2002expected,ahmadi2012entropic,bensaid1992derivative,jorion2000value. VaR, in particular, has been widely used to estimate the level of risk by measuring the amount of loss for investments during a given period under normal market conditions jorion2000value. Especially, the literature documents the development of risk models with VaR to accommodate non-normal market conditions, for example, asymmetric, heavy-tailed stock returns, and volatility clustering (see, e.g., huisman1998var,natarajan2008incorporating,wu2007value,zoia2018value for non-normal stock returns; billio2000value,broda2009chicago,engle2004caviar,fantazzini2008dynamic,mcneil2000estimation for volatility clustering).
However, the 2008 financial crisis demonstrates that the development of risk models with dynamics for non-normal conditions does not guarantee financial firms to be successful in managing extreme systematic risks. One of the supporting arguments can be that risk managers still consider only stock data in their target portfolio fan2003semiparametric,hendricks1996evaluation,patton2019dynamic. They tend to overlook the risks of the assets that are not included in the portfolio; hence, the portfolio size for risk analysis can be relatively small. This so-called small portfolio risk analysis may not sufficiently capture portfolio risk dynamics, particularly systematic risks inherent in the market.
In this regard, this study highlights a clear message that the use of financial big data can help risk managers take into consideration latent risks that cannot be captured in the portfolios of interest. The contributions of this paper are threefold. First, we propose a dynamic risk model to resolve the above-mentioned problem by capturing systematic risk dynamics. At the core of the proposed model is the use of financial big data to analyze portfolio risks. The use of financial big data can help risk managers take into account a universe of assets larger than the size of their portfolios. This usefulness can lead to a higher accuracy of risk prediction, particularly, when it comes to capturing the systematic risk inherent in the markets. Consequentially, financial regulators can benefit from a high accuracy of risk prediction for market players to maintain the stability of financial markets and thus reduce the probability of a systemic risk occurrence.
One may be concerned about the fact that the use of financial big data can lead multivariate VaR methods to be computationally intensive in parameter estimation, which is called the curse of dimensionality.\footnote{For instance, when the BEKK model engle1995multivariate is conducted on $p$-dimensional log return data, the number of parameters increases with the $p^2$ order. As the dimension $p$ increases, the parameter estimation becomes computationally demanding due to the exploding of computation time, which is the NP hard problem. Even if it is possible to estimate the parameters, they are not consistent estimators bickel2008covariance,engle2017large.} Our second contribution is to address this point in that the proposed model overcomes the curse of dimensionality with a two-step estimation. The first step is to use the principal orthogonal complement thresholding (POET) procedure fan2013large to estimate the latent factor and idiosyncratic component of the approximate latent factor model ait2017using, fan2018robust, fan2019robust, kim2018large, li2018embracing. The second step is a maximum quasi-likelihood estimation for the Generalized AutoRegressive Conditional Heteroskedasticity (GARCH) parameter with the latent factor estimator from the first step. The estimated GARCH parameters are used to predict one-step ahead portfolio volatility, with which the one-step ahead VaR is predicted under the parametric and non-parametric $\sigma$-based approaches brooks2003volatility,giot2004modelling,hull1998incorporating.
Lastly, our contribution is the incorporation of financial big data featuring the blessing of dimensionality donoho2000high, fan2019structured,fan2013large, li2018embracing by discussing the asymptotic properties of $\sigma$-based approaches. Upon the risk measurement procedure with integrated econometric approaches, our study is the first to suggest how a big data approach can help financial risk management. The proposed model can be particularly helpful for financial firms and regulators, who may be concerned about the stability of financial markets and question how big data can be optimally used for an improved prediction of risks.
The rest of the paper is organized as follows. In Section (ref), we introduce the concept of VaR and develop a dynamic risk model. In Section 3, we propose a maximum quasi-likelihood estimation method and study its asymptotic behavior. Section (ref) shows how the large volatility matrix can be predicted, how the proposed method can be applied to measure VaR, and how the curse of dimensionality can be solved. In Section (ref), we implement the Monte Carlo simulation to check the finite sample performance of the proposed methods and apply the proposed VaR measure to the empirical data. We conclude in Section (ref).\footnote{All technical proofs are provided in the online supplement (see Appendix A).}
Consider a market consisting of $p$ stocks. Let $\bfm y_t$ denote the $\mathbb{R}^p$-valued random vector of the assets' log returns at time $t$. Then, the portfolio is characterized by a vector of asset weights $\bfm w\in \mathbb{R}^p$, for example, $r_t = \bfm w ^{\top} \bfm y_t$. We assume that the log returns admit the factor model ait2017using,bai2003inferential,carhart1997persistence,fan2013large,fan2019robust,fama1992cross,fama1993common,fama2015five as follows:
where $\ensuremath{\boldsymbol{\mu}}_t$ is a mean vector, $\bfm f_t$ is an $r$-dimensional factor, $\bfm V$ is a $p\times r$ factor loading matrix, and $\bfm u_t$ is an idiosyncratic component. We assume that the factor $\bfm f_t$ and idiosyncratic component $\bfm u_t$ are independent, and without loss of generality, we assume that the mean of $\bfm f_t$ and $\bfm u_t$ are zero. Usually, the number of market factors is much smaller than the number of assets, so we assume that $r$ is finite. Moreover, in this paper, we consider the latent factor model ait2017using,fan2013large,fan2019robust,li2018embracing, so $\bfm f_t$ and $\bfm u_t$ are not observable. In this section, under the latent factor model, we propose a VaR estimation procedure.
One of the most popular risk measures is VaR, which estimates how much a portfolio of assets might lose under normal market conditions. Specifically, VaR at level $\alpha$ is defined as the $\alpha$-percentile of a portfolio log return distribution as follows:
where $\mathbb{P}_{t-1}$ denotes the conditional distribution of the portfolio log returns $r_t$ given a filtration representing information to time $t-1$. When the portfolio log return distribution follows a location-scale family, such as normal distribution and t-distribution, we can obtain VaR from the portfolio conditional mean $\mu_t$ and conditional volatility $\sigma_t$ given a filtration information to time $t-1$ as follows:
where $c_{\alpha}$ is an $\alpha$-quantile value of the distributions of the standardized $r_t$. This method is called a $\sigma$-based approach jorion2000value. The performance of the $\sigma$-based approach depends on three components: the conditional mean $\mu_t$, conditional volatility $\sigma_t$, and quantile $c_{\alpha}$. Quantile $c_{\alpha}$ can be determined using the parametric and non-parametric methods. The parametric method involves the assumption of the parametric distribution for the standardized log return $(r_t - \mu_t) / \sigma_t$, such as normal distribution or student-t distribution, and $c_{\alpha}$ is determined based on the $\alpha$-quantile of each distribution. The non-parametric method uses a sample quantile of the historical standardized portfolio log returns $(r_t - \mu_t)/\sigma_t$. Detailed implementations are described in Section (ref). However, several empirical studies indicate that the portfolio conditional mean $\mu_t$ does not have strong time series patterns cajueiro2004hurst,fama1998market,malkiel2003efficient,narayan2004south, which makes it difficult to predict the mean. In this paper, we simply assume that the mean process is constant over time in order to reduce the model misspecification error. Meanwhile, in the stock market, we observe volatility time series structures, such as volatility clustering mandelbrot1963variation. From this point of view, it is crucial to determine the dynamic structure of the conditional volatility $\sigma_t$. Thus, in this paper, we focus on how the dynamic volatility structure can be incorporated in VaR. We note that the mean structure would have a significant effect on VaR; hence, a parametric structure of the mean process should be developed. We leave this for future study.
When analyzing portfolio volatility, we usually investigate only the stock log return data that belong to the portfolio. For example, let $\bfm y_{t}^s$ be the $s$-dimensional log return vector whose stocks belong to the portfolio. We also denote the conditional co-volatility matrix of $\bfm y_{t}^s$ by $\bfsym \Sigma_{t}^s$ and the $s$-dimensional portfolio weight vector by $\bfm w^s$. Then, the portfolio conditional volatility can be obtained from $\bfsym \Sigma_{t}^s$ as follows:
To account for the market dynamics of the portfolio return, researchers have introduced several dynamic models, such as the BEKK engle1995multivariate, CCC bollerslev1990modelling, and portfolio univariate GARCH bollerslev1986generalized. See also engle2002dynamic, kawakatsu2006matrix, paolella2021non, paolella2015comfort, and van2002go. The BEKK model with the conditional mean vector $\ensuremath{\boldsymbol{\mu}}_{t}^s$ has the following structure:
where $\bfm C$ is an $s\times s$ lower triangular matrix, $\(\bfsym \Sigma_{t}^s\)^{\frac{1}{2}}$ is the square root of the matrix $\bfsym \Sigma_{t}^s$, $\bfm A$ and $\bfm B$ are $s\times s$ matrices, and $F(0,1)$ is some arbitrary distribution with mean zero and variance one. As in the above BEKK model, the dynamic volatility models have the autoregressive form of the historical squared returns. Empirical studies shows that the volatility is heterogeneous, and these models can explain the market dynamics bollerslev1986generalized, bollerslev1990modelling, engle2002dynamic,engle1995multivariate, kawakatsu2006matrix,van2002go. Moreover, the dynamic volatility model helps to improve the measurement of VaR engle2004caviar,giot2004modelling,kuester2006value. That is, the performance of measuring VaR can be improved by identifying the market dynamics. In financial markets, we often observe that much of the market dynamic structure comes from the systematic risk, which is related to the entire stock market; hence, the systematic risk information is contained in the entire stock. From this point of view, even if we consider only the small number of assets in the portfolio, to account for the market dynamics, we need to investigate the whole stock market and extract the systematic risk information from the stock returns. One of the naive extensions is to use a multivariate volatility model, such as the BEKK and CCC models, for the whole stock market. However, when applying these volatility models to the incorporation of a large number of assets, we face the curse of dimensionality bickel2008covariance, pakel2020fitting,engle2017large. Specifically, let $p$ be the total number of assets, and then the number of parameters quadratically increases with $p$. Hence, with a large $p$, the estimated model is statistically inconsistent. Furthermore, in practice, the parameter optimization takes exponential computation time, which is known as the non-deterministic polynomial-time (NP) hard problem. Thus, applying the usual multivariate volatility model to a large number of assets is not feasible. In this paper, we discuss how the curse of dimensionality issue can be handled while accounting for the market dynamics.
In this section, we propose a dynamic factor model under the latent factor structure in (ref). The factor model implies that the risk of assets stems from factor and idiosyncratic components, which are called systematic and idiosyncratic risks, respectively. An idiosyncratic risk is related to the firm-specific risk; hence, it does not affect the entire market. Moreover, it can be mitigated by the portfolio diversification. By contrast, a systematic risk arises from common market factor, such as interest rate, inflation, oil price, etc. Because systematic risk affects the entire market, we cannot mitigate this risk. From this point of view, we impose a dynamic structure on the factor to account for the market risk. For the idiosyncratic risk, we employ some martingale assumption, for example, the idiosyncratic volatility is constant over time. Then, the conditional expected co-volatility matrix of $\bfm y_t$ given the current available information $\mathcal{F}_{t}$ is expressed as follows:
where $\bfsym \Sigma_{f,t+1}= \mathbb{E} \[ \bfm f_{t+1} \bfm f_{t+1} ^{\top} \middle | \mathcal{F}_{t} \] $, and $\bfsym \Sigma_u$ represents the idiosyncratic volatility matrix. Under the framework (ref), the market dynamics can be explained by the factor volatility dynamics.
The latent factor model has an identification problem in $(\bfm V, \bfm f_t)$. For example, the pair $(\bfm V, \bfm f_t)$ is not distinguishable from $(\bfm V \bfm H^\top, \bfm H \bfm f_t)$ for any orthonormal matrix $\bfm H$. To uniquely define the latent factor model, we often impose the following identifiability condition bai2012statistical,bai2013principal, fan2013large,kim2019factor,li2018embracing:
The identification condition (ref) implies that the scaled factor loading matrix $p^{-1/2}\bfm V$ and elements of $p \bfsym \Sigma_{f,t}$ are the eigenmatrix and eigenvalues of the conditional variance of the factor part. Thus, the market dynamics are explained by the dynamic structure of the eigenvalues, whereas its factor loading matrix is constant over time.
To account for the factor dynamics, we employ the famous GARCH model bauwens2006multivariate,engle1995multivariate,engle2002dynamic,engle2016dynamic,van2002go. The conditional expected volatility of the factor is modeled by the following GARCH structure:
where $\epsilon_{it}$'s are $\text{i.i.d.}$ random variables with mean zero and unit variance; $\bfm h_{t}(\ensuremath{\boldsymbol{\theta}})$ is an $r$-dimensional vector of the conditional expected factor volatility $\bfsym \Sigma_{f,t}$, that is, $\bfm h_{t}= \big(h_{1t} (\ensuremath{\boldsymbol{\theta}}), \ldots,\allowbreak h_{rt} (\ensuremath{\boldsymbol{\theta}})\big) = \operatorname{diag}(\bfsym \Sigma_{f,t})$, where $\operatorname{diag}(\bfm X)$ is a vector whose elements are diagonal entries of $\bfm X$, $\bfm f_t^2$ is an element-wise squared vector of the factor log return $\bfm f_t$, and $\ensuremath{\boldsymbol{\theta}} =( \ensuremath{\boldsymbol{\omega}}, \mathrm{vec}(\bfm A), \mathrm{vec}(\bfm B))$ is the GARCH parameters. The GARCH model in (ref) indicates that the conditional expected volatility of the factor is the autoregressive form of the historical squared factor log returns. Thus, we expect that this model can capture the stylized market features, such as the volatility clustering and heavy tail. The empirical study in Section (ref) also supports this.
The idiosyncratic component comes from the firm-specific risk, so their risk is not strongly connected. Empirical studies show that there are some local factors that affect few other idiosyncratic components ait2017using, boivin2006more. In light of these, we allow a weak relationship between idiosyncratic components so that the idiosyncratic volatility $\bfsym \Sigma_u =\( (\Sigma_{u,ij})_{i,j=1,\ldots, p} \)$ satisfies the following sparse condition:
where $ q \in [0, 1)$, and the sparsity measure $s_p$ diverges slowly with the dimension $p$, for example, $\log{p}$. We define $0^0 =0$. This sparsity condition is widely employed in the large convariance matrix inferences ait2017using,bickel2008covariance, cai2011adaptive, fan2013large, fan2019robust, kim2019factor. When the idiosyncratic volatility satisfies the exact sparsity, that is, $q=0$, the sparsity condition indicates that each asset has at most $s_p$ non-zero idiosyncratic correlations with other assets.
Under the volatility structure (ref) and (ref), the VaR of portfolios in (ref) can be calculated as follows:
To evaluate the above VaR value, we need to estimate the unobserved factor components and the idiosyncratic volatility. However, to incorporate the entire market asset information, we consider a large number of assets and consequently run into the curse of dimensionality problem. In the following section, we discuss how to overcome this issue and introduce an estimation procedure to incorporate the financial big data for risk analysis.
First, we define the notations. We denote $\|\cdot\|,\|\cdot\|_F$, and $\|\cdot\|_{\max}$ by the matrix spectral norm, Frobenius norm, and max norm, respectively. We use $O_p$ as a big-O in probability, $\lambda_k(\bfm A)$ as the $k^{th}$ largest eigenvalue of the square matrix $\bfm A$, and $\mathrm{vec}(\bfm A)$ as the vectorization of $\bfm A$. We also denote the true parameters by $\ensuremath{\boldsymbol{\theta}}_0 = (\ensuremath{\boldsymbol{\omega}}_0, \mathrm{vec} (\bfm A_0), \mathrm{vec}(\bfm B_0) )$.
To estimate the model parameter $\ensuremath{\boldsymbol{\theta}}_0 = (\ensuremath{\boldsymbol{\omega}}_0, \mathrm{vec} (\bfm A_0), \mathrm{vec}(\bfm B_0) )$, we first need to estimate the latent factor components $\bfm f_t$ and $\bfm V$. We recall that the conditional volatility matrix is
and the mean conditional volatility matrix is
where $\mkern 1.5mu\overline{\mkern-1.5mu\bfsym \Sigma\mkern-1.5mu}\mkern 1.5mu_{f} = T^{-1} \sum_{t=1}^T \bfsym \Sigma_{f,t}$. Under the identification and sparsity conditions (ref) and (ref), the mean conditional volatility matrix has the low-rank plus sparse structure that is widely used in the high-dimensional factor analysis ait2017using,fan2013large, fan2019robust,kim2019factor. With the low-rank plus sparse structure, fan2013large introduced the POET estimation procedure to estimate the latent factor volatility and idiosyncratic volatility matrices. To harness the POET procedure, we need a good proxy of the mean conditional volatility matrix. We use the sample covariance matrix $\widehat{\mkern 1.5mu\overline{\mkern-1.5mu\bfsym \Sigma\mkern-1.5mu}\mkern 1.5mu} = T^{-1} \sum_{t=1}^T (\bfm y_t - \mkern 1.5mu\overline{\mkern-1.5mu\bfm y\mkern-1.5mu}\mkern 1.5mu ) (\bfm y_t - \mkern 1.5mu\overline{\mkern-1.5mu\bfm y\mkern-1.5mu}\mkern 1.5mu) ^\top$ with sample mean $\mkern 1.5mu\overline{\mkern-1.5mu\bfm y\mkern-1.5mu}\mkern 1.5mu=T^{-1}\sum_{t=1}^T{\bfm y_t}$, and under some mild condition, the martingale convergence theorem implies that the sample covariance matrix converges to $\mkern 1.5mu\overline{\mkern-1.5mu\bfsym \Sigma\mkern-1.5mu}\mkern 1.5mu$. Thus, we apply the POET procedure with the sample covariance matrix to estimate the latent factor components. Specifically, the eigenvalue decomposition admits
where $\widehat{\lambda}_k$ and $\widehat{\bfm q}_k$ are the $k^{th}$ largest eigenvalues and eigenvectors of the sample covariance matrix $\widehat{\mkern 1.5mu\overline{\mkern-1.5mu\bfsym \Sigma\mkern-1.5mu}\mkern 1.5mu}$, respectively. Then, the factor loading matrix estimator $\widehat{\bfm V}$ is $\sqrt{p}( \widehat{\bfm q}_1, \ldots, \widehat{\bfm q}_r)$, and the mean factor volatility matrix estimator $\widehat{\mkern 1.5mu\overline{\mkern-1.5mu\bfsym \Sigma\mkern-1.5mu}\mkern 1.5mu}_{f}$ is $p^{-1}\mathrm{Diag} \big( (\widehat{\lambda}_1, \ldots, \widehat{\lambda}_r )\big)$, where $\mathrm{Diag}(\bfm x)$ is a diagonal matrix whose diagonal entries are $\bfm x$. Note that $\widehat{\bfm V}$ has a multiplier $\sqrt{p}$ so that $\widehat{\bfm V}$ satisfies the condition $\bfm V^\top\bfm V = p\bfm I_r$. Then, the latent factors can be obtained as follows:
Theorem (ref) in the online supplement shows that the convergence rate of $\widehat{\bfm f}_t^2$ is $1/\sqrt{T} + \sqrt{s_p/p}$. We note that when the factor is not observable and the number of assets is finite, the latent factors are impossible to estimate. By contrast, Theorem (ref) shows the blessing of financial big data in the latent factor estimation. Specifically, every asset contains the common market information. Thus, as the number of assets increases, the information for the latent factors also increases. Therefore, more stock samples exhibit a clearer latent signal, and we can estimate the latent factors consistently with the convergence rate $\sqrt{s_p/p}$.
In this section, we propose a model parameter estimation procedure. Specifically, we adopt the following quasi-maximum likelihood estimation (QMLE) method:
where $\ensuremath{\boldsymbol{\Theta}}$ is a compact parameter space. Then, the well-developed asymptotic theorems for the QMLE provide the consistency comte2003asymptotic. However, the factors are not observable; hence, to evaluate the quasi-likelihood function, we use the non-parametric factor estimator $\widehat{\bfm f_t^2}$. For example, the conditional co-volatility $\bfm h_t(\ensuremath{\boldsymbol{\theta}})$ is estimated by
and we use $\widehat{\bfm h}_1(\ensuremath{\boldsymbol{\theta}}) = (\bfm I_r - \bfm A - \bfm B)^{-1} \ensuremath{\boldsymbol{\omega}}$ as the initial value. Note that $\widehat{\bfm h}_1(\ensuremath{\boldsymbol{\theta}})$ is the unconditional volatility of the factor. Then, we calculate the quasi-likelihood function with the conditional co-volatility estimator $\widehat{\bfm h}_t(\ensuremath{\boldsymbol{\theta}})$ and obtain the maximum quasi-likelihood estimator $\widehat{\ensuremath{\boldsymbol{\theta}}}$ as follows:
Theorem (ref) in the online supplement shows the consistency of $\widehat{\ensuremath{\boldsymbol{\theta}}}$.
In the previous section, we find the blessing of dimensionality in the latent factor estimation. However, when estimating large volatility matrices, we still suffer from the curse of dimensionality. For example, sample volatility matrix estimators are inconsistent when both the number of assets and sample size go to infinity bickel2008covariance, cai2011adaptive. To overcome the curse of dimensionality, we impose the sparse structure (ref) on the idiosyncratic volatility matrix. As discussed in Section (ref), in the stock market, the co-movement of stocks can be explained by the common factor, and the remaining idiosyncratic co-volatilities are weakly correlated. Thus, the sparsity condition is realistic. To estimate the sparse idiosyncratic volatility matrix, we employ the POET procedure fan2013large as follows. First, we estimate the input idiosyncratic volatility matrix estimator by using the non-pervasive eigen-components as follows: $$ \widehat{\bfsym \Sigma}_u = \sum_{i=r+1}^p {\widehat{\lambda}_i \widehat{\bfm q}_i {\widehat{\bfm q}_i}^\top}, $$ where the eigenvalues $\widehat{\lambda}_i$'s and eigenvectors $\widehat{\bfm q}_i$'s are defined in (ref). Then, we apply the thresholding method to the input idiosyncratic volatility matrix estimator as follows:
where $\mathbbm{1}_{(\cdot)}$ is an indicator function, and $s_{ij}(\cdot)$ is a shrinkage function satisfying $|s_{ij}(x) - x| \le \tau_{T} \sqrt{\widehat{\Sigma}_{u,ii} \widehat{\Sigma}_{u,jj}}$. The examples of the shrink function $s_{ij}(x)$ are the soft thresholding function $s_{ij} (x) = x - sign(x)\tau_T \sqrt{ \widehat{\Sigma}_{u,ii} \widehat{\Sigma}_{u,jj}}$ and the hard thresholding function $s_{ij} (x) = x$. The thresholding level $\tau_T$ will be given in Theorem (ref). The working principle of the thresholding method is that the co-volatility is zero if the estimated correlation is weak. This makes the estimated idiosyncratic volatility matrix estimator sparse, so the estimated idiosyncratic volatility satisfies the sparse condition.
With the estimated factor volatility $\widehat{\bfsym \Sigma}_{f,t+1}$ and idiosyncratic volatility matrix ${\cal T}(\widehat{\bfsym \Sigma}_u)$, we estimate the conditional volatility matrix as follows:
We call the conditional volatility matrix estimator $ \widehat{\bfsym \Sigma}_{t+1}$ the P-GARCH estimator.
The following theorem shows the asymptotic behaviors for the P-GARCH estimator.
With the realistic low-rank plus sparse structure, we can enjoy the blessing of dimensionality for estimating the factor volatility matrix. Moreover, using the regularization method, as shown in Theorem (ref), we can overcome the curse of dimensionality.
In this section, we discuss how the VaR value can be measured with the P-GARCH estimator $\widehat{\bfsym \Sigma}_{t+1}$ in (ref). Using the plug-in method, we estimate the one-step ahead VaR as follows:
where $c_\alpha$ is an $\alpha$-quantile, and $\mkern 1.5mu\overline{\mkern-1.5mu\bfm y\mkern-1.5mu}\mkern 1.5mu$ is a sample mean vector. To evaluate the VaR value, we need to determine the $\alpha$-quantile value. To do this, we assume that the standardized portfolio log returns are i.i.d. Then, when the standardized portfolio log returns follow the standard normal distribution, $c_{\alpha}$ is the $\alpha$-quantile of the standard normal $z_{\alpha}$. When the standardized portfolio log returns follow the multivariate t-distribution with the degrees of freedom $\nu$, then $c_{\alpha}=t_{\nu,\alpha}\sqrt{(\nu-2)/\nu}$, where $t_{\nu,\alpha}$ is an $\alpha$-quantile of the t-distribution with the degrees of freedom $\nu$ glasserman2002portfolio. We call them the parametric $\sigma$-based VaR estimator. The performance of the parametric $\sigma$-based VaR estimator depends on the distribution assumption. On the other hand, to obtain distribution robust VaR estimators, we use the non-parametric sample quantile method, which we call the non-parametric $\sigma$-based VaR estimator. For example, $c_{\alpha}$ is set to be $\lceil \alpha T \rceil$-th smallest value of $\{(\bfm w^\top (\bfm y_t - \mkern 1.5mu\overline{\mkern-1.5mu\bfm y\mkern-1.5mu}\mkern 1.5mu)/ (\bfm w^\top \widehat{\bfsym \Sigma}_{t}\bfm w)^{1/2}\}_{t=1}^T$, where $\lceil \cdot \rceil$ is a ceiling function.
Then, the following theorem shows the convergence rates of VaR.
In this paper, we focus on investigating the effect of the VaR estimator with the financial big data for a small portfolio. When we do not incorporate financial big data, that is, $p$ is finite, the absolute VaR error is not consistent for all parametric and non-parametric estimators. This is because the term $\sqrt{s_p/p}$ does not converge and so the latent factor estimation error is dominant. Therefore, Theorem (ref) supports that incorporating financial big data leads to the blessing in the VaR forecast by capturing common factor dynamics.
In this section, we conducted Monte Carlo simulations to check the finite sample performances of the proposed P-GARCH and corresponding VaR model. The data-generating process is analogous to the models (ref) and (ref). The number of factors was chosen to be $r=3$; the mean of $\bfm y_t$ is set to be zero $\ensuremath{\boldsymbol{\mu}}=\mathbf{0}$; and the factor loading matrix $\bfm V$ was randomly sampled from the first $r$ right singular vector of random matrix, which has the elements $\text{i.i.d.}\,\text{Unif}(0,1)$. The latent factors $\bfm f_t$ were generated by the multivariate normal distribution with the conditional expected volatility $\bfsym \Sigma_{f,t}$ as follows:
where $\bfm h_{1}(\ensuremath{\boldsymbol{\theta}}_0) = \left(\bfm I_r - \bfm A_0 -\bfm B_0 \right)^{-1} \ensuremath{\boldsymbol{\omega}}_0$,
The idiosyncratic risk $\bfm u_t$ was generated from the multivariate normal distribution with the following sparse co-volatility:
We varied $p$ from 20 to 500 and $T$ from 500 to 10,000, and employed the QMLE procedure in Section (ref) to estimate the GARCH parameters. We repeated the entire procedure 500 times.
Table (ref) shows the mean absolute error (MAE) of the QMLE estimate $\widehat{\ensuremath{\boldsymbol{\theta}}}$ for various $p$ and $T$. For the sake of spatial efficiency, we only documented estimation results with nine parameters.\footnote{The estimations with the rest of parameters provide similar results, which are available upon request.} The result illustrates that MAEs decreases as the parameter as $T$ or $p$ increases, which supports the theoretical findings in Theorem (ref).
With the estimated parameter $\widehat{\ensuremath{\boldsymbol{\theta}}}$, we validated the one-step ahead volatility prediction $\widehat{\bfsym \Sigma}_{t+1}$ with several matrix norms. We implemented the thresholding procedure for the idiosyncratic volatility in (ref) with the tuning parameters $C_{\tau}=1$ and $s_p=1$. Figure (ref) displays the prediction errors of $\widehat{\bfsym \Sigma}_{t+1} - \bfsym \Sigma_{t+1}$ with the Frobenius, spectral, max, and relative Frobenius norms. We find that the errors are large at $p=500$ for the Frobenius and spectral norms that depend on the dimensionality $p$, whereas the error of the relative Frobenius norm decreases as the dimension p increases. This result thus highlights that the relative Frobenius norm can lead to the blessing of dimensionality, thereby helping risk managers enjoy the use of financial big data for volatility prediction. The max norm shows no significant difference for different $p$, which can be explained by the fact that a large volatility matrix is more likely to have a large max norm error, despite a decrease in the element-wise error. In all cases, the errors decrease as the sample size $T$ increases, which supports the theoretical findings in Theorem (ref).
We compared the performance of predicting one-step ahead volatility matrices of the assets in the portfolio to clarify the benefits of using financial big data for small portfolio risk analysis We set the portfolio size as $s=5$ and the total number of assets as $p=500$. To capture the market dynamics within the portfolios of interest and compare with our proposed model (P-GARCH), we considered the CCC model bollerslev1990modelling, and BEKK with diagonal-constraint and variance targeting engle1995multivariate,pedersen2014multivariate. Additionally, in the empirical study, we often impose the static or slow time varying covariance assumption, and under this condition, we adopted POET fan2013large and historical sample covariance (Hist-Vol). Figure (ref) plots the average predicted volatility errors of the portfolio volatility matrix $\widehat{\bfsym \Sigma}_{s,t+1}-\bfsym \Sigma_{s,t+1}$ against the sample size $T$ under the Frobenius, spectral, max, and relative Frobenius norms. We observe that the P-GARCH model is superior to the other multivariate volatility matrix models by showing much lower errors across different sample sizes. This finding can be explained by the fact that the P-GARCH can address the market dynamics with small portfolios, whereas the others cannot fully address dynamics.
Finally, we compared the VaR forecasting performance with the small portfolio risk analysis methods. We set $\alpha$ as 1% and fixed $p=500$ and varied the portfolio size from 1 to 20. We also calculated the small portfolio matrix estimators $\widehat{\bfsym \Sigma}_{s,t+1}$ based on the P-GARCH, CCC, BEKK, POET, and Hist-Vol for the volatility forecast. We compared the considered models with the portfolio univariate GARCH(1,1) model, which is a widely-used approach to modeling a portfolio return bollerslev1986generalized. For the $\alpha$-quantile, we considered the normal and student-t distribution with 6 degrees of freedom ($\nu=6$) for the parametric method, whereas we used the non-parametric method with $\lceil \alpha T \rceil$-th smallest values of the standardized portfolio log returns, as in Section (ref). The portfolio was set to be equally weighted. The true VaR is $z_\alpha (\bfm w^\top \bfsym \Sigma_{t+1}\bfm w)^{1/2}$, where $z_\alpha$ is the $\alpha$-quantile for the standard normal distribution.
Figure (ref) shows the MAEs of the 1%-level VaR forecasts for $p=500$ with 500 replications. We identify that the P-GARCH outperforms the other models for any $\sigma$-based methods and portfolio size. When comparing the quantile estimation methods, given that the true model is based on the normal distribution, the normal distribution-based method shows the best performance. However, the sample quantile-based method shows similar performance as the sample size $T$ increases. This is because the sample quantile is a non-parametric consistent estimation procedure; thus, with a large $T$, we can estimate the quantile consistently. Finally, the parametric-based methods, such as P-GARCH, CCC, BEKK, and port-GARCH, usually perform better than the non-parametric methods. However, as the portfolio size increases, the BEKK method performs worse. One possible explanation is that as the portfolio size increases, the BEKK model becomes too complex to estimate model parameters.
In this section, we examined the performance of one-step ahead VaR forecasting with the empirical data. The data comprise 18 years CRSP daily percentage log returns of S&P 500 constituents from January 1, 2000, to December 31, 2017 (4,523 days). The log returns are calculated from the percentage return of the last sale or closing bid/ask price including dividends. We selected firms that have been ever a constituent of the S&P 500 during this period and filtered out the stocks that have any missing data. This selection process led us to include 492 firms ($p=492$) for the empirical study.
To employ the proposed P-GARCH model, we should determine the number of factors ($r$). Figure 4 shows that the optimal $r$ is less than $3$, and the rank $\widehat{r}$ estimated by (ref) in the online supplement is also 3. Thus, we determined the possible values of $r$ as 1, 2, and 3, which is in line with previous empirical studies chan1998risk,fama1992cross,fama1993common. For the thresholding step, we used the global industry classification standard (GICS) sector fan2016incorporating. For example, we maintained within-sector volatilities but set others to zero. The equally weighted portfolios of sizes 5 and 20 are randomly sampled from the 492 stocks with 500 repetitions, and the single-asset portfolio contains only one stock from 492 assets. With these portfolios, we predicted VaR with a significance level of $\alpha=10\%$, 5%, 2%, and 1% using the parametric and non-parametric methods, as discussed in Section (ref). For the model fitting, we employed a rolling window scheme with the window size $T=252$ days (1 year). The parameter is updated every 10 days, whereas the VaR forecasting is done every day with the updated log return data. The total forecasting number $N$ is 4,270. The estimated parameters of P-GARCH(r=3) for the entire period are $\widehat{\ensuremath{\boldsymbol{\theta}}}=(\widehat{\ensuremath{\boldsymbol{\omega}}},\mathrm{vec}(\widehat{\bfm A}),\mathrm{vec}(\widehat{\bfm B}))=10^{-3}\times($16.32, 0.32, 0.42, 83.99, 0.00, 0.08, 6.65, 48.23, 0.90, 87.04, 5.88, 62.87, 892.7, 0.00, 0.00, 0.01, 944.7, 0.23, 0.01, 0.00, 934.99$)$. The estimated parameter matrix $\widehat{\bfm B}$ is almost diagonal, which indicates that the recursive structure mainly comes from its own volatility. Similarly, the estimated parameter matrix $\widehat{\bfm A}$ is also almost diagonal, except for $A_{13}$. However, when considering the magnitudes of the first and third eigenvalues, $A_{13}$ does not affect the first eigenvalue significantly. Thus, the dynamics of the three factors are mainly explained by their own past volatilities. Figure (ref) depicts one sample path of the estimated one-step ahead 1%-level VaR under t-distribution with a portfolio size of 5. It shows that predicted VaR of the P-GARCH model can capture the dynamics of extreme loss.
We compared the out-of-sample VaR forecast performances with the small portfolio risk analysis methods considered in Section (ref). For example, we used the P-GARCH, CCC, BEKK, portfolio GARCH, Hist-Vol, and POET for the volatility prediction and the standard normal quantile, t-quantile with degrees of freedom 6, and sample quantile for the $\alpha$-quantile. To test whether the predicted VaR is correct, we used the unconditional coverage test ($LR_{uc}$) kupiec1995techniques, the conditional coverage test ($LR_{cc}$) christoffersen1998evaluating, and the dynamic quantile test (DQ$_{\text{VaR}}$) kuester2006value. Table (ref) reports the averaged hit rates ($N_{\alpha}/N$) and the average $p$-values of the $LR_{uc}, LR_{cc}$, DQ$_{\text{Hit}}$, and DQ$_{\text{VaR}}$ for portfolio sizes of 1, 5, 20 and $\alpha=1\%$ with the normal distribution, student-t distribution, and sample quantile. Figure (ref) presents box plots of individual portfolio's LR$_{cc}$ $p$-values for portfolio sizes of 1, 5, and 20 and $\alpha=1\%$ with the normal distribution, t-distribution, and sample quantile. From Table (ref) and Figure (ref), we find that when comparing the dynamic models and static models, the dynamic models, such as the P-GARCH, BEKK, CCC, and portfolio GARCH, show better performance than the non-parametric models. This indicates that the financial market is not static and that the GARCH-type models can account for the market dynamics. When the dynamic models are compared, the P-GARCH has the highest $p$-values for most of the VaR tests. From this empirical result, we can conjecture that the P-GARCH accounts for the market dynamics by incorporating financial big data, and this leads to improved performance in VaR forecasting for the relatively small portfolio. Finally, when the quantile methods are compared, the t-distribution shows the best performance. Moreover, the normal distribution has the smallest $p$-values for all the portfolio sizes. This pattern coincides with the stylized fact that the log return follows the conditional heavy-tailed distribution cont2001empirical.
In this paper, we aim to develop a dynamic process of portfolio risk measurement with the use of financial big data. The proposed model benefits from the use of financial big data in that it leads risk managers to take into account latent risk factors that are not included in their portfolios. One may be concerned about the curse of dimensionality with regard to the use of financial big data in parameter estimation; however, the proposed approach can overcome this problem with the POET method fan2013large under the dynamic factor model in a GARCH form. We showed that a large number of assets aids in increasing the accuracy of factor volatility estimation. This finding results from the fact that every asset contains common market information so that an increase in the number of assets enables one to obtain more information on latent risk factors. Thus, the proposed model achieves the blessing of dimensionality in estimating latent common factors.
We tested a finite sample performance of the proposed dynamic volatility model and the corresponding VaR model with Monte Carlo simulations and showed that the results support our theoretical findings. We further investigated the performance of our one-step ahead volatility matrices in the asset portfolio and concluded that our dynamic process with P-GARCH outperforms other competitive models for any $\sigma$-based methods and portfolio size. To guarantee the performance of our predictive model with financial big data, we used S&P 500 constituents with parametric and non-parametric methods. The results illustrated that the predicted VaR of our dynamic volatility model can capture the dynamics of extreme loss, which can be extreme cases from systematic factors in the market. Finally, the empirical study shows that our dynamic model with P-GARCH outperforms the other competitive dynamic models in terms of forecasting the VaR with financial big data for a relatively small portfolio.
The lesson from the 2008 financial crisis and similar systemic events in the financial market of history clearly argues that risk managers in the financial industry may have to take into account latent risks possibly accounting for systematic factors. The portfolios that they aim to analyze may consist of idiosyncratic factors with a limited number of risk factors (or assets); hence, the predicted level of risk in their target portfolios may fail to address the entire risk landscape. This problem can be also of importance for financial regulators in that underestimation of risk capitals based on a small portfolio risk measurement would potentially prevent the financial markets from being stabilized, thereby leading to negative externality to the economy, as observed in 2008 kaserer2019systemic\footnote{The negative externality was realized in 2008 as a huge bailout to, e.g., AIG as described in the introduction. Most of the \$374 billion total bailout was provided to financial firms, e.g., AIG, Citigroup and Bank of America, and the government rationale of this financial support was to avoid the collapse of world financial markets potentially creating the risk of another Great Depression harrington2009financial. This negative externality to the economy clearly burdened tax-payers for not only the size of bailout from taxes, but also the negative spillover effect on real economy.}. This study contributes to this regard by providing an innovative approach to using financial big data in improve the accuracy of portfolio risks.
In this paper, we only consider the factor dynamics, but several studies have also shown that the idiosyncratic volatilities are also affected by common market factors barigozzi2016generalized, connor2006common, herskovic2016common, rangel2012factor, shin2021factor. Thus, it is interesting but difficult to develop idiosyncratic dynamic models. In addition, we consider the $\sigma$-based approach to estimate the VaR values with the parametric and non-parametric quantile structures. However, the parametric structure cannot account for some stylized features, such as the skewness of the asset returns, whereas the empirical quantile approach often has large errors in the extreme tail. Thus, it is important and interesting to develop a quantile estimation procedure that can accomodate the stylized features of the asset returns. Finally, we assume that the mean process is constant over time, but this assumption is often violated in finance practice. Unlike the volatility process, the mean process does not have a strong linear time series structure, which makes the modeling of a dynamic mean process difficult. Thus, a parametric structure of the mean process, which can account for the heterogeneity of the mean process, should be developed. We leave these interesting problems for future study.
Theorems and technical proofs are provided online at Wiley in the Supplementary Material of this article. Readers may refer to the supplementary material associated with this article, available at Wiley (\url{https://onlinelibrary.wiley.com/journal/15396975}).
\let