Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
83,999 characters · 20 sections · 72 citation commands
Statistical Estimation for Covariance Structures with Tail Estimates using Nodewise Quantile Predictive Regression Models
\pagenumbering{roman}
I wish to thank my main PhD supervisor Prof. Jose Olmo for guidance and continuous encouragement throughout the PhD programme as well as Prof. Peter W. Smith for helpful discussion and for stimulating the further development of the current study in the direction of node exclusion and feature selection. In addition, I wish to thank Zudi Lu, Jean-Yves Pitarakis, Tassos Magdalinos, Hector Calvo Pardo, Zacharias Maniadis, Ruben Sanchez Garcia and Tullio Mancini as well as Departmental Seminar Series speakers at the University of Southampton: Luis E. Candelaria, Juan Carlos Escanciano, Marcelo C. Medeiros, Sebastian Engelke, Markus Pelger and Loriano Mancini for helpful conversations. \\
Lecturer in Economics, University of Exeter Business School, Exeter EX4 4PU, United Kingdom. Email: \textcolor{blue}{[email removed]} } } }
\setcounter{page}{1} \pagenumbering{arabic}
Network analysis has seen in recent decades a growing attention in the statistics, econometrics and quantitative finance literature; with a common aspect of consideration the role of node centrality in model estimation and portfolio optimization problems. Specifically, in the statistics literature related applications include graphical models and node exclusion tests (see, roverato1998isserlis and salgueiro2005power), while in the econometrics literature recent applications investigate the modelling of interconnectedness and financial spillover effects based on graph structures that capture interactions among economic agents (see, hong2009granger, billio2012econometric, diebold2012better, hardle2016tenet, blasques2016spillover, barigozzi2017network, barunik2018measuring). In the finance and risk management literature, a relevant aspect of interest is the optimal portfolio allocation problem using financial networks\footnote{A specific research question of these studies is the relation between stock centrality and portfolio risk. peralta2016network examine the minimum variance and mean-variance asset allocation problems by connecting the optimal portfolio weights of each investment strategy to the eigenvector centrality. The authors use the correlation decomposition of the covariance matrix to represent the network topology. Thus, the variance-covariance matrix is decomposed into a quadratic form with the correlation matrix as a quadratic matrix and a diagonal matrix with the variances of the assets. By defining a correlation matrix as a matrix which has $1'$s on its main diagonal and the pairwise correlations on its off diagonal elements then this matrix is considered as an adjacency matrix.} (see, peralta2016network, olmo2021optimal, katsouris2021optimal and lin2022portfolio). All aforementioned studies consider suitable statistical dependence measures that are robust to deformations of the network topology, a commonly occurring phenomenon during a financial crisis. Our main research objective is to investigate methodologies for modelling extreme events in graphs combining both tail risk and the induced centrality measures.
Our first contribution to the literature is the construction of a regression-based covariance-type matrix for modelling tail dependency in graphs using pairwise quantile regression models\footnote{The econometric specification for obtaining the risk measures of VaR and CoVaR are presented in the studies of adrian2016covar and hardle2016tenet (see also, white2015var).}; which to the best our knowledge is a novel estimation approach. Specifically, the proposed risk matrix captures higher-order dependency induced by the tail behaviour of the underline distributions of a fixed graph\footnote{A graph representation of a financial network with the proposed risk matrix is proposed by katsouris2021optimal who show that this covariance-type induces a positive-definite quadratic form which permits to consider optimal portfolio allocation problems under tail events. In this paper, the focus is on the proposed feature selection algorithm.}, under the assumption of time series stationarity\footnote{In this paper, we consider that the pairwise quantile predictive regression models are driven only by the predictive regression equation without using a nonstationary specification for regressors. An extension to the case in which the proposed tail-risk matrix depends nonlinearly on the degree of persistence is currently work in progress by the author.}. Although we do not directly model the multivariate extreme dependence as in the study of engelke2020graphical, the constructed large covariance-type matrix captures a general form of network interconnectedness through pairwise graph causality. In particular, the conditional tail (extremal) dependency across the nodes of the graph is measured in relation to the risk measures of VaR and CoVaR obtained as tail estimates (see, Section (ref)).
Our second contribution to the literature, is a novel feature selection algorithm based on the proposed risk matrix, that we call Feature Ordering by Centrality Exclusion (FOCE). This node selection algorithm provides an ordering mechanism using the parametrically specified covarariance type matrix that corresponds to a fixed quantile level, denoted with $\boldsymbol{\Gamma}_n ( \uptau )$ for some $\uptau \in (0,1)$, which implies that the covariance structure is approximated with tail estimates obtained from pairwise quantile (predictive) regressions fitted on a sample size of $n$ observations. Recently, caner2022sharpe present a framework for nodewise regression when the residuals from a fitted factor model are used and a feasible residual-based nodewise regression is proposed to estimate the precision matrix of errors.
In this paper we focus on the introduction of the novel risk matrix as well as in the introduction of the novel graph topology measure. In particular, the proposed risk matrix captures tail dependence and causality in graphs and thus the proposed algorithm of feature ordering is based on this conditional tail dependence measure providing this way a methodology to order characteristic statistics of the graph topology such as centrality measures based on the tail risk matrix. Specifically, the proposed algorithm has an oracle property in the sense that if we plug the true population parameters then our algorithm gives the true feature selection by centrality exclusion of the population variables. Using the proposed algorithm and risk matrix, we aim to model multivariate tail behaviour by considering the tail behaviour of each pair of nodes in the graph. More importantly, it appears to have good performance in simulated and real data sets. Therefore, we focus on the development of the FSCE as well as the proofs of its consistency. The proposed algorithm is not possible to be implemented in low-dimensional models due to convergence and consistency issues based on the proposed stopping rule of the algorithm. We aim to investigate whether our procedure can be estimated with respect to convergence rate $\mathcal{O} \left( n \ \mathsf{log} n \right)$ where $n$ is the sample size.
Usually in the literature the main concern is to show that for $p >> n$ the Sharpe Ratio estimators are consistent in global minimum-variance and mean-variance portfolios. However, in this paper our main concern is to demonstrate the use of the proposed tail-risk matrix in portfolio optimization problems. Specifically, covariance matrices estimated using high dimensional statistical models are used in asset allocation settings as a precision matrix (inverse of the covariance matrix). On the other hand, in high dimensional settings a sparsity condition on the precision matrix of the errors from a statistical model (such as a factor model, a nodewise quantile predictive regression model etc.), can be justified in empirical finance studies. Furthermore, the statistical interpretation of of zero entries on a precision matrix can be attributed into conditional independence. Then, our proposed methodology focuses on estimating a covariance-type matrix using the tail forecasts rather than based on the estimated residuals from a factor model as in the study of caner2022sharpe.
A different stream of literature considers shrinkage methods such as the framework proposed by ledoit2012nonlinear, who consider a nonlinear shrinkage estimator in which small eigenvalues of the sample covariance matrix are increased and large eigenvalues are decreased. Although methodologies for shrinkaging the eigenvalues of covariance matrices which are used for asset allocation purposes ensure that the optimization is not ill-conditioned, this aspect is of independent interest since our main contribution for this paper is to study the main properties of this novel tail-risk matrix. A conjecture is that similar to the case when the covariance matrix is estimated from the errors of factor models, a larger number of included factors will result on a slower convergence rate of the estimation error. Similarly, when the tail-based covariance matrix is estimated with nodewise quantile predictive regressions, an increased number of regressors (either stationary or nonstationary) will result to a slower convergence of the estimation error. Notice that with estimation error here we mean the error in estimating the true structure of the covariance matrix. Moreover, the conventional approach in the literature is to separate between the number of units (e.g., $p$ nodes in the network) versus the number of time series observations. In our case we have another relevant parameter which is the number of regressors included in each quantile predictive regression model (similar to the case of factors when considering factor models). In this paper, we focus in the case in which the number of nodes in the network is not necessarily much larger than the number of time series observations, that is, $p < n$ (or $N < T$ in an equivalent notation).
In the conventional factor structure model case, the following theorem holds for the precision matrix.
Notice that Example (ref) illustrates the decomposition of the precision matrix when the vector of stock returns have a factor structure representation. On the other hand, in this article we assume that the there exists an underline network structure that describes the stochastic behaviour and comovements of these nodes in the graph. Some key points are summarized below:
Extreme events (such as financial crises) are commonly employed for causal discovery since the relationships between variables may manifest themselves especially in the largest observations. In particular, a modelling methodology based on extreme value theory is presented in the study of gnecco2021causal\footnote{In the particular framework the authors use $X_k := \sum_{ k \in pa (j, G) } \beta_{jk} X_k + \epsilon_j, \ j \in V$ under the assumption that the noise sequence $\epsilon_j, j \in V$, are heavy-tailed with extreme value index $\gamma > 0$ such that $P ( \epsilon_j > x ) \sim \ \ell(x)x^{-2/ \gamma}, \ \ x \to \infty$.}. A discussion about tail dependence in graphs is provided in the paper of engelke2020graphical (see also discussion in dombry2016asymptotic). On the other hand, our regression-based risk matrix takes into utilized tail estimates rather than relying on fitting a parametric model to heavy-tailed data, which implies that the corresponding matrix estimator is robust to the presence of potential outliers in financial and macroeconomic data.
The estimation of covariance matrices is widely used in various applications such as optimal portfolio choice problems and other statistical problems such as multivariate analysis, principal components analysis and graphical modeling among others. The commonly used assumption for the estimation of large covariance matrices is the assumption of multivariate normality. For instance, brown1975techniques and browne1984asymptotically proposed statistical methodologies for the analysis of covariance matrices such as asymptotically distribution-free methods. In particular, the majority of estimators of parameters in structural models for covariance matrices are members of the class of generalized least squares (GLS) estimators. Specifically, this type of estimators include the maximum Wishart likelihood estimators (MWL) in the sense that a generalized least squares discrepancy function can be constructed which will in general attain its minimum at the point defined by the MWL estimator (see, browne1984asymptotically). A recent study who consider the MLE estimation of covariance type matrices when both rows and columns can be correlated is presented by drton2021existence.
The statistical literature on methodologies for the robust estimation of large covariance matrices has been thrived the last decades. In particular, assumptions such as the Wishart distribution for the unknown covariance matrix allows to implement asymptotically distribution-free methods for the robust estimation. However, features such as heavy-tailed time series, temporal dependence or even the presence of structural instability in time-series affects the robustness of the estimation procedure. For instance, dendramis2021estimation accommodate the time variation, dependence and heavy-tailedness of distributions with implementation of a time-varying covariance matrix estimation procedure that includes thresholding. Therefore, the novelty of our estimation procedure is the use of quantile predictive regression models for capturing these features, which implies a regression-based tail risk matrix. Moreover, the particular proposition allows to take into consideration the time series properties via the estimation methodology rather than implementing ex-ante corrections to the estimation procedure which implies implementing bias corrections for the possible effect of the particular features to the robustness and degree of accuracy of the large covariance matrix.
A different stream of literature investigates the use of centrality measures (see, a discussion in aamari2021graph) for the development of graph-based modelling methodologies. Our approach is different than portfolio sorting procedures such as the studies of mcgee2020optimal and ledoit2018efficient who propose methodologies for sorting portfolios based on variable characteristics and nonlinear time series models respectively, however our motivation is driven by the construction of an endogenous to the system sorting mechanism. Furthermore, the aspect of bridging centrality and extremity is examined in the study of einmahl2015bridging. In terms of rank centrality measures a related approach is presented in the study of negahban2017rank, while a formal asymptotic theory framework based on the random matrix theory is given by Li2021central. Furthermore, negahban2017rank propose a rank centrality measure which is an iterative rank aggregation algorithm for discovering scores for objects (or items) from pairwise comparisons in a statistical sense.
A third line of literature which is related to our study is the use of the concept of conditional independence. Specifically, a novel conditional independence measure is proposed by azadkia2021simple which corresponds to the case of conditional mean preserving measures. Lastly, although we mainly focus with the aspects of statistical estimation, relevant inference procedures include testing for misspecification of the underline covariance structure (see, guo2021specification and chang2022testing). Further limit theorems on covariance structures are given in the recent study of liu2021robust while chang2021central establish central limit theorems for high-dimensional dependent data.
In summary, our proposed modelling approach focuses on two key aspects: (i) the construction of the risk matrix based on tail dependence measures and (ii) the implementation of the feature selection algorithm. Although our feature selection algorithm is applied to the tail risk matrix, we conjecture that this estimation procedure can be also used in alternative structures of covariance matrices.
Specifically, the construction of the proposed risk matrix captures causality\footnote{A discussion related to statistical causality can be found in schenone2018causality.} for extreme values based on nodewise quantile predictive regression models which capture the quantile behaviour of the underline distributions. Our novel risk matrix is constructed using nodewise regression models in the sense that the elements of the risk matrix are estimated using pairwise quantile regression models.
Given the singular value decomposition $\boldsymbol{\Gamma} = \boldsymbol{U} \boldsymbol{D} \boldsymbol{V}^{\prime}$. Moreover, denote with $\big\{ \boldsymbol{\Gamma}_t \big\}_{t=0}^T$ be the sequence obtained by the iteration of Algorithm 1.
Furthermore, in this section we derive some useful oracle inequalities such as provide suitable error bounds for the difference between the sequence of matrices $\boldsymbol{\Gamma}_n$ generated by Algorithm 1 and the true matrix $\boldsymbol{\Gamma}$. These results are determined by the strong convexity of the optimization function $\rho_{\uptau}$.
Notice that Assumption A3 is required in order to obtain the bounds on the tail probabilities of the estimation error. Furthermore, we consider the non-asymptotic bound for $\left\lVert \boldsymbol{\Gamma}_t - \boldsymbol{\Gamma} \right\rVert_F$ in the general situation that the number of factors can be increasing with $n$.
Our risk matrix is employed to model tail causality for the nodewise regression models. In particular, these nodewise regressions correspond to quantile predictive regression models as defined by adrian2016covar which are used to obtain estimates for the risk measures of VaR and CoVaR. Consider that we have the vector $\underline{Y} = ( Y_{1,t},..., Y_{p,t} )$ such that $i \in \left\{ 1,..., p \right\}$ and $j \in \left\{ 1,..., p \right\}$. Consider the single index model for estimating the $\mathsf{CoVaR}$ and $\mathsf{VaR}$ risk measures
The given specification is motivated from the literature of financial connectedness and tail risk and in particular is based on the seminal works of adrian2016covar and hardle2016tenet who examine statistical estimation methodologies and their properties. We firstly consider the following covariance-type structure
where $\mathbf{\widetilde{\Sigma}}_{ij} \in \mathbb{R}^{ N \times N }$ represents the financial connectedness matrix and provides a robust representation of idiosyncratic and systemic risk within the financial network. Our proposed risk matrix belongs to the family of dynamic high dimensional covariance matrices and captures network tail risk dependence, since is based on tail risk measures. We estimate the time-varying VaRs and CoVaRs conditional on a vector of lagged state variables $\boldsymbol{X}_{t-1}$. Both VaR and CoVaR risk measures represent quantiles of the distribution of returns under different conditioning sets; hence, to obtain one-period ahead predicted values for these risk quantities we employ the quantile regression specifications proposed by the seminal study of Koenker1978regression (see, also KoenkerXiao02). The construction of one-period ahead forecasts for the dynamic portfolio allocation, based on a graph representation, consists one of the main contributions of the empirical application of the paper.
Consider the conditional quantile function estimated via $Q_{y_i} \left( \tau | x_i \right) = F_{y_i}^{-1} \left( \tau | x_i \right)$. Then, the optimization function to obtain the model estimates is expressed as below
where $\tau \in (0,1)$ is a specific quantile level, and $\rho_{\tau}( u ) = u \left( \tau - \mathbf{1}_{ \left\{ u < 0 \right\} } \right)$ is the check function (see, newey1987asymmetric). Suppose $1 \leq t \leq T$, then based on the estimation method given by (ref) the VaR and CoVaR are generated by the following econometric specifications
where $\{ R_{(i),t} \}$ for $i \in \left\{ 1,...,N \right\}$ is the vector of portfolio returns at time $t$, and $\boldsymbol{X}_{t-1}$ is a vector of exogenous regressors containing macroeconomic characteristics common across assets and denote with $\Theta = \{ \nu_{(i)}, \boldsymbol{\xi}_{(i)}, \nu_{(j)}, \boldsymbol{ \xi }_{(j)}, \beta_{(j)} \}$ to be a compact parameter space.
Let $Q_{\tau} \left( \cdot\ | \ \mathcal{F}_{t-1} \right)$ denote the quantile operator for $\tau \in (0,1)$ conditional on an information set $\mathcal{F}_{t-1}$. Then, the error terms satisfy $Q_{\tau} \left( u_{(i),t} \ | \ \boldsymbol{X}_{t-1} \right)=0$ for the quantile regression model (ref), and $Q_{\tau} \left( v_{(j),t} \ | \ R_{(j),t}, \boldsymbol{X}_{t-1} \right) = 0$ for the quantile regression model given by expression (ref). Moreover, denote with $\boldsymbol{\theta}_{(1)} = \left( \nu_{(i)}, \boldsymbol{\xi}_{(i)}^{\prime} \right)$ and $\widetilde{\boldsymbol{X} }_{t-1}^{(1)} = \left( 1, \boldsymbol{X}_{t-1}^{\prime} \right)^{\prime}$ for model (ref), and with $\boldsymbol{\theta}_{(2)} = \left( \nu_{(j)}, \boldsymbol{\xi}_{(j)}^{\prime}, \beta_{(j)} \right)$ and $\widetilde{\boldsymbol{X}}_{t-1}^{(2)} = \left( 1, \boldsymbol{X}_{t-1}^{\prime}, R_{(i),t} \right)^{\prime}$ for model (ref).
Then, the model estimate
Similarly, the quantile estimator of (ref) is obtained via the optimization function below
The state variables, that is, the macroeconomic and financial variables, included in the econometric model, capture time variation in the conditional moments of asset returns. Therefore, the $\mathsf{VaR}$ and $\mathsf{CoVaR}$ of each financial institution are estimated as vectors from a time-varying distribution.
This approach allows to explain the optimal portfolio allocation in terms of changes in the systemic risk and financial connectedness in the network. To implement this, we use a rolling window, within which the quantile regressions given by (ref) and (ref) are fitted and the $\boldsymbol{\Gamma}$ risk matrix is constructed based on one-period ahead forecasts. Specifically, the one-period ahead $\mathsf{VaR}_{t+1}$ for a coverage probability $\tau \in (0,1)$ is obtained by regressing $R_{(i),t}$ on $\boldsymbol{X}_{t-1}$ as given by the first step regression ((ref)) which gives the parameter estimates $\widehat{\nu}_{(i)}$ and $\widehat{ \boldsymbol{\xi} }_{(i)}$. The forecasted one-period ahead VaR is computed as below
Similarly, the one-period ahead forecast of $\mathsf{CoVaR}_{t+1}$ is obtained by regressing $R_{(j),t}$ on $\boldsymbol{X}_{t-1}$ and $R_{(i),t}$ as shown in the second step regression ((ref)) using the quantile estimation as defined by (ref). To do this, we collect the parameter estimates $\widehat{\nu}_{(j)}, \widehat{\boldsymbol{\xi} }_{(j)}, \widehat{\beta}_{(j)}$ from the second regression, and construct the forecasted one-period ahead $\mathsf{CoVaR}_{t+1}$ measure using the following expression:
where $\widehat{\mathsf{VaR}}_{i,t+1}( \tau )$ replaces the actual $\mathsf{VaR}_{i,t+1} ( \tau )$ measure. Then, the $\mathsf{\Delta \text{CoVaR}}$ measure is computed as the difference of the two risk measures
An important aspect of our proposed estimation methodology is the assumption of a graph representation with nodes being the stock returns and edges being pairwise tail forecasts, which can be thought as the level of simultaneous Granger causality due to the existence of tail and graph dependence. Therefore, these tail forecasts are estimated using the aforementioned predictive regression models based on the bipartite graph structure, which implies an estimation implementation for each pair $(i,j)$ of nodes for the graph $\mathcal{G} = (\mathcal{V}, \mathcal{E})$.
In this section, we derive results for establishing the asymptotic theory for the estimator of the proposed risk matrix that correspond to the population risk measures. Define with $\mathtt{Upper} ( \mathbf{\Gamma} ) \in \mathbb{R}^{ N \times N}$ the upper triangular part of the matrix, excluding the diagonal entries and with $\mathtt{Lower} ( \mathbf{\Gamma} ) \in \mathbb{R}^{ N \times N}$ the lower triangular part of the matrix, excluding the diagonal entries. Then, the risk matrix $\mathbf{\Gamma} \in \mathbb{R}^{ N \times N}$ is given by the following expression:
where the off-diagonal elements consist of the term:
Since the main diagonal of the risk matrix includes only the VaR tail measure which correspond to the underline node specific distribution, then we suppose that is generated by a common cumulative distribution function $F(.)$. An important condition for our risk matrix to be utilized within the optimal portfolio problem is to ensure that it converges in probability to a positive definite matrix for large sample size, that is, in our case, as the number of nodes in the network grows to infinity, i.e., $N \to \infty$.
Then, we can observe that the proposed covariance structure has a similar decomposition to the covariance matrix as below:
Moreover, we are interested in estimating the precision matrix of the $\mathsf{VaR}-\Delta \mathsf{CoVaR}$ risk matrix, $\boldsymbol{\Omega} := \boldsymbol{\Gamma}^{-1} \equiv \displaystyle \frac{ \omega_{j_1, j_2} }{ \omega_{j_1,j_1} \omega_{j_2,j_2} }$ where $\omega_{j_1, j_2}$ correspond to the elements of the precision matrix of the $\mathsf{VaR}-\Delta \mathsf{CoVaR}$ matrix for some $(j_1, j_2) \in \left\{ 1,..., N \right\}$.
The asymptotic behaviour of the risk matrix can be determined by considering the limiting distribution of the expression $ \sqrt{N} \left( \widehat{ \mathbf{\Gamma} } - \mathbf{\Gamma}_0 \right)$. Denote with
The term $\mathcal{A}_3$ converges to a Gaussian random variable, with a non-stochastic covariance matrix under the assumption of stationary and ergodic regressors in the quantile predicitve regression model.
We begin with the traditional optimal portfolio allocation problem. We extend this framework by proposing a novel risk matrix that captures tail events. We present regularity conditions which ensure the existence of a closed-form solution for the associated quadratic risk function and optimization functions. Our proposed risk matrix induces a well defined quadratic form which permits to be employed for optimal portfolio choice problems\footnote{A rational investor who holds a set of securities aims to minimize the portfolio risk based on the portfolio optimization framework introduced by markowitz1956optimization. The specific linear optimization mechanism has seen growing attention in various fields of financial economics such as portfolio management, risk management, capital investment and asset pricing. Furthermore, the seminal paper of sharpe1964capital introduced the fundamental principles of investment under conditions of uncertainty in financial markets. Furthermore, various studies examined the effects of constraints on the portfolio performance measures such as the expected return and the portfolio risk (see fishburn1977mean, jagannathan2003risk, and references therein).}.
The optimal portfolio theory proposed by Markowitz52 considers the minimization of the portfolio variance without any further restriction on the corresponding expected return. More formally, let $\mathbf{R}_t = ( R_{1,t} ,..., R_{N,t} )^{\prime}$ be the $N-$dimensional vector of assets returns at time $t$ and denote with $\text{\boldmath$\mu$} := \mathbb{E} \left( \mathbf{R}_t \right) = \left( \mathbb{E} \left( R_{1,t} \right) ,..., \mathbb{E} \left( R_{N,t} \right) \right)^{\prime}$ the vector of expected returns and $\mathbf{\Sigma} := \text{Cov}\left( \mathbf{R}_t \right) = \mathbb{E} \left( \mathbf{R}_t \mathbf{R}_t^{\prime} \right) - \text{\boldmath$\mu$} \text{\boldmath$\mu$}^{\prime}$ the associated variance-covariance matrix. Let $r_t^{p} = \mathbf{w}_i^{\prime} \mathbf{R}_t$ denote the return on a portfolio of $N$ assets, with $\mathbf{w}_i= \left( w_{1},...,w_{N} \right)^{\prime}$ the vector of portfolio weights representing the financial position of the investor as a weight allocation of assets to the portfolio. The variance of this portfolio is $V \left( r_t^{p} \right) = \mathbf{w}^{\prime} \mathbf{\Sigma} \mathbf{w}$. Then, the Markowitz's minimum variance optimization problem is expressed as below
The first order conditions to the minimization problem yield the optimal portfolio allocation given by the following vector of weights
The symmetry of $\widehat{ \mathbf{\Sigma} }$ yields the spectral decomposition $\widehat{ \mathbf{\Sigma} } = \mathbf{U} \mathbf{D}_{\sigma} \mathbf{U}^{\prime}$, with $\mathbf{U}$ is an $N \times N$ orthonormal matrix such that $\mathbf{U}^{\prime} = \mathbf{U}^{-1}$ which includes as columns the linearly independent eigenvectors of $\widehat{ \mathbf{\Sigma} }$, and $\mathbf{D}_{\sigma}$ an $N \times N$ diagonal matrix with the corresponding eigenvalues which is defined as $\mathbf{D}_{ \sigma } = diag \left\{ \lambda^{ \sigma }_1,..., \lambda^{ \sigma }_N \right\}$. Moreover, we can rearrange the diagonal eigenvalue matrix as $\mathbf{D}_{\sigma} =\mathbf{U}^{\prime} \widehat{ \mathbf{\Sigma} } \mathbf{U}$.
This allow us to express the eigenvalues of the risk matrix $\widehat{ \mathbf{\Sigma} }$ in terms of the elements of the variance-covariance matrix as below
where $\left\{ u_{ij} \right\}_{i,j = 1,...,N}$ denotes the eigenvectors of the risk matrix $\widehat{ \mathbf{\Sigma} } \in \mathbb{R}^{N \times N}$.
The condition in Assumption (ref) guarantees that $\widetilde{ \mathbf{\Gamma} }$ is positive definite, and as a result implies that the corresponding quadratic form $\mathbf{w}^{\prime} \widetilde{ \mathbf{\Gamma} } \mathbf{w}$ has certain desirable properties as summarized by the following propositions.
Lastly, if $\boldsymbol{w}_n \boldsymbol{n} = 1$ with $\widehat{\boldsymbol{w}}_n \overset{p}{\to} \boldsymbol{w}$, then the composite estimator $\widehat{\gamma}_n ( \widehat{\boldsymbol{w}}_n )$ is $\sqrt{k}-$asymptotically equivalent to $\widehat{\gamma}_n ( \boldsymbol{w}_n )$ in the sense that $\sqrt{k} \big( \widehat{\gamma}_n ( \widehat{\boldsymbol{w}}_n ) - \widehat{\gamma}_n ( \widehat{\boldsymbol{w}}_n ) \big) = o_p(1)$.
We denote with $ \mathcal{G} = \{ \mathcal{V} , \mathcal{E} \}$ a graph structure, representing a financial network, which consists of a set of nodes, $\mathcal{V}$, and a set of edges, $\mathcal{E}$, connecting the pairs of nodes. A full characterization of the information in the network is provided by the $N \times N$ adjacency matrix $\mathbf{\Omega}^0$ whose element $\{ \Omega_{ij}^0 \}_{i,j=1,...,N}$ determine whether there is a link connecting node $i$ and $j$ for the graph $\mathcal{G}$. The link between two elements can be characterized by a binary variable which determines the existence of a connection or by a real value that corresponds to the pairwise weight between nodes.
We propose the following adjacency matrix $\mathbf{\Omega}^0 = \widetilde{ \mathbf{\Gamma} } - diag \left( \widetilde{ \mathbf{\Gamma} } \right)$, where $diag(\cdot)$ is an $N \times N$ diagonal matrix, implying that $diag \left( \widetilde{ \mathbf{\Gamma} } \right) := \text{diag} \{ \gamma^{0}_1, \ldots, \gamma^{0}_N \}$. Therefore, within our framework we consider a weighted network, where the connections between stocks are determined by the $\widetilde{ \gamma}_{i|j}$ measures as defined by the off-diagonal elements of the $\widetilde{ \mathbf{\Gamma} }$ matrix. The adjacency matrix $\mathbf{\Omega}^0$ is symmetric by definition, since the main diagonal has a vector of zeros and the off-diagonal terms are equivalent to the symmetric risk matrix $\widetilde{ \mathbf{\Gamma} }$, that is, $\Omega_{ij}^0= \widetilde{ \gamma}_{i|j}$ for $i \neq j$.
The symmetry of the adjacency matrix entails the spectral decomposition $\mathbf{\Omega}^0 = \mathbf{Z}_{ \Omega^0 } \mathbf{D}_{ \Omega^0 } \mathbf{Z}_{ \Omega^0 }^{\prime}$, with $\mathbf{Z}_{ \Omega^0 }$ an $N \times N$ orthonormal matrix such that $\mathbf{Z}_{ \Omega^0 }^{\prime} = \mathbf{Z}_{ \Omega^0 }^{-1}$ that contains the linearly independent eigenvectors of $\mathbf{\Omega}^0$, and $\mathbf{D}_{\Omega^0}$ is an $N \times N$ diagonal matrix with the corresponding eigenvalues. Moreover, we express the eigenvalues of the adjacency matrix $\mathbf{\Omega}^0$ using the spectral vector decomposition as below
with $z_{ij}$ the eigenvectors of the matrix $\mathbf{\Omega}^0 \in \mathbb{R}^{N \times N}$. We decompose the loss function $\mathbf{w}' \widetilde{ \mathbf{\Gamma} } \mathbf{w}$ as the sum of a quadratic loss function given by the adjacency matrix and a quadratic loss function given by a diagonal risk matrix with elements given by $\gamma^{0}_i$. More formally, the quadratic form becomes
The study of the contribution of the centrality of assets to the loss function $Q( \widetilde{ \mathbf{\Gamma}} , \mathbf{w})$ is examined by katsouris2021optimal and olmo2021optimal. To do so, we further examine the components of expression $\mathbf{w}' \mathbf{ \Omega }^0 \mathbf{w}$ and provide a definition of asset centrality with respect to the eigenvalue decomposition of the adjacency matrix. The notion of centrality quantifies the influence of certain nodes in a given network. There are several measurements in the literature each corresponding to a specific definition of centrality. We focus on the eigenvector centrality, see also Bonacich72. This measure of centrality is closely related to the concept of Katz centrality, see katz1953new. The eigenvector centrality of asset $i$ is defined as the proportional sum of its neighbors' centrality, see Bonacich72.
A centrality measure although has no formal definition in the literature in this paper we propose some regularity properties for these measures to correspond to centrality ranking. Motivated from the above findings we propose a novel centrality measure. In particular, in this paper we define the context of recursive monotonicity. In mathematical terms recursive monotonicity means that if vertex $i$ has more neighbors than vertex $j$, and $i$'s neighbours are more central than those of $j$, then $i$ must be more central than $j$. Therefore, we focus on two different network topologies, a topology of low connectivity and a topology of strong connectivity (see, sadler2022ordinal). \color{black}
A measure of centrality is a function $C$ that takes a node $i \in V$ and the graph $G = ( V, E)$ it belongs to, and returns a non-negative real number $C(i ; G) \geq 0$ which translates to a quantity which specifies how central $i$ is in $G$. Currently, there is no agreement on a formal definition of a centrality measure the following properties should hold
Furthermore and more precisely, given a preorder on the vertex set of a graph, the neighborhood of vertex $i$ dominates that of another vertex $j$ if there exists an injective function from the neighbourhood of $j$ to that of $i$ such that every neighbor of $j$ is mapped to a higher ranked neighbour of $i$.
The preorder satisfies recursive monotonicity, and is therefore an ordinal centrality, if it ranks $i$ higher than $j$ whenever $i$'s neighbourhood dominates that of $j$. Consider for example commonly used centrality measures in the literature such as degree centrality, Katz-Bonacich centrality and eigenvector centrality, all of these centrality measures satisfy assumption 1 (recursive monotonicity). Therefore, following our preliminary empirical findings of the optimal portfolio performance measures with respect to the low centrality topology as well as the high centrality topology, we obtain some useful insights regarding the how the degree of connectedness in the network and the strength of the centrality association can affect the portfolio allocation. Building on this idea, in this section we focus on the proposed algorithm which provides a statistical methodology for constructing a centrality ranking as well as a characterization of the strength of the centrality association between nodes.
More specifically, we assume that a node exhibits "strong centrality" when all the rank centrality statistics assign a higher preference to vertex $i$ above $j$ (in terms of graph topology), which implies that in a sense $i$ is considered to be more central than the node $j$ for all $i,j \in \left\{ 1,..., N \right\}$ such that $i \neq j$. On the other hand, we assume that a node exhibits "weak centrality" when vertex $i$ is weakly more central than vertex $j$ if there exists some strict ordinal centrality that ranks $i$ higher than $j$.
According to sadler2022ordinal weak centrality itself is a strict ordinal centrality, so it is the maximal such order, and it extends strong centrality. Moreover, in this paper we employ these terminologies for the development of an algorithm for feature ordering by centrality exclusion. We define some useful mathematical definitions such as the binary relation $\succsim$ on $V$ is a set of ordered pairs of vertices, and we write $ i \succsim j$ if $(i,j) \in \succsim$.
We follow the example presented in janssen2003bootstrap. Given $k(n)$ let $\big( R_1,..., R_{k(n)} \big)$ be uniformly distributed ranks, that is, a random variable with uniform distribution on a set of permutations $\mathcal{P}_{k(n)}$ of $\left\{ 1,..., k(n) \right\}$. Furthermore, an additional index $n$ concerning the ranks is suppressed throughout. Consider a sequence of simple linear permutation statistics such that
where $c_{n_i}$ are regression coefficients, $1 \leq i \leq k(n)$, with
and random scores $d_n(i): \tilde{\Omega} \mathbb{R}$, $1 \leq i \leq k(n)$ with
The $c'$s are here considered to be fixed and the $d'$ are allowed to be random variables independent of the ranks $R_j$. Without restrictions we may assume ordered regression coefficients
Then, we have equality in distribution for $\big( d_n (R_i) \big)_{ i \leq k(n) } \overset{ D }{ = } \big( d_{ R_{i : k(n) } } \big)_{ i \leq k(n) }$, which is a consequence of the ex-changeability of these variables. Therefore, in these cases ranks and order statistics are independent. Thus, without restriction we may assume that the score functions are also ordered
Therefore, since $\mathcal{E} \big( S_n | W \big) = 0$ holds always convergence subsequences of $S_n$ exist and under some regularity conditions we can classify all possible limit distributions of subsequences with respect to the distributional convergence. We define the normalized random variables as below
The proposed algorithm of exclusion centrality is constructed using the proposed high-dimensional tail forecast risk matrix. However, the feature ordering by centrality exclusion (FOCE) procedure in practise can be applicable for covariance matrices for graphical models. On the other hand, when the FOCE procedure is implemented based on the VaR-$\Delta$CoVaR risk matrix, then the feature selection is based on the tail dependence of the response variable $\boldsymbol{Y}_t$ given the predictors $\boldsymbol{X}_{t-1}$. Furthermore, since the construction of the proposed risk matrix is based on the quantile predictive regression models, then we specifically focus on a type of granger causality in the tails of the underline distributions. Thus, our proposed algorithm provides a methodology for ordering the vertices of the graph based on the effect of their centrality in relation to the other nodes in the graph when we exclude the most central node in the graph. It produces an ordering of the predictors according to their predictive power. This ordering is used for variable selection without putting any assumption on the distribution of the data or assuming any particular underlying model.
The simplicity of the estimation of the conditional dependence coefficient makes it an efficient method for variable ordering and variable selection that can be used for high dimensional settings. In this section, motivated by the study of centrality measures and the proposed risk matrix of the paper, we propose a novel feature selection algorithm for multivariate regression models using the centrality exclusion procedure based on our novel regression-based risk matrix. Notice that other methodologies currently proposed in the literature include the model-based methods.
Let $Y$ be the response variable and let $\boldsymbol{X} = \left( X_j \right)_{ 1 \leq j \leq p }$ be the set of predictors. The data consists of $n$ i.i.d copies of $\left( Y, \boldsymbol{X} \right)$. Similar to the framework proposed by azadkia2021simple, first, choose $j_1$ to be the index $j$ that maximizes $T_n( Y, X_j )$. Then, having obtained $j_1,..., j_k$, choose $j_{k+1}$ to be the index $j \not\in \left\{ j_1,..., j_k \right\}$ that maximizes $T_n \left( Y, X_j | X_{j_1},...., X_{j_k} \right)$. Furthermore, continue like this until arriving at the first $k$ such that $T_n \left( Y, X_{j_{k+1} } | X_{j_1},...., X_{j_k} \right) \leq 0$, and then declare the chosen subset to be $\hat{S} := \big\{ j_1,..., j_k \big\}$. If there is no such $k$, define $\hat{S}$ to be the whole set of variables. Furthermore, this may also happen that $T_n \big( Y, X_{j_1} \big) \leq 0$.
In that case, we declare $\hat{S}$ to be the empty set. In the next section, we prove the consistency of FSCE under a set of assumptions on the law of $\left( Y, \boldsymbol{X} \right)$. Furthermore, we focus on proposing on a well-defined stopping rule. Our stopping rule seems to work well in practise, and we are able to prove consistency of variable selection for this rule.
Suppose that $\boldsymbol{X}$ is a normal random vector with zero mean and arbitrary covariance structure,
Then, for any non-empty $S \subset \left\{ 1,..., p \right\}$ and any $j \in \left\{ 1,..., p \right\} \backslash S$. Notice that if $S$ is a sufficient set of predictors, then $\rho( S, j ) = 0$ for any $j \not\in S$. To the best of our knowledge the proposed algorithm for higher-order conditional risk measures is novel and is based on distribution-free inference.
See for reference the paper: chao2018multivariate. In the particular paper the authors introduce the following estimation procedure. Let $\big\{ \left( \boldsymbol{X}_i, Y_{i1},..., Y_{im} \right) \big\}_{ 1 \leq i \leq n }$ where $Y_{ij}$ represents the value observed from the response $j$ at the time point $i$, and $\left\{ \boldsymbol{X}_i \right\}_{i=1}^n$ are the covariates. Furthermore, we assume that the samples are i.i.d over $i$. For $\uptau \in (0,1)$, the conditional expectile $e_j \left( \uptau | \boldsymbol{X}_i \right)$ of $Y_{ij}$ given $\boldsymbol{X}_i$ is defined as below
where
The proposed feature selection by node exclusion algorithm is considered as a dimension reduction methodology for the graph. Below we present some related algorithms from the economic theory and econometrics literature.
We propose a novel feature ordering algorithm for a multivariate regression with a fixed number of covariates, using a backward stepwise algorithm which is applied to our network driven tail risk matrix. In other words, given a high dimensional response vector $\boldsymbol{Y} = ( Y_j )_{1 \leq j \leq p}$ and a common set of predictors $\boldsymbol{X} = ( X_j )_{1 \leq j \leq k}$ then, the high dimensional response vector is reduced to a lower dimensional space based on the same number of predictors. In other words, given a high dimensional response vector and a set of covariates, we conjecture that the dimension of the response vector can be reduced without having to extend the set of predictors in order to better explain variation in the high-dimensional response vector.
The main idea of the proposed algorithm is a way of selecting features or reducing the number of nodes in a graph, since the algorithm chooses the most central node and then removes it from the graph iteratively using a stopping rule. In other words, the proposed procedure can be considered to be a methodology that bridge the gap between a node exclusion statistical methodology and a procedure for feature selection in graphs without employing Lasso regularization techniques.
In this section, we apply our method on the simulated data to evaluate the estimation performance on the factors and loadings, as the number of factor varies.
We set with $n = m = p = 100$. For $i = 1,..., n, j = 1,...,m$, and let $\boldsymbol{X}_i \sim \mathcal{N} \big( \boldsymbol{0}_{p \times 1}, \boldsymbol{\Sigma}_{p \times p} \big)$ with $\boldsymbol{\Sigma}_{ij} = 0.5^{ | j - k | }$ and $\boldsymbol{\epsilon}_i \overset{ i.i.d }{ \sim } \mathcal{N} \big( \boldsymbol{0}_{m \times 1}, \boldsymbol{I}_{ m \times m} \big)$. Then, the response variables are generated by
where $r = \mathsf{rank} \left( \boldsymbol{\Gamma} \right)$.
\color{black}
Let $\Gamma$ be the an $N \times N$ matrix whose $(i,j)$ entry is $\upgamma_{ij}$. Furthermore, let $\displaystyle \Gamma = \sum_{j=1}^N s_j u_j v_j^{\prime}$. Furthermore, notice that since the elements of the risk matrix $\Gamma$ are constructed based from population moments of the pairwise conditional distribution functions. Considering the fact that the entries of the risk matrix $\Gamma$ are estimated by the nodewise quantile predictive regression models. Thus, our matrix is considered to be a matrix with pairwise tail forecasts.
Let $\Gamma$ be the $N \times N$ matrix whose $(i,j)-$th element is $f( \beta_i, \beta_j )$. Then, our data matrix $X$ has the following form
where $\epsilon_{ij}$ are independent errors with zero mean, satisfying the restriction that $| x_{ij} | < 1$ almost surely. For example, X may be the adjacency matrix of a random graph where the probability of an edge existing between vertices $i$ and $j$ is $f( \beta_i, \beta_j )$.
Our contributions to the literature are twofold: (i) we propose a novel regression-based risk matrix for modelling tail dependence based on graph structures; and (ii) we propose a statistical procedure for feature selection in graphical models. Furthermore, we focus on the interpretation of the risk matrix for optimal portfolio allocation problems as well as a mechanism to introduce our novel variable selection algorithm. Our study discusses important aspects related to the robust estimation of large tail forecast covariance-type matrices as well as the implementation of feature ordering procedures using the proposed graph-based matrix.
Obviously, the drawback of our proposition is the related challenges in evaluating the predictive accuracy of the tail risk matrix since the elements of the matrix consist of risk measures such as the Value-at-Risk and Conditional-Value-at-Risk which are tail estimates of the underline distributions. Secondly, the formal study of our tail risk matrix as a suitable representation for optimal portfolio choice problems is crucial since this is a novel aspect not previously proposed in the literature. Our motivation in this paper although is from the optimal portfolio allocation perspective in financial networks, we focus on the aspect of the induced centrality measure for the proposed novel risk matrix.
Some future research worth mentioning include the formal study of the asymptotic properties of the proposed risk matrix. Although we assume that we employ stationary time series observations our framework can be extended to accommodate features of non-stationary time series. We leave the particular aspect for future research. A second aspect of interest are the time series properties of regressors when estimating the risk measures of VaR and CoVaR. Although the risk management literature usually operates under the assumption of stationarity, an interesting aspect for future research is the effect of non-stationarity when estimating our novel risk matrix. We leave the particular aspect for future research. Although in this paper, we do not consider a formal statistical methodology for edge exclusion this is certainly an aspect which worth future research in subsequent studies. Furthermore, our proposed framework can be employed as a methodology for covariate screening in high dimensional environments. We leave the particular aspect for future research.
\paragraph{Acknowledgements}