EconBase
← Back to paper

Statistical Estimation for Covariance Structures with Tail Estimates using Nodewise Quantile Predictive Regression Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

83,999 characters · 20 sections · 72 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Statistical Estimation for Covariance Structures with Tail Estimates using Nodewise Quantile Predictive Regression Models

\pagenumbering{roman}

I wish to thank my main PhD supervisor Prof. Jose Olmo for guidance and continuous encouragement throughout the PhD programme as well as Prof. Peter W. Smith for helpful discussion and for stimulating the further development of the current study in the direction of node exclusion and feature selection. In addition, I wish to thank Zudi Lu, Jean-Yves Pitarakis, Tassos Magdalinos, Hector Calvo Pardo, Zacharias Maniadis, Ruben Sanchez Garcia and Tullio Mancini as well as Departmental Seminar Series speakers at the University of Southampton: Luis E. Candelaria, Juan Carlos Escanciano, Marcelo C. Medeiros, Sebastian Engelke, Markus Pelger and Loriano Mancini for helpful conversations. \\

Lecturer in Economics, University of Exeter Business School, Exeter EX4 4PU, United Kingdom. Email: \textcolor{blue}{[email removed]} } } }

abstractThis paper considers the specification of covariance structures with tail estimates. We focus on two aspects: (i) the estimation of the $\mathsf{VaR-\Delta CoVaR}$ risk matrix in the case of larger number of time series observations than assets in a portfolio using quantile predictive regression models without assuming the presence of nonstationary regressors and; (ii) the construction of a novel variable selection algorithm, so-called Feature Ordering by Centrality Exclusion (FOCE), which is based on an assumption-lean regression framework, has no tuning parameters and is proved to be consistent under general sparsity assumptions. We illustrate the usefulness of our proposed methodology with numerical studies of real and simulated datasets when modelling systemic risk in a network. \\ Keywords: feature ordering, centrality measures, covariance matrix, quantile regression. AMS 2000 Classification: 62F07, 62H05, 62H12, 62H20.

\setcounter{page}{1} \pagenumbering{arabic}

Introduction.

Network analysis has seen in recent decades a growing attention in the statistics, econometrics and quantitative finance literature; with a common aspect of consideration the role of node centrality in model estimation and portfolio optimization problems. Specifically, in the statistics literature related applications include graphical models and node exclusion tests (see, roverato1998isserlis and salgueiro2005power), while in the econometrics literature recent applications investigate the modelling of interconnectedness and financial spillover effects based on graph structures that capture interactions among economic agents (see, hong2009granger, billio2012econometric, diebold2012better, hardle2016tenet, blasques2016spillover, barigozzi2017network, barunik2018measuring). In the finance and risk management literature, a relevant aspect of interest is the optimal portfolio allocation problem using financial networks\footnote{A specific research question of these studies is the relation between stock centrality and portfolio risk. peralta2016network examine the minimum variance and mean-variance asset allocation problems by connecting the optimal portfolio weights of each investment strategy to the eigenvector centrality. The authors use the correlation decomposition of the covariance matrix to represent the network topology. Thus, the variance-covariance matrix is decomposed into a quadratic form with the correlation matrix as a quadratic matrix and a diagonal matrix with the variances of the assets. By defining a correlation matrix as a matrix which has $1'$s on its main diagonal and the pairwise correlations on its off diagonal elements then this matrix is considered as an adjacency matrix.} (see, peralta2016network, olmo2021optimal, katsouris2021optimal and lin2022portfolio). All aforementioned studies consider suitable statistical dependence measures that are robust to deformations of the network topology, a commonly occurring phenomenon during a financial crisis. Our main research objective is to investigate methodologies for modelling extreme events in graphs combining both tail risk and the induced centrality measures.

Our first contribution to the literature is the construction of a regression-based covariance-type matrix for modelling tail dependency in graphs using pairwise quantile regression models\footnote{The econometric specification for obtaining the risk measures of VaR and CoVaR are presented in the studies of adrian2016covar and hardle2016tenet (see also, white2015var).}; which to the best our knowledge is a novel estimation approach. Specifically, the proposed risk matrix captures higher-order dependency induced by the tail behaviour of the underline distributions of a fixed graph\footnote{A graph representation of a financial network with the proposed risk matrix is proposed by katsouris2021optimal who show that this covariance-type induces a positive-definite quadratic form which permits to consider optimal portfolio allocation problems under tail events. In this paper, the focus is on the proposed feature selection algorithm.}, under the assumption of time series stationarity\footnote{In this paper, we consider that the pairwise quantile predictive regression models are driven only by the predictive regression equation without using a nonstationary specification for regressors. An extension to the case in which the proposed tail-risk matrix depends nonlinearly on the degree of persistence is currently work in progress by the author.}. Although we do not directly model the multivariate extreme dependence as in the study of engelke2020graphical, the constructed large covariance-type matrix captures a general form of network interconnectedness through pairwise graph causality. In particular, the conditional tail (extremal) dependency across the nodes of the graph is measured in relation to the risk measures of VaR and CoVaR obtained as tail estimates (see, Section (ref)).

Our second contribution to the literature, is a novel feature selection algorithm based on the proposed risk matrix, that we call Feature Ordering by Centrality Exclusion (FOCE). This node selection algorithm provides an ordering mechanism using the parametrically specified covarariance type matrix that corresponds to a fixed quantile level, denoted with $\boldsymbol{\Gamma}_n ( \uptau )$ for some $\uptau \in (0,1)$, which implies that the covariance structure is approximated with tail estimates obtained from pairwise quantile (predictive) regressions fitted on a sample size of $n$ observations. Recently, caner2022sharpe present a framework for nodewise regression when the residuals from a fitted factor model are used and a feasible residual-based nodewise regression is proposed to estimate the precision matrix of errors.

In this paper we focus on the introduction of the novel risk matrix as well as in the introduction of the novel graph topology measure. In particular, the proposed risk matrix captures tail dependence and causality in graphs and thus the proposed algorithm of feature ordering is based on this conditional tail dependence measure providing this way a methodology to order characteristic statistics of the graph topology such as centrality measures based on the tail risk matrix. Specifically, the proposed algorithm has an oracle property in the sense that if we plug the true population parameters then our algorithm gives the true feature selection by centrality exclusion of the population variables. Using the proposed algorithm and risk matrix, we aim to model multivariate tail behaviour by considering the tail behaviour of each pair of nodes in the graph. More importantly, it appears to have good performance in simulated and real data sets. Therefore, we focus on the development of the FSCE as well as the proofs of its consistency. The proposed algorithm is not possible to be implemented in low-dimensional models due to convergence and consistency issues based on the proposed stopping rule of the algorithm. We aim to investigate whether our procedure can be estimated with respect to convergence rate $\mathcal{O} \left( n \ \mathsf{log} n \right)$ where $n$ is the sample size.

Usually in the literature the main concern is to show that for $p >> n$ the Sharpe Ratio estimators are consistent in global minimum-variance and mean-variance portfolios. However, in this paper our main concern is to demonstrate the use of the proposed tail-risk matrix in portfolio optimization problems. Specifically, covariance matrices estimated using high dimensional statistical models are used in asset allocation settings as a precision matrix (inverse of the covariance matrix). On the other hand, in high dimensional settings a sparsity condition on the precision matrix of the errors from a statistical model (such as a factor model, a nodewise quantile predictive regression model etc.), can be justified in empirical finance studies. Furthermore, the statistical interpretation of of zero entries on a precision matrix can be attributed into conditional independence. Then, our proposed methodology focuses on estimating a covariance-type matrix using the tail forecasts rather than based on the estimated residuals from a factor model as in the study of caner2022sharpe.

Illustrative Examples

exampleConsider the conventional covariance structure that corresponds to the errors of the factor model (see, fan2015risks, caner2022sharpe among others) such that \begin{align} \hat{\sigma}_{ij}^o = \frac{1}{T} \sum_{t=1}^T u_{it} u_{jt} \ \ \ and \ \ \ \hat{\omega}_{ij}^o = \frac{ \hat{\sigma}_{ij}^o }{ \big( \hat{\sigma}_{ii}^o \hat{\sigma}_{jj}^o \big)^{1/2} } \end{align} where the $\sigma^o$ notation stands for oracle, which indicates that these estimators are not feasible because the true model errors are required. Clearly, the covariance structure given by (ref), implies that all related statistical concepts such as deriving probability bounds for the estimation error, deriving consistent estimators of portfolio performance measures (such as sharpe ratios) in a global minimum-variance or mean-variance portfolio, will depend on the unknown error of the underline factor model. On the other hand, the novelty of our proposed framework is that the covariance matrix under consideration corresponds to the true values of the tail risk measures, $\mathsf{VaR}-\mathsf{CoVaR}$, which are estimated as forecasts\footnote{Notice that for the purpose of this paper, we assume that these risk measures have good elicitability properties. } from nodewise quantile predictive regression models rather than their errors as it is the common practice in the literature (see, callot2021nodewise and caner2022sharpe).

A different stream of literature considers shrinkage methods such as the framework proposed by ledoit2012nonlinear, who consider a nonlinear shrinkage estimator in which small eigenvalues of the sample covariance matrix are increased and large eigenvalues are decreased. Although methodologies for shrinkaging the eigenvalues of covariance matrices which are used for asset allocation purposes ensure that the optimization is not ill-conditioned, this aspect is of independent interest since our main contribution for this paper is to study the main properties of this novel tail-risk matrix. A conjecture is that similar to the case when the covariance matrix is estimated from the errors of factor models, a larger number of included factors will result on a slower convergence rate of the estimation error. Similarly, when the tail-based covariance matrix is estimated with nodewise quantile predictive regressions, an increased number of regressors (either stationary or nonstationary) will result to a slower convergence of the estimation error. Notice that with estimation error here we mean the error in estimating the true structure of the covariance matrix. Moreover, the conventional approach in the literature is to separate between the number of units (e.g., $p$ nodes in the network) versus the number of time series observations. In our case we have another relevant parameter which is the number of regressors included in each quantile predictive regression model (similar to the case of factors when considering factor models). In this paper, we focus in the case in which the number of nodes in the network is not necessarily much larger than the number of time series observations, that is, $p < n$ (or $N < T$ in an equivalent notation).

In the conventional factor structure model case, the following theorem holds for the precision matrix.

theorem[caner2022sharpe] Under the main assumptions it holds that \begin{align} \underset{ 1 \leq j \leq p }{ \mathsf{max} } \ \left\lVert \widehat{\boldsymbol{\Gamma}}_j - \boldsymbol{\Gamma}_j \right\rVert_1 \equiv \mathcal{O}_p \big( \bar{m} \ell_n \big) = o_p(1). \end{align} and secondly, \begin{align} \left\lVert \widehat{\boldsymbol{\mu}} - \boldsymbol{\mu} \right\rVert_{\infty} = \mathcal{O}_p \left( \mathsf{max} \left\{K \sqrt{ \frac{ ln(n) }{n} }, \sqrt{ \frac{ ln(p) }{n} } \right\} \right) = o_p(1). \end{align}
remarkUsing a direct estimation procedure (i.e., without a pseudo-inverse approximation) for the precision matrix provides a faster convergence to the true population values of the portfolio performance metrics. However, regardless of the direct or indirect estimation of the precision matrix, a larger number of $p$ (nodes in the network) affects the error by a logarithmic factor. On the other hand, the estimation error also increases with the non-sparsity of the precision matrix, especially as the dimensionality of the matrix increases. Thus, in the case of a non-sparse precision matrix, we can only get consistency when $p << n$, that is, the number of nodes in the network is much smaller than the time series observations on which the statistical model is fitted on. A key assumption in the conventional framework is that, $\big( \boldsymbol{Y}_t, \boldsymbol{X}_t \big)_{t=1}^n$ are both stationary and ergodic vectors.
lemmaThe following probability bound holds \begin{align} \mathbb{P} \left( \underset{ 1 \leq j \leq p }{ \mathsf{max} } \ \underset{ 1 \leq \ell \leq p }{ \mathsf{max} } \left| \frac{1}{n} \sum_{t=1}^n u_{\ell, t} u_{j,t} - \mathbb{E} \big[ u_{\ell,t} u_{j,t} \big] \right| > C \sqrt{ \mathsf{ln} (p) / n } \right) = \mathcal{O} \left( \frac{1}{p^2} \right) \end{align} when the covariance structure is based on a prespecified factor model.
exampleConsider a given factor structure for stock returns such that \begin{align} \boldsymbol{Y} = \boldsymbol{B} \boldsymbol{X} + \boldsymbol{U} \end{align} Then, the covariance matrix of $\boldsymbol{Y}$ is expressed as below \begin{align} \boldsymbol{\Sigma}_y = \boldsymbol{B} \mathsf{Cov} \left( \boldsymbol{f}_t \right) \boldsymbol{B}^{\prime} + \boldsymbol{\Sigma}_n. \end{align} while the precision matrix, defined by $\boldsymbol{\Gamma}_y := \boldsymbol{\Sigma}_y^{-1}$, such that \begin{align} \boldsymbol{\Gamma}_y := \boldsymbol{\boldsymbol{\Omega}}_n - \boldsymbol{\boldsymbol{\Omega}}_n^{-1} \boldsymbol{B} \left[ \mathsf{Cov} \left( \boldsymbol{f}_t \right)^{-1} + \boldsymbol{B}^{\prime} \right] \end{align} where $\boldsymbol{\Omega}_n = \boldsymbol{\boldsymbol{\Sigma}}_n^{-1} \equiv \left\{ \mathbb{E} \left[ \boldsymbol{u}_t \boldsymbol{u}_t^{\prime} \right] \right\}^{-1}$.

Notice that Example (ref) illustrates the decomposition of the precision matrix when the vector of stock returns have a factor structure representation. On the other hand, in this article we assume that the there exists an underline network structure that describes the stochastic behaviour and comovements of these nodes in the graph. Some key points are summarized below:

itemize• Our proposed risk matrix is equivalent to the covariance matrix $\boldsymbol{\Sigma}_y$, after fitting the nodewise quantile predictive regression models in the system, while $\boldsymbol{\Omega} := \left\{ \mathbb{E} \left[ \boldsymbol{u}_t \boldsymbol{u}_t^{\prime} \right] \right\}^{-1}$, i.e., the precision matrix corresponds to the inverse of the covariance matrix of errors. However, the advantage of our approach is that we focus directly on the inversion of this tail risk matrix by assuming that $\boldsymbol{\Gamma}_{j_1 j_2 }^{-1} = \frac{ \omega_{j_1 j_2} }{ \omega_{j_1 j_1} \omega_{j_2 j_2 } }$. A novelty in the construction of this risk matrix is that the estimation of these risk measures is conditioned in the presence of covariates (i.e., stationary or nonstationary regressors) which is more informative than based solely on conditional distribution functionals. • In terms of the relevant testing methodology this involves finding statistical significant blocks based on the tail interdependence can be interpreted as revealing blocks which are significantly important using a form of pooling these extreme events together based on the risk measures of $\mathsf{VaR}$ and $\mathsf{CoVaR}$. Although, we do not study this here, the role of the degree of persistence of these covariates included in the models, is an interesting future research question.

Related Literature.

Extreme events (such as financial crises) are commonly employed for causal discovery since the relationships between variables may manifest themselves especially in the largest observations. In particular, a modelling methodology based on extreme value theory is presented in the study of gnecco2021causal\footnote{In the particular framework the authors use $X_k := \sum_{ k \in pa (j, G) } \beta_{jk} X_k + \epsilon_j, \ j \in V$ under the assumption that the noise sequence $\epsilon_j, j \in V$, are heavy-tailed with extreme value index $\gamma > 0$ such that $P ( \epsilon_j > x ) \sim \ \ell(x)x^{-2/ \gamma}, \ \ x \to \infty$.}. A discussion about tail dependence in graphs is provided in the paper of engelke2020graphical (see also discussion in dombry2016asymptotic). On the other hand, our regression-based risk matrix takes into utilized tail estimates rather than relying on fitting a parametric model to heavy-tailed data, which implies that the corresponding matrix estimator is robust to the presence of potential outliers in financial and macroeconomic data.

The estimation of covariance matrices is widely used in various applications such as optimal portfolio choice problems and other statistical problems such as multivariate analysis, principal components analysis and graphical modeling among others. The commonly used assumption for the estimation of large covariance matrices is the assumption of multivariate normality. For instance, brown1975techniques and browne1984asymptotically proposed statistical methodologies for the analysis of covariance matrices such as asymptotically distribution-free methods. In particular, the majority of estimators of parameters in structural models for covariance matrices are members of the class of generalized least squares (GLS) estimators. Specifically, this type of estimators include the maximum Wishart likelihood estimators (MWL) in the sense that a generalized least squares discrepancy function can be constructed which will in general attain its minimum at the point defined by the MWL estimator (see, browne1984asymptotically). A recent study who consider the MLE estimation of covariance type matrices when both rows and columns can be correlated is presented by drton2021existence.

The statistical literature on methodologies for the robust estimation of large covariance matrices has been thrived the last decades. In particular, assumptions such as the Wishart distribution for the unknown covariance matrix allows to implement asymptotically distribution-free methods for the robust estimation. However, features such as heavy-tailed time series, temporal dependence or even the presence of structural instability in time-series affects the robustness of the estimation procedure. For instance, dendramis2021estimation accommodate the time variation, dependence and heavy-tailedness of distributions with implementation of a time-varying covariance matrix estimation procedure that includes thresholding. Therefore, the novelty of our estimation procedure is the use of quantile predictive regression models for capturing these features, which implies a regression-based tail risk matrix. Moreover, the particular proposition allows to take into consideration the time series properties via the estimation methodology rather than implementing ex-ante corrections to the estimation procedure which implies implementing bias corrections for the possible effect of the particular features to the robustness and degree of accuracy of the large covariance matrix.

A different stream of literature investigates the use of centrality measures (see, a discussion in aamari2021graph) for the development of graph-based modelling methodologies. Our approach is different than portfolio sorting procedures such as the studies of mcgee2020optimal and ledoit2018efficient who propose methodologies for sorting portfolios based on variable characteristics and nonlinear time series models respectively, however our motivation is driven by the construction of an endogenous to the system sorting mechanism. Furthermore, the aspect of bridging centrality and extremity is examined in the study of einmahl2015bridging. In terms of rank centrality measures a related approach is presented in the study of negahban2017rank, while a formal asymptotic theory framework based on the random matrix theory is given by Li2021central. Furthermore, negahban2017rank propose a rank centrality measure which is an iterative rank aggregation algorithm for discovering scores for objects (or items) from pairwise comparisons in a statistical sense.

A third line of literature which is related to our study is the use of the concept of conditional independence. Specifically, a novel conditional independence measure is proposed by azadkia2021simple which corresponds to the case of conditional mean preserving measures. Lastly, although we mainly focus with the aspects of statistical estimation, relevant inference procedures include testing for misspecification of the underline covariance structure (see, guo2021specification and chang2022testing). Further limit theorems on covariance structures are given in the recent study of liu2021robust while chang2021central establish central limit theorems for high-dimensional dependent data.

In summary, our proposed modelling approach focuses on two key aspects: (i) the construction of the risk matrix based on tail dependence measures and (ii) the implementation of the feature selection algorithm. Although our feature selection algorithm is applied to the tail risk matrix, we conjecture that this estimation procedure can be also used in alternative structures of covariance matrices.

The regression-based tail risk matrix.

Specifically, the construction of the proposed risk matrix captures causality\footnote{A discussion related to statistical causality can be found in schenone2018causality.} for extreme values based on nodewise quantile predictive regression models which capture the quantile behaviour of the underline distributions. Our novel risk matrix is constructed using nodewise regression models in the sense that the elements of the risk matrix are estimated using pairwise quantile regression models.

Modelling Environment.

Given the singular value decomposition $\boldsymbol{\Gamma} = \boldsymbol{U} \boldsymbol{D} \boldsymbol{V}^{\prime}$. Moreover, denote with $\big\{ \boldsymbol{\Gamma}_t \big\}_{t=0}^T$ be the sequence obtained by the iteration of Algorithm 1.

remark(Initialization and the Stopping Rule) We suggest to initialize the algorithm with $\boldsymbol{\Gamma}_0$ in Algorithm 1, but because the optimization problem is convex, this can be replaced by any matrix. Furthermore, the algorithm converges faster when it is initialized with a matrix that is close to the minimizer. Furthermore, we suggest to stop the algorithm at iteration $T$ satisfying the condition $\big| F( \boldsymbol{\Gamma}_{T+1} ) - F( \boldsymbol{\Gamma}_{T} ) \big| \leq \epsilon$, for some small $\epsilon > 0$.

Furthermore, in this section we derive some useful oracle inequalities such as provide suitable error bounds for the difference between the sequence of matrices $\boldsymbol{\Gamma}_n$ generated by Algorithm 1 and the true matrix $\boldsymbol{\Gamma}$. These results are determined by the strong convexity of the optimization function $\rho_{\uptau}$.

assumptionWe impose the following conditions. \begin{itemize} • There exists $C > 0$ such that for $u_{ij} = Y_{ij} - \boldsymbol{X}_{i}^{\prime} \boldsymbol{\Gamma}_{.j}$ such that it holds that \begin{align} \mathbb{P} \big( | u_{ij} | > s \big) \leq \mathsf{exp} \left( 1 - \frac{s^2}{C^2} \right), \forall \ s \geq 0 \end{align} with sub-gaussian norm given by \begin{align} \left\lVert u_{ij} \right\rVert_{ \psi_2 } \overset{ \mathsf{def} }{ = } \underset{ p \geq 1 }{ \mathsf{sup} } \ p^{-1/2} \big( \mathbb{E} | u_{ij} |^p \big)^{1 / p} \ \ and \ \ K_u \overset{ \mathsf{def} }{ = } \underset{ 1 \leq j \leq m }{ \mathsf{max} } \ \left\lVert u_{ij} \right\rVert_{ \psi_2 }. \end{align} • Conditional on $\boldsymbol{X}_i$, $Y_{ij}$ are independent over $j$. \end{itemize}

Notice that Assumption A3 is required in order to obtain the bounds on the tail probabilities of the estimation error. Furthermore, we consider the non-asymptotic bound for $\left\lVert \boldsymbol{\Gamma}_t - \boldsymbol{\Gamma} \right\rVert_F$ in the general situation that the number of factors can be increasing with $n$.

Estimation Methodology.

Our risk matrix is employed to model tail causality for the nodewise regression models. In particular, these nodewise regressions correspond to quantile predictive regression models as defined by adrian2016covar which are used to obtain estimates for the risk measures of VaR and CoVaR. Consider that we have the vector $\underline{Y} = ( Y_{1,t},..., Y_{p,t} )$ such that $i \in \left\{ 1,..., p \right\}$ and $j \in \left\{ 1,..., p \right\}$. Consider the single index model for estimating the $\mathsf{CoVaR}$ and $\mathsf{VaR}$ risk measures

align[align omitted — 141 chars of source]

The given specification is motivated from the literature of financial connectedness and tail risk and in particular is based on the seminal works of adrian2016covar and hardle2016tenet who examine statistical estimation methodologies and their properties. We firstly consider the following covariance-type structure

align[align omitted — 241 chars of source]

where $\mathbf{\widetilde{\Sigma}}_{ij} \in \mathbb{R}^{ N \times N }$ represents the financial connectedness matrix and provides a robust representation of idiosyncratic and systemic risk within the financial network. Our proposed risk matrix belongs to the family of dynamic high dimensional covariance matrices and captures network tail risk dependence, since is based on tail risk measures. We estimate the time-varying VaRs and CoVaRs conditional on a vector of lagged state variables $\boldsymbol{X}_{t-1}$. Both VaR and CoVaR risk measures represent quantiles of the distribution of returns under different conditioning sets; hence, to obtain one-period ahead predicted values for these risk quantities we employ the quantile regression specifications proposed by the seminal study of Koenker1978regression (see, also KoenkerXiao02). The construction of one-period ahead forecasts for the dynamic portfolio allocation, based on a graph representation, consists one of the main contributions of the empirical application of the paper.

Consider the conditional quantile function estimated via $Q_{y_i} \left( \tau | x_i \right) = F_{y_i}^{-1} \left( \tau | x_i \right)$. Then, the optimization function to obtain the model estimates is expressed as below

align[align omitted — 174 chars of source]

where $\tau \in (0,1)$ is a specific quantile level, and $\rho_{\tau}( u ) = u \left( \tau - \mathbf{1}_{ \left\{ u < 0 \right\} } \right)$ is the check function (see, newey1987asymmetric). Suppose $1 \leq t \leq T$, then based on the estimation method given by (ref) the VaR and CoVaR are generated by the following econometric specifications

align[align omitted — 248 chars of source]

where $\{ R_{(i),t} \}$ for $i \in \left\{ 1,...,N \right\}$ is the vector of portfolio returns at time $t$, and $\boldsymbol{X}_{t-1}$ is a vector of exogenous regressors containing macroeconomic characteristics common across assets and denote with $\Theta = \{ \nu_{(i)}, \boldsymbol{\xi}_{(i)}, \nu_{(j)}, \boldsymbol{ \xi }_{(j)}, \beta_{(j)} \}$ to be a compact parameter space.

Let $Q_{\tau} \left( \cdot\ | \ \mathcal{F}_{t-1} \right)$ denote the quantile operator for $\tau \in (0,1)$ conditional on an information set $\mathcal{F}_{t-1}$. Then, the error terms satisfy $Q_{\tau} \left( u_{(i),t} \ | \ \boldsymbol{X}_{t-1} \right)=0$ for the quantile regression model (ref), and $Q_{\tau} \left( v_{(j),t} \ | \ R_{(j),t}, \boldsymbol{X}_{t-1} \right) = 0$ for the quantile regression model given by expression (ref). Moreover, denote with $\boldsymbol{\theta}_{(1)} = \left( \nu_{(i)}, \boldsymbol{\xi}_{(i)}^{\prime} \right)$ and $\widetilde{\boldsymbol{X} }_{t-1}^{(1)} = \left( 1, \boldsymbol{X}_{t-1}^{\prime} \right)^{\prime}$ for model (ref), and with $\boldsymbol{\theta}_{(2)} = \left( \nu_{(j)}, \boldsymbol{\xi}_{(j)}^{\prime}, \beta_{(j)} \right)$ and $\widetilde{\boldsymbol{X}}_{t-1}^{(2)} = \left( 1, \boldsymbol{X}_{t-1}^{\prime}, R_{(i),t} \right)^{\prime}$ for model (ref).

Then, the model estimate

align[align omitted — 282 chars of source]

Similarly, the quantile estimator of (ref) is obtained via the optimization function below

align[align omitted — 283 chars of source]

The state variables, that is, the macroeconomic and financial variables, included in the econometric model, capture time variation in the conditional moments of asset returns. Therefore, the $\mathsf{VaR}$ and $\mathsf{CoVaR}$ of each financial institution are estimated as vectors from a time-varying distribution.

This approach allows to explain the optimal portfolio allocation in terms of changes in the systemic risk and financial connectedness in the network. To implement this, we use a rolling window, within which the quantile regressions given by (ref) and (ref) are fitted and the $\boldsymbol{\Gamma}$ risk matrix is constructed based on one-period ahead forecasts. Specifically, the one-period ahead $\mathsf{VaR}_{t+1}$ for a coverage probability $\tau \in (0,1)$ is obtained by regressing $R_{(i),t}$ on $\boldsymbol{X}_{t-1}$ as given by the first step regression ((ref)) which gives the parameter estimates $\widehat{\nu}_{(i)}$ and $\widehat{ \boldsymbol{\xi} }_{(i)}$. The forecasted one-period ahead VaR is computed as below

align[align omitted — 159 chars of source]

Similarly, the one-period ahead forecast of $\mathsf{CoVaR}_{t+1}$ is obtained by regressing $R_{(j),t}$ on $\boldsymbol{X}_{t-1}$ and $R_{(i),t}$ as shown in the second step regression ((ref)) using the quantile estimation as defined by (ref). To do this, we collect the parameter estimates $\widehat{\nu}_{(j)}, \widehat{\boldsymbol{\xi} }_{(j)}, \widehat{\beta}_{(j)}$ from the second regression, and construct the forecasted one-period ahead $\mathsf{CoVaR}_{t+1}$ measure using the following expression:

align[align omitted — 214 chars of source]

where $\widehat{\mathsf{VaR}}_{i,t+1}( \tau )$ replaces the actual $\mathsf{VaR}_{i,t+1} ( \tau )$ measure. Then, the $\mathsf{\Delta \text{CoVaR}}$ measure is computed as the difference of the two risk measures

equation[equation omitted — 129 chars of source]

An important aspect of our proposed estimation methodology is the assumption of a graph representation with nodes being the stock returns and edges being pairwise tail forecasts, which can be thought as the level of simultaneous Granger causality due to the existence of tail and graph dependence. Therefore, these tail forecasts are estimated using the aforementioned predictive regression models based on the bipartite graph structure, which implies an estimation implementation for each pair $(i,j)$ of nodes for the graph $\mathcal{G} = (\mathcal{V}, \mathcal{E})$.

Main Results

Asymptotic theory for the VaR-$\Delta$CoVaR risk matrix

In this section, we derive results for establishing the asymptotic theory for the estimator of the proposed risk matrix that correspond to the population risk measures. Define with $\mathtt{Upper} ( \mathbf{\Gamma} ) \in \mathbb{R}^{ N \times N}$ the upper triangular part of the matrix, excluding the diagonal entries and with $\mathtt{Lower} ( \mathbf{\Gamma} ) \in \mathbb{R}^{ N \times N}$ the lower triangular part of the matrix, excluding the diagonal entries. Then, the risk matrix $\mathbf{\Gamma} \in \mathbb{R}^{ N \times N}$ is given by the following expression:

small\begin{align*} \mathbf{\Gamma} := \begin{bmatrix} \mathsf{VaR}^{+}_{1} & (\mathsf{VaR}^{+}_{1} \mathsf{VaR}^{+}_{2})^{1/2} \Delta \mathsf{CoVaR}_{(1,2)} & \dots (\mathsf{VaR}^{+}_{1} \mathsf{VaR}^{+}_{N})^{1/2} \Delta \mathsf{CoVaR}_{(1,N)} \\ \vdots & \vdots & \vdots \\ \vdots & \ddots & \vdots \\ (\mathsf{VaR}^{+}_{N} \mathsf{VaR}^{+}_{1})^{1/2} \Delta \mathsf{CoVaR}_{(1,N)} & \dots & \mathsf{VaR}^{+}_{N} \\ \end{bmatrix} \end{align*}

where the off-diagonal elements consist of the term:

align[align omitted — 90 chars of source]

Since the main diagonal of the risk matrix includes only the VaR tail measure which correspond to the underline node specific distribution, then we suppose that is generated by a common cumulative distribution function $F(.)$. An important condition for our risk matrix to be utilized within the optimal portfolio problem is to ensure that it converges in probability to a positive definite matrix for large sample size, that is, in our case, as the number of nodes in the network grows to infinity, i.e., $N \to \infty$.

Then, we can observe that the proposed covariance structure has a similar decomposition to the covariance matrix as below:

small\begin{align*} \mathbf{\Gamma} \equiv \begin{pmatrix} \sqrt{ \mathsf{VaR}^{+}_{1} } & & & \\ & \sqrt{ \mathsf{VaR}^{+}_{2} } \\ & & \ddots \\ & & & \sqrt{ \mathsf{VaR}^{+}_{N} } \end{pmatrix} \begin{pmatrix} 1 & \Delta \mathsf{CoVaR}_{1,2} & & \\ \Delta \mathsf{CoVaR}_{2,1} & 1 \\ & & \ddots \\ & & & 1 \end{pmatrix} \begin{pmatrix} \sqrt{ \mathsf{VaR}^{+}_{1} } & & & \\ & \sqrt{ \mathsf{VaR}^{+}_{2} } \\ & & \ddots \\ & & & \sqrt{ \mathsf{VaR}^{+}_{N} } \end{pmatrix} \end{align*}

Moreover, we are interested in estimating the precision matrix of the $\mathsf{VaR}-\Delta \mathsf{CoVaR}$ risk matrix, $\boldsymbol{\Omega} := \boldsymbol{\Gamma}^{-1} \equiv \displaystyle \frac{ \omega_{j_1, j_2} }{ \omega_{j_1,j_1} \omega_{j_2,j_2} }$ where $\omega_{j_1, j_2}$ correspond to the elements of the precision matrix of the $\mathsf{VaR}-\Delta \mathsf{CoVaR}$ matrix for some $(j_1, j_2) \in \left\{ 1,..., N \right\}$.

The asymptotic behaviour of the risk matrix can be determined by considering the limiting distribution of the expression $ \sqrt{N} \left( \widehat{ \mathbf{\Gamma} } - \mathbf{\Gamma}_0 \right)$. Denote with

align[align omitted — 494 chars of source]

The term $\mathcal{A}_3$ converges to a Gaussian random variable, with a non-stochastic covariance matrix under the assumption of stationary and ergodic regressors in the quantile predicitve regression model.

exampleAs an illustration of the required derivations to obtain asymptotic theory results, consider the case in which the graph matrix depends on the unknown error terms of $\boldsymbol{\Gamma} := [u_{ij}]$. We follow similar derivations as in the framework of guo2021specification. Define with $\mathcal{G}_1 := \big\{ i : (i,j) \in \mathcal{G}, \ \text{for some} \ j \leq p \big\}$. For any $(k \times k)$ matrix $\boldsymbol{A}$, denote by $\lambda_i (\boldsymbol{A})$ the $i-$th largest eigenvalue of $\boldsymbol{A}$ for $i = 1,..., k$. Then, some useful results are given below.
lemmaUnder the assumptions of the Theorem, there exists $\zeta_1 > 0$ and $\zeta_2 > 0$, such that \begin{align} \mathbb{P} \left( \underset{ (i,j) \in \mathcal{G} }{ \mathsf{max} } \ T^{1/2} \left| \frac{1}{T} \sum_{t=1}^T \hat{u}_{it} \hat{u}_{jt} - \frac{1}{T} \sum_{t=1}^T u_{it} u_{jt} \right| \geq \zeta_1 \right) < \zeta_2 \end{align} where $\zeta_1 = o \big( ( \mathsf{log}(p) )^{-1/2} \big)$ and $\zeta_1 = o(1)$.
proofSince it holds that \begin{align} \sum_{t=1}^T \big( \hat{u}_{it} \hat{u}_{jt} - u_{it} u_{jt} \big) = \sum_{t=1}^T \bigg\{ \big( \hat{u}_{it} - u_{it} \big) \big( \hat{u}_{jt} - u_{jt} \big) + u_{it} \big( \hat{u}_{jt} - u_{jt} \big) + \big( \hat{u}_{it} - u_{it} \big) u_{jt} \bigg\} \end{align} Therefore, it suffices to show that \begin{align} \underset{ (i,j) \in \mathcal{G} }{ \mathsf{max} } \left| \sum_{t=1}^T \big( \hat{u}_{it} - u_{it} \big) \big( \hat{u}_{jt} - u_{jt} \big) \right| = o_p \left( T^{1/2} \zeta_1 \right) \end{align} In particular, it holds that \begin{align*} \underset{ (i,j) \in \mathcal{G} }{ \mathsf{max} } \left| \sum_{t=1}^T \big( \hat{u}_{it} - u_{it} \big) \big( \hat{u}_{jt} - u_{jt} \big) \right| &\leq \underset{ (i,j) \in \mathcal{G} }{ \mathsf{max} } \sum_{t=1}^T \left| \big( \hat{u}_{it} - u_{it} \big) \big( \hat{u}_{jt} - u_{jt} \big) \right| \\ &\leq \underset{ (i,j) \in \mathcal{G} }{ \mathsf{max} } \left\{ \sum_{t=1}^T \big( \hat{u}_{it} - u_{it} \big)^2 . \sum_{t=1}^T \big( \hat{u}_{jt} - u_{jt} \big)^2 \right\}^{1/2} \\ & = \underset{ i \in \mathcal{G}_1 }{ \mathsf{max} } \left\{ \sum_{t=1}^T \big( \hat{u}_{it} - u_{it} \big)^2 \right\}^{1/2} \underset{ j \in \mathcal{G}_2 }{ \mathsf{max} } \left\{ \sum_{t=1}^T \big( \hat{u}_{jt} - u_{jt} \big)^2 \right\}^{1/2} \\ & = \underset{ i \in \mathcal{G}_1 }{ \mathsf{max} } \left\{ \sum_{t=1}^T \left[ \big( b_i^{\prime} - \hat{b}_i^{\prime} \big) z_t \right]^2 \right\}^{1/2} \underset{ j \in \mathcal{G}_2 }{ \mathsf{max} } \left\{ \sum_{t=1}^T \left[ \big( b_j^{\prime} - \hat{b}_j^{\prime} \big) z_t \right]^2 \right\}^{1/2} \end{align*} Consider the following expression \begin{align*} M_{ij} := \frac{1}{T} \sum_{t=1}^T \left\{ u_{it} u_{jt} - \phi_{ij} ( \theta_0) - \frac{ \partial \phi_{ij} ( \theta) }{ \partial \theta^{\top} } \bigg|_{\theta = \theta_0 } \left[ \sum_{(k, \ell) \in \mathcal{H} } \frac{ \partial \phi_{k \ell} (\theta) }{ \partial \theta } \frac{ \partial \phi_{k \ell} (\theta) }{ \partial \theta^{\top} } \right]^{-1} \times \sum_{(k, \ell) \in \mathcal{H} } \bigg[ u_{kt} u_{\ell t} - \phi_{k \ell } (\theta_0) \bigg] \frac{ \partial \phi_{k \ell} (\theta) }{ \partial \theta } \bigg|_{\theta = \theta_0 } \right\} \end{align*} where $\theta^{*}$ lies between $\hat{\theta}$ and $\theta_0$. Specifically, we will show that \begin{align} \mathbb{P} \left( \underset{ (i,j) \in \mathcal{H} }{ \mathsf{max} } \ T^{1/2} \left| \phi_{ij} \left( \hat{\theta} \right) - \phi_{ij} ( \theta_0 ) - M_{ij}^* \right| \geq \zeta_1 \right) < \zeta_2, \end{align} where \begin{align*} M_{ij}^{*} = \frac{1}{T} \sum_{t=1}^T \left( \frac{ \partial \phi_{ij} (\theta) }{ \partial \theta^{\top} } \bigg|_{\theta = \theta_0 } \left[ \sum_{(k, \ell) \in \mathcal{H} } \frac{ \partial \phi_{k \ell} (\theta) }{ \partial \theta } \frac{ \partial \phi_{k \ell} (\theta) }{ \partial \theta^{\top} } \right]^{-1} \sum_{ ( k,\ell ) \in \mathcal{H} } \left\{ \bigg[ u_{kt} u_{\ell t} - \phi_{k \ell } (\theta_0) \bigg] \frac{ \partial \phi_{k \ell} (\theta) }{ \partial \theta } \bigg|_{\theta = \theta_0 } \right\} \right) \end{align*} We know that \begin{align} \underset{ (i,j) \in \mathcal{H} }{ \mathsf{max} } \left\lVert \frac{1}{2} \left( \hat{\theta} - \theta_0 \right)^{\top} \frac{ \partial^2 \phi_{ij} (\theta) }{ \partial \theta \partial \theta^{\top} } \bigg|_{ \theta = \theta^* } \left( \hat{\theta} - \theta_0 \right) \right\rVert_2 \leq \left\lVert \hat{\theta} - \theta_0 \right\rVert_2^2 = \mathcal{O}_p \left( \frac{ \mathsf{log} (p) }{T} \right). \end{align} because of the definition of $\hat{\theta}$, which implies that \begin{align} \sum_{ (i,j) \in \mathcal{H} } \big( \hat{\sigma}_{ij} - \phi_{ij} (\hat{\theta} ) \big) \frac{ \partial \phi_{ij} (\theta) }{ \partial \theta} \bigg|_{ \theta = \hat{\theta} } = 0_q. \end{align} From Taylor expansion, we have that \begin{align*} 0_q &= \sum_{ (i,j) \in \mathcal{H} } \left[ \left( \hat{\sigma}_{ij} - \phi_{ij} (\hat{\theta} ) \right) \frac{ \partial \phi_{ij} (\theta) }{ \partial \theta} \bigg|_{ \theta = \theta_0 } \right] \\ &+ \left( \sum_{ (i,j) \in \mathcal{H} } \left[ \left( \hat{\sigma}_{ij} - \phi_{ij} ( \theta^* ) \right) \frac{ \partial^2 \phi_{ij} (\theta) }{ \partial \theta \partial \theta^{\top} } \bigg|_{ \theta = \theta^* } \right] - \sum_{ (i,j) \in \mathcal{H} } \left[ \frac{ \partial \phi_{ij} (\theta) }{ \partial \theta} \bigg|_{ \theta = \theta^* } \frac{ \partial \phi_{ij} (\theta) }{ \partial \theta^{\top} } \bigg|_{ \theta = \theta^* } \right] \right) \left( \hat{\theta} - \theta_0 \right) \end{align*} where $\theta^*$ lies between $\hat{\theta}$ and $\theta_0$. Therefore, it holds that \begin{align} \hat{\theta} - \theta_0 = N^{-1} \sum_{ (i,j) \in \mathcal{H} } \left[ \left( \hat{\sigma}_{ij} - \phi_{ij} ( \theta_0 ) \right) \frac{ \partial \phi_{ij} (\theta) }{ \partial \theta} \bigg|_{ \theta = \theta_0 } \right] \end{align}
remarkNotice that it is important to distinguish the source of high-dimensionality. In particular, for $\beta_j, j \in \left\{ 1,..., p \right\}$, we can consider asymptotic expressions in the limit where both $n, p \to \infty$ and focus on the high-dimensional regime with $c^{*} = \mathsf{lim} \ p / n \in (0, \infty]$. Moreover, our estimation method differs from various current methodologies in the literature. For example, in the study of gonzalo2021spurious the authors are concerned with the interactions of persistence and dimensionality in the context of the eigenvalue estimation problem of large covariance matrices arising in cointegration and principal components analysis. In other words, we are not concerned about the presence of spurious relationships since our covariance matrix is constructed based tail estimates and not the error terms of cointegrating systems. On the other hand, there is a clear interplay between the dimensionality of the statistical problem of interest and granger causality.

Optimal Portfolio Optimization.

We begin with the traditional optimal portfolio allocation problem. We extend this framework by proposing a novel risk matrix that captures tail events. We present regularity conditions which ensure the existence of a closed-form solution for the associated quadratic risk function and optimization functions. Our proposed risk matrix induces a well defined quadratic form which permits to be employed for optimal portfolio choice problems\footnote{A rational investor who holds a set of securities aims to minimize the portfolio risk based on the portfolio optimization framework introduced by markowitz1956optimization. The specific linear optimization mechanism has seen growing attention in various fields of financial economics such as portfolio management, risk management, capital investment and asset pricing. Furthermore, the seminal paper of sharpe1964capital introduced the fundamental principles of investment under conditions of uncertainty in financial markets. Furthermore, various studies examined the effects of constraints on the portfolio performance measures such as the expected return and the portfolio risk (see fishburn1977mean, jagannathan2003risk, and references therein).}.

Traditional Minimum-Variance Portfolio Optimization.

The optimal portfolio theory proposed by Markowitz52 considers the minimization of the portfolio variance without any further restriction on the corresponding expected return. More formally, let $\mathbf{R}_t = ( R_{1,t} ,..., R_{N,t} )^{\prime}$ be the $N-$dimensional vector of assets returns at time $t$ and denote with $\text{\boldmath$\mu$} := \mathbb{E} \left( \mathbf{R}_t \right) = \left( \mathbb{E} \left( R_{1,t} \right) ,..., \mathbb{E} \left( R_{N,t} \right) \right)^{\prime}$ the vector of expected returns and $\mathbf{\Sigma} := \text{Cov}\left( \mathbf{R}_t \right) = \mathbb{E} \left( \mathbf{R}_t \mathbf{R}_t^{\prime} \right) - \text{\boldmath$\mu$} \text{\boldmath$\mu$}^{\prime}$ the associated variance-covariance matrix. Let $r_t^{p} = \mathbf{w}_i^{\prime} \mathbf{R}_t$ denote the return on a portfolio of $N$ assets, with $\mathbf{w}_i= \left( w_{1},...,w_{N} \right)^{\prime}$ the vector of portfolio weights representing the financial position of the investor as a weight allocation of assets to the portfolio. The variance of this portfolio is $V \left( r_t^{p} \right) = \mathbf{w}^{\prime} \mathbf{\Sigma} \mathbf{w}$. Then, the Markowitz's minimum variance optimization problem is expressed as below

align[align omitted — 212 chars of source]

The first order conditions to the minimization problem yield the optimal portfolio allocation given by the following vector of weights

align[align omitted — 136 chars of source]

Portfolio Allocation and Asset Centrality.

align[align omitted — 170 chars of source]

The symmetry of $\widehat{ \mathbf{\Sigma} }$ yields the spectral decomposition $\widehat{ \mathbf{\Sigma} } = \mathbf{U} \mathbf{D}_{\sigma} \mathbf{U}^{\prime}$, with $\mathbf{U}$ is an $N \times N$ orthonormal matrix such that $\mathbf{U}^{\prime} = \mathbf{U}^{-1}$ which includes as columns the linearly independent eigenvectors of $\widehat{ \mathbf{\Sigma} }$, and $\mathbf{D}_{\sigma}$ an $N \times N$ diagonal matrix with the corresponding eigenvalues which is defined as $\mathbf{D}_{ \sigma } = diag \left\{ \lambda^{ \sigma }_1,..., \lambda^{ \sigma }_N \right\}$. Moreover, we can rearrange the diagonal eigenvalue matrix as $\mathbf{D}_{\sigma} =\mathbf{U}^{\prime} \widehat{ \mathbf{\Sigma} } \mathbf{U}$.

This allow us to express the eigenvalues of the risk matrix $\widehat{ \mathbf{\Sigma} }$ in terms of the elements of the variance-covariance matrix as below

align[align omitted — 280 chars of source]

where $\left\{ u_{ij} \right\}_{i,j = 1,...,N}$ denotes the eigenvectors of the risk matrix $\widehat{ \mathbf{\Sigma} } \in \mathbb{R}^{N \times N}$.

assumptionThe eigenvectors $\{ u_{ij} \}_{i,j=1,...,N}$ of the risk matrix $\mathbf{\Gamma}$ satisfy the condition \begin{equation} \overset{N}{\underset{i=1}{\sum}} u_{ik}^2 > - \overset{N}{\underset{i=1}{\sum}} \overset{N}{\underset{\underset{j \neq i}{j=1}}{\sum}} u_{ik} u_{jk} \frac{ \widetilde{ \gamma}_{i|j} }{ \gamma_{k}^{0} }, \end{equation} for all $k=1,\ldots,N$.

The condition in Assumption (ref) guarantees that $\widetilde{ \mathbf{\Gamma} }$ is positive definite, and as a result implies that the corresponding quadratic form $\mathbf{w}^{\prime} \widetilde{ \mathbf{\Gamma} } \mathbf{w}$ has certain desirable properties as summarized by the following propositions.

propositionLet $Q( \widetilde{ \mathbf{\Gamma} } ,\mathbf{w}) = \mathbf{w}^{\prime} \widetilde{ \mathbf{\Gamma} } \mathbf{w}$ denote the portfolio risk, where the risk matrix $\widetilde{ \mathbf{\Gamma} }$ satisfying Assumption (ref). Then the following conditions hold \begin{itemize} • $Q( \widetilde{ \mathbf{\Gamma} },\mathbf{w})= 0$ if and only if $\mathbf{w}=0$. • $Q( \widetilde{ \mathbf{\Gamma} },\mathbf{w}) > 0$ for $\mathbf{w} \neq 0$. • The associated quadratic form defined by the Lagrangian of the objective function (ref) given by $Q( \widetilde{ \mathbf{\Gamma} },\mathbf{w}) + \zeta (\mathbf{w}^{\prime} \mathbf{1} - 1)$, for some $\zeta > 0$, has a global minimum on $\mathbf{w}$. \end{itemize}
remarkNotice that the optimal choice of weights goes beyond the portfolio allocation optimization problem. For example, the choice of $\boldsymbol{w}$ can be based on the vector of naive weights or a vector estimated through an optimization procedure. For example, the covariance matrix used in portfolio allocation problems is positive-definite which allow to estimate the vector of optimal weights. We provide some examples below:
itemize$\textbf{(Variance-Optimal Weights)}$. There is a unique solution to the minimization problem of $\boldsymbol{w}^{\top} \boldsymbol{V} \boldsymbol{w}$ subject to the constraint that $\boldsymbol{w}^{\top} \boldsymbol{1} = 1$, which is \begin{align} \boldsymbol{w}_{gmvp} = \frac{ \boldsymbol{V}_c^{-1} \boldsymbol{1} }{ \boldsymbol{1}^{\top} \boldsymbol{V}_c^{-1} \boldsymbol{1} }, \ \sqrt{k} \big( \widehat{\gamma}_n \left( \boldsymbol{w}_{gmvp} \right) - \gamma \big) \overset{d}{\to} \mathcal{N} \left( \frac{ \boldsymbol{1}^{\top} \boldsymbol{V}_c^{-1} \boldsymbol{B}_c }{ \boldsymbol{1}^{\top} \boldsymbol{V}_c^{-1} \boldsymbol{1} }, \frac{1}{ \boldsymbol{1}^{\top} \boldsymbol{V}_c^{-1} \boldsymbol{1} } \right). \end{align} • $\textbf{(AMSE-Optimal Weights)}$. There is a unique solution to the minimization problem of $\mathsf{AMSE}(\boldsymbol{w}) = \displaystyle \frac{1}{k} \left[ \left( \boldsymbol{w}^{\top} \boldsymbol{B}_c \right)^2 + \boldsymbol{w}^{\top} \boldsymbol{V}_c \boldsymbol{w} \right]$ subject to the constraint $\boldsymbol{w}^{\top} \boldsymbol{1} = 1$.

Lastly, if $\boldsymbol{w}_n \boldsymbol{n} = 1$ with $\widehat{\boldsymbol{w}}_n \overset{p}{\to} \boldsymbol{w}$, then the composite estimator $\widehat{\gamma}_n ( \widehat{\boldsymbol{w}}_n )$ is $\sqrt{k}-$asymptotically equivalent to $\widehat{\gamma}_n ( \boldsymbol{w}_n )$ in the sense that $\sqrt{k} \big( \widehat{\gamma}_n ( \widehat{\boldsymbol{w}}_n ) - \widehat{\gamma}_n ( \widehat{\boldsymbol{w}}_n ) \big) = o_p(1)$.

A Novel Centrality Measure.

Graph Topology and Node Centrality.

We denote with $ \mathcal{G} = \{ \mathcal{V} , \mathcal{E} \}$ a graph structure, representing a financial network, which consists of a set of nodes, $\mathcal{V}$, and a set of edges, $\mathcal{E}$, connecting the pairs of nodes. A full characterization of the information in the network is provided by the $N \times N$ adjacency matrix $\mathbf{\Omega}^0$ whose element $\{ \Omega_{ij}^0 \}_{i,j=1,...,N}$ determine whether there is a link connecting node $i$ and $j$ for the graph $\mathcal{G}$. The link between two elements can be characterized by a binary variable which determines the existence of a connection or by a real value that corresponds to the pairwise weight between nodes.

We propose the following adjacency matrix $\mathbf{\Omega}^0 = \widetilde{ \mathbf{\Gamma} } - diag \left( \widetilde{ \mathbf{\Gamma} } \right)$, where $diag(\cdot)$ is an $N \times N$ diagonal matrix, implying that $diag \left( \widetilde{ \mathbf{\Gamma} } \right) := \text{diag} \{ \gamma^{0}_1, \ldots, \gamma^{0}_N \}$. Therefore, within our framework we consider a weighted network, where the connections between stocks are determined by the $\widetilde{ \gamma}_{i|j}$ measures as defined by the off-diagonal elements of the $\widetilde{ \mathbf{\Gamma} }$ matrix. The adjacency matrix $\mathbf{\Omega}^0$ is symmetric by definition, since the main diagonal has a vector of zeros and the off-diagonal terms are equivalent to the symmetric risk matrix $\widetilde{ \mathbf{\Gamma} }$, that is, $\Omega_{ij}^0= \widetilde{ \gamma}_{i|j}$ for $i \neq j$.

The symmetry of the adjacency matrix entails the spectral decomposition $\mathbf{\Omega}^0 = \mathbf{Z}_{ \Omega^0 } \mathbf{D}_{ \Omega^0 } \mathbf{Z}_{ \Omega^0 }^{\prime}$, with $\mathbf{Z}_{ \Omega^0 }$ an $N \times N$ orthonormal matrix such that $\mathbf{Z}_{ \Omega^0 }^{\prime} = \mathbf{Z}_{ \Omega^0 }^{-1}$ that contains the linearly independent eigenvectors of $\mathbf{\Omega}^0$, and $\mathbf{D}_{\Omega^0}$ is an $N \times N$ diagonal matrix with the corresponding eigenvalues. Moreover, we express the eigenvalues of the adjacency matrix $\mathbf{\Omega}^0$ using the spectral vector decomposition as below

align[align omitted — 166 chars of source]

with $z_{ij}$ the eigenvectors of the matrix $\mathbf{\Omega}^0 \in \mathbb{R}^{N \times N}$. We decompose the loss function $\mathbf{w}' \widetilde{ \mathbf{\Gamma} } \mathbf{w}$ as the sum of a quadratic loss function given by the adjacency matrix and a quadratic loss function given by a diagonal risk matrix with elements given by $\gamma^{0}_i$. More formally, the quadratic form becomes

equation[equation omitted — 218 chars of source]

The study of the contribution of the centrality of assets to the loss function $Q( \widetilde{ \mathbf{\Gamma}} , \mathbf{w})$ is examined by katsouris2021optimal and olmo2021optimal. To do so, we further examine the components of expression $\mathbf{w}' \mathbf{ \Omega }^0 \mathbf{w}$ and provide a definition of asset centrality with respect to the eigenvalue decomposition of the adjacency matrix. The notion of centrality quantifies the influence of certain nodes in a given network. There are several measurements in the literature each corresponding to a specific definition of centrality. We focus on the eigenvector centrality, see also Bonacich72. This measure of centrality is closely related to the concept of Katz centrality, see katz1953new. The eigenvector centrality of asset $i$ is defined as the proportional sum of its neighbors' centrality, see Bonacich72.

Main Theory.

itemize• Let $\succsim$ be a non-trivial binary relation over $X$. We say that $\succsim$ has a weak representation if there exists a non-constant real valued function $v : X \to \mathbb{R}$ such that for all $x, y \in X, v(x) > v(y)$ implies $x \succ y$. • We say that $v$ is a partial representation for $\succsim$ if for all $x,y \in X, x \succ y$ implies $v(x) > v(y)$. • The function $v$ is called strong representation of $\succsim$ if for all $x, y \in X, x \succ y$ if and only if $v(x) > v(y)$, that is, $v$ is a strong representation if it is both a weak and partial representation.

A centrality measure although has no formal definition in the literature in this paper we propose some regularity properties for these measures to correspond to centrality ranking. Motivated from the above findings we propose a novel centrality measure. In particular, in this paper we define the context of recursive monotonicity. In mathematical terms recursive monotonicity means that if vertex $i$ has more neighbors than vertex $j$, and $i$'s neighbours are more central than those of $j$, then $i$ must be more central than $j$. Therefore, we focus on two different network topologies, a topology of low connectivity and a topology of strong connectivity (see, sadler2022ordinal). \color{black}

A measure of centrality is a function $C$ that takes a node $i \in V$ and the graph $G = ( V, E)$ it belongs to, and returns a non-negative real number $C(i ; G) \geq 0$ which translates to a quantity which specifies how central $i$ is in $G$. Currently, there is no agreement on a formal definition of a centrality measure the following properties should hold

itemize• Invariance. • Monotonicity. Adding an edge between $i$ and another node $j$, then the centrality of $i$ does not increase.
assumption[Centrality measures] Under the assumption of a graph $\mathcal{G} = \left( E, V \right)$ it holds: \begin{itemize} • Linear Homogeneity. Removing one node from the graph, then the properties of the centrality measure remain valid since it provides a ranking statistic in relation to the graph topology, that is, the ranked association of the node $i \in \left\{ 1,..., N \right\}$ with respect to the remaining nodes. • Invariance. The centrality measure remains invariant under a linear transformation. A linear (affine) transformation of the adjacency matrix does not affect the ranking statistics in relation to the graph topology. Furthermore, the centrality measures are invariant under graph automorphisms, that is, nodes re-labeling. • \textit{Recursive Monotonicity}. A relevant definition of recursive monotonicity is given by bommier2017monotone. • \textit{Sub-Additivity}. \end{itemize}

Furthermore and more precisely, given a preorder on the vertex set of a graph, the neighborhood of vertex $i$ dominates that of another vertex $j$ if there exists an injective function from the neighbourhood of $j$ to that of $i$ such that every neighbor of $j$ is mapped to a higher ranked neighbour of $i$.

The preorder satisfies recursive monotonicity, and is therefore an ordinal centrality, if it ranks $i$ higher than $j$ whenever $i$'s neighbourhood dominates that of $j$. Consider for example commonly used centrality measures in the literature such as degree centrality, Katz-Bonacich centrality and eigenvector centrality, all of these centrality measures satisfy assumption 1 (recursive monotonicity). Therefore, following our preliminary empirical findings of the optimal portfolio performance measures with respect to the low centrality topology as well as the high centrality topology, we obtain some useful insights regarding the how the degree of connectedness in the network and the strength of the centrality association can affect the portfolio allocation. Building on this idea, in this section we focus on the proposed algorithm which provides a statistical methodology for constructing a centrality ranking as well as a characterization of the strength of the centrality association between nodes.

More specifically, we assume that a node exhibits "strong centrality" when all the rank centrality statistics assign a higher preference to vertex $i$ above $j$ (in terms of graph topology), which implies that in a sense $i$ is considered to be more central than the node $j$ for all $i,j \in \left\{ 1,..., N \right\}$ such that $i \neq j$. On the other hand, we assume that a node exhibits "weak centrality" when vertex $i$ is weakly more central than vertex $j$ if there exists some strict ordinal centrality that ranks $i$ higher than $j$.

According to sadler2022ordinal weak centrality itself is a strict ordinal centrality, so it is the maximal such order, and it extends strong centrality. Moreover, in this paper we employ these terminologies for the development of an algorithm for feature ordering by centrality exclusion. We define some useful mathematical definitions such as the binary relation $\succsim$ on $V$ is a set of ordered pairs of vertices, and we write $ i \succsim j$ if $(i,j) \in \succsim$.

definition[recursive monotonicity] Given a graph $G = ( V, E)$ and a binary relation $\succsim$ on $V$, the subset $S \subseteq V$ $\mathbf{\succsim}-$dominates $S^{\prime} \subseteq V$, and we write $S \succsim S^{\prime}$, if there exists an injective function $f: S^{\prime} \to S$ such that $f(i) \succsim i$ for all $i \in S^{\prime}$. Furthermore, the binary relation $\succsim$ satisfies recursive monotonicity if $i \succsim j$ whenever $G_i \succsim G_j$.
remarkAn ordinal centrality $\succsim$ on $G = (V, E)$ is a preorder on $V$ that satisfies recursive monotonicity. Recursive monotonicity captures the essence of a board class of centrality measures. Thus a vertex $i$ is more central than vertex $j$ if $i$ has more connections to more central nodes than $j$.
definition[preorder] The function $c: V \to \mathbb{R}$ represents the preorder $\succsim$ if we have $i \succsim j$ if and only if $c(i) \succsim c(j)$.
definition[strongly more central] Given a graph $G = ( V, E )$, node $i$ is strongly more central than node $j$, and we write $i \succsim_s j$, if and only if $i \succsim j$ for every ordinal centrality on $G$.

Centrality Rank Statistic.

We follow the example presented in janssen2003bootstrap. Given $k(n)$ let $\big( R_1,..., R_{k(n)} \big)$ be uniformly distributed ranks, that is, a random variable with uniform distribution on a set of permutations $\mathcal{P}_{k(n)}$ of $\left\{ 1,..., k(n) \right\}$. Furthermore, an additional index $n$ concerning the ranks is suppressed throughout. Consider a sequence of simple linear permutation statistics such that

align[align omitted — 77 chars of source]

where $c_{n_i}$ are regression coefficients, $1 \leq i \leq k(n)$, with

align[align omitted — 130 chars of source]

and random scores $d_n(i): \tilde{\Omega} \mathbb{R}$, $1 \leq i \leq k(n)$ with

align[align omitted — 199 chars of source]

The $c'$s are here considered to be fixed and the $d'$ are allowed to be random variables independent of the ranks $R_j$. Without restrictions we may assume ordered regression coefficients

align[align omitted — 59 chars of source]

Then, we have equality in distribution for $\big( d_n (R_i) \big)_{ i \leq k(n) } \overset{ D }{ = } \big( d_{ R_{i : k(n) } } \big)_{ i \leq k(n) }$, which is a consequence of the ex-changeability of these variables. Therefore, in these cases ranks and order statistics are independent. Thus, without restriction we may assume that the score functions are also ordered

align[align omitted — 67 chars of source]

Therefore, since $\mathcal{E} \big( S_n | W \big) = 0$ holds always convergence subsequences of $S_n$ exist and under some regularity conditions we can classify all possible limit distributions of subsequences with respect to the distributional convergence. We define the normalized random variables as below

align[align omitted — 164 chars of source]

Feature Sorting by Centrality Exclusion.

The proposed algorithm of exclusion centrality is constructed using the proposed high-dimensional tail forecast risk matrix. However, the feature ordering by centrality exclusion (FOCE) procedure in practise can be applicable for covariance matrices for graphical models. On the other hand, when the FOCE procedure is implemented based on the VaR-$\Delta$CoVaR risk matrix, then the feature selection is based on the tail dependence of the response variable $\boldsymbol{Y}_t$ given the predictors $\boldsymbol{X}_{t-1}$. Furthermore, since the construction of the proposed risk matrix is based on the quantile predictive regression models, then we specifically focus on a type of granger causality in the tails of the underline distributions. Thus, our proposed algorithm provides a methodology for ordering the vertices of the graph based on the effect of their centrality in relation to the other nodes in the graph when we exclude the most central node in the graph. It produces an ordering of the predictors according to their predictive power. This ordering is used for variable selection without putting any assumption on the distribution of the data or assuming any particular underlying model.

The simplicity of the estimation of the conditional dependence coefficient makes it an efficient method for variable ordering and variable selection that can be used for high dimensional settings. In this section, motivated by the study of centrality measures and the proposed risk matrix of the paper, we propose a novel feature selection algorithm for multivariate regression models using the centrality exclusion procedure based on our novel regression-based risk matrix. Notice that other methodologies currently proposed in the literature include the model-based methods.

Let $Y$ be the response variable and let $\boldsymbol{X} = \left( X_j \right)_{ 1 \leq j \leq p }$ be the set of predictors. The data consists of $n$ i.i.d copies of $\left( Y, \boldsymbol{X} \right)$. Similar to the framework proposed by azadkia2021simple, first, choose $j_1$ to be the index $j$ that maximizes $T_n( Y, X_j )$. Then, having obtained $j_1,..., j_k$, choose $j_{k+1}$ to be the index $j \not\in \left\{ j_1,..., j_k \right\}$ that maximizes $T_n \left( Y, X_j | X_{j_1},...., X_{j_k} \right)$. Furthermore, continue like this until arriving at the first $k$ such that $T_n \left( Y, X_{j_{k+1} } | X_{j_1},...., X_{j_k} \right) \leq 0$, and then declare the chosen subset to be $\hat{S} := \big\{ j_1,..., j_k \big\}$. If there is no such $k$, define $\hat{S}$ to be the whole set of variables. Furthermore, this may also happen that $T_n \big( Y, X_{j_1} \big) \leq 0$.

In that case, we declare $\hat{S}$ to be the empty set. In the next section, we prove the consistency of FSCE under a set of assumptions on the law of $\left( Y, \boldsymbol{X} \right)$. Furthermore, we focus on proposing on a well-defined stopping rule. Our stopping rule seems to work well in practise, and we are able to prove consistency of variable selection for this rule.

Suppose that $\boldsymbol{X}$ is a normal random vector with zero mean and arbitrary covariance structure,

align[align omitted — 49 chars of source]

Then, for any non-empty $S \subset \left\{ 1,..., p \right\}$ and any $j \in \left\{ 1,..., p \right\} \backslash S$. Notice that if $S$ is a sufficient set of predictors, then $\rho( S, j ) = 0$ for any $j \not\in S$. To the best of our knowledge the proposed algorithm for higher-order conditional risk measures is novel and is based on distribution-free inference.

See for reference the paper: chao2018multivariate. In the particular paper the authors introduce the following estimation procedure. Let $\big\{ \left( \boldsymbol{X}_i, Y_{i1},..., Y_{im} \right) \big\}_{ 1 \leq i \leq n }$ where $Y_{ij}$ represents the value observed from the response $j$ at the time point $i$, and $\left\{ \boldsymbol{X}_i \right\}_{i=1}^n$ are the covariates. Furthermore, we assume that the samples are i.i.d over $i$. For $\uptau \in (0,1)$, the conditional expectile $e_j \left( \uptau | \boldsymbol{X}_i \right)$ of $Y_{ij}$ given $\boldsymbol{X}_i$ is defined as below

align[align omitted — 120 chars of source]

where

align[align omitted — 257 chars of source]

The proposed feature selection by node exclusion algorithm is considered as a dimension reduction methodology for the graph. Below we present some related algorithms from the economic theory and econometrics literature.

itemize• Heuristic Pricing Algorithm: The particular algorithm iteratively creates an ordering of the choice situations, which can correspond to a rational choice type. First, we explain the link between orderings of the choice situations and rational choice types. Next, we explain how to build an ordering that provides a good solution. • Generation of Choice Types for Tightening: To tighten the set based on a subset of the rational choice types, we generate the subset in a semi-random way. First, we generate (likely irrational) choice types by randomly choosing one patch in each choice situation. If this choice type is rational, we add it to the subset for tightening. If it is not, we identify the subsets of choice situations for which preference cycle exists. For each such subset, we randomly pick one choice situation. For that choice situation, we look for a patch which (i) removes at least one preference relation within the subset, (ii) is as close as possible to the currently selected patch in that choice situation, and (iii) removes (rathern than adds) revealed preference relations. In this way, we slightly change the choice type, while increasing the probability that it is a rational choice type. If after these changes the choice type is not yet rational, the procedure is repeated until a rational choice type is found. The algorithm below contains the pseudo-code to generate these rational choice types in a semi-random way. In the algorithm, we again define $\mathcal{T} = \left\{ t | 1 \leq t \leq T \right\}$ as the set of all choice situations.

Algorithm.

We propose a novel feature ordering algorithm for a multivariate regression with a fixed number of covariates, using a backward stepwise algorithm which is applied to our network driven tail risk matrix. In other words, given a high dimensional response vector $\boldsymbol{Y} = ( Y_j )_{1 \leq j \leq p}$ and a common set of predictors $\boldsymbol{X} = ( X_j )_{1 \leq j \leq k}$ then, the high dimensional response vector is reduced to a lower dimensional space based on the same number of predictors. In other words, given a high dimensional response vector and a set of covariates, we conjecture that the dimension of the response vector can be reduced without having to extend the set of predictors in order to better explain variation in the high-dimensional response vector.

exampleConsider a class of linear additive models \begin{align} y_t = B z_t + u_t, \ \ \ t = 1,..., T \end{align} with $y_t = \big( y_{1t},..., y_{pt} \big)^{\prime}$ and $u_t = \big( u_{1t},..., u_{pt} \big)^{\prime}$
itemize• Both $y_t$, the high-dimensional response vector, and $z_t$ are observable. Usually, the literature considers the estimation of $\Sigma$ associated with the unobserved error term $u_t$, which is a measure of uncertainty. On the other hand we consider the predictive ability using the one-ahead period forecasts. Our methodology reduces the dimensionality of the response vector for a fixed number of covariates at the quantile level. • Notice that in practise we examined two different procedures here. The first procedure utilizes a suitable stopping rule with respect to a portfolio optimization problem. The second procedure, considers an alternative way of ordering the importance (with respect to the degree of centrality) of the high-dimensional response vector. • A relevant notion from the statistical theory perspective is the aspect of $\textit{sufficient dimension reduction in classical statistics}$. In the classical statistics setting, if one can find a small subset of predictors that is sufficient, then we can assume that these predictors contain all the relevant predictive information about $Y$ among the given set of predictors, and the statistician can then fit a predictive model based on this small subset of predictors. • On the other hand, if the reduction of the dimensionality of the high dimensional response vector is what the statistician is interested to, then our approach can work regardless of the dimensionality of the predictors. This implies that a question of constructing a statistical procedure for testing the inclusion of high dimensional controls is replaced on whether a multivariate regression model should include or not a high dimensional response vector.

The main idea of the proposed algorithm is a way of selecting features or reducing the number of nodes in a graph, since the algorithm chooses the most central node and then removes it from the graph iteratively using a stopping rule. In other words, the proposed procedure can be considered to be a methodology that bridge the gap between a node exclusion statistical methodology and a procedure for feature selection in graphs without employing Lasso regularization techniques.

Numerical Studies.

Simulation study.

In this section, we apply our method on the simulated data to evaluate the estimation performance on the factors and loadings, as the number of factor varies.

We set with $n = m = p = 100$. For $i = 1,..., n, j = 1,...,m$, and let $\boldsymbol{X}_i \sim \mathcal{N} \big( \boldsymbol{0}_{p \times 1}, \boldsymbol{\Sigma}_{p \times p} \big)$ with $\boldsymbol{\Sigma}_{ij} = 0.5^{ | j - k | }$ and $\boldsymbol{\epsilon}_i \overset{ i.i.d }{ \sim } \mathcal{N} \big( \boldsymbol{0}_{m \times 1}, \boldsymbol{I}_{ m \times m} \big)$. Then, the response variables are generated by

align[align omitted — 288 chars of source]

where $r = \mathsf{rank} \left( \boldsymbol{\Gamma} \right)$.

\color{black}

Let $\Gamma$ be the an $N \times N$ matrix whose $(i,j)$ entry is $\upgamma_{ij}$. Furthermore, let $\displaystyle \Gamma = \sum_{j=1}^N s_j u_j v_j^{\prime}$. Furthermore, notice that since the elements of the risk matrix $\Gamma$ are constructed based from population moments of the pairwise conditional distribution functions. Considering the fact that the entries of the risk matrix $\Gamma$ are estimated by the nodewise quantile predictive regression models. Thus, our matrix is considered to be a matrix with pairwise tail forecasts.

Let $\Gamma$ be the $N \times N$ matrix whose $(i,j)-$th element is $f( \beta_i, \beta_j )$. Then, our data matrix $X$ has the following form

align[align omitted — 59 chars of source]

where $\epsilon_{ij}$ are independent errors with zero mean, satisfying the restriction that $| x_{ij} | < 1$ almost surely. For example, X may be the adjacency matrix of a random graph where the probability of an edge existing between vertices $i$ and $j$ is $f( \beta_i, \beta_j )$.

Conclusion.

Our contributions to the literature are twofold: (i) we propose a novel regression-based risk matrix for modelling tail dependence based on graph structures; and (ii) we propose a statistical procedure for feature selection in graphical models. Furthermore, we focus on the interpretation of the risk matrix for optimal portfolio allocation problems as well as a mechanism to introduce our novel variable selection algorithm. Our study discusses important aspects related to the robust estimation of large tail forecast covariance-type matrices as well as the implementation of feature ordering procedures using the proposed graph-based matrix.

Obviously, the drawback of our proposition is the related challenges in evaluating the predictive accuracy of the tail risk matrix since the elements of the matrix consist of risk measures such as the Value-at-Risk and Conditional-Value-at-Risk which are tail estimates of the underline distributions. Secondly, the formal study of our tail risk matrix as a suitable representation for optimal portfolio choice problems is crucial since this is a novel aspect not previously proposed in the literature. Our motivation in this paper although is from the optimal portfolio allocation perspective in financial networks, we focus on the aspect of the induced centrality measure for the proposed novel risk matrix.

Some future research worth mentioning include the formal study of the asymptotic properties of the proposed risk matrix. Although we assume that we employ stationary time series observations our framework can be extended to accommodate features of non-stationary time series. We leave the particular aspect for future research. A second aspect of interest are the time series properties of regressors when estimating the risk measures of VaR and CoVaR. Although the risk management literature usually operates under the assumption of stationarity, an interesting aspect for future research is the effect of non-stationarity when estimating our novel risk matrix. We leave the particular aspect for future research. Although in this paper, we do not consider a formal statistical methodology for edge exclusion this is certainly an aspect which worth future research in subsequent studies. Furthermore, our proposed framework can be employed as a methodology for covariate screening in high dimensional environments. We leave the particular aspect for future research.

\paragraph{Acknowledgements}

smallI wish to thank my main PhD supervisor Prof. Jose Olmo for guidance and continuous encouragement throughout the PhD programme as well as Prof. Peter W. Smith for helpful discussion and for stimulating the further development of the current study in the direction of node exclusion and feature selection. In addition, I wish to thank Zudi Lu, Jean-Yves Pitarakis, Tassos Magdalinos, Hector Calvo Pardo, Zacharias Maniadis, Ruben Sanchez Garcia and Tullio Mancini as well as Departmental Seminar Series speakers: Luis E. Candelaria, Juan Carlos Escanciano, Marcelo C. Medeiros, Sebastian Engelke, Markus Pelger and Loriano Mancini for helpful conversations. The author acknowledge the use of Iridis 5 HPC Facility and associated support services at the University of Southampton in the completion of this work. \paragraph{Funding} Financial support from the Vice-Chancellor's PhD scholarship of the University of Southampton is gratefully acknowledged. The author declares no conflicts of interests.