Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
79,359 characters · 14 sections · 54 citation commands
Network Regression and Supervised Centrality Estimation
\def\spacingset#1{ {#1}} \spacingset{1}
\if11 \fi
\if01 {
} \fi
{\it Keywords:} Hub centrality, Authority centrality, {Network effect,} Global trade network, Currency risk premium
\spacingset{1.8}
{0pt} {0pt} {0pt} {0pt} {0pt}
\newdimen\origiwspc \newdimen\origiwstr \origiwspc=\fontdimen2\font \origiwstr=\fontdimen3\font \if11 \fi
In many disciplines such as economics, finance, and sociology, there has been great interest in studying the network effect, that is, the effect of a network on certain outcomes of interest due to relationships among agents (e.g., individuals, firms, industries, and countries). One popular approach is to bridge the outcome and network via an intermediary or a sufficient statistics -- the centrality of the network.
As a low-rank summary of a network, centrality is a common metric to measure agents' importance in the network, which in turn induces a wide range of agent behaviors that consequently shapes certain outcomes of {theirs}. A strong motivation for centrality is that many real-world networks exhibit a low-rank structure, i.e., the leading singular value dominates the rest in magnitude allen2019ownership,zhu2020networks1,liu2020dynamical. Centrality itself has rich implications for studying human capital investment jackson2017economic, information sharing and advertising banerjee2019using, breza2019social, firms' investment decision-making allen2019ownership, the identification of banks that are too-connected-to-fail gofman2017efficiency, and stock returns ahern2013network, richmond2019trade1, among many others.
To be specific, researchers often regress the outcome of interest on the network centrality to study the network effect. This approach has been implemented in many fields including portfolio management, finance, and social networks. In portfolio management, hochberg2007whom, ahern2013network and richmond2019trade1 demonstrated that, for a trade network of firms or countries, a strategy that shorts portfolios with high centralities and longs those with low centralities yields a significant excess return, and regressing risk metrics on the centrality of the financial institutions {helps to} understand the amplification of severe adversarial shocks to the central institutions in the network. liu2019industrial examined the effect of centrality in the production network on the government's investment in strategic industries to illustrate the effectiveness of industrial policies. For social networks, ozsoylev2014investor and rossi2018network regressed the excess returns of investment managers on the centrality of their social networks to study trading behaviors; kornienko2018peer and mojzisch2021interactive studied the network effect on mental health by regressing the stress level on the network centrality.
\fontdimen3\font=0.15em Network centrality, however, is not directly observable. In practice, researchers often follow a two-stage procedure: in Stage 1, they compute the centrality from a given network adjacency matrix using an algorithm; in Stage 2, the computed centrality is then used as an input in the regression analysis. Such a practice will be referred to as the two-stage procedure throughout. \fontdimen3\font=\origiwspc
The validity of the two-stage procedure, however, hinges upon one critical assumption that the centrality is computed from a noiseless observed adjacency matrix in Stage 1 so that it is accurate. In reality, a network is often observed with noise due to the cost of data collection lakhina2003sampling. There are numerous examples of such noise: the friendship network on Facebook or Twitter is far from a perfect measure of real-life social connections; using self-reported friendships to measure social ties suffers from subjective biases banerjee2013diffusion; using patent citations to measure the knowledge flow between companies neglects the communication among workers or executives zhu2020networks1. Overlooking noise in networks has demonstrable consequences for network analysis borgatti2006robustness, frantz2009robustness, wang2012measurement, martin2019influence, candelaria2022identification.
Given a noisy observed network, one has two goals in understanding the network effect:
The two-stage procedure attempts to achieve these two goals in a sequential manner, yet it has the following drawbacks. First, Stage 1 only uses the information from the noisy network to estimate centrality without incorporating the auxiliary information from the regression on the centrality, which can result in inaccurate estimation of the centrality due to large observational errors in the network. Second, Stage 2 is contingent upon Stage 1 -- regressing the outcome on the inaccurately estimated centrality exacerbates an inaccurate estimation of the regression coefficients, thereby invalidating the follow-up statistical inference.
To remedy the shortcomings of the two-stage procedure, we first propose a unified framework that fuses two models to achieve the two goals: one network generation model based on the centralities for (ref) and one network regression model for the dependency of the outcome on the centralities for (ref). We then propose a novel supervised network centrality estimation (SuperCENT) methodology that accomplishes both (ref) and (ref) simultaneously, instead of sequentially.
SuperCENT exploits information from the two models -- the network regression model contains auxiliary information on the centrality in addition to the network, and thus provides supervision to the centrality estimation. The supervision effect improves the centrality estimation, which in turn benefits the network regression. Therefore, the centrality estimation and the network regression complement and empower each other. Under the unified framework, we derive the theoretical convergence rates and asymptotic distributions of the centralities and regression coefficients estimators, for both the two-stage and SuperCENT methods, which can be used to construct confidence intervals.
\fontdimen3\font=0em We summarize our contributions as follows. First, to the best of our knowledge, despite the popular adoption of the two-stage procedure, we are the first to provide a unified framework to study properties of centrality estimation and inference, and the subsequent network regression analysis when the observed network is noisy.
{ Second, we are the first to study the properties of the common practice of the two-stage procedure and demonstrate that it can be problematic when the network noise is large. The accuracy of the two-stage centrality estimates in Stage 1 depends on the network noise. When the network noise is large, the centrality estimates are inaccurate, which results in inaccurate centrality coefficient estimates with invalid ad-hoc inference in Stage 2. }
\fontdimen3\font=\origiwspc
Thirdly, we show theoretically and empirically that the proposed SuperCENT dominates the two-stage procedure universally. Specifically, for (ref), SuperCENT yields a more accurate centrality estimation, especially under large network noise; {for (ref), SuperCENT boosts the accuracy of the regression coefficient estimation and provides confidence intervals that are valid and narrower than the ad-hoc two-stage confidence intervals.}
\fontdimen3\font=0.1em Lastly, we apply both SuperCENT and the two-stage procedure to predict the currency risk premium, based on an economic theory that links a country's currency risk premium with its importance within the global trade network richmond2019trade1. We show that a long-short {trading} strategy based on SuperCENT centrality estimates {yields} a return {double} that of the two-stage procedure. Furthermore, SuperCENT can verify the economic theory via a rigorous statistical test while the two-stage fails. \fontdimen3\font=\origiwspc
Our paper contributes to several strands of literature, including network modeling, network regression with centralities, covariate-assisted network modeling, and network effect modeling. First, the proposed unified framework bridges the gap between research on noisy networks and network regression with centralities. Most of the existing network literature focuses on one of these two aspects. On one hand, in studies involving noisy networks, many empirical works have estimated the true network without incorporating centrality measures (e.g., lakhina2003sampling, handcock2010modeling, banerjee2013diffusion, le2018estimating, rohe2019critical, breza2020using). On the other hand, numerous work, including those mentioned earlier, have focused on the network regression model with centralities while ignoring the estimation error of the centralities inherited from the noise of the network.
Our unified framework also relates to the line of research on networks with covariates supervision zhang2016community, li2016supervised, fan2016projected, binkiewicz2017covariate, yan2019statistical, ma2020universal. One major difference is that SuperCENT uses both the covariates and the response to supervise the estimation, instead of only the covariates. In addition, the existing literature has focused mostly on network formation or community detection.
In econometrics, there has been significant effort to model the network effect on an outcome of interest through regression de2017econometrics. One popular approach follows the pioneering work of manski1993identification and his “reflection model” lee2007identification, bramoulle2009identification, lee2010specification, hsieh2016social,zhu2017network. This approach models the network effect through the observed adjacency matrix itself, not through the centralities like ours. There has also been a recent surge of literature in network recovery based on the reflection model de2019identifying, battaglini2021endogenous. This literature focuses on the issue of identifiability of the network effect, while our work attends to both estimation and inference of the network effect. Another popular approach assumes that the outcome depends on individual fixed effects, and casts the role of the network through the Laplacian matrix, such that connected nodes share similar individual fixed effects li2019prediction, le2020linear. This approach emphasizes network homophily, while ours concentrates on the nodes' position or importance in the network using the centralities.
{ It is worth mentioning that there is an extensive body of literature discussing the concept of centrality in networks with negative-weight edges. The foundational work on these networks stems from social balance theory in sociology harary1953notion, cartwright1956structural. bonacichCalculatingStatusNegative2004 extends the concept of centrality to such networks, providing interpretations grounded in balance theory. Several subsequent studies have built on this foundation chiang2014prediction, everettNetworksContainingNegative2014, singh2019eigenvector, ma2019clusters, gromov2025social. Empirical research has also explored these networks, including studies on workplace dynamics labiancaExploringSocialLedger2006 and alliance-enemy networks in wartime konig2017networks. Our method is applicable to these networks, and centrality can be interpreted within the framework established by the literature. }
The rest of this article is organized as follows. Section (ref) provides the background and formally introduces the unified framework. Descriptions of the two-stage procedure and SuperCENT are given in Section (ref). Theoretical properties are studied in Section (ref) and the simulation study is shown in Section (ref). Section (ref) presents the case study of the relationship between currency risk premiums and the global trade network centralities. Section (ref) concludes with a summary and future work. The supplementary materials contain additional background information on network and centralities, detailed descriptions of the algorithms for undirected networks, more simulation results, additional information of the case study, some concrete mathematical expressions, and the proofs. We developed an R package, SuperCENT, that implements the methods \if01 (\url{https://cccfran.github.io/SuperCENT}). \fi \if11 (\url{https://jh-cai.com/SuperCENT}). \fi
We observe a sample of $n$ observations $(\x_1, y_1), \, (\x_2, y_2),\ldots, (\x_n, y_n)$ where $y_i \in \mathbb{R}$ is the response and $\x_i\in\mathbb{R}^{p-1}$ is the vector of $p-1$ covariates for the $i$-th observation as in the multivariate regression setting. Let $\by \in \mathbb{R}^n$ denote the column vector of outcome and $\bX \in \mathbb{R}^{n\times p}$ denote the design matrix including the intercept, which is assumed to be fixed.
In {a} network, the nodes are agents and the edges represent relationships between the agents. The edges can be directed or undirected depending on whether the relationships are reciprocal. This article focuses on directed networks; the Supplement provides the results for undirected ones. A weighted directed network with $n$ nodes can be represented by an asymmetric adjacency matrix $\bA \in \mathbb{R}^{n \times n}$ where $a_{ij}$'s represent the weighted edges.
Researchers have used multiple versions of network centrality. We refer to Chapter 2 of jackson2010social for a comprehensive introduction to centrality. We focus on the hub and authority centralities kleinberg1999authoritative, which extend the well-known {eigenvector centrality} associated with the undirected network to the directed network.
For directed networks, there is a distinction between the giver and the recipient, such as the citee-citor in citation networks or web-page networks, the exporter-importer in trade networks, and the investor-investee in investment networks. The hub and authority centralities take into account the different roles of the giver and the recipient, and thus measure the importance of nodes from these two different perspectives. The concept of “hubs and authorities” {originated} from web searching. Intuitively, the {hub centrality} of a web page depends on the total level of authority centrality of the web pages it links to, while the {authority centrality} of a web page depends on the total level of hub centrality of the web pages it receives links from. Supplement (ref) provides an example to further illuminate this intuition.
Let $u_i$ denote the hub centrality and $v_i$ denote the authority centrality for node $i$, and let $\bu=(u_1,u_2,\ldots,u_n)^\top,~~\bv=(v_1,v_2,\ldots,v_n)^\top$. Their relationship hence satisfies $ \bu =\A \bv,~ \bv=\A^\top \bu. $ Given $\A$, to calculate the centralities, kleinberg1999authoritative proposes iterating with proper normalization as follows until convergence, for $k = 1, 2, 3, \ldots,$ \be \bu^{(k)} \leftarrow \A \bv^{(k-1)}, \bv^{(k)} \leftarrow \A^\top \bu^{(k)}. \ee This iterative algorithm is also well known as the power method to compute the singular value decomposition (SVD) of $\bA$ van1996matrix. Therefore, the hub and authority centralities are the leading left and right singular vectors of $\bA$ respectively. It is worth mentioning that such definition of centrality and the algorithm essentially assume that the adjacency matrix $\bA$ is noiseless.
\ifdoublespacing \fi
\fontdimen3\font=0em We propose the following unified modelling framework that encapsulates (ref)-(ref), {
where $\bD$ is a diagonal matrix of dimension $r\times r$ with the singular values $d > d_2 \geq \ldots \geq d_r \geq 0$ as the diagonal entries, and $\bU = (\bu, \bu_2, \ldots, \bu_r)$ and $\bV = (\bv, \bv_2, \ldots, \bv_r)$ are two matrices of size $n\times r$ with orthogonal columns of length $\sqrt{n}$. } The intuitions of the unified framework are as follows. The hub and authority centralities are calculated as the leading left and right singular vectors of the observed adjacency matrix. As such, it is natural to consider the generative model (ref) for the observed adjacency matrix, where $\bA_0$ is the true adjacency matrix, the true centralities $\bu,\bv\in \mathbb{R}^n$ are the parameters of interest to be estimated, {$(\bu_2, \ldots, \bu_r)$ and $(\bv_2, \ldots, \bv_r)$ are the non-leading singular vectors orthogonal to $\bu, \bv$,} and $\bE$ is the additive noise of mean zero. Then, (ref) naturally models the relationship between the centralities and the response variable. Here, $\bbeta_x \in \mathbb{R}^p$ is the vector of the regression coefficients, $\betau, \betav \in \mathbb{R}$ are the coefficients of the hub and authority centralities, and the regression error $\bepsilon$ has mean zero. Note that in (ref) it is the true centralities, not the estimated ones, that have direct impacts on the response {and only $\bu$ and $\bv$ are included instead of the entire $\bU$ and $\bV$ because it is common practice to consider the network effect via only the centralities.}
Under the unified framework (ref) with observed data $\{\bA, \bX, \by\}$, our original two goals (ref)-(ref) become concrete: (i) estimate the true centralities $\bu,\bv$; (ii) estimate the regression coefficients $\bbeta_x,\beta_u,\beta_v$; and (iii) construct valid confidence intervals (CIs) for the centralities and the regression coefficients.
\fontdimen3\font=0em The low-rank mean plus noise model (ref) has been commonly adopted for matrix estimation or denoising shabalin2013reconstruction1, yang2016rate, cai2018rate1, matrix completion candes2010matrix, and network community detection with slight modifications rohe2011spectral, zhao2012consistency, lei2015consistency, le2016optimization, gao2021minimax. There is a strand of literature on latent variables network models that can be rewritten as (ref) hoff2009multiplicative, soufiani2012graphlet, fosdick2015testing. \fontdimen3\font=\origiwstr
The unified framework unites our estimation goals and provides a theoretical framework to study the behaviors of the two-stage procedure and motivates our new methodology. Under Model (ref) and some extra assumptions on the noise, shabalin2013reconstruction1 proves that if the noise-to-signal ratio is large, the leading singular vector of $\bA$ and that of $\bA_0$ converge to orthogonal as $n$ goes to infinity. This implies that the naive estimation of the centralities by implementing SVD on the observed network will fail in the presence of large noise, which invalidates the common practice of two-stage. Furthermore, unifying the two models motivates our supervised network centrality estimation (SuperCENT) methodology, which we will describe formally in the next section. We name it the “supervised” centrality estimation because $(\bX, \by)$ in the regression (ref) can be thought of as the supervisors that offer additional supervision to the centrality estimation. It is expected that if the centralities indeed have strong predictive power (that is, the centrality regression coefficients $\betau,\betav$ are large compared with the regression noise level), the estimation of the centralities will be better when considering both (ref) and (ref) instead of only (ref). With the improved estimation of the centralities, {SuperCENT can further improve the estimation and inference of the regression model.}
Note that $\bu,\bv$ are only identifiable up to a scalar. SVD assumes $\bu$ and $\bv$ have unit length. However, we assume $\|\bu\|_2=\|\bv\|_2=\sqrt{n}$, because the network can grow and consequently the centralities should roughly be on the same scale with the network. This prevents the centrality regression coefficients from exploding as the network grows.
\fontdimen3\font=0em \jcM{ Sections (ref) and (ref) formally introduce the two-stage procedure and SuperCENT, respectively. In Supplement (ref), we will derive SuperCENT algorithm and prove its algorithmic convergence, discuss the prediction procedure and tuning parameter selection, and provide an algorithm for undirected networks with eigenvector centrality. } \fontdimen3\font=\origiwstr
As mentioned in the introduction, given the unified framework (ref) and the observed data $\{\bA, \bX, \by\}$, a natural and ad-hoc procedure is the two-stage estimator, which can serve as a benchmark. In view of (ref), the first stage is to perform SVD on the observed adjacency matrix $\bA$ and take its leading left and right singular vectors and rescale them to have length $\sqrt{n}$, denoted as $\huts$ and $\hvts$, as the estimates for the centralities $\bu$ and $\bv$, respectively. The superscript ts stands for {two-stage}. In view of (ref), given the estimates $\huts$ and $\hvts$, the second stage performs {the} ordinary least square (OLS) regression of $\by$ on $\bX$ and $\huts,\hvts$, treating $\huts,\hvts$ as fixed covariates.
Hence, the two-stage procedure solves the following two optimizations sequentially,
It follows that $\hbbetats = (\hW^\top \hW)^{-1}\hW^\top\y$, where $\hW = (\X, \huts, \hvts)$.
In the two-stage procedure, the estimation of the regression model in Step 2 depends on the centrality estimation in Step 1. The more accurate the centrality estimates are, the better we are able to make inference in the regression model. On the other hand, the centralities are incorporated into the regression model as regressors, so $(\bX, \by)$ can supervise centrality estimation and thus boost the estimation accuracy.
\fontdimen3\font=\origiwspc Motivated by the intuition above, we propose to optimize the following objective function $\mathcal L(\bu,\bv,\bbeta, d)$, where $\bbeta = (\bbeta_x^\top,\betau,\betav)^\top$, to obtain the SuperCENT estimates, \be (\hu,\hv, \hbbeta,\hd) := \argmin_{\substack{\|\bu\|_2=\|\bv\|_2=\sqrt{n}, \bbeta, d}} \frac{1}{n}\|\by - \bX\bbeta_x - \bu\beta_u - \bv\beta_v\|_2^2 + \frac{\lambda}{n^2} \|\bA - d\bu\bv^\top\|_F^2, \ee where $\hbbeta = (\hbetax^\top,\hbetau,\hbetav)^\top$ and $\|\cdot\|_F$ is the Frobenius norm of a matrix. The above objective function combines the residual sum of squares in (ref) and the rank-one approximation error of the observed network in (ref). The connection between the two terms is the centralities. The trade-off between them can be tuned through {a} proper selection of the hyper-parameter $\lambda$.
To solve (ref), we use a block gradient descent algorithm by updating $(\hu,\hv, \hbbeta,\hd)$ iteratively until convergence. The initialization is based upon the two-stage estimation, $(\bu^{(0)}, \bv^{(0)}) = (\huts, \hvts)$. {For any matrix $\bA$, let} $\bP_\bA$ {denote} the projection matrix that projects onto the column space of $\bA$, and $\|\bA\|_2$ {be} the matrix operator norm. We use {$(\bu^{(t)}, \bv^{(t)}, \bbeta^{(t)}, d^{(t)})$} to denote the estimation in the $t$-th iteration. The complete algorithm with a given tuning parameter $\lambda$ is shown in Algorithm (ref).
\jcM{
The tuning parameter can be chosen via cross-validation, for which prediction procedure has to be introduced. The prediction is a non-trivial task, since it involves the network expansion and justification to use the same regression coefficients for centralities with more nodes introduced into the model. Given the length constraint, the detailed prediction procedure and why it works are given in Supplement (ref), and the cross-validation procedure is provided in Supplement (ref). }
We investigate the statistical properties of SuperCENT and compare it with the two-stage procedure in this section. \jcM{The proofs are deferred to Supplement (ref).} We start with introducing notations and assumptions.
Denote SuperCENT estimators, \jcM{i.e., the minimizer of the objective function (ref) with a given tuning parameter $\lambda$,} as $\hd$, $\hu$, $\hv$, and $\hbbeta = ((\hbetax)^\top,\hbetau,\hbetav)^\top$ and the two-stage counterparts as $\hdts$, $\huts$, $\hvts$, and $\hbbetats$. We further denote $\bAp = \bUp \bDp \bVp^\top$ where $\bUp = (\bu_2, \ldots, \bu_r)$, $\bVp = (\bv_2, \ldots, \bv_r)$ and $\bDp = diag(d_2, \ldots, d_r)$, $\bm\Omega = \left(
\right)$, $\pxuv$ as the projection matrix that projects onto the column space of $(\bX, \bu,\bv)$, and similarly for $\bP_{\bX}$, $\pu$ and $\pv$. Define $\tu = (\bI-\bP_{\bX})\bu$, $\tv = (\bI-\bP_{\bX})\bv$, which are the centralities projected onto the orthogonal space of $\bX$, and $ \bCuv = \left(\tu, \tv\right)^\top \left(\tu, \tv\right) $. In the following, the theorems are for $\argmax_{\mathbf{h}\in \{\hu,-\hu\}} \mathrm{sign}(\mathbf{h}^\top\bu) $ and $\argmax_{\mathbf{g}\in \{\hv,-\hv\}} \mathrm{sign}(\mathbf{g}^\top\bv) $, and we continue to use $\hu,\hv$ to denote them. While these notions are a bit of an abuse of notation, it is reasonable since both the objective function and algorithm are sign-invariant (i.e., proper flipping of signs gives the same value or another valid iteration sequence). The same notation applies in the simulations as well.
In Assumption (ref), the independence is assumed for simplicity. If the network noises $e_{ij}$'s or the regression noises $\epsilon_i$'s are dependent with known covariance, the theorems and the corollaries below still hold with slight modifications {by simply plugging their covariance matrices into appropriate places}; if they are dependent with unknown covariance, extra assumptions on the covariance structure need to be made and new methodologies and theories should be developed. Assumption (ref) simply states that the regression is in the conventional low-dimensional fixed-design regime. Assumption (ref) is required for the consistency of the two-stage and SuperCENT, which essentially requires the signal-to-noise ratio (SNR) of the network and the gap between the leading and second singular values of the network to be large enough.
\jcM{
}
\jcM{
}
{Similarly, we derive the asymptotic distribution of the two-stage estimator in Theorem (ref).} Comparing the covariance matrices with those of the two-stage in Theorem (ref), all $\bSigma_\bu, \bSigma_\bv$ and $ \bSigma_\bbeta$ involve both $\sigma_a^2$ and $\sigma_y^2$ due to the simultaneous estimation, while $\bsigmats_\bu$ and $\bsigmats_\bv$ of the two-stage only involve $\sigma_a^2$ and $\bsigmats_\bbeta$ involves both. Specifically, $\bC_\bu$ and $\bC_\bv$ are functions of $(\sigma_a, \bD, \bU, \bV, \sigmay, \bX, \betau, \betav, \lambda)$, while the two-stage counterparts $\bCts_\bu$ and $\bCts_\bv$ only involve $(\sigma_a, \bD, \bU, \bV)$. Therefore, the difference between SuperCENT and the two-stage estimators of $\bu$ and $\bv$ lies in the tuning parameter $\lambda$ as well as the SNRs of the network and regression. Following Theorem (ref), Proposition (ref) provides the convergence rates of $\hu$ and $\hv$. We focus on the rank-one scenario where $\bA_0 = d\bu\bv^\top$ in model (ref) to provide clearer insights for understanding the difference between SuperCENT and the two-stage estimators.
{
}
\jcM{
}
In this section, we investigate the empirical performances, including the estimation and inference properties of the two-stage and SuperCENT estimators under various settings. Section (ref) describes the simulation setups and Section (ref) shows the results. Additional simulations, including a phase-transition experiment, are deferred to Supplement (ref).
We generate the network following model (ref). We consider the case of $r = 10$ where the leading singular value $d = 1$ and the non-leading ones as $d_2 = \ldots = d_r = 2^{-1}$. All entries of $\bU$ are first generated from i.i.d. $N(0,1)$ and $\bV = 0.5\bU+\bepsilon_{\bV}$ where $\bepsilon_{\bV}$ are generated from i.i.d. $N(0,1)$. We then apply Gram–Schmidt to ensure orthogonality between columns of $\bU$ and $\bV$, and finally rescale each column to have length $\sqrt{n}$. For the regression model (ref), the regression coefficients are $\bbeta_x = (1,3,5)^\top$, the design matrix $\X$ consists of a column of 1's and $p-1$ columns whose entries follow $N(0,1)$ independently.
\fontdimen3\font=0.15em For the properties of the estimators and inference, only the network SNR $\netsnr$ and the regression SNR $(\frac{\beta_u}{\sigma_y}, \frac{\beta_v}{\sigma_y})$ matter. Hence, we fix $n=2^{8},~d=1$, {$d_2=\ldots=d_r=2^{-1}$,} and $\betav=1$ and vary $\sigma_a, \sigma_y$, and $\beta_u$. To study the effect of the regression SNR, we consider $\sigmay \in 2^{-4, -2, 0 }$ and $\betau \in 2^{0, 2, 4}$, while ensuring {$\frac{\beta_u^2}{\sigma_y^2} \geq \frac{{\beta_v^2}}{\sigma_y^2}$.} As the network SNR is controlled by $\sigmaa$, we vary $\sigmaa\in 2^{ 0, 2 }$. {We study the effects of the non-leading singular values $d_2, \ldots, d_r$ in additional simulations in Supplement (ref).} \fontdimen3\font=\origiwstr
For estimation property, we compare the following procedures: 1. Two-stage; 2. $\supercentoracle$, which implements Algorithm (ref) with oracle $\lambda_0 = n\sigma_y^2/\sigma_a^2$ using the true $\sigmay,\sigmaa$ and serves as the benchmark; 3. $\supercentplugin$ is SuperCENT with estimated tuning parameter $\hat\lambda_0 = n(\hat\sigma_y^{ts})^2/(\hat\sigma_a^{ts})^2$, where $(\hat\sigma_y^{ts})^2 = \frac{1}{n-p-2}\|\hy^{ts}-\by\|_2^2$ and $(\hat\sigma_a^{ts})^2 = \frac{1}{n^2} \|\hAts - \bA_0\|_F^2$ are estimated from the two-stage procedure; and 4. $\supercentcv$ is SuperCENT with tuning parameter $\hat\lambda_{cv}$ chosen by cross-validation as in Algorithm (ref).
For inference property, we consider the following procedures to construct the confidence intervals (CIs) for the regression coefficient: 1. $\ts$-{adhoc}: $\hbeta^{ts} \pm z_{1-\alpha/2} \hat\sigma^{OLS}(\hbeta^{ts})$, where $z_{1-\alpha/2}$ denote the $(1-\alpha/2)$-quantile of the standard normal distribution, $\hbeta^{ts}$ is the two-stage estimate of $\beta$ and $\hat\sigma^{OLS}(\hbeta^{ts})$ is the standard error from OLS, assuming $\huts,\hvts$ are fixed predictors; 2. $\ts$-oracle: $\hbeta^{ts} \pm z_{1-\alpha/2} \sigma(\hbeta^{ts})$, where $\sigma(\hbeta^{ts})$ is the standard error of $\hbeta^{ts}$, whose mathematical expressions are given in (ref)-(ref) or (ref)-(ref) with the true parameters plugged in; 3. $\ts$-plugin: $\hbeta^{ts} \pm z_{1-\alpha/2} \hat\sigma(\hbeta^{ts})$, where $\hat\sigma(\hbeta^{ts})$ is the standard error of $\hbeta^{ts}$ by plugging all the two-stage estimators into (ref)-(ref) or (ref)-(ref); 4. $\supercentoracle$-oracle: $\hbetaoracle \pm z_{1-\alpha/2} \sigma(\hbetaoracle)$, where $\hbetaoracle$ is the estimate of $\beta$ by $\supercentoracle$ and $\sigma(\hbetaoracle)$ follows (ref) with the true parameters plugged in; and 5. $\supercentcv$: $\hbetacv \pm z_{1-\alpha/2} \hat\sigma(\hbetacv)$, where $\hbetacv$ is the estimate of $\beta$ by $\supercentcv$ and $\hat\sigma(\hbetacv)$ is obtained by plugging the $\supercentcv$ estimates into (ref). {Note that for the $\ts$-plugin and \textbf{$\supercentcv$}, $\hat\sigma(\hbeta^{ts})$ and $\hat\sigma(\hbetacv)$ involve estimation for $\bAp = \bUp \bDp \bVp^\top$. For the \textbf{$\ts$-plugin}, we plug in the estimate from SVD; for \textbf{$\supercentcv$}, we perform SVD on $\bA - \hd \hu \hv^\top$ and then plug in the estimates.} The experiments are repeated 500 times.
From the perspective of estimation, we compare the following metrics: the estimation accuracy for the centralities, the network, and the regression coefficients. Let $\bP$ denote the projection matrix. Figure (ref) shows the loss $l(\hu,{\bu}) = \| \bP_{\hu} - \bP_{\bu}\|^2_2$ and $\hbetau-\betau$, respectively, across different $\sigmaa$, $\sigmay$ and $\betau$ with $d = 1$ and $\betav = 1$. Losses such as $l(\hv,{\bv}) = \| \bP_{\hv} - \bP_{\bv}\|^2_2$, $l(\hA,{\bA_0}) = \|\hA - \bA_0\|_F^2/\|\bA_0\|_F^2$, $l(\hbetau,{\betau}) = (\hbetau - \betau)^2/\beta_u^2$, $l(\hbetav,{\betav}) = (\hbetav - \betav)^2/\beta_v^2$, and $\hbetav-\betav$ are given in Supplement (ref).
Figure (ref) shows the boxplot of $\log_{10}(l(\hu,\bu))$. The rows correspond to $\log_2(\sigmaa)$ and the columns correspond to $\log_2(\betau)$. For each panel, the x-axis is $\log_2(\sigmay)$ and the y-axis is $\log_{10}(l(\hu,\bu))$. The super-imposed red symbols show the theoretical rates of $\huts$ {and $\hu$ calculated from Theorems (ref) and (ref) respectively.} As expected, three SuperCENT-based methods estimate $\bu$ much more accurately than the two-stage procedure. In particular, the supervision effect of $(\bX, \by)$ is more pronounced when the noise of the outcome regression ($\sigmay$) is small, or when the signal of the outcome regression ($\betau$) is large, or when the network noise-to-signal ($\frac{\sigmaa}{d}=\sigmaa$) is large. The numerical comparison validates Remarks (ref) and (ref) on the theoretical comparison of the estimators. Comparing the three SuperCENT-based methods, $\supercentcv$ and $\supercentplugin$ are sometimes worse than the benchmark $\supercentoracle$, but still better than the two-stage. $\supercentplugin$ is typically comparable to or worse than $\supercentcv$, because $\supercentplugin$ fails to locate the optimal $\lambda_0$ due to inaccurate estimate of $\sigma_a$ and $\sigma_y$ from the two-stage procedure.
Figure (ref) shows $\hbetau - \betau$. With large $\sigmaa$ or large $\betau$, the two-stage estimates are inaccurate, while SuperCENT estimates remain accurate. { In particular, the two-stage estimates becomes more inaccurate as $\sigmaa$ or $\betau$ increases. The three SuperCENT-based methods all outperform the two-stage and $\supercentcv$ and $\supercentplugin$ are comparable with the benchmark $\supercentoracle$. }
From the perspective of inference property, Figure (ref) shows the empirical coverage probability (CP) and the average width of the 95% confidence interval for $\betau$ respectively. The CP and width for the centralities, the network, and $\betav$ are given in the Supplement.
Figure (ref) shows how the inaccurate estimation of $\betau$ by the two-stage further affects its confidence interval. Regarding empirical coverage, when $\betau$ is small (leftmost column), all methods are above the nominal level. As $\betau$ increases and $\sigmaa$ remains small (top right two panels), most methods (except for two-stage-oracle) remain valid, but for different reasons: the two SuperCENT-based methods remain valid due to the accurate estimation of both $\betau$ and the standard error, whereas two-stage and two-stage-{ad-hoc} remain valid mainly because they over-estimate ${\sigma_y^2}$, and this conservativeness masks the issue of the inaccurate estimation. Two-stage-oracle uses the true ${\sigma_y^2}$ and the issue of the inaccurate estimate cannot be hidden, hence the corresponding intervals undercover. When $\betau$ increases and $\sigmaa$ gets large too (bottom right panel), the over-estimation of ${\sigma_y^2}$ can no longer hide the issue of inaccurate estimation, causing all two-stage-related methods to become invalid. On the other hand, the empirical coverage of SuperCENT remains closer to the nominal level.
As for the width of $CI_\betau$, Figure (ref) shows that the confidence intervals by the SuperCENT-based methods have better coverage and are narrower than those by the two-stage methods. The improvement in width becomes more pronounced with larger $\betau$ and $\sigmaa$.
In this case study, we demonstrate that SuperCENT can provide a more accurate estimation of the centralities using the global trade network. This has a profound and lucrative implication on portfolio management because the centrality is closely related to currency risk premium, i.e., the excess return from holding foreign currency compared to the US dollar. We further show the advantage of SuperCENT over the two-stage in the inference of regression coefficients, and thus strengthens a related economic theory.
In international finance literature, economists have studied extensively the currency risk premium and remain puzzled by its driving forces. One recent theory, developed by richmond2019trade1 using a general equilibrium, shows that countries' positions in the trade network can explain the difference in currency premiums and countries that are central in the trade network exhibit lower currency risk premiums. This theory has two implications: {(i)} the regression coefficients for the centralities should be negative; and {(ii)} international investors can leverage and profit from a long-short strategy for foreign exchange by taking a long position in currencies of countries with low centralities and a short position in currencies of countries with high centralities. Therefore, if the centralities can be estimated accurately, one can yield a significant investment return based on the strategy.
Motivated by richmond2019trade1, we investigate how the global trade network drives the currency risk premium by regressing the currency risk premium on the centrality of the international trade network. To be specific, we consider a triplet of $\{\bA, \bX, \by\}$, where $\bA$ is the country-level trade network, $\by$ is the currency risk premium, and $\bX$ is the share of the world's GDP. Since all these quantities are not directly available, we compute them following richmond2019trade1. It is worth mentioning that the trade linkage in $\bA$ is defined as the trade volume normalized by the pair-wise total GDP, which represents the relative trade (export/import) intensity between two countries. We use a five-year moving average: when considering year $t$, the average is taken from year $t-4$ to year $t$. More details are provided in Supplement (ref). We focus on the period between 1999 and 2013 and include the 24 countries/regions whose exchange rates are available during this period.\footnote{The euro was first adopted in 1999. The bilateral trade data is available untill 2013.}\footnote{The list of country abbreviations is provided in Supplement (ref).} In Figure (ref), the dotted line shows the time series plot of the rank of the five-year moving average of risk premium from 2003 to 2012 for the 24 countries/regions.\footnote{We leave the last available year 2013 for the validation purposes.} In each year, we rank the 24 countries/regions' risk premiums from the largest to the smallest as the 1st to 24th. We show a circular plot to visualize the average trade volume (2003-2012) in Figure (ref).
\paragraph{Centrality estimation.} Since neither the two-stage nor SuperCENT is applicable for panel data, we will repeat the analysis for each year from 2003 to 2012. Besides the network and the response variable, we also include the GDP share as a predictor, which is defined as the percentage of country/region GDP among the total GDP of all available countries in the sample for that year. In summary, the unified framework is, for each $t$, \be {a_{ijt}} &=& d \cdotHub_{it} \times Authority_{jt} + e_{ijt},\nonumber\\ {y_{it}} &=& \alpha + {\beta_{ut}} \cdot Hub_{it} + {\beta_{vt}} \cdot Authority_{it} + {\beta_{xt}} \cdot GDP share_{it} + \epsilon_{it}.\nonumber \ee
{In Sections (ref) and (ref), we have demonstrated that the two-stage is problematic under large network noise. In this case study, the observational error of the network comes from two sources: GDPs and the trade volumes, because each entry of the observed network $a_{ijt}$ is defined as the trade volume normalized by their GDPs. The accounting of GDP has been a challenge in macroeconomics landefeld2008taking. For the trade volume, measurement errors are mostly due to (i) underground or illegal import and export; (ii) excluding service trade; (iii) trade cost like transportation or taxes lipsey20091. Consequently, the observed trade network can be very noisy and the two-stage will perform badly. }
\plotfig[.94]{\trade}{plot/hub_gap5}{ Time series of ranking of risk premium in descending order and ranking of hub centrality estimated by two-stage and SuperCENT in ascending order from 2003 to 2012. The vertical dashed line indicates 2008, the year of the financial crisis. }{fig:trade-hub}
On the other hand, SuperCENT can significantly improve over the two-stage when the network noise is large. In what follows, we focus on $\supercentcv$ using $10$-fold cross-validation. We will refer to $\supercentcv$ as SuperCENT for simplicity and use the superscript $sc$ for all the $\supercentcv$-related estimates. We determine the signs of the centrality estimates by the empirical rule described at the end of Section (ref). Figure (ref) shows the time series plots of the ranking of the hub centrality estimated by two-stage and SuperCENT for the 24 countries/regions, together with the ranking of the currency risk premium. Figure (ref) is for the authority centrality. {We rank the centrality in ascending order and the risk premium in descending order. Based on the negative relationship between centralities and risk premium established in richmond2019trade1, the closer the trends of rankings between centralities and risk premium are, the better the centralities capture the time variation in the risk premium.} The centrality estimated by the two-stage procedure is relatively more stable over time compared to SuperCENT. This is because SuperCENT incorporates information of both the GDP share and the currency risk premium, which is more volatile than the trade network itself. Asian trade hubs such as Hong Kong (HKG) and Singapore (SGP) are the most central; while countries like South Africa (ZAF) and New Zealand (NZL) are peripheral. Comparing the ranking of risk premium, the time variation is not reflected in the centrality estimated by the two-stage procedure, while it is well captured by SuperCENT. For the 2008 financial crisis, the SuperCENT centralities fluctuate together with the risk premium while the two-stage centralities mostly {remain unchanged}.
\plotfig[.9]{\trade}{plot/excess_return_next_gap5_top3}{Time series of the next-year return {from 2004 to 2013} based on a strategy {that takes a long position on the currencies with the lowest 3 centralities and a short position on the currencies with the highest 3 centralities {estimated from 2003 to 2012 respectively}.} }{fig:excess-return}
To emphasize the importance of accurate centrality estimation for portfolio management, we examine whether a long-short strategy based on SuperCENT's estimated centrality can significantly boost investment performance based on two-stage. For either two-stage or SuperCENT, we take a long position on the currencies with the lowest 3 centralities (bottom 10%) and a short position on the currencies with the highest 3 centralities (top 10%). We obtain a return based on the estimated centrality of the period between year $t-4$ and $t$. {Similarly, we include a naive long-short strategy based on the return of year $t-1$ as a baseline.} Figure (ref) shows the year $t+1$ return based on this strategy. {The centrality-based portfolios both outperform the naive strategy.} The return based on the centrality estimated by SuperCENT is much higher than that of the two-stage procedure. Table (ref) shows the 10-year average {annualized return and Sharpe ratio sharpe1994sharpe} based on this strategy with the top and bottom 3, 4, and 5 currencies, respectively. The 10-year average return based on SuperCENT centralities {doubles} that of the two-stage procedure. SuperCENT achieves Sharpe ratios ranging from 0.27 to 0.39, compared to the Sharpe ratios of 0.28 for the Dow Jones, 0.42 for the S&P 500, and 0.39 for the NASDAQ over the sample period (2004-2013).
{
}
\paragraph{Inference of regression.} We further demonstrate the superiority of SuperCENT in inference. {Again since our method is not directly applicable to longitudinal data, we take the 10-year average of trade volume and GDP to construct a 10-year trade network and GDP share. Similarly, we take the 10-year average of risk premium as the response.}
To better understand the behavior of the two-stage and SuperCENT estimators and how much improvement SuperCENT can potentially achieve, we compare the noise-to-signal ratio $\netsnr$ of the trade network and the SNR of the regression. Since both quantities are unknown, we estimate using results from SuperCENT. For the noise-to-signal ratio, $\hat\netsnr^{\lambdacv} = 0.36 \approx 2^{-1.5}$, which is larger than $\netsnr=2^{-8}$ in the simulation where the accuracy of the two-stage estimators is already low. For the SNR of the regression: $(\hbetau^\lambdacv/\widehat{\sigma}_y^{\lambdacv})^2 = 1.8 \times 10^{7} \approx 2^{24}$ and $(\hbetav^\lambdacv/\widehat{\sigma}_y^{\lambdacv})^2 = 6.1\times 10^{5} \approx 2^{19}$. Compared with the simulation settings where $\netsnr = 2^{-4}$, $\beta_u^2/\sigma_y^2 \leq 2^{16} $ and $\beta_v^2/\sigma_y^2 \leq 2^{8}$, We expect SuperCENT to significantly outperform the two-stage method in both the estimation and inference of $\betau$, due to the large value of $(\hbetau^\lambdacv/\widehat{\sigma}_y^{\lambdacv})^2$ and the fact that $|\hbetau^\lambdacv| \gg |\hbetav^\lambdacv|$ under a relatively large $\hat\netsnr^{\lambdacv}$, while the improvement of $\betav$ is less pronounced.
\tradetbl{1999}{10}{gdp}
Table (ref) shows the coefficient estimation, the standard error, and the significant level for the two-stage-adhoc, two-stage, and $\supercent$, respectively. { The standard errors of two-stage and $\supercent$ are based on the trade network being rank-one as we tested using the rank inference by han2023universal in Supplement (ref). } For the hub centrality $\betau$, (i) the estimate from the two-stage methods is $-0.0011$, while the estimate from $\supercent$ is $-0.0020$, {which is consistent with the inaccuracy we observed in the simulation;} {(ii) the standard errors from the two-stage methods are close to $0.0007$, much larger than $0.0001$ from $\supercent$, which reinforces the problem of overestimation of $\sigma_y^2$ in two-stage; } (iii) {the above two facts combined} make the confidence intervals by two-stage-adhoc and two-stage unnecessarily wide, yet still invalid: {consequently the hub centrality $\betau$ is barely significant at level $0.1$ using two-stage and is insignificant using two-stage-adhoc;} (iv) the two facts in (i) and (ii) also lead to a valid but narrower confidence interval for $\supercent$, making the hub centrality a significant factor {at} level $0.01$ for the currency risk premium; and (v) {conclusions drawn from the} two-stage-adhoc and two-stage methods contradict the theory in richmond2019trade1, while $\supercent$ supports the theory. Other regression coefficients' significance can be also explained by Remark (ref); the details are given in Supplement (ref).
Motivated by the rising use of centrality in empirical literature, we examined centrality estimation and inference on a noisy network (ref) as well as network effect through the centralities in the subsequent network regression (ref). We proposed a unified framework that incorporates the network generation model and the network regression model to achieve both goals. { Under the unified framework, we showed that the properties of the commonly used two-stage procedure and that it could yield inaccurate centrality estimates and regression coefficient estimates, as well as invalid inference when the noise-to-signal ratio of the network is large. } We proposed SuperCENT which incorporates the two models and simultaneously estimates the centralities and the effects of the centralities on the outcome. We further derived the convergence rate and the distribution of the SuperCENT {estimator} and provided valid confidence intervals for all the parameters of interest. { We showed that SuperCENT dominates the two-stage universally and improves over the two-stage in terms of centrality estimation, regression coefficient estimations, and inference. } The theoretical results are corroborated with extensive simulations and a real case study in predicting currency risk premiums from the global trade network.
The {unified framework and SuperCENT methodology} can be extended in multiple directions. One can consider a generalized linear model for the outcome model and \jcM{a generalized network model for networks with noncontinuous edges via link functions to generalize SuperCENT. } In the case when only a subset of covariates and outcomes are observed, semi-supervised SuperCENT can be developed. In the case when the network is partially observed, we can perform matrix completion with supervision. SuperCENT can also be extended to a longitudinal model with additional assumptions by using techniques from tensor decomposition as well as functional data analysis to obtain centralities that are smooth over time. For ultra-high-dimensional problems, sparsity can be imposed on centralities due to the existence of abundant peripheral nodes.
\if11
Shen's research is supported in part by Hong Kong CRF C7162-20GF, the Ministry of Science and Technology Major Project of China 2017YFC1310903, University of Hong Kong (HKU) Stanley Ho Alumni Challenge Fund, and HKU BRC Grant. Yang's research is supported in part by NSF grant IIS-1741390, Hong Kong GRF 17301620 and CRF C7162-20GF. Zhao's research is supported in part by Wharton Global Initiative Fund. Zhu's research is supported in part by Tsinghua University Initiative Scientific Research Program and Tsinghua University School of Economics and Management Research Grant. \fi
{\linespread{1}\selectfont
{.6em plus 0.3ex}
\fontdimen3\font=0em \fontdimen2\font=0.7ex
}
{ \linespread{1.2}\selectfont