EconBase
← Back to paper

Tractable Unified Skew-t Distribution and Copula for Heterogeneous Asymmetries

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

91,143 characters · 33 sections · 63 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

myblue Tractable Unified Skew-t Distribution and Copula for Heterogeneous Asymmetries

{0.15cm} {0.15cm}

\pagestyle{empty}

titlepage{ Lin Deng is a PhD student and Michael Smith is Professor of Management (Econometrics), both at the Melbourne Business School, University of Melbourne, Australia. Worapree Maneesoothorn is Associate Professor at the Department of Econometrics and Statistics at Monash University, Australia. Correspondence should be directed to Michael Smith at {\tt [email removed]}. \\ Acknowledgments: This research was supported by The University of Melbourne’s Research Computing Services and the Petascale Campus Initiative. Michael Stanley Smith research has been partially supported by the Australian Research Council (ARC) Discovery Project grant DP250101069, and that of Worapree Maneesoonthorn by ARC Discovery Project Grant DP200101414.}\\ \begin{center} \\ \\ { Abstract \\} \end{center} \onehalfspacing { Multivariate distributions that allow for asymmetry and heavy tails are important building blocks in many econometric and statistical models. The Unified Skew-t (UST) is a promising choice because it is both scalable and allows for a high level of flexibility in the asymmetry in the distribution. However, it suffers from parameter identification and computational hurdles that have to date inhibited its use for modeling data. In this paper we propose a new tractable variant of the unified skew-t (TrUST) distribution that addresses both challenges. Moreover, the copula of this distribution is shown to also be tractable, while allowing for greater heterogeneity in asymmetric dependence over variable pairs than the popular skew-t copula. We show how Bayesian posterior inference for both the distribution and its copula can be computed using an extended likelihood derived from a generative representation of the distribution. The efficacy of this Bayesian method, and the enhanced flexibility of both the TrUST distribution and its implicit copula, is first demonstrated using simulated data. Applications of the TrUST distribution to highly skewed regional Australian electricity prices, and the TrUST copula to intraday U.S. equity returns, demonstrate how our proposed distribution and its copula can provide substantial increases in accuracy over the popular skew-t and its copula in practice. } {\bf Keywords}: Asymmetric Dependence, Bayesian Data Augmentation, Electricity Prices, Implicit Copulas, Intraday Equity Returns, Markov chain Monte Carlo, Skew-Elliptical Distributions.

\pagestyle{plain} \setcounter{equation}{0}

Introduction

Skew-elliptical distributions are models of multivariate asymmetry that scale well. While there are different variants (for some of these see genton2004skew) those formed by conditioning on a latent truncated variable as suggested by branco2001general are prominent. This approach is often called “hidden truncation” and is used to form the skew-t distribution of Azzalini_Capitanio_2003 (AC hereafter). The AC skew-t distribution is a particularly popular choice in applied data analysis, with diverse applications in renewable energy Hering01032010, actuarial studies eling2012, macroeconomic modeling adrian2019vulnerable, atmospheric science morris2017space, and in flow cytometric analysis PyneetalPNAS2009,fruhwirth2010bayesian among others. However, in $d$ dimensions skew-elliptical distributions typically have only $d$ parameters to control asymmetry across $d(d-1)/2$ variable pairs, which can limit their effectiveness for modeling asymmetry. Unified skew-elliptical (USE) \footnote{arellano2006unification use the abbreviation SUN for the Unified Skew-Normal distribution, later adopting SUE for the Unified Skew-Elliptical family and SUT for the Unified Skew-t. In this paper, we break with this convention and adopt the word-ordered acronyms USN, USE and UST for these three distributions.} distributions arellano2010multivariate are an extension that introduces more latent variables and parameters to allow for a much more flexible form of asymmetry. In particular, the unified skew-t (UST) discussed by wang2024multivariate is a promising generalization of the impactful AC skew-t distribution. However, despite the strong potential of USE distributions, they have not been adopted for data analysis because (i) the parameters are unidentified wang2023non, and (ii) likelihood evaluation and optimization is computationally difficult in even low dimensions.

The current paper addresses these problems by specifying a novel subclass of the USE that is tractable. For this proposed subclass, it is shown that the USE parameters are identified under an ordering of the eigenvalues of the conditional scale matrix of the latent variables. The special case of a tractable UST distribution, which we label a TrUST distribution, is considered in detail. A generative representation is given for the proposed TrUST distribution that is used to define an extended likelihood that is easier to evaluate than the regular likelihood. From this, a Bayesian augmented posterior can be defined and computed using a proposed Markov chain Monte Carlo (MCMC) sampler that is both fast and efficient, and for which the identifying eigenvalue constraint is easy to impose. The proposed TrUST distribution nests the popular AC skew-t as a special case, but generalizes it to allow for richer multivariate asymmetry. An application of the TrUST distribution to both simulated data and highly skewed regional daily Australian electricity prices shows it captures asymmetry more accurately than the popular AC skew-t, and that this improves the quality of fit considerably.

However, the main advantage of the proposed TrUST distribution is that its implicit copula is also tractable and has parameters that are identified. Every multivariate continuous distribution has a unique implicit copula; see smith2021implicit for an introduction to this class of copulas. Currently, implicit copulas of differing skew-t distributions are used to capture the strong asymmetric dependence often found in financial data; see smith_gan_kohn_2010, Christoffersen_Errunza_Jacobs_Langlois_2012, creal2015high, lucas2017, opschoor2021, oh_patton_2023 and deng2024large for examples. Of these, deng2024large show that the implicit copula of the AC skew-t distribution allows for the greatest level of asymmetric dependence. But the implicit copula of the TrUST distribution (hereafter the TrUST copula) allows for a much greater level of heterogeneity in asymmetric dependencies over variable pairs than the AC skew-t copula. We show here that by doing so, the TrUST copula can better capture the dependence between financial returns data. To estimate the TrUST copula parameters a Bayesian augmented posterior is specified using a modification of the extended likelihood of the TrUST distribution, which is then evaluated using an MCMC sampler. Expressions for the Kendall and Spearman rank correlations for this TrUST copula are also derived. As far as we are aware, ours is the first paper to construct the implicit copula of any USE distribution and show how to compute statistical inference for its parameters.

deng2024large use the implicit copula of the AC skew-t distribution to capture dependence between intraday equity returns. In our main application, we extend their analysis to show that the TrUST copula is more effective at capturing such dependence than the AC skew-t copula. This is done for five large equities and an equity market volatility index (VIX) using data collected both before and during the COVID-19 pandemic. The performance of the TrUST copula is compared with that of both the symmetric Student-t copula and the AC skew-t copula using the Deviance Information Criterion (DIC) and cumulative log-score to assess accuracy. The results indicate that the TrUST copula better captures heterogeneity in asymmetric dependence across variables. High levels of heterogeneity is also found by le2021covid and ando2022quantile using network models of tail dependencies.

The rest of this paper is organized as follows. Section (ref) briefly introduces the USE distribution and the special case of the UST, after which our proposed tractable USE variant is outlined. Section (ref) considers the case of the TrUST distribution in detail. Section (ref) outlines estimation methodology for the TrUST distribution, and illustrates its efficacy through a simulated example and application to electricity prices. Section (ref) presents the TrUST copula, derivation of its rank correlations, Bayesian inference and a simulation example. Section (ref) shows how our TrUST copula captures heterogeneity in asymmetric dependence between high-frequency financial variables; Section (ref) concludes. Some extra results are given in the Appendix, while an Online Appendix contains additional information, figures, tables and all proofs.

Unified Skew-Elliptical Distributions

In this section, the unified skew-elliptical (USE) distribution is introduced, including the special case of the unified skew-t (UST). Our tractable variant of the USE distribution is then outlined in detail. We consider the standardized case (i.e. with zero mean and unit scale) because this is used in forming the implicit copula in Section (ref). Generalization of the distribution by multiplying the random vector by a diagonal scaling matrix and adding a location parameter is straightforward.

Unified Skew-Elliptical Distribution

arellano2006unification extend the popular skew-elliptical distribution to allow for additional skew parameters. It can be specified for the standardized case as follows. Let $\text{\boldmath$X$}\in \text{$\mathds{R}$}^d$ and $\text{\boldmath$L$} \in \text{$\mathds{R}$}^q$, for $q\geq 1$, be jointly (radially symmetric) elliptically distributed as

equation[equation omitted — 273 chars of source]

where $R$ is a $(d+q) \times (d+q)$ correlation matrix partitioned to be consistent with $(\text{\boldmath$X$}^\top,\text{\boldmath$L$}^\top)^\top$, $\mathbf{0}$ is a zero mean vector\footnote{Throughout this paper we use $\mathbf{0}$ to denote a conformable vector or matrix of zeros without explicitly denoting its dimension.}, and $g: [0, \infty) \to [0, \infty)$ is the density generator function; see Fang_Kotz_Ng_1990 for specification of an elliptical distribution. Then the USE distribution can be defined as

equation[equation omitted — 154 chars of source]

where “$\overset{\mathrm{d}}{=}$” denotes equality in distribution, and the inequality $\text{\boldmath$L$}>\mathbf{0}$ holds element-wise. The vector $\text{\boldmath$L$}$ are latent variables and the specification at (ref) is often called “hidden truncation” arnold2000hidden,arnold2004elliptical. The $(q\times d)$ matrix $\Delta$ contains parameters that control the level and direction of skew, where $q$ is typically low, such as $q\in\{1,2,3,4\}$.

Let $f_{\mathrm{EL}}(\text{\boldmath$y$};V,g)$ and $F_{\mathrm{EL}}(\text{\boldmath$y$};V,g)$ denote the density and distribution functions, respectively, of a zero-mean elliptically distributed random variable $\text{\boldmath$Y$}\sim \mathrm{EL} {\left(\mathbf{0}, V, g\right)}$. Also, for the partition $\text{\boldmath$Y$}=(\text{\boldmath$Y$}_1^\top,\text{\boldmath$Y$}_2^\top)^\top$, the conditional $\text{\boldmath$Y$}_1 \mid \text{\boldmath$Y$}_2$ is elliptically distributed Fang_Kotz_Ng_1990 with density generator denoted here as \( g_{\text{\boldmath$Y$}_2} \). Then, the density of $\text{\boldmath$Z$}$ is obtained by using Bayes theorem to evaluate the conditional density on the righthand side of (ref), giving

equation[equation omitted — 383 chars of source]

If a random variable $\text{\boldmath$Z$}$ has this density, then we write $\text{\boldmath$Z$} \sim \mathrm{USE}_{q}{\left( \Omega, \Delta, \Sigma, g\right)} $.

The number of elements in $\Delta$ increases linearly with $q$. In this way, $q$ plays a role for multivariate asymmetry that is analogous to that of the number of factors in a traditional factor model for the covariance. When \(q = 1\), the USE distribution reduces to the skew-elliptical distribution introduced by branco2001general and Azzalini_Capitanio_2003. In this case there is only a single skew parameter for each dimension, so that \(\Delta \equiv \text{\boldmath$\delta$}^\top = {\left(\delta_1,\ldots,\delta_d\right)}^\top \). Because $\text{\boldmath$\delta$}$ is constrained, for computing inference it is common to transform it to an unconstrained parameterization via the one-to-one transformation $\text{\boldmath$\alpha$} = (1 - \text{\boldmath$\delta$}^{\top} \Omega^{-1} \text{\boldmath$\delta$})^{-1 / 2} \Omega^{-1} \text{\boldmath$\delta$}$ with inverse $\text{\boldmath$\delta$} = (1+\text{\boldmath$\alpha$}^{\top} \Omega \text{\boldmath$\alpha$})^{-1 / 2} \Omega \text{\boldmath$\alpha$}$.

Unified Skew-t Distribution

The elliptical distribution with greatest applied potential is the t distribution, which has density generator function $g(x)=(1+x/\nu)^{-(\nu+d)/2}$ that introduces the parameter $\nu>0$. Assuming $(\text{\boldmath$X$}^\top,\text{\boldmath$L$}^\top)^\top$ is multivariate t with location zero, scale matrix $R$ and $\nu$ degrees of freedom, the density at (ref) is given by

equation[equation omitted — 345 chars of source]

Here, $Q(\text{\boldmath$z$}) = \text{\boldmath$z$}^\top\Omega^{-1}\text{\boldmath$z$}$, while \( t_d{\left(\text{\boldmath$y$}; V, \upsilon\right)} \) and \( T_d(\text{\boldmath$y$}; V, \upsilon) \) denote the density and distribution functions, respectively, of a zero mean $d$-dimensional multivariate t random variable with scale matrix $V$ and degrees of freedom \(\upsilon\) evaluated at $\text{\boldmath$y$}$. If a random variable $\text{\boldmath$Z$}$ has density at (ref), then we write $\text{\boldmath$Z$} \sim \mathrm{UST}_{q}{\left(\Omega, \Delta, \Sigma, \nu\right)} $.

Three special cases of the UST distribution are: (i) when $\nu\rightarrow \infty$, the UST converges to the unified skew-normal of arellano2006unification; (ii) when $\Delta$ is a matrix with all elements equal to zero, then $\text{\boldmath$Z$}$ is distributed symmetric multivariate t with location zero, scale matrix $\Omega$ and $\nu$ degrees of freedom; and (iii) when \(q=1\), the UST distribution is the AC skew-t distribution. wang2024multivariate provide a comprehensive overview of the UST distribution and its properties.

Tractable Unified Skew-Elliptical Distribution

The novel subclass of the USE distribution is now outlined, that we call here a Tractable Unified Skew-Elliptical (TrUSE) distribution. It employs the following assumption of conditional independence in the latent variables.

assumption[Conditional Independence] Let $(\text{\boldmath$X$}^\top,\text{\boldmath$L$}^\top)^\top$ follow the joint elliptical distribution at (ref). Then assume each component of $\text{\boldmath$L$}$ is independent, conditional on $\text{\boldmath$X$}$, so that $L_i \perp\!\!\!\!\perp L_j | \text{\boldmath$X$} $, for all ${\left(i,j\right)} \subset {\left\{1,\ldots,q\right\}}$ and $q > 1$.

Under this assumption, the scale matrix $\Sigma$ is a deterministic function of $\{\Omega,\Delta\}$, and the density at (ref) is simplified, as summarized by the following two lemmas.

lemmaIf Assumption (ref) holds, then the scale matrix of the marginal elliptical distribution of $\text{\boldmath$L$}$ in (ref) is given by $\Sigma = I_q + \left(M-\operatorname{diag}(M)\right)$, with $M=\Delta \Omega^{-1}\Delta^\top$.
lemmaIf Assumption (ref) holds, then \[H=\Sigma-\Delta \Omega^{-1} \Delta^\top = \mbox{diag}\left((1 - \text{\boldmath$\delta$}_1^\top \Omega^{-1} \text{\boldmath$\delta$}_1),\ldots,(1 - \text{\boldmath$\delta$}_q^\top \Omega^{-1} \text{\boldmath$\delta$}_q)\right)\,,\] is a diagonal matrix, and the distribution function of $\text{\boldmath$L$}|\text{\boldmath$X$}$ is given by \[ F_{\mathrm{EL}}{\left(\Delta \Omega^{-1} \text{\boldmath$z$}; \Sigma - \Delta \Omega^{-1} \Delta^\top, g_{\text{\boldmath$X$}} \right)}= \prod_{k=1}^{q} F_{\mathrm{EL}} {\left(\text{\boldmath$\alpha$}_k^\top \text{\boldmath$z$};1, g_\text{\boldmath$X$}\right)} \] where $\Delta^\top = [\text{\boldmath$\delta$}_1|\text{\boldmath$\delta$}_2| \cdots| \text{\boldmath$\delta$}_q]$ and $\text{\boldmath$\alpha$}_k = {\left(1 - \text{\boldmath$\delta$}_k^\top \Omega^{-1} \text{\boldmath$\delta$}_k\right)}^{-1/2} \Omega^{-1} \text{\boldmath$\delta$}_k$ for $k=1,\ldots,q$.

Proofs of both lemmas are found in Appendix (ref), and together they can be used to define our proposed tractable USE distribution as follows.

definition[TrUSE Distribution] If $\text{\boldmath$Z$} \sim \mathrm{USE}_{q}{\left( \Omega, \Delta, \Sigma, g \right)}$ and Assumption (ref) holds, then $\text{\boldmath$Z$}$ is said to be distributed Tractable Unified Skew-Elliptical (TrUSE) with density function \begin{equation} f_{\mathrm{TrUSE}, q}{\left(\boldmath$z$; \Omega, A, g \right)} = f_{\mathrm{EL}}{\left(\boldmath$z$; \Omega, g\right)} \frac{ \prod_{k=1}^{q} F_{\mathrm{EL}} {\left(\boldmath$\alpha$_k^\top \boldmath$z$;1, g_\boldmath$X$ \right)} }{ F_{\mathrm{EL}} {\left(\mathbf{0}; \Sigma, g \right)} }. \end{equation} where $\Sigma$ is given in Lemma (ref), $\text{\boldmath$\alpha$}_k$ is given in Lemma (ref), $g$ is the density generator function at (ref), and $g_\text{\boldmath$X$}$ is the density generator function of the elliptical distribution of $\text{\boldmath$L$}|\text{\boldmath$X$}$. We write $\text{\boldmath$Z$} \sim \mathrm{TrUSE}_{q}{\left(\Omega, A, g \right)}$ where $A=[\text{\boldmath$\alpha$}_1|\text{\boldmath$\alpha$}_2|\cdots|\text{\boldmath$\alpha$}_q]$.

Parameterization of the TrUSE distribution in terms of $A$ is attractive because each $\text{\boldmath$\alpha$}_k\in \mathbb{R}^d$ is unbounded given $\{\Omega,\text{\boldmath$\alpha$}_{j\neq k}\}$, whereas $\text{\boldmath$\delta$}_k$ has complex nonlinear bounds given $\{\Omega,\text{\boldmath$\delta$}_{j\neq k}\}$. We show later that this property is useful for constructing MCMC schemes to generate $\text{\boldmath$\alpha$}_k$, rather than $\text{\boldmath$\delta$}_k$, from its posterior.

The TrUSE distribution is more tractable than the general USE for three reasons. First, it simplifies the parameter space because $\Sigma$ is a deterministic function of $\{\Omega,A\}$, whereas for a general USE the off-diagonal elements of $\Sigma$ have complex nonlinear bounds given $\{\Omega,\Delta\}$. Second, it is more computationally tractable to compute inference because (i) the density at (ref) is simplified, and (ii) efficient MCMC sampling is possible from the augmented posterior constructed using an extended likelihood based on the joint of $(\text{\boldmath$X$},\text{\boldmath$L$})$. Third, a parameter identification constraint that is straightforward to impose exists as outlined below. In contrast, it is hard to compute statistical inference for the unconstrained parameters of the general USE distribution. The tractability of TrUST distribution is illustrated in Section (ref).

Adopting Assumption (ref) alone is insufficient to identify the USE distribution under permutation of $\text{\boldmath$L$}$. wang2023non discuss this problem in detail, which is characterized by the lemma below.

lemma[Permutation Un-identification] Let \(G(q) = \{\pi : \pi \text{ is a bijection from } \{1, \ldots, q\} \text{ to itself}\}\) represents the set of all possible orderings of the indices \(\{1, \ldots, q\}\). Then, the distribution of $ \text{\boldmath$Z$} \stackrel{\mathrm{d}}{=} {\left(\text{\boldmath$X$} | \text{\boldmath$L$} > \mathbf{0}\right)} \sim \mathrm{TrUSE}{\left( \Omega, A, g\right)}$ is unidentified under permutation of latent variable $\text{\boldmath$L$}$. That is, for every permutation $\pi \in G(q) $, $\text{\boldmath$Z$} \overset{\mathrm{d}}{=} \text{\boldmath$Z$}_{\pi}$, where $ \text{\boldmath$Z$}_{\pi} \overset{\mathrm{d}}{=} \text{\boldmath$X$} | \text{\boldmath$L$}_\pi > \mathbf{0}$ and $\text{\boldmath$L$}_\pi = {\left(L_{\pi(1)},\ldots,L_{\pi(q)}\right)}^\top$.

To identify the parameters a constraint based on an ordering of the eigenvalues of the scale matrix $H$ of the conditional distribution of $\text{\boldmath$L$}|\text{\boldmath$X$}$ is used.

assumption[Latent Permutation Constraint] Let \(\text{\boldmath$Z$} \overset{\mathrm{d}}{=} {\left(\text{\boldmath$X$} | \text{\boldmath$L$} > \mathbf{0}\right)}\) follow a USE distribution, and $H =\Sigma-\Delta \Omega^{-1}\Delta^\top$ be the scale matrix of the conditional distribution of $\text{\boldmath$L$}|\text{\boldmath$X$}$ with eigenvalues $\lambda_1,\ldots,\lambda_q$. Then we assume the permutation \(\pi^*\in G(q)\) of the latent vector $\text{\boldmath$L$}$ and associated columns of $\Delta^\top$, solves the problem \begin{equation*} \pi^* = \operatorname*{{arg\,max}}_{\pi \in G(q)} \sum_{k=1}^q \pi(k) \lambda_{\pi(k)}. \end{equation*}

This assumption, when applied to the TrUSE distribution under Assumption (ref), results in a permutation of the latent vector $\text{\boldmath$L$}$ that identifies the columns of \(A\) and \(\Delta^\top\), as below.

theorem[Latent Permutation Identification] Under Assumption (ref) , $H=\Sigma-\Delta \Omega^{-1} \Delta^\top=\mbox{diag}(h_1,\ldots,h_q)$ is a diagonal matrix, so that $\lambda_k=h_k$ for $k=1,\ldots,q$. Under the additional Assumption (ref), the permutation of $\text{\boldmath$L$}$ (and therefore also the columns of $A$ and $\Delta^\top$) in the TrUSE distribution is based on the optimal ordering: \begin{equation*} \pi^*: 0 \leq h_{\pi^*(1)} \leq \cdots \leq h_{\pi^*(q)} \leq 1, \end{equation*} where \( h_{\pi^*(k)} = 1 - \text{\boldmath$\delta$}_{\pi^*(k)}^\top \Omega^{-1} \text{\boldmath$\delta$}_{\pi^*(k)}\), for \(k = 1, \ldots, q\). Moreover, when these inequalities are strict, the ordering is unique.

In practice, because the posterior distribution of $(h_1,\ldots,h_q)^\top$ is continuous, the inequalities in Theorem (ref) are strict, so that the ordering is unique. In the remainder of the paper Assumption (ref) is adopted to identify the parameters in the TrUSE (and TrUST) distributions, so that $\Delta, A, \Sigma$ are all defined with respect to the unique ordering $\pi^*$. Imposing this constraint is straightforward within an MCMC scheme as discussed in Section (ref).

We finish this section with some comments on the appropriateness of the TrUSE distribution. We stress that Assumption (ref) is on the distribution of $\bm{L}$, not on that of $\bm{X}$, so that the TrUSE distribution can capture the same degree of rank correlation as the USE distribution. Moreover, Assumptions (ref) and (ref) are identifying assumptions that are necessary for the USE distribution to be well-defined for data analysis. In comparison, wang2024multivariate suggest (but do not implement) some other ways to identify the parameters in a UST distribution that involve much stronger restrictions on $\Delta$ and/or $\Omega$. However, these reduce substantially the ability to capture variation in pairwise correlations and/or asymmetry across variable pairs, which is the key advantage of the UST over skew-t distributions. wang2023non suggest identification is secured by fixing the order of the eigenvalues of \(\Sigma\), but this condition is insufficient when \(q\geq 2\) because interchanging rows of \(\Delta\) does not alter the order of the eigenvalues of $\Sigma$.

Tractable Unified Skew-t Distribution

The special case of the tractable unified skew-t (TrUST) distribution that is the focus of our paper is now outlined in detail.

Joint Distribution

The TrUST distribution results from selecting a t distribution with $\nu$ degrees of freedom as the choice of elliptical distribution in Section (ref). In this case, the density at Definition (ref) is

equation[equation omitted — 353 chars of source]

If a random vector $\text{\boldmath$Z$}$ has the above density, then we write $\text{\boldmath$Z$} \sim \mathrm{TrUST}_{q}{\left(\Omega, A, \nu\right)}$. If $\nu \rightarrow \infty$, then this reduces to the tractable unified skew-normal (TrUSN) distribution, with density

equation*[equation* omitted — 262 chars of source]

where $\phi_d{\left(\text{\boldmath$x$}; \Omega\right)}$ and $\Phi_d{\left(\text{\boldmath$x$}; \Omega\right)}$ are density and distribution functions, respectively, for a $N{\left(\mathbf{0}, \Omega\right)}$ distribution. If a random vector $\text{\boldmath$Z$}$ has the above density, then we write $\text{\boldmath$Z$} \sim \mathrm{TrUSN}_{q}{\left(\Omega,A\right)}$. If $q=1$ then the TrUST, UST and AC skew-t all coincide.

Figure (ref) presents the bivariate density contours of the TrUST distribution, when $d=2$, $q=2$, $\nu = 5$ and the first skew parameter vector is \(\text{\boldmath$\alpha$}_1 = (5, 5)^\top\). Each row corresponds to a different correlation parameter \(\omega \in {\left\{-0.5, 0, 0.5\right\}}\), ordered from top to bottom. The columns correspond to \(\text{\boldmath$\alpha$}_2 \in {\left\{ (5, 5)^\top, (0, 5)^\top, (-5, 5)^\top \right\}}\) and show the impact of varying the second skew parameter vector.

figure[figure omitted — 623 chars of source]

Marginal Distribution

Because the UST distribution is closed under marginalization, the marginals of a TrUST distribution are UST with parameters derived from those of the joint as below.

lemmaLet $\text{\boldmath$Z$} \sim \mathrm{TrUST}_{q}{\left(\Omega, A, \nu\right)}$ with $\Delta$ computed from $A$, and $\Sigma$ computed from $\{\Delta,\Omega\}$ as in Lemma (ref). If \( J \subset {\left\{1,\ldots,d\right\}} \) is a subset of $d_J<d$ indices, then the marginal distribution of the random vector $\text{\boldmath$Z$}_J = (Z_{J(1)},Z_{J(2)},\ldots,Z_{J(d_J)})$ is \( \text{\boldmath$Z$}_J \sim \mathrm{UST}_{ q} {\left(\Omega_J, \Delta_J, \Sigma, \nu\right)} \), where $\Delta_{J} = {\left(\Delta_{J(1)}, \ldots, \Delta_{J(d_J)}\right)}$ contains the columns of $\Delta$ indexed by $J$, and $\Omega_J=\{\omega_{ij}\}_{i,j\in J}$ is the sub-matrix of $\Omega=\{\omega_{ij}\}$ comprising the elements indexed by set $J$.

The case where $d_J=1$ is of particular importance when constructing the implicit copula of a TrUST distribution in Section (ref). Here, the marginal for the $j$-th element is $Z_j \sim \mathrm{UST}_{q}( 1, \Delta_{j}, \Sigma, \nu)$ for $j \in {\left\{1,\ldots,d\right\}}$ with density function

equation[equation omitted — 299 chars of source]

where $\Delta_j$ is the $j$th column of $\Delta$ and $Q(z_j)=z_j^2$. The distribution and quantile functions are computed numerically from (ref) using the method given in Yoshiba_2018 and Smith_Maneesoonthorn_2018.

Generative Representation

For generating from the TrUST distribution, we use a scale mixture of normals representation of a t distribution at (ref). If $W \sim \mathrm{Gamma} {\left(\nu/2, \nu/2\right)}$, then $(\text{\boldmath$X$}^\top,\text{\boldmath$L$}^\top)^\top|W\sim N(\mathbf{0},W^{-1}\odot R)$, where `$\odot$' denotes the product of a scalar with each element of a vector or matrix. From this representation, a draw of $\text{\boldmath$Z$}$ from the TrUST distribution can be obtained by completing the following three sequential steps: (i) generate $W \sim \mathrm{Gamma} {\left(\nu/2, \nu/2\right)}$, (ii) draw $\text{\boldmath$L$}|W \sim N(\mathbf{0},W^{-1}\odot\Sigma)$ constrained so that $\text{\boldmath$L$}>\mathbf{0}$, and (iii) draw from the conditional distribution

equation*[equation* omitted — 226 chars of source]

Because the $(q\times 1)$ vector $\text{\boldmath$L$}$ is low dimensional (e.g. $1\leq q \leq 3$ in our empirical work) generating from the constrained distribution at step (ii) is both easy and fast using the approach of botev2017normal or another method.

From the generative algorithm, the joint density of $(\text{\boldmath$Z$},\text{\boldmath$L$},W)$ is

equation[equation omitted — 337 chars of source]

where $f_{Gam}(w;\alpha,\beta)$ denotes the density of a $\mbox{Gamma}(\alpha,\beta)$ distribution, and $\phi_{\text{\boldmath$x$}>\mathbf{0}}(\text{\boldmath$x$} ;\mathbf{0},V)$ is the density of a $N(\mathbf{0},V)$ distribution constrained so that $\text{\boldmath$x$}>\mathbf{0}$. Equation (ref) is used to define an extended likelihood when computing Bayesian inference in Section (ref) below. Two other alternative generative representations of the TrUST distribution are given in Part A of the Online Appendix. However, that above is preferred because it provides for a numerically stable extended likelihood.

Discussion of Distribution

We make four additional comments on the proposed TrUST distribution. First, it is defined with zero mean and correlation matrix $R$ because it is convenient for formation of the implicit copula in Section (ref), where location and scale are unidentified in the copula. Second, generalization of the distribution is straightforward by adding a location parameter and multiplying by scale parameters, as we do when modeling electricity prices in Section (ref). Third, Assumptions (ref) and (ref) can also be adopted in the “extended UST”, where an additional parameter $\text{\boldmath$\tau$} \in \text{$\mathds{R}$}^q$ is introduced and $\text{\boldmath$Z$}\overset{d}{=}\text{\boldmath$X$}|\text{\boldmath$L$}+\text{\boldmath$\tau$}>\mathbf{0}$ Azzalini_Capitanio_2003,arellano2006unification,wang2024multivariate. Doing so further generalizes the TrUST distribution; see Appendix (ref) for details. Fourth, the conditional distribution of a TrUST distribution is of this extended form, as discussed in Appendix (ref).

Bayesian Inference and Application

This section outlines how to compute Bayesian inference for the parameters of the TrUST distribution, and applies the approach to both simulated data and highly skewed Australian regional electricity prices.

Extended Likelihood

The inferential problem is for the location-scaled version of the TrUST distribution, where \(\text{\boldmath$Y$} = {\left(\text{\boldmath$\mu$} + S \text{\boldmath$Z$}\right)}\), \( \text{\boldmath$Z$} \sim \mathrm{TrUST}_{q}{\left(\Omega, A, \nu\right)} \), $\text{\boldmath$\mu$} = {\left(\mu_1, \ldots, \mu_d\right)}^\top$ denotes the location vector, and $S = \operatorname{diag}{\left(\text{\boldmath$s$}\right)}$ is a diagonal scale matrix with leading diagonal $\text{\boldmath$s$} = {\left(s_{1}, \ldots, s_{d}\right)}^\top$. The vector \(\text{\boldmath$Y$}\) has density

equation[equation omitted — 309 chars of source]

The likelihood based on (ref) can exhibit a complex geometry, so that both its direct optimization and evaluation of the resulting Bayesian posterior can be difficult. To simplify the problem the following extended likelihood is employed. Let $\text{\boldmath$y$}_{\mbox{\tiny obs}}=\{\text{\boldmath$y$}_i\}_{i=1}^n$ be the observed data, with $\text{\boldmath$y$}_i=(y_{i1},\ldots,y_{id})^\top $ the $i$th observation of $\text{\boldmath$Y$}$. Also let $\text{\boldmath$w$} = \{w_i\}_{i=1}^n$ and $\text{\boldmath$l$} = \{\text{\boldmath$l$}_i\}_{i=1}^n$, with $\text{\boldmath$l$}_i = {\left(l_{i1}, \ldots, l_{iq}\right)}^\top$, be the corresponding $n$ values of latent $W$ and $\text{\boldmath$L$}$. Then if $\text{\boldmath$\theta$}$ denotes the model parameters, an extended likelihood based on the joint of $(\text{\boldmath$Y$},\text{\boldmath$L$},W)$ is \[ {\cal L}(\text{\boldmath$\theta$},\text{\boldmath$l$},\text{\boldmath$w$};\text{\boldmath$y$}_{\mbox{\tiny obs}})=\mbox{det}(S)^{-n}\prod_{i=1}^n f_{Z,L,W}(\text{\boldmath$z$}_i,\text{\boldmath$l$}_i,w_i;\text{\boldmath$\theta$})\,. \] Here, $\text{\boldmath$z$}_i=S^{-1}(\text{\boldmath$y$}_i-\text{\boldmath$\mu$})$, while the density $f_{Z,L,W}$ is given at (ref) and written here as a function of $\text{\boldmath$\theta$}$. Another advantage of the extended likelihood is that its evaluation does not involve computation of the distribution functions in (ref), unlike with the conventional likelihood based on (ref).

Parameterization, Prior and Augmented Posterior

An effective parameterization of a correlation matrix is in terms of hyper-spherical angles as in Rebonato_Jäckel_1999. We adopt this for \(\Omega = B B^\top\) by setting $B = \{b_{ij}\}$ to a lower triangular Choleskey factor with elements

equation*[equation* omitted — 309 chars of source]

that are functions of the angles $\psi_{i(i-1)}\in(0,2\pi]$ and $\psi_{ij}\in (0,\pi]$ for $j<i-1$. The bounds on each angle $\psi_{ij}$ are unchanged when conditioning on the other angles, making them an attractive parameterization for MCMC sampling as in creal2011dynamic. Denoting the $d(d-1)/2$ unique angles as $\Psi=\{\psi_{ij}\}$, the parameters \(\text{\boldmath$\theta$}= \{\Psi, A, \nu,\text{\boldmath$\mu$},\text{\boldmath$s$}\}\).

The angles have uniform independent priors

equation[equation omitted — 243 chars of source]

with \(\epsilon = 0.03\) chosen to bound $\Omega$ away from singular values. Independent proper Gaussian priors are assigned to each component of \(A\), such that \(\alpha_{jk} \sim N_1(0, 25)\). The prior for \((\nu - 2) \sim \mathrm{Gamma}(3, 0.2)\), where the location shift ensures \(\nu > 2\), and we adopt the reference prior $p(\text{\boldmath$\mu$},\text{\boldmath$s$})\propto \prod_{j=1}^d s_j^{-1}$ as in liseo2006note and others. The prior density is $p(\text{\boldmath$\theta$})\propto p_0(\text{\boldmath$\theta$})\mathds{1}(\Delta\in {\cal P}(\Omega))$, where $p_0(\text{\boldmath$\theta$})$ is the product of the prior densities outlined above, and the indicator function $\mathds{1}(X)=1$ if $X$ is true, and zero otherwise. This indicator function constrains $\Delta$ to a set ${\cal P}(\Omega)$, which contains all values of $\Delta$ that satisfy Assumption (ref). This is where the elements of the diagonal matrix $H=\Sigma-\Delta \Omega^{-1}\Delta^\top$ (which is a deterministic function of $\{\Delta,\Omega\}$) are in ascending order so that ${\cal P}$ will depend on the value of $\Omega$ and we write it as such.

The Bayesian augmented posterior \[ p(\text{\boldmath$\theta$},\text{\boldmath$l$},\text{\boldmath$w$}|\text{\boldmath$y$}_{\mbox{\tiny obs}})\propto {\cal L}(\text{\boldmath$\theta$},\text{\boldmath$l$},\text{\boldmath$w$};\text{\boldmath$y$}_{\mbox{\tiny obs}}) p_0(\text{\boldmath$\theta$})\mathds{1}(\Delta\in {\cal P}(\Omega))\,, \] is computed using the MCMC sampler below.

{\bf Algorithm 1} {\em (MCMC Sampler for TrUST Distribution Parameters)}\\ Step 1: Generate from $p(\text{\boldmath$l$}|\text{\boldmath$w$},\text{\boldmath$y$}_{\mbox{\tiny obs}})$ as a block.\\ Step 2: Generate from $p(\text{\boldmath$w$}|\text{\boldmath$l$},\text{\boldmath$y$}_{\mbox{\tiny obs}})$ as a block.\\ Step 3: Generate from $p(\text{\boldmath$\theta$}|\text{\boldmath$l$},\text{\boldmath$w$},\text{\boldmath$y$}_{\mbox{\tiny obs}})$.

At Step 1, $\text{\boldmath$l$}$ is sampled first as a block by exploiting the fact that each element is independently constrained Gaussian, where its conditional posterior density is \[ p(\text{\boldmath$l$}|\text{\boldmath$w$},\text{\boldmath$y$}_{\mbox{\tiny obs}})=\prod_{i=1}^n \prod_{k=1}^q p(l_{ik}|\text{\boldmath$z$},w,\text{\boldmath$\theta$}) =\prod_{i=1}^n \prod_{k=1}^q\phi_{l_{ik}>0}(m_{ik},w_i^{-1}r_k)\,, \] $m_{ik}= \text{\boldmath$\delta$}_k^\top \Omega^{-1}\text{\boldmath$z$}_i$ and $r_k = 1 - \text{\boldmath$\delta$}_k^\top \Omega^{-1} \text{\boldmath$\delta$}_k$. The conditional independence of the elements of $\text{\boldmath$l$}$ is a key feature of the TrUST distribution, and does not apply to the general UST distribution.

At Step 2, $\text{\boldmath$w$}$ is sampled next, also as a block from its posterior $p(\text{\boldmath$w$}|\text{\boldmath$l$},\text{\boldmath$y$}_{\mbox{\tiny obs}})=\prod_{i=1}^n f_{Gam}(w_i;a,b_i)$, with $a = (d+q+\nu)/2$ and \[ b_i = \frac{1}{2}{\left[ {\left(\text{\boldmath$z$}_i - \Delta^\top \Sigma^{-1}\text{\boldmath$l$}_i\right)}^\top{\left(\Omega - \Delta^\top \Sigma^{-1} \Delta\right)}^{-1}{\left(\text{\boldmath$z$}_i - \Delta^\top \Sigma^{-1}\text{\boldmath$l$}_i\right)} + \text{\boldmath$l$}_i^\top \Sigma^{-1} \text{\boldmath$l$}_i + \nu \right]}. \]

At Step 3, each element of \(\text{\boldmath$\theta$}\) is sampled conditional on $\text{\boldmath$l$}$ and $\text{\boldmath$w$}$. A range of popular methods can be used here, and we simply sample each element one-at-a-time\footnote{Each constrained or bounded element of $\text{\boldmath$\theta$}$ is first transformed to the real line using a one-to-one transformation (e.g. logarithm for positive valued parameters) to simplify sampling.} (in random order to improve mixing) using adaptive random walk Metropolis-Hastings. The parameter constraint in Assumption (ref) is imposed after generating each element of $\Psi$ and $A$ as follows. Compute the diagonal matrix $H=\Sigma-\Delta \Omega^{-1} \Delta^\top$, where recall that $\Delta, \Omega$ and $\Sigma$ are all known functions of $\{\Psi,A\}$. Then simply reorder the columns of $\Delta$ to ensure the leading diagonal elements of $H$ are ordered as permutation $\pi^*$; i.e. in ascending order. To see that the order of the leading diagonal elements of $H$ corresponds to the order of the columns of $\Delta$, recall that $\Sigma$ is a function of $\Delta$ as in Lemma (ref).

Simulation Example

The Bayesian method and increased flexibility of the TrUST distribution over the AC skew-t are first illustrated using a simulation, where we assume \(\text{\boldmath$\mu$}=\mathbf{0}\) and \(\text{\boldmath$s$}=(1,\ldots,1)^\top\) for simplicity. We generate \(n=1040\) observations from two data generating processes: DGP1 with $q=1$ (i.e. a skew-t), and DGP2 with \(q=2\). For both cases \(d=3\), \(\nu=10\), and elements of \(\Omega\) set to $\omega_{21}=0.5$, $\omega_{31}=0.3$, and \( \omega_{32}=0.811 \). The values of $A$ are:

itemize• DGP1: \(A = (-5, 3, 5)^\top\), implying \(\Delta = {\left(-0.271, 0.618, 0.805\right)}^\top\). • DGP2: \(A = [\text{\boldmath$\alpha$}_1|\text{\boldmath$\alpha$}_2]\), where \(\text{\boldmath$\alpha$}_1 = (5, 3, 5)^\top\), \(\text{\boldmath$\alpha$}_2 = (-10, 0, 5)^\top\), implying $\Delta^\top = [\text{\boldmath$\delta$}_1| \text{\boldmath$\delta$}_2]$, where \(\text{\boldmath$\delta$}_1 = {\left(0.748, 0.894, 0.835\right)}^\top\), and \(\text{\boldmath$\delta$}_2 = {\left(-0.868, -0.097, 0.204\right)}^\top\).
table[table omitted — 2,558 chars of source]

For each DGP, TrUST distributions with \(q =1 \) and \(q = 2\) are estimated and the quality of fit measured using the Deviance Information Criterion (DIC) \[ \operatorname{DIC} = -4\,\mathbb{E}_{\text{\boldmath$\theta$}}[\log p(\text{\boldmath$y$}_{\mbox{\tiny obs}} \mid \text{\boldmath$\theta$})] + 2\log p(\text{\boldmath$y$}_{\mbox{\tiny obs}} \mid \text{\boldmath$\hat \theta$}), \] where \(p(\text{\boldmath$y$}_{\mbox{\tiny obs}}|\text{\boldmath$\theta$})=\prod_{i=1}^n f{\left(\text{\boldmath$y$}_i; \text{\boldmath$\mu$}, S, \Omega, A, \nu\right)} \) is evaluated as in (ref). The expectation \(\mathbb{E}_{\text{\boldmath$\theta$}}\) is computed from the MCMC samples after burn-in and \(\text{\boldmath$\hat \theta$}\) denotes the posterior mean.

Table (ref) presents the results. For DGP1 the misspecified TrUST distribution with $q=2$ (Panel B) is almost as accurate at the correctly specified skew-t distribution with $q=1$ (Panel A). In contrast, when the model is misspecified with too few latent variables, the impact is substantial with a much higher DIC value in Panel C compared to Panel D. Figure (ref) displays the major discrepancy in the two fitted densities for the case of DGP2, with the contours in the first row (Panels (a)–(c) for $q=1$) failing to align well with those in the second row (Panels (d)–(f) for $q=2$).

figure[figure omitted — 405 chars of source]

Daily Australian Electricity Prices

Over \$17.5bn of electricity is traded annually in the Australian National Electricity Market (NEM). This wholesale market comprises the five interconnected regions of New South Wales (NSW), Victoria (VIC), Queensland (QLD), South Australia (SA), and Tasmania (TAS). There are separate prices in each region, which exhibit complex dependencies and extreme positive skew; see panagiotelis2008bayesian for a discussion of this market and the distribution of prices.

The data comprise the $n=2070$ daily prices from 1 December 2018 to 31 July 2024 for the \(d = 5\) regions. We consider TrUST distributions with values of $q\in\{0,1,2,3\}$ and for both Gaussian (i.e. $\nu \rightarrow \infty$) and unconstrained degrees of freedom $\nu$. Eight distributions are fit to the logarithm of prices using the Bayesian method, and Table (ref) lists these, their dimension of $\text{\boldmath$\theta$}$ and DIC values. The posteriors means of $\nu$ are also given, indicating the extreme kurtosis in this data, and the means of the other parameters can be found in the Online Appendix. Overall, the TrUST distributions with \(q=2\) and $q=3$ have the lowest DIC values, and clearly dominate the other benchmarks, including the simpler skew-t distribution with $q=1$.

table[table omitted — 1,364 chars of source]
figure[figure omitted — 477 chars of source]

Figure (ref) gives bivariate marginal density contour plots for the fitted TrUST distributions with \(q \in \{0,1,2,3\}\) and variable pairs (a--d) QLD, TAS, and (e--h) SA, VIC. These are overlayed with scatterplots of the data, trimmed for extreme outliers for improved visualization only. Increasing $q$ impacts the asymmetry in the distribution substantially. Further contour plots for all pairwise slices of the TrUST distribution with $q=2$, along with five univariate marginals, are given in the Online Appendix. Overall, this example highlights that the TrUST distribution fits this five-dimensional data more accurately than the skew-t.

TrUST Copula

This section outlines the copula of the TrUST distribution, which extends the AC skew-t copula of Yoshiba_2018 and deng2024large to allow for greater variability in asymmetric dependence across variable pairs.

Copula Specification

The copula of a continuous random vector $\text{\boldmath$Z$}$ is known as an implicit copula and is obtained by inversion of Sklar's theorem; see nelsen2006introduction and smith2021implicit. Let $\text{\boldmath$Z$} = {\left( Z_1, \ldots, Z_d\right)}^\top \sim \mathrm{TrUST}_q {\left(\Omega, A, \nu\right)}$, then from (ref) and (ref), the TrUST copula density is

equation[equation omitted — 280 chars of source]

Here, $\text{\boldmath$u$}=(u_1,\ldots,u_d)^\top \in [0,1]^d$, $z_j=F_{\mathrm{UST},q}^{-1}(u_j;1, \Delta_j, \Sigma, \nu)$ is the marginal quantile function of $Z_j$ evaluated at $u_j$ for $j=1,\ldots,d$, $\Delta_j$ is the $j$th column of $\Delta$,\footnote{Note that this is different than $\text{\boldmath$\delta$}_j$ which denotes the $j$th row of $\Delta$.} and $\text{\boldmath$z$}=(z_1,\ldots,z_d)^\top$. The copula parameters are $\text{\boldmath$\theta$}=\{\Omega,A,\nu \}$, and $\Sigma, \Delta$ are both known functions of $\text{\boldmath$\theta$}$ as in Section (ref).

To visualize the copula, Figure (ref) in the Online Appendix plots contours of $c$ for the bivariate TrUST copulas of the bivariate distributions with $q=2$ in Figure (ref). This illustrates how the parameter $\text{\boldmath$\alpha$}_2$ (which is an additional copula parameter over that of the skew-t copula where $q=1$) affects the dependence asymmetries greatly.

Rank Correlations

We now present expressions for both Kendall's correlation \(\rho_K\) and Spearman's correlation \(\rho_S\) for the UST distribution, which includes the TrUST subclass. As far as we are aware, these general expressions are given for the first time in the literature.

lemma[Kendall's Correlation of UST] Let \( {\left(Z_1, Z_2\right)}^\top \sim \mathrm{UST}_{q}{\left(\Omega, \Delta, \Sigma, \nu\right)} \), then the Kendall's correlation for this pair is \begin{equation} \rho_{K, \mathrm{UST}}{\left(Z_1, Z_2\right)} = 4 \frac{ T_{2+2q} {\left(\mathbf{0}; R_K, \nu \right)}}{ {\left[T_{q} {\left(\mathbf{0}; \Sigma, \nu\right)}\right]}^2 } - 1, \end{equation} where \begin{equation*} R_K = \left(\begin{array}{lc} 2\Omega & \Delta_K^\top \\ \Delta_K & I_2 \otimes \Sigma \\ \end{array}\right)\,, \quad \quad \Delta_K = \left(\begin{array}{c} \Delta \\ -\Delta \\ \end{array}\right)\,, \end{equation*} and \( \otimes \) denotes the Kronecker product.

For the specific case of $q=2$, $\rho_{K, \mathrm{UST}}{\left(Z_1, Z_2\right)} = 4 \frac{T_{6}{\left(\mathbf{0}; R_K, \nu\right)}}{\left[1/4 + {\left(2\pi\right)}^{-1} \arcsin(\sigma)\right]^2} - 1$, where \(\sigma\) is the off-diagonal of \(\Sigma\). For the specific case of $q=1$ (i.e. the skew-t distribution) $\rho_{K, \mathrm{UST}}{\left(Z_1, Z_2\right)} = 16\,T_{4}{\left(\mathbf{0}; R_K, \nu\right)} - 1$, where the skew parameter is \(\Delta = \text{\boldmath$\delta$}^\top\).

lemma[Spearman's Correlation of UST] Let ${\left(Z_1, Z_2\right)}^\top \sim \mathrm{UST}_{q}{\left(\Omega, \Delta, \Sigma, \nu\right)}$, then Spearman's correlation for this pair has is: \begin{equation*} \rho_{S}{\left(Z_1, Z_2\right)} = 12 \frac{T_{2+3q} {\left(\mathbf{0}; R_S, \nu \right)}}{{\left[T_q {\left( \mathbf{0}; \Sigma, \nu \right)}\right]}^3} - 3, \end{equation*} where \begin{equation*} R_S = \left(\begin{array}{cc} \Omega + I_2 & \Delta_S^\top \\ \Delta_S & I_3 \otimes \Sigma \\ \end{array}\right) \,, \quad \quad \Delta_S = \left(\begin{array}{c} \Delta \\ -{\left(\Delta_1 \oplus \Delta_2\right)} \end{array}\right)\,, \end{equation*} and $\oplus$ denotes the direct sum operation that stacks each vector in \(\Delta\) diagonally to construct a block diagonal matrix.

For the specific case of $q=2$, $\rho_{S,\mathrm{UST}}{\left(Z_1, Z_2\right)} = 12\,\frac{T_{8}{\left(\mathbf{0}; R_S, \nu\right)}}{\left[1/4 + (2\pi)^{-1}\arcsin(\sigma)\right]^3} - 3$, where \(\sigma\) is the off-diagonal element of \(\Sigma\). For the specific case of $q=1$ (i.e. the skew-t distribution) $\rho_{S,\mathrm{UST}}{\left(Z_1, Z_2\right)} = 96\,T_{5}{\left(\mathbf{0}; R_S, \nu\right)} - 3$, where the skew parameter is \(\Delta = \text{\boldmath$\delta$}^\top\). When $q=0$, Lemmas (ref) and (ref) correspond to the usual expressions for an elliptical distribution.

For the TrUST distribution, these expressions can be computed for a pair of elements in $\text{\boldmath$Z$}$ using their bivariate marginal. When \(q=1\) the expressions have been presented previously in heinen2022kendall and lu2024kendall. Table (ref) reports the the rank correlations for the nine bivariate distributions with $q=2$ shown in Figure (ref) and their implied copulas.

table[table omitted — 1,427 chars of source]

Asymmetric Dependence Measurements

To measure asymmetric dependence between any pair of variables $Y_1,Y_2$, differences between quantile dependence metrics are computed as follows. Let $U_1=F_1(Y_1)$ and $U_2=F_2(Y_2)$, then for quantile $\kappa\in(0,0.5]$ the quantile dependence metrics in the four quadrants are:

align*[align* omitted — 465 chars of source]

These probabilities are computed from the bivariate copula $C(U_1,U_2)$ of the joint distribution of $(Y_1,Y_2)$ as in Appendix (ref). Then we measure the level of asymmetry in the dependence along the major and minor diagonals of the unit square support of $C$ as

equation[equation omitted — 316 chars of source]

In most empirical applications, including when capturing dependence between equity returns, $\Lambda_{\scalebox{0.6}{Major}}(\kappa)$ is the more important of these two measures.

Bayesian Inference for the Copula Parameters

We now outline the Bayesian inference for the TrUST copula parameters $\text{\boldmath$\theta$}=\{\Omega,A,\nu\}$. For simplicity, the marginals of data $Y_{ij}\sim F_{Y_j}$ are assumed known, although joint estimation of these with $\text{\boldmath$\theta$}$ is possible by extending the MCMC scheme discussed here.

Let $U_j:=F_{\mathrm{UST},q}(Z_j;1, \Delta_j, \Sigma, \nu)$ for $j=1,\ldots,d$, where $\text{\boldmath$Z$}\sim \mbox{TrUST}_q(\Omega,A,\nu)$, and $f_{\mathrm{UST},q}=\frac{d}{dz_j} F_{\mathrm{UST},q}$ is the marginal density of $Z_j$ given at (ref). Then the joint density of $(\text{\boldmath$U$},\text{\boldmath$L$},W)$ can be evaluated using a change of variables from $\text{\boldmath$Z$}$ to $\text{\boldmath$U$}=(U_1,\ldots,U_d)^\top$ to obtain \[ f_{U,L,W}(\text{\boldmath$u$},\text{\boldmath$l$},w;\text{\boldmath$\theta$})=f_{Z,L,W}(\text{\boldmath$z$},\text{\boldmath$l$},w)/\prod_{j=1}^{d} f_{\mathrm{UST}, q}{\left(z_j; 1, \Delta_j, \Sigma, \nu\right)}\,, \] where $f_{Z,L,W}$ is defined previously at (ref), and the denominator arises from the Jacobian of the transformation. This density can be used to define an extended likelihood and augmented posterior for $\text{\boldmath$\theta$}$ and the latents $\text{\boldmath$l$},\text{\boldmath$w$}$.

For the observed copula data \(\text{\boldmath$u$}_{obs}= \{\text{\boldmath$u$}_i\}_{i=1}^n\), where \(\text{\boldmath$u$}_i=(u_{i1},\dots,u_{id})^\top\) and \(u_{ij}=F_{Y_j}(y_{ij})\), the extended likelihood is simply $\mathcal{L}{\left(\text{\boldmath$\theta$},\text{\boldmath$l$},\text{\boldmath$w$};\text{\boldmath$u$}_{obs}\right)} = \prod_{i=1}^n f_{U,L,W}(\text{\boldmath$u$}_i,\text{\boldmath$l$}_i,w_i;\text{\boldmath$\theta$})$. When combined with the same prior for $\text{\boldmath$\theta$}$ in Section (ref), the resulting Bayesian augmented posterior is

equation[equation omitted — 325 chars of source]

where the identification constraint for $\Delta$ is imposed through the prior as before.

The MCMC scheme at Algorithm 1 with some minor adjustments can be used to evaluate this augmented posterior. Generating $\text{\boldmath$l$}$ and $\text{\boldmath$w$}$ in Steps 1 and 2 is unchanged, except that $\text{\boldmath$z$}=\{\text{\boldmath$z$}_i\}_{i=1}^n$, with $\text{\boldmath$z$}_i=(z_{i1},\ldots,z_{id})^\top$, is computed from the copula data using the $d$ quantile functions \[ z_{ij}=F_{\mathrm{UST},q}^{-1}(u_{ij};1, \Delta_j, \Sigma, \nu)\,,\;\;j=1,\ldots,d\,;\quad i=1,\ldots,n\,. \] Evaluating $\text{\boldmath$z$}$ is the most demanding computation of the sampler. The quantile functions above are unavailable in closed form, so that we use the efficient and accurate numerical method in Yoshiba_2018 and Smith_Maneesoonthorn_2018. In Step 3 of Algorithm 1 each element of $\text{\boldmath$\theta$}$ is generated one element at a time using an adaptive random walk Metropolis-Hastings scheme. We use the same approach here, except that the target conditional posterior $p(\text{\boldmath$\theta$}|\text{\boldmath$l$},\text{\boldmath$w$},\text{\boldmath$u$}_{obs})$ is now given by (ref) (up to proportionality).

Simulation Example Continued

The flexibility of the TrUST copula and the efficacy of the Bayesian inference method are first illustrated using the data simulated from the two DGPs in Section (ref). The copula data $\text{\boldmath$u$}_{obs}$ was computed using the expressions for the univariate marginals $u_{ij}=F_{\mathrm{UST},q}(z_{ij};1, \Delta_j, \Sigma, \nu)$. The two DGPs have different levels of rank correlation and asymmetric dependence over variable pair. For the data from each DGP, we estimate two TrUST copulas: Copula 1 with \(q=1\) (i.e. skew-t copula) and Copula 2 with \(q=2\).

table[table omitted — 2,653 chars of source]

For DGP1, both Copula 1 and Copula 2 provide very similar and accurate estimates of the dependence metrics (see Table (ref) in the Online Appendix). Therefore, we focus on DGP2 where greater heterogeneity in dependence over variable pairs is apparent. Table (ref) reports the true rank correlations and asymmetric dependence metrics for DGP2 across the three variable pairs. The posterior mean and posterior standard deviations for these dependence metrics from the fitted copulas are also given. Copula 2 (the TrUST copula with $q=2$) captures these well, whereas Copula 1 (the skew-t copula with $q=1$) is unable to capture the variability in these metrics across variable pairs. Overall, the results suggest that a TrUST copula with $q\geq 2$ is preferable whenever dependence is heterogeneous over variable pairs, as the next section shows is the case for intraday equity returns.

Asymmetric Dependence in Intraday Equity Returns

Tail dependence between the returns on equity pairs is often asymmetric patton2004out,harvey2010 and recent studies suggest it also exhibits high levels of heterogeneity over asset pairs, including during the recent COVID-19 pandemic le2021covid,ando2022quantile. deng2024large found evidence of this in 15-minute equity returns using an AC skew-t copula model, and we extend their study here. We show that the TrUST copula with $q>1$ is more effective than the skew-t copula at capturing the dependence between the market volatility VIX index published by the Chicago Board of Exchange (CBOE) and returns of leading US stocks.

Data and Marginal Models

Observations on 15-minute equity returns and the VIX index were obtained from the LSEG DataScope Select Database. The intraday GARCH(1,1) model of Engle_Sokalska_2012 with student t innovations is used as a marginal for the equity returns. For the VIX index, the marginal is a first-order autoregression with a nonparametric disturbance given by a kernel density estimator. The copula data are computed from these marginal models. The same filtering approach was also employed by deng2024large, to whom we refer for further details.

Banking Stocks Example (3-dim)

The first study is of dependence during the COVID-19 crash between two key banking stocks, Bank of America (BAC) and JP Morgan (JPM), and the VIX. TrUST copulas with \(q=1\) and \(q=2\) of dimension $d=3$ are fit to the $n=1040$ intraday observations from February 11 to April 15, 2020. The posterior estimates of the copula parameters $\text{\boldmath$\theta$}$, pairwise rank correlations and asymmetric dependencies are given in Table (ref) in the Online Appendix. They show that increasing the number of latent variables in the TrUST copula from $q=1$ to $q=2$ has a strong impact on the estimated dependence structure. In particular, the DIC when $q=2$ is 5742, while for the skew-t copula with $q=1$ it is 6653, suggesting an improved fit even though the example is low dimensional.

figure[figure omitted — 552 chars of source]

To visualize the asymmetric dependence between each pair of variables we use a “quantile dependence plot” as in patton2012review. This is where \(\lambda_{\mathrm{LL}}(\kappa)\) (blue line) and \(\lambda_{\mathrm{UL}}(\kappa)\) (purple line) are plotted against quantile $\kappa$, and \(\lambda_{\mathrm{UR}}(\kappa)\) (red line), \(\lambda_{\mathrm{LR}}(\kappa)\) (yellow line) are plotted against $1-\kappa$. Asymmetric dependence in either the minor or major diagonal results in visual asymmetry around $\kappa=0.5$ in the plot. Moreover, greater consistency with the noisy empirical quantiles (black line) suggest an improved fit. Figure (ref) gives the quantile dependence plots for the pairs (BAC, VIX), (JPM, VIX), and (JPM, BAC) for the fitted TrUST copulas with $q=1$ and $q=2$. The quantile estimates with $q=2$ presented in panels (d)-(f) more closely align with the empirical quantiles for all three pairs, reinforcing the improvement also indicated by the DIC values.

Cross Industry Example (6-dim)

Cross industry dependence in U.S. equity returns is estimated using $d=6$ dimensional TrUST copulas. The filtered copula data was computed for the VIX and leading stocks from five key sectors of the economy: Apple (AAPL) from Electronics, JPMorgan (JPM) from Banking, Visa (V) from Business Services, Exxon Mobil (XOM) from Energy, and United Parcel Service (UPS) from Freight/Transport. We consider two periods each with $n=1040$ observations: a pre-COVID period from 18 December 2018 to 15 February 2019, and the COVID period from 11 February 2020 to 15 April 2020. Market volatility was high during both periods, allowing investigation of how extreme market conditions influence dependence between the variables.

table[table omitted — 2,374 chars of source]

The TrUST copula model with \(q=1\) and \(q=2\) is estimated for both periods. Table (ref) reports the posterior means of the copula parameters for the pre-COVID period. The DIC for both copulas is also reported, indicating that the fit with $q=2$ provides a substantial improvement over the AC skew-t copula with $q=1$. This is also the case for the COVID period data, for which the estimates are given in Table (ref) in the Online Appendix.

sidewaysfigure[htbp] \begin{minipage}[t]{1\linewidth} \\[3ex] \end{minipage} \caption{Pairwise asymmetric dependencies $\Lambda_{\scalebox{0.6}{Major}}(0.05)$ and $\Lambda_{\scalebox{0.6}{Minor}}(0.05)$ of the estimated TrUST copula models for 6-dimensional cross-industry example. Panels (a,b,e,f) give results for the pre-COVID period data, and panels (c,d,g,h) for the COVID period. The panels in the top row correspond to $q=1$ and panels in the the bottom row when $q=2$.}

Estimates of the asymmetric dependence measures $\Lambda_{\scalebox{0.6}{Major}}(0.05)$ and $\Lambda_{\scalebox{0.6}{Minor}}(0.05)$ for all variable pairs are presented as heatmaps in Figure (ref). The level of asymmetry varies across variable pairs in both periods and for both fitted copulas, particularly for $\Lambda_{\scalebox{0.6}{Major}}(0.05)$ which is the more important measure. However, there are substantive differences in direction and degree between the two copulas for some variable pairs due to the improved fit of the TrUST copula with $q=2$. Equivalent results are given for quantile $\tau=0.01$ in the Online Appendix.

To measure whether these differences are meaningful, we consider their impact on density forecasts. To do so, the out-of-sample one-step-ahead predictive distributions were computed from both copula models over 40 trading day evaluation periods (equivalent to 1040 intraday observations) following the two estimation periods. Figure (ref) presents the cumulative log-score (LS) differences between the TrUST copula models and a Student-t copula model with the same marginals. Panel (a) gives the results for the earlier pre-COVID estimation period, and panel (b) for the later COVID estimation period. Positive values of the cumulative LS difference indicate improved performance, and the TrUST copula with $q=2$ outperforms both the AC skew-t copula and the benchmark Student-t copula in both evaluation periods.

figure[figure omitted — 306 chars of source]

Discussion

The UST distribution can capture richer asymmetry than the popular AC skew-t distribution which it nests. However, this potential has not been exploited in data analysis to date because inference is difficult. The current paper addresses this by specifying a novel tractable subclass for which likelihood-based inference can be computed using data augmentation methods. The flexibility of this TrUST distribution is illustrated empirically using both simulated and regional Australian electricity price data.

The implicit copulas of skew-elliptical distributions demarta_mcneil_2005,smith_gan_kohn_2010,lucas2017,Yoshiba_2018,oh_patton_2023,deng2024large are popular because they are both scalable and invariant to permutation of the random vector. A major contribution of the current paper is to show how the implicit copula of the TrUST distribution is also tractable, and how Bayesian data augmentation method can also be used to compute the posterior of its parameters. The TrUST copula with $q\geq 2$ allows for greater heterogeneity in asymmetric dependence across variable pairs than the AC skew-t copula. An important application is the modeling of dependence between equity returns, where an improvement is demonstrated using a TrUST copula model for the intraday returns of large U.S. equities.

We finish by discussing three possible future extensions of our work. First, in this paper we compute exact posterior inference for both the TrUST distribution and its copula using MCMC methods. However, approximate Bayesian inference methods---notably variational inference---are faster and have the potential to allow estimation for higher dimensions $d$. Second, the prior for $\Omega$ is uniform on its $d(d-1)/2$ hyper-spherical angles. For high values of $d$ these angles can be regularized, such as through Bayesian global-local priors. Alternatively, a factor decomposition can be used for $\Omega$, as in Murray_Dunson_Carin_Lucas_2013 or deng2024large, so that parameterization of $\Omega$ only increases linearly with $d$. Interestingly, this is very similar to the approach used in USE distributions to parameterize skew using $\Delta$. Here, $\Delta$ is a $(d\times q)$ matrix where $q$ is fixed, so that the number of skew parameters only increases linearly with $d$. Last, adopting a dynamic specification of the TrUST copula, as in the models of creal2015high,oh2017 and oh_patton_2023, would introduce time variation in the dependence structure.

\oldappendix {\appendixname Part \Alph{section}\quad}

Extended Expression

Extended TrUST Distribution

The term “extended” follows from Azzalini_Capitanio_2003 that has hidden truncation on $ L + \tau > 0$.

corollary[Extended TrUST Distribution] For joint random vector \( {\left(\text{\boldmath$X$}^\top, \text{\boldmath$L$}^\top\right)}^\top \sim \mathrm{T}{\left(\mathbf{0}, R, \nu\right)}\). We denote the Extended-TrUST (ETrUST) distribution for random vector \(\text{\boldmath$Z$} \overset{\mathrm{d}}{=} \text{\boldmath$X$} | \text{\boldmath$L$} + \text{\boldmath$\tau$} > \mathbf{0}\), where with $\text{\boldmath$\tau$} = {\left(\tau_1,\ldots,\tau_q\right)}^\top $. Then $\text{\boldmath$Z$} \sim \mathrm{ETrUST}_q {\left(\mathbf{0}, \Omega, A, \nu, \text{\boldmath$\tau$}\right)}$ has density function \begin{equation} f_{\mathrm{ETrUST}, q}{\left(\boldmath$z$; \Omega, A, \nu, \boldmath$\tau$\right)} = t {\left(\boldmath$z$; \Omega, \nu\right)} \frac{\prod_{k=1}^{q} T {\left( \sqrt{\frac{\nu+d}{\nu+Q {\left(\boldmath$z$\right)}}} {\left($\tilde \tau$_k + \boldmath$\alpha$_k^\top \text{\boldmath$z$}\right)}; 1, \nu +d \right)}ght)}; 1, \nu +d } }{ T {\left(\text{\boldmath$\tau$}; \Sigma, \nu\right)} }, \end{equation} where $\text{$\tilde \tau$}_k = \tau_k {\left(1-\text{\boldmath$\delta$}_k^\top \Omega^{-1}_k \text{\boldmath$\delta$}_k\right)}^{-1/2}$, for $k=1,\ldots,q$.

Conditional Distribution

corollary[Conditional TrUST Distribution] Let \(\text{\boldmath$Z$} \sim \mathrm{TrUST}_{q}(\Omega, A, \nu)\). Partition \[ \text{\boldmath$Z$} \;=\; (\text{\boldmath$Z$}_1,\text{\boldmath$Z$}_2)^\top \;\overset{d}{=}\; (\text{\boldmath$X$}_1,\text{\boldmath$X$}_2 \mid \text{\boldmath$L$}>\mathbf{0}), \] where \(\text{\boldmath$Z$}_1\in\mathbb{R}^{d_1}\) and \(\text{\boldmath$Z$}_2\in\mathbb{R}^{d_2}\), and \[ \Omega = \begin{pmatrix} \Omega_1 & \Omega_{12} \\ \Omega_{21} & \Omega_2 \end{pmatrix}, A = \begin{pmatrix} A_1 \\ A_2 \end{pmatrix} \] with \( A_1 = {\left(\text{\boldmath$\alpha$}_{k(1)}\right)}_{d_1 \times q},\; A_2 = {\left(\text{\boldmath$\alpha$}_{k(2)}\right)}_{d_2 \times q} \). Define the conditional random vector \[ \text{\boldmath$Z$}_{1\mid 2} \;:=\; (\text{\boldmath$Z$}_1 \mid \text{\boldmath$Z$}_2 = \text{\boldmath$z$}_2)\,\overset{d}{=}\, (\text{\boldmath$X$}_1 \mid \text{\boldmath$X$}_2 = \text{\boldmath$x$}_2,\;\text{\boldmath$L$}>\mathbf{0})\,. \] Then \(\text{\boldmath$Z$}_{1\mid 2}\) admits an “extended” TrUST distribution, in which the latent location is shifted by conditioning \(\text{\boldmath$Z$}_2\). Its density is given by: \[ p_{\text{\boldmath$Z$}_{1|2}}(\text{\boldmath$z$}_1) = t_{d_1}{\left(\text{\boldmath$z$}_1 - \text{\boldmath$\mu$}_*; \frac{\nu+Q(\text{\boldmath$z$}_2)}{\nu+d_2} \Omega_*, \nu+d_2\right)} \frac{ \prod_{k=1}^{q} T_1{\left(\sqrt{\frac{\nu+d_2+d_1}{\nu + Q(\text{\boldmath$z$})} } (\text{\boldmath$\alpha$}_{k(1)}^\top \text{\boldmath$z$}_1 + \text{\boldmath$\alpha$}_{k(2)}^\top \text{\boldmath$z$}_2); 1, \nu+d_2+d_1\right)} }{ T_q{\left(\Delta_2 \Omega_{2}^{-1} \text{\boldmath$z$}_2; \frac{\nu+Q(\text{\boldmath$z$}_2)}{\nu+d_2} (\Sigma - \Delta_2 \Omega_{2}^{-1} \Delta_2^\top), \nu+d_2 \right)} }, \] where \( \text{\boldmath$\mu$}_* = \Omega_{12}\Omega_2^{-1} \text{\boldmath$z$}_2\), \( \Omega_{*} = \Omega_1 - \Omega_{12} \Omega_2^{-1} \Omega_{21}\), \(Q(\text{\boldmath$z$}_2) = \text{\boldmath$z$}_2^\top \Omega_{2}^{-1} \text{\boldmath$z$}_2 \).

Quantile Dependence

The pairwise quantile dependencies of the TrUST copula are computed from the bivariate marginal copula function of \( (U_1, U_2) \), which is \[ C(u_1, u_2) = F_{\mathrm{UST}, q}{\left((z_1, z_2)^\top; \Omega, \Delta, \Sigma, \nu\right)} = \frac{T_{2+q}{\left((z_1, z_2, \mathbf{0}_q^\top)^\top; R^*, \nu \right)} }{ T_q{\left(\mathbf{0}^\top; \Sigma, \nu \right)} } \] with \[ R^* = \left(

array[array omitted — 64 chars of source]

\right), \] which can be derived from Equations (ref) and (ref) that \[ \text{$\mathds{P}$}{\left(\text{\boldmath$Z$} \leq \text{\boldmath$z$}\right)} = \text{$\mathds{P}$}{\left(\text{\boldmath$X$} \leq \text{\boldmath$z$} | \text{\boldmath$L$} > \mathbf{0}\right)} = \frac{\text{$\mathds{P}$}{\left(\text{\boldmath$X$} \leq \text{\boldmath$z$}, -\text{\boldmath$L$} \leq \mathbf{0}\right)} }{\text{$\mathds{P}$}{\left( -\text{\boldmath$L$} \leq \mathbf{0}\right)} }. \] Each pseudo observation \( z_{ij}=F_{\mathrm{UST},q}^{-1}(u_{ij};1, \Delta_j, \Sigma, \nu) \) is obtained via numerical optimization.

\singlespacing

center[center omitted — 131 chars of source]

\spacing{1.5} \setcounter{page}{1} \setcounter{figure}{0} \setcounter{table}{0} \setcounter{section}{0} \setcounter{equation}{0} \setcounter{algorithm}{1} \makeatletter \makeatother

This Online Appendix has four parts:

itemize• {\bf Part A}: Alternative Generative Representations \begin{itemize} • A.1: Generative Representation 1 • A.2: Generative Representation 2 \end{itemize} • {\bf Part B}: Supplementary Figures • {\bf Part C}: Supplementary Tables • {\bf Part D}: Proofs

Alternative Generative Representations

Generative Representation 1

The augmentation leverages the hidden truncation mechanism, as described in Eq. (ref), within the framework of the joint Student-t distribution without scale mixture of normals. Specifically, the joint distribution \( {\left( \text{\boldmath$X$}^\top, \text{\boldmath$L$}^\top \right)}^\top \sim \mathrm{T}_{d+q}{\left(\mathbf{0}, R, \nu\right)} \). Then from this expression, we can derive the \(\text{\boldmath$Z$}\) in the conditional distribution

equation*[equation* omitted — 288 chars of source]

The augmented density of \({\left\{ \text{\boldmath$Z$}, \text{\boldmath$L$} \right\}} \) is

equation[equation omitted — 390 chars of source]

Still, the latent variable $\text{\boldmath$L$}$ can be sampled as a block by:

equation*[equation* omitted — 210 chars of source]

for \(k = 1,\ldots,q\), where $m_k= \text{\boldmath$\delta$}_k^\top \Omega^{-1}\text{\boldmath$Z$}$, $ r_k = 1 - \text{\boldmath$\delta$}_k^\top \Omega^{-1} \text{\boldmath$\delta$}_k$.

Generative Representation 2

The generative representation in the Section (ref) using the scaling then hidden truncation ordering. The other method follows a sequence where hidden truncation is applied first, then scaling. First, we have the joint distribution \( (\text{$\widetilde{\text{\boldmath$X$}}$}^\top, \text{$\widetilde{\text{\boldmath$L$}}$}^\top )^\top \sim \mathrm{N}_{d+q}{\left(\mathbf{0}, R\right)}, \) then expressed the scaling as:

equation*[equation* omitted — 208 chars of source]

where $\odot$ is the element-wise multiplication.

Then the augmented likelihood is expressed as:

equation*[equation* omitted — 290 chars of source]

with the posterior distribution for latent variable $\text{\boldmath$L$}$ given by:

equation*[equation* omitted — 124 chars of source]

and the posterior of $W$ is not in closed form.

Both (GR1) and (GR2) provide alternative augmentations for the model, offering flexibility in computational implementation while adhering consistency with the underlying hidden truncation framework. The choice between representations, including the (GR1) without scaling, depends on the specific modeling objectives and computational requirements, ensuring adaptability to diverse data contexts.