Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
91,143 characters · 33 sections · 63 citation commands
myblue Tractable Unified Skew-t Distribution and Copula for Heterogeneous Asymmetries
{0.15cm} {0.15cm}
\pagestyle{empty}
\pagestyle{plain} \setcounter{equation}{0}
Skew-elliptical distributions are models of multivariate asymmetry that scale well. While there are different variants (for some of these see genton2004skew) those formed by conditioning on a latent truncated variable as suggested by branco2001general are prominent. This approach is often called “hidden truncation” and is used to form the skew-t distribution of Azzalini_Capitanio_2003 (AC hereafter). The AC skew-t distribution is a particularly popular choice in applied data analysis, with diverse applications in renewable energy Hering01032010, actuarial studies eling2012, macroeconomic modeling adrian2019vulnerable, atmospheric science morris2017space, and in flow cytometric analysis PyneetalPNAS2009,fruhwirth2010bayesian among others. However, in $d$ dimensions skew-elliptical distributions typically have only $d$ parameters to control asymmetry across $d(d-1)/2$ variable pairs, which can limit their effectiveness for modeling asymmetry. Unified skew-elliptical (USE) \footnote{arellano2006unification use the abbreviation SUN for the Unified Skew-Normal distribution, later adopting SUE for the Unified Skew-Elliptical family and SUT for the Unified Skew-t. In this paper, we break with this convention and adopt the word-ordered acronyms USN, USE and UST for these three distributions.} distributions arellano2010multivariate are an extension that introduces more latent variables and parameters to allow for a much more flexible form of asymmetry. In particular, the unified skew-t (UST) discussed by wang2024multivariate is a promising generalization of the impactful AC skew-t distribution. However, despite the strong potential of USE distributions, they have not been adopted for data analysis because (i) the parameters are unidentified wang2023non, and (ii) likelihood evaluation and optimization is computationally difficult in even low dimensions.
The current paper addresses these problems by specifying a novel subclass of the USE that is tractable. For this proposed subclass, it is shown that the USE parameters are identified under an ordering of the eigenvalues of the conditional scale matrix of the latent variables. The special case of a tractable UST distribution, which we label a TrUST distribution, is considered in detail. A generative representation is given for the proposed TrUST distribution that is used to define an extended likelihood that is easier to evaluate than the regular likelihood. From this, a Bayesian augmented posterior can be defined and computed using a proposed Markov chain Monte Carlo (MCMC) sampler that is both fast and efficient, and for which the identifying eigenvalue constraint is easy to impose. The proposed TrUST distribution nests the popular AC skew-t as a special case, but generalizes it to allow for richer multivariate asymmetry. An application of the TrUST distribution to both simulated data and highly skewed regional daily Australian electricity prices shows it captures asymmetry more accurately than the popular AC skew-t, and that this improves the quality of fit considerably.
However, the main advantage of the proposed TrUST distribution is that its implicit copula is also tractable and has parameters that are identified. Every multivariate continuous distribution has a unique implicit copula; see smith2021implicit for an introduction to this class of copulas. Currently, implicit copulas of differing skew-t distributions are used to capture the strong asymmetric dependence often found in financial data; see smith_gan_kohn_2010, Christoffersen_Errunza_Jacobs_Langlois_2012, creal2015high, lucas2017, opschoor2021, oh_patton_2023 and deng2024large for examples. Of these, deng2024large show that the implicit copula of the AC skew-t distribution allows for the greatest level of asymmetric dependence. But the implicit copula of the TrUST distribution (hereafter the TrUST copula) allows for a much greater level of heterogeneity in asymmetric dependencies over variable pairs than the AC skew-t copula. We show here that by doing so, the TrUST copula can better capture the dependence between financial returns data. To estimate the TrUST copula parameters a Bayesian augmented posterior is specified using a modification of the extended likelihood of the TrUST distribution, which is then evaluated using an MCMC sampler. Expressions for the Kendall and Spearman rank correlations for this TrUST copula are also derived. As far as we are aware, ours is the first paper to construct the implicit copula of any USE distribution and show how to compute statistical inference for its parameters.
deng2024large use the implicit copula of the AC skew-t distribution to capture dependence between intraday equity returns. In our main application, we extend their analysis to show that the TrUST copula is more effective at capturing such dependence than the AC skew-t copula. This is done for five large equities and an equity market volatility index (VIX) using data collected both before and during the COVID-19 pandemic. The performance of the TrUST copula is compared with that of both the symmetric Student-t copula and the AC skew-t copula using the Deviance Information Criterion (DIC) and cumulative log-score to assess accuracy. The results indicate that the TrUST copula better captures heterogeneity in asymmetric dependence across variables. High levels of heterogeneity is also found by le2021covid and ando2022quantile using network models of tail dependencies.
The rest of this paper is organized as follows. Section (ref) briefly introduces the USE distribution and the special case of the UST, after which our proposed tractable USE variant is outlined. Section (ref) considers the case of the TrUST distribution in detail. Section (ref) outlines estimation methodology for the TrUST distribution, and illustrates its efficacy through a simulated example and application to electricity prices. Section (ref) presents the TrUST copula, derivation of its rank correlations, Bayesian inference and a simulation example. Section (ref) shows how our TrUST copula captures heterogeneity in asymmetric dependence between high-frequency financial variables; Section (ref) concludes. Some extra results are given in the Appendix, while an Online Appendix contains additional information, figures, tables and all proofs.
In this section, the unified skew-elliptical (USE) distribution is introduced, including the special case of the unified skew-t (UST). Our tractable variant of the USE distribution is then outlined in detail. We consider the standardized case (i.e. with zero mean and unit scale) because this is used in forming the implicit copula in Section (ref). Generalization of the distribution by multiplying the random vector by a diagonal scaling matrix and adding a location parameter is straightforward.
arellano2006unification extend the popular skew-elliptical distribution to allow for additional skew parameters. It can be specified for the standardized case as follows. Let $\text{\boldmath$X$}\in \text{$\mathds{R}$}^d$ and $\text{\boldmath$L$} \in \text{$\mathds{R}$}^q$, for $q\geq 1$, be jointly (radially symmetric) elliptically distributed as
where $R$ is a $(d+q) \times (d+q)$ correlation matrix partitioned to be consistent with $(\text{\boldmath$X$}^\top,\text{\boldmath$L$}^\top)^\top$, $\mathbf{0}$ is a zero mean vector\footnote{Throughout this paper we use $\mathbf{0}$ to denote a conformable vector or matrix of zeros without explicitly denoting its dimension.}, and $g: [0, \infty) \to [0, \infty)$ is the density generator function; see Fang_Kotz_Ng_1990 for specification of an elliptical distribution. Then the USE distribution can be defined as
where “$\overset{\mathrm{d}}{=}$” denotes equality in distribution, and the inequality $\text{\boldmath$L$}>\mathbf{0}$ holds element-wise. The vector $\text{\boldmath$L$}$ are latent variables and the specification at (ref) is often called “hidden truncation” arnold2000hidden,arnold2004elliptical. The $(q\times d)$ matrix $\Delta$ contains parameters that control the level and direction of skew, where $q$ is typically low, such as $q\in\{1,2,3,4\}$.
Let $f_{\mathrm{EL}}(\text{\boldmath$y$};V,g)$ and $F_{\mathrm{EL}}(\text{\boldmath$y$};V,g)$ denote the density and distribution functions, respectively, of a zero-mean elliptically distributed random variable $\text{\boldmath$Y$}\sim \mathrm{EL} {\left(\mathbf{0}, V, g\right)}$. Also, for the partition $\text{\boldmath$Y$}=(\text{\boldmath$Y$}_1^\top,\text{\boldmath$Y$}_2^\top)^\top$, the conditional $\text{\boldmath$Y$}_1 \mid \text{\boldmath$Y$}_2$ is elliptically distributed Fang_Kotz_Ng_1990 with density generator denoted here as \( g_{\text{\boldmath$Y$}_2} \). Then, the density of $\text{\boldmath$Z$}$ is obtained by using Bayes theorem to evaluate the conditional density on the righthand side of (ref), giving
If a random variable $\text{\boldmath$Z$}$ has this density, then we write $\text{\boldmath$Z$} \sim \mathrm{USE}_{q}{\left( \Omega, \Delta, \Sigma, g\right)} $.
The number of elements in $\Delta$ increases linearly with $q$. In this way, $q$ plays a role for multivariate asymmetry that is analogous to that of the number of factors in a traditional factor model for the covariance. When \(q = 1\), the USE distribution reduces to the skew-elliptical distribution introduced by branco2001general and Azzalini_Capitanio_2003. In this case there is only a single skew parameter for each dimension, so that \(\Delta \equiv \text{\boldmath$\delta$}^\top = {\left(\delta_1,\ldots,\delta_d\right)}^\top \). Because $\text{\boldmath$\delta$}$ is constrained, for computing inference it is common to transform it to an unconstrained parameterization via the one-to-one transformation $\text{\boldmath$\alpha$} = (1 - \text{\boldmath$\delta$}^{\top} \Omega^{-1} \text{\boldmath$\delta$})^{-1 / 2} \Omega^{-1} \text{\boldmath$\delta$}$ with inverse $\text{\boldmath$\delta$} = (1+\text{\boldmath$\alpha$}^{\top} \Omega \text{\boldmath$\alpha$})^{-1 / 2} \Omega \text{\boldmath$\alpha$}$.
The elliptical distribution with greatest applied potential is the t distribution, which has density generator function $g(x)=(1+x/\nu)^{-(\nu+d)/2}$ that introduces the parameter $\nu>0$. Assuming $(\text{\boldmath$X$}^\top,\text{\boldmath$L$}^\top)^\top$ is multivariate t with location zero, scale matrix $R$ and $\nu$ degrees of freedom, the density at (ref) is given by
Here, $Q(\text{\boldmath$z$}) = \text{\boldmath$z$}^\top\Omega^{-1}\text{\boldmath$z$}$, while \( t_d{\left(\text{\boldmath$y$}; V, \upsilon\right)} \) and \( T_d(\text{\boldmath$y$}; V, \upsilon) \) denote the density and distribution functions, respectively, of a zero mean $d$-dimensional multivariate t random variable with scale matrix $V$ and degrees of freedom \(\upsilon\) evaluated at $\text{\boldmath$y$}$. If a random variable $\text{\boldmath$Z$}$ has density at (ref), then we write $\text{\boldmath$Z$} \sim \mathrm{UST}_{q}{\left(\Omega, \Delta, \Sigma, \nu\right)} $.
Three special cases of the UST distribution are: (i) when $\nu\rightarrow \infty$, the UST converges to the unified skew-normal of arellano2006unification; (ii) when $\Delta$ is a matrix with all elements equal to zero, then $\text{\boldmath$Z$}$ is distributed symmetric multivariate t with location zero, scale matrix $\Omega$ and $\nu$ degrees of freedom; and (iii) when \(q=1\), the UST distribution is the AC skew-t distribution. wang2024multivariate provide a comprehensive overview of the UST distribution and its properties.
The novel subclass of the USE distribution is now outlined, that we call here a Tractable Unified Skew-Elliptical (TrUSE) distribution. It employs the following assumption of conditional independence in the latent variables.
Under this assumption, the scale matrix $\Sigma$ is a deterministic function of $\{\Omega,\Delta\}$, and the density at (ref) is simplified, as summarized by the following two lemmas.
Proofs of both lemmas are found in Appendix (ref), and together they can be used to define our proposed tractable USE distribution as follows.
Parameterization of the TrUSE distribution in terms of $A$ is attractive because each $\text{\boldmath$\alpha$}_k\in \mathbb{R}^d$ is unbounded given $\{\Omega,\text{\boldmath$\alpha$}_{j\neq k}\}$, whereas $\text{\boldmath$\delta$}_k$ has complex nonlinear bounds given $\{\Omega,\text{\boldmath$\delta$}_{j\neq k}\}$. We show later that this property is useful for constructing MCMC schemes to generate $\text{\boldmath$\alpha$}_k$, rather than $\text{\boldmath$\delta$}_k$, from its posterior.
The TrUSE distribution is more tractable than the general USE for three reasons. First, it simplifies the parameter space because $\Sigma$ is a deterministic function of $\{\Omega,A\}$, whereas for a general USE the off-diagonal elements of $\Sigma$ have complex nonlinear bounds given $\{\Omega,\Delta\}$. Second, it is more computationally tractable to compute inference because (i) the density at (ref) is simplified, and (ii) efficient MCMC sampling is possible from the augmented posterior constructed using an extended likelihood based on the joint of $(\text{\boldmath$X$},\text{\boldmath$L$})$. Third, a parameter identification constraint that is straightforward to impose exists as outlined below. In contrast, it is hard to compute statistical inference for the unconstrained parameters of the general USE distribution. The tractability of TrUST distribution is illustrated in Section (ref).
Adopting Assumption (ref) alone is insufficient to identify the USE distribution under permutation of $\text{\boldmath$L$}$. wang2023non discuss this problem in detail, which is characterized by the lemma below.
To identify the parameters a constraint based on an ordering of the eigenvalues of the scale matrix $H$ of the conditional distribution of $\text{\boldmath$L$}|\text{\boldmath$X$}$ is used.
This assumption, when applied to the TrUSE distribution under Assumption (ref), results in a permutation of the latent vector $\text{\boldmath$L$}$ that identifies the columns of \(A\) and \(\Delta^\top\), as below.
In practice, because the posterior distribution of $(h_1,\ldots,h_q)^\top$ is continuous, the inequalities in Theorem (ref) are strict, so that the ordering is unique. In the remainder of the paper Assumption (ref) is adopted to identify the parameters in the TrUSE (and TrUST) distributions, so that $\Delta, A, \Sigma$ are all defined with respect to the unique ordering $\pi^*$. Imposing this constraint is straightforward within an MCMC scheme as discussed in Section (ref).
We finish this section with some comments on the appropriateness of the TrUSE distribution. We stress that Assumption (ref) is on the distribution of $\bm{L}$, not on that of $\bm{X}$, so that the TrUSE distribution can capture the same degree of rank correlation as the USE distribution. Moreover, Assumptions (ref) and (ref) are identifying assumptions that are necessary for the USE distribution to be well-defined for data analysis. In comparison, wang2024multivariate suggest (but do not implement) some other ways to identify the parameters in a UST distribution that involve much stronger restrictions on $\Delta$ and/or $\Omega$. However, these reduce substantially the ability to capture variation in pairwise correlations and/or asymmetry across variable pairs, which is the key advantage of the UST over skew-t distributions. wang2023non suggest identification is secured by fixing the order of the eigenvalues of \(\Sigma\), but this condition is insufficient when \(q\geq 2\) because interchanging rows of \(\Delta\) does not alter the order of the eigenvalues of $\Sigma$.
The special case of the tractable unified skew-t (TrUST) distribution that is the focus of our paper is now outlined in detail.
The TrUST distribution results from selecting a t distribution with $\nu$ degrees of freedom as the choice of elliptical distribution in Section (ref). In this case, the density at Definition (ref) is
If a random vector $\text{\boldmath$Z$}$ has the above density, then we write $\text{\boldmath$Z$} \sim \mathrm{TrUST}_{q}{\left(\Omega, A, \nu\right)}$. If $\nu \rightarrow \infty$, then this reduces to the tractable unified skew-normal (TrUSN) distribution, with density
where $\phi_d{\left(\text{\boldmath$x$}; \Omega\right)}$ and $\Phi_d{\left(\text{\boldmath$x$}; \Omega\right)}$ are density and distribution functions, respectively, for a $N{\left(\mathbf{0}, \Omega\right)}$ distribution. If a random vector $\text{\boldmath$Z$}$ has the above density, then we write $\text{\boldmath$Z$} \sim \mathrm{TrUSN}_{q}{\left(\Omega,A\right)}$. If $q=1$ then the TrUST, UST and AC skew-t all coincide.
Figure (ref) presents the bivariate density contours of the TrUST distribution, when $d=2$, $q=2$, $\nu = 5$ and the first skew parameter vector is \(\text{\boldmath$\alpha$}_1 = (5, 5)^\top\). Each row corresponds to a different correlation parameter \(\omega \in {\left\{-0.5, 0, 0.5\right\}}\), ordered from top to bottom. The columns correspond to \(\text{\boldmath$\alpha$}_2 \in {\left\{ (5, 5)^\top, (0, 5)^\top, (-5, 5)^\top \right\}}\) and show the impact of varying the second skew parameter vector.
Because the UST distribution is closed under marginalization, the marginals of a TrUST distribution are UST with parameters derived from those of the joint as below.
The case where $d_J=1$ is of particular importance when constructing the implicit copula of a TrUST distribution in Section (ref). Here, the marginal for the $j$-th element is $Z_j \sim \mathrm{UST}_{q}( 1, \Delta_{j}, \Sigma, \nu)$ for $j \in {\left\{1,\ldots,d\right\}}$ with density function
where $\Delta_j$ is the $j$th column of $\Delta$ and $Q(z_j)=z_j^2$. The distribution and quantile functions are computed numerically from (ref) using the method given in Yoshiba_2018 and Smith_Maneesoonthorn_2018.
For generating from the TrUST distribution, we use a scale mixture of normals representation of a t distribution at (ref). If $W \sim \mathrm{Gamma} {\left(\nu/2, \nu/2\right)}$, then $(\text{\boldmath$X$}^\top,\text{\boldmath$L$}^\top)^\top|W\sim N(\mathbf{0},W^{-1}\odot R)$, where `$\odot$' denotes the product of a scalar with each element of a vector or matrix. From this representation, a draw of $\text{\boldmath$Z$}$ from the TrUST distribution can be obtained by completing the following three sequential steps: (i) generate $W \sim \mathrm{Gamma} {\left(\nu/2, \nu/2\right)}$, (ii) draw $\text{\boldmath$L$}|W \sim N(\mathbf{0},W^{-1}\odot\Sigma)$ constrained so that $\text{\boldmath$L$}>\mathbf{0}$, and (iii) draw from the conditional distribution
Because the $(q\times 1)$ vector $\text{\boldmath$L$}$ is low dimensional (e.g. $1\leq q \leq 3$ in our empirical work) generating from the constrained distribution at step (ii) is both easy and fast using the approach of botev2017normal or another method.
From the generative algorithm, the joint density of $(\text{\boldmath$Z$},\text{\boldmath$L$},W)$ is
where $f_{Gam}(w;\alpha,\beta)$ denotes the density of a $\mbox{Gamma}(\alpha,\beta)$ distribution, and $\phi_{\text{\boldmath$x$}>\mathbf{0}}(\text{\boldmath$x$} ;\mathbf{0},V)$ is the density of a $N(\mathbf{0},V)$ distribution constrained so that $\text{\boldmath$x$}>\mathbf{0}$. Equation (ref) is used to define an extended likelihood when computing Bayesian inference in Section (ref) below. Two other alternative generative representations of the TrUST distribution are given in Part A of the Online Appendix. However, that above is preferred because it provides for a numerically stable extended likelihood.
We make four additional comments on the proposed TrUST distribution. First, it is defined with zero mean and correlation matrix $R$ because it is convenient for formation of the implicit copula in Section (ref), where location and scale are unidentified in the copula. Second, generalization of the distribution is straightforward by adding a location parameter and multiplying by scale parameters, as we do when modeling electricity prices in Section (ref). Third, Assumptions (ref) and (ref) can also be adopted in the “extended UST”, where an additional parameter $\text{\boldmath$\tau$} \in \text{$\mathds{R}$}^q$ is introduced and $\text{\boldmath$Z$}\overset{d}{=}\text{\boldmath$X$}|\text{\boldmath$L$}+\text{\boldmath$\tau$}>\mathbf{0}$ Azzalini_Capitanio_2003,arellano2006unification,wang2024multivariate. Doing so further generalizes the TrUST distribution; see Appendix (ref) for details. Fourth, the conditional distribution of a TrUST distribution is of this extended form, as discussed in Appendix (ref).
This section outlines how to compute Bayesian inference for the parameters of the TrUST distribution, and applies the approach to both simulated data and highly skewed Australian regional electricity prices.
The inferential problem is for the location-scaled version of the TrUST distribution, where \(\text{\boldmath$Y$} = {\left(\text{\boldmath$\mu$} + S \text{\boldmath$Z$}\right)}\), \( \text{\boldmath$Z$} \sim \mathrm{TrUST}_{q}{\left(\Omega, A, \nu\right)} \), $\text{\boldmath$\mu$} = {\left(\mu_1, \ldots, \mu_d\right)}^\top$ denotes the location vector, and $S = \operatorname{diag}{\left(\text{\boldmath$s$}\right)}$ is a diagonal scale matrix with leading diagonal $\text{\boldmath$s$} = {\left(s_{1}, \ldots, s_{d}\right)}^\top$. The vector \(\text{\boldmath$Y$}\) has density
The likelihood based on (ref) can exhibit a complex geometry, so that both its direct optimization and evaluation of the resulting Bayesian posterior can be difficult. To simplify the problem the following extended likelihood is employed. Let $\text{\boldmath$y$}_{\mbox{\tiny obs}}=\{\text{\boldmath$y$}_i\}_{i=1}^n$ be the observed data, with $\text{\boldmath$y$}_i=(y_{i1},\ldots,y_{id})^\top $ the $i$th observation of $\text{\boldmath$Y$}$. Also let $\text{\boldmath$w$} = \{w_i\}_{i=1}^n$ and $\text{\boldmath$l$} = \{\text{\boldmath$l$}_i\}_{i=1}^n$, with $\text{\boldmath$l$}_i = {\left(l_{i1}, \ldots, l_{iq}\right)}^\top$, be the corresponding $n$ values of latent $W$ and $\text{\boldmath$L$}$. Then if $\text{\boldmath$\theta$}$ denotes the model parameters, an extended likelihood based on the joint of $(\text{\boldmath$Y$},\text{\boldmath$L$},W)$ is \[ {\cal L}(\text{\boldmath$\theta$},\text{\boldmath$l$},\text{\boldmath$w$};\text{\boldmath$y$}_{\mbox{\tiny obs}})=\mbox{det}(S)^{-n}\prod_{i=1}^n f_{Z,L,W}(\text{\boldmath$z$}_i,\text{\boldmath$l$}_i,w_i;\text{\boldmath$\theta$})\,. \] Here, $\text{\boldmath$z$}_i=S^{-1}(\text{\boldmath$y$}_i-\text{\boldmath$\mu$})$, while the density $f_{Z,L,W}$ is given at (ref) and written here as a function of $\text{\boldmath$\theta$}$. Another advantage of the extended likelihood is that its evaluation does not involve computation of the distribution functions in (ref), unlike with the conventional likelihood based on (ref).
An effective parameterization of a correlation matrix is in terms of hyper-spherical angles as in Rebonato_Jäckel_1999. We adopt this for \(\Omega = B B^\top\) by setting $B = \{b_{ij}\}$ to a lower triangular Choleskey factor with elements
that are functions of the angles $\psi_{i(i-1)}\in(0,2\pi]$ and $\psi_{ij}\in (0,\pi]$ for $j<i-1$. The bounds on each angle $\psi_{ij}$ are unchanged when conditioning on the other angles, making them an attractive parameterization for MCMC sampling as in creal2011dynamic. Denoting the $d(d-1)/2$ unique angles as $\Psi=\{\psi_{ij}\}$, the parameters \(\text{\boldmath$\theta$}= \{\Psi, A, \nu,\text{\boldmath$\mu$},\text{\boldmath$s$}\}\).
The angles have uniform independent priors
with \(\epsilon = 0.03\) chosen to bound $\Omega$ away from singular values. Independent proper Gaussian priors are assigned to each component of \(A\), such that \(\alpha_{jk} \sim N_1(0, 25)\). The prior for \((\nu - 2) \sim \mathrm{Gamma}(3, 0.2)\), where the location shift ensures \(\nu > 2\), and we adopt the reference prior $p(\text{\boldmath$\mu$},\text{\boldmath$s$})\propto \prod_{j=1}^d s_j^{-1}$ as in liseo2006note and others. The prior density is $p(\text{\boldmath$\theta$})\propto p_0(\text{\boldmath$\theta$})\mathds{1}(\Delta\in {\cal P}(\Omega))$, where $p_0(\text{\boldmath$\theta$})$ is the product of the prior densities outlined above, and the indicator function $\mathds{1}(X)=1$ if $X$ is true, and zero otherwise. This indicator function constrains $\Delta$ to a set ${\cal P}(\Omega)$, which contains all values of $\Delta$ that satisfy Assumption (ref). This is where the elements of the diagonal matrix $H=\Sigma-\Delta \Omega^{-1}\Delta^\top$ (which is a deterministic function of $\{\Delta,\Omega\}$) are in ascending order so that ${\cal P}$ will depend on the value of $\Omega$ and we write it as such.
The Bayesian augmented posterior \[ p(\text{\boldmath$\theta$},\text{\boldmath$l$},\text{\boldmath$w$}|\text{\boldmath$y$}_{\mbox{\tiny obs}})\propto {\cal L}(\text{\boldmath$\theta$},\text{\boldmath$l$},\text{\boldmath$w$};\text{\boldmath$y$}_{\mbox{\tiny obs}}) p_0(\text{\boldmath$\theta$})\mathds{1}(\Delta\in {\cal P}(\Omega))\,, \] is computed using the MCMC sampler below.
{\bf Algorithm 1} {\em (MCMC Sampler for TrUST Distribution Parameters)}\\ Step 1: Generate from $p(\text{\boldmath$l$}|\text{\boldmath$w$},\text{\boldmath$y$}_{\mbox{\tiny obs}})$ as a block.\\ Step 2: Generate from $p(\text{\boldmath$w$}|\text{\boldmath$l$},\text{\boldmath$y$}_{\mbox{\tiny obs}})$ as a block.\\ Step 3: Generate from $p(\text{\boldmath$\theta$}|\text{\boldmath$l$},\text{\boldmath$w$},\text{\boldmath$y$}_{\mbox{\tiny obs}})$.
At Step 1, $\text{\boldmath$l$}$ is sampled first as a block by exploiting the fact that each element is independently constrained Gaussian, where its conditional posterior density is \[ p(\text{\boldmath$l$}|\text{\boldmath$w$},\text{\boldmath$y$}_{\mbox{\tiny obs}})=\prod_{i=1}^n \prod_{k=1}^q p(l_{ik}|\text{\boldmath$z$},w,\text{\boldmath$\theta$}) =\prod_{i=1}^n \prod_{k=1}^q\phi_{l_{ik}>0}(m_{ik},w_i^{-1}r_k)\,, \] $m_{ik}= \text{\boldmath$\delta$}_k^\top \Omega^{-1}\text{\boldmath$z$}_i$ and $r_k = 1 - \text{\boldmath$\delta$}_k^\top \Omega^{-1} \text{\boldmath$\delta$}_k$. The conditional independence of the elements of $\text{\boldmath$l$}$ is a key feature of the TrUST distribution, and does not apply to the general UST distribution.
At Step 2, $\text{\boldmath$w$}$ is sampled next, also as a block from its posterior $p(\text{\boldmath$w$}|\text{\boldmath$l$},\text{\boldmath$y$}_{\mbox{\tiny obs}})=\prod_{i=1}^n f_{Gam}(w_i;a,b_i)$, with $a = (d+q+\nu)/2$ and \[ b_i = \frac{1}{2}{\left[ {\left(\text{\boldmath$z$}_i - \Delta^\top \Sigma^{-1}\text{\boldmath$l$}_i\right)}^\top{\left(\Omega - \Delta^\top \Sigma^{-1} \Delta\right)}^{-1}{\left(\text{\boldmath$z$}_i - \Delta^\top \Sigma^{-1}\text{\boldmath$l$}_i\right)} + \text{\boldmath$l$}_i^\top \Sigma^{-1} \text{\boldmath$l$}_i + \nu \right]}. \]
At Step 3, each element of \(\text{\boldmath$\theta$}\) is sampled conditional on $\text{\boldmath$l$}$ and $\text{\boldmath$w$}$. A range of popular methods can be used here, and we simply sample each element one-at-a-time\footnote{Each constrained or bounded element of $\text{\boldmath$\theta$}$ is first transformed to the real line using a one-to-one transformation (e.g. logarithm for positive valued parameters) to simplify sampling.} (in random order to improve mixing) using adaptive random walk Metropolis-Hastings. The parameter constraint in Assumption (ref) is imposed after generating each element of $\Psi$ and $A$ as follows. Compute the diagonal matrix $H=\Sigma-\Delta \Omega^{-1} \Delta^\top$, where recall that $\Delta, \Omega$ and $\Sigma$ are all known functions of $\{\Psi,A\}$. Then simply reorder the columns of $\Delta$ to ensure the leading diagonal elements of $H$ are ordered as permutation $\pi^*$; i.e. in ascending order. To see that the order of the leading diagonal elements of $H$ corresponds to the order of the columns of $\Delta$, recall that $\Sigma$ is a function of $\Delta$ as in Lemma (ref).
The Bayesian method and increased flexibility of the TrUST distribution over the AC skew-t are first illustrated using a simulation, where we assume \(\text{\boldmath$\mu$}=\mathbf{0}\) and \(\text{\boldmath$s$}=(1,\ldots,1)^\top\) for simplicity. We generate \(n=1040\) observations from two data generating processes: DGP1 with $q=1$ (i.e. a skew-t), and DGP2 with \(q=2\). For both cases \(d=3\), \(\nu=10\), and elements of \(\Omega\) set to $\omega_{21}=0.5$, $\omega_{31}=0.3$, and \( \omega_{32}=0.811 \). The values of $A$ are:
For each DGP, TrUST distributions with \(q =1 \) and \(q = 2\) are estimated and the quality of fit measured using the Deviance Information Criterion (DIC) \[ \operatorname{DIC} = -4\,\mathbb{E}_{\text{\boldmath$\theta$}}[\log p(\text{\boldmath$y$}_{\mbox{\tiny obs}} \mid \text{\boldmath$\theta$})] + 2\log p(\text{\boldmath$y$}_{\mbox{\tiny obs}} \mid \text{\boldmath$\hat \theta$}), \] where \(p(\text{\boldmath$y$}_{\mbox{\tiny obs}}|\text{\boldmath$\theta$})=\prod_{i=1}^n f{\left(\text{\boldmath$y$}_i; \text{\boldmath$\mu$}, S, \Omega, A, \nu\right)} \) is evaluated as in (ref). The expectation \(\mathbb{E}_{\text{\boldmath$\theta$}}\) is computed from the MCMC samples after burn-in and \(\text{\boldmath$\hat \theta$}\) denotes the posterior mean.
Table (ref) presents the results. For DGP1 the misspecified TrUST distribution with $q=2$ (Panel B) is almost as accurate at the correctly specified skew-t distribution with $q=1$ (Panel A). In contrast, when the model is misspecified with too few latent variables, the impact is substantial with a much higher DIC value in Panel C compared to Panel D. Figure (ref) displays the major discrepancy in the two fitted densities for the case of DGP2, with the contours in the first row (Panels (a)–(c) for $q=1$) failing to align well with those in the second row (Panels (d)–(f) for $q=2$).
Over \$17.5bn of electricity is traded annually in the Australian National Electricity Market (NEM). This wholesale market comprises the five interconnected regions of New South Wales (NSW), Victoria (VIC), Queensland (QLD), South Australia (SA), and Tasmania (TAS). There are separate prices in each region, which exhibit complex dependencies and extreme positive skew; see panagiotelis2008bayesian for a discussion of this market and the distribution of prices.
The data comprise the $n=2070$ daily prices from 1 December 2018 to 31 July 2024 for the \(d = 5\) regions. We consider TrUST distributions with values of $q\in\{0,1,2,3\}$ and for both Gaussian (i.e. $\nu \rightarrow \infty$) and unconstrained degrees of freedom $\nu$. Eight distributions are fit to the logarithm of prices using the Bayesian method, and Table (ref) lists these, their dimension of $\text{\boldmath$\theta$}$ and DIC values. The posteriors means of $\nu$ are also given, indicating the extreme kurtosis in this data, and the means of the other parameters can be found in the Online Appendix. Overall, the TrUST distributions with \(q=2\) and $q=3$ have the lowest DIC values, and clearly dominate the other benchmarks, including the simpler skew-t distribution with $q=1$.
Figure (ref) gives bivariate marginal density contour plots for the fitted TrUST distributions with \(q \in \{0,1,2,3\}\) and variable pairs (a--d) QLD, TAS, and (e--h) SA, VIC. These are overlayed with scatterplots of the data, trimmed for extreme outliers for improved visualization only. Increasing $q$ impacts the asymmetry in the distribution substantially. Further contour plots for all pairwise slices of the TrUST distribution with $q=2$, along with five univariate marginals, are given in the Online Appendix. Overall, this example highlights that the TrUST distribution fits this five-dimensional data more accurately than the skew-t.
This section outlines the copula of the TrUST distribution, which extends the AC skew-t copula of Yoshiba_2018 and deng2024large to allow for greater variability in asymmetric dependence across variable pairs.
The copula of a continuous random vector $\text{\boldmath$Z$}$ is known as an implicit copula and is obtained by inversion of Sklar's theorem; see nelsen2006introduction and smith2021implicit. Let $\text{\boldmath$Z$} = {\left( Z_1, \ldots, Z_d\right)}^\top \sim \mathrm{TrUST}_q {\left(\Omega, A, \nu\right)}$, then from (ref) and (ref), the TrUST copula density is
Here, $\text{\boldmath$u$}=(u_1,\ldots,u_d)^\top \in [0,1]^d$, $z_j=F_{\mathrm{UST},q}^{-1}(u_j;1, \Delta_j, \Sigma, \nu)$ is the marginal quantile function of $Z_j$ evaluated at $u_j$ for $j=1,\ldots,d$, $\Delta_j$ is the $j$th column of $\Delta$,\footnote{Note that this is different than $\text{\boldmath$\delta$}_j$ which denotes the $j$th row of $\Delta$.} and $\text{\boldmath$z$}=(z_1,\ldots,z_d)^\top$. The copula parameters are $\text{\boldmath$\theta$}=\{\Omega,A,\nu \}$, and $\Sigma, \Delta$ are both known functions of $\text{\boldmath$\theta$}$ as in Section (ref).
To visualize the copula, Figure (ref) in the Online Appendix plots contours of $c$ for the bivariate TrUST copulas of the bivariate distributions with $q=2$ in Figure (ref). This illustrates how the parameter $\text{\boldmath$\alpha$}_2$ (which is an additional copula parameter over that of the skew-t copula where $q=1$) affects the dependence asymmetries greatly.
We now present expressions for both Kendall's correlation \(\rho_K\) and Spearman's correlation \(\rho_S\) for the UST distribution, which includes the TrUST subclass. As far as we are aware, these general expressions are given for the first time in the literature.
For the specific case of $q=2$, $\rho_{K, \mathrm{UST}}{\left(Z_1, Z_2\right)} = 4 \frac{T_{6}{\left(\mathbf{0}; R_K, \nu\right)}}{\left[1/4 + {\left(2\pi\right)}^{-1} \arcsin(\sigma)\right]^2} - 1$, where \(\sigma\) is the off-diagonal of \(\Sigma\). For the specific case of $q=1$ (i.e. the skew-t distribution) $\rho_{K, \mathrm{UST}}{\left(Z_1, Z_2\right)} = 16\,T_{4}{\left(\mathbf{0}; R_K, \nu\right)} - 1$, where the skew parameter is \(\Delta = \text{\boldmath$\delta$}^\top\).
For the specific case of $q=2$, $\rho_{S,\mathrm{UST}}{\left(Z_1, Z_2\right)} = 12\,\frac{T_{8}{\left(\mathbf{0}; R_S, \nu\right)}}{\left[1/4 + (2\pi)^{-1}\arcsin(\sigma)\right]^3} - 3$, where \(\sigma\) is the off-diagonal element of \(\Sigma\). For the specific case of $q=1$ (i.e. the skew-t distribution) $\rho_{S,\mathrm{UST}}{\left(Z_1, Z_2\right)} = 96\,T_{5}{\left(\mathbf{0}; R_S, \nu\right)} - 3$, where the skew parameter is \(\Delta = \text{\boldmath$\delta$}^\top\). When $q=0$, Lemmas (ref) and (ref) correspond to the usual expressions for an elliptical distribution.
For the TrUST distribution, these expressions can be computed for a pair of elements in $\text{\boldmath$Z$}$ using their bivariate marginal. When \(q=1\) the expressions have been presented previously in heinen2022kendall and lu2024kendall. Table (ref) reports the the rank correlations for the nine bivariate distributions with $q=2$ shown in Figure (ref) and their implied copulas.
To measure asymmetric dependence between any pair of variables $Y_1,Y_2$, differences between quantile dependence metrics are computed as follows. Let $U_1=F_1(Y_1)$ and $U_2=F_2(Y_2)$, then for quantile $\kappa\in(0,0.5]$ the quantile dependence metrics in the four quadrants are:
These probabilities are computed from the bivariate copula $C(U_1,U_2)$ of the joint distribution of $(Y_1,Y_2)$ as in Appendix (ref). Then we measure the level of asymmetry in the dependence along the major and minor diagonals of the unit square support of $C$ as
In most empirical applications, including when capturing dependence between equity returns, $\Lambda_{\scalebox{0.6}{Major}}(\kappa)$ is the more important of these two measures.
We now outline the Bayesian inference for the TrUST copula parameters $\text{\boldmath$\theta$}=\{\Omega,A,\nu\}$. For simplicity, the marginals of data $Y_{ij}\sim F_{Y_j}$ are assumed known, although joint estimation of these with $\text{\boldmath$\theta$}$ is possible by extending the MCMC scheme discussed here.
Let $U_j:=F_{\mathrm{UST},q}(Z_j;1, \Delta_j, \Sigma, \nu)$ for $j=1,\ldots,d$, where $\text{\boldmath$Z$}\sim \mbox{TrUST}_q(\Omega,A,\nu)$, and $f_{\mathrm{UST},q}=\frac{d}{dz_j} F_{\mathrm{UST},q}$ is the marginal density of $Z_j$ given at (ref). Then the joint density of $(\text{\boldmath$U$},\text{\boldmath$L$},W)$ can be evaluated using a change of variables from $\text{\boldmath$Z$}$ to $\text{\boldmath$U$}=(U_1,\ldots,U_d)^\top$ to obtain \[ f_{U,L,W}(\text{\boldmath$u$},\text{\boldmath$l$},w;\text{\boldmath$\theta$})=f_{Z,L,W}(\text{\boldmath$z$},\text{\boldmath$l$},w)/\prod_{j=1}^{d} f_{\mathrm{UST}, q}{\left(z_j; 1, \Delta_j, \Sigma, \nu\right)}\,, \] where $f_{Z,L,W}$ is defined previously at (ref), and the denominator arises from the Jacobian of the transformation. This density can be used to define an extended likelihood and augmented posterior for $\text{\boldmath$\theta$}$ and the latents $\text{\boldmath$l$},\text{\boldmath$w$}$.
For the observed copula data \(\text{\boldmath$u$}_{obs}= \{\text{\boldmath$u$}_i\}_{i=1}^n\), where \(\text{\boldmath$u$}_i=(u_{i1},\dots,u_{id})^\top\) and \(u_{ij}=F_{Y_j}(y_{ij})\), the extended likelihood is simply $\mathcal{L}{\left(\text{\boldmath$\theta$},\text{\boldmath$l$},\text{\boldmath$w$};\text{\boldmath$u$}_{obs}\right)} = \prod_{i=1}^n f_{U,L,W}(\text{\boldmath$u$}_i,\text{\boldmath$l$}_i,w_i;\text{\boldmath$\theta$})$. When combined with the same prior for $\text{\boldmath$\theta$}$ in Section (ref), the resulting Bayesian augmented posterior is
where the identification constraint for $\Delta$ is imposed through the prior as before.
The MCMC scheme at Algorithm 1 with some minor adjustments can be used to evaluate this augmented posterior. Generating $\text{\boldmath$l$}$ and $\text{\boldmath$w$}$ in Steps 1 and 2 is unchanged, except that $\text{\boldmath$z$}=\{\text{\boldmath$z$}_i\}_{i=1}^n$, with $\text{\boldmath$z$}_i=(z_{i1},\ldots,z_{id})^\top$, is computed from the copula data using the $d$ quantile functions \[ z_{ij}=F_{\mathrm{UST},q}^{-1}(u_{ij};1, \Delta_j, \Sigma, \nu)\,,\;\;j=1,\ldots,d\,;\quad i=1,\ldots,n\,. \] Evaluating $\text{\boldmath$z$}$ is the most demanding computation of the sampler. The quantile functions above are unavailable in closed form, so that we use the efficient and accurate numerical method in Yoshiba_2018 and Smith_Maneesoonthorn_2018. In Step 3 of Algorithm 1 each element of $\text{\boldmath$\theta$}$ is generated one element at a time using an adaptive random walk Metropolis-Hastings scheme. We use the same approach here, except that the target conditional posterior $p(\text{\boldmath$\theta$}|\text{\boldmath$l$},\text{\boldmath$w$},\text{\boldmath$u$}_{obs})$ is now given by (ref) (up to proportionality).
The flexibility of the TrUST copula and the efficacy of the Bayesian inference method are first illustrated using the data simulated from the two DGPs in Section (ref). The copula data $\text{\boldmath$u$}_{obs}$ was computed using the expressions for the univariate marginals $u_{ij}=F_{\mathrm{UST},q}(z_{ij};1, \Delta_j, \Sigma, \nu)$. The two DGPs have different levels of rank correlation and asymmetric dependence over variable pair. For the data from each DGP, we estimate two TrUST copulas: Copula 1 with \(q=1\) (i.e. skew-t copula) and Copula 2 with \(q=2\).
For DGP1, both Copula 1 and Copula 2 provide very similar and accurate estimates of the dependence metrics (see Table (ref) in the Online Appendix). Therefore, we focus on DGP2 where greater heterogeneity in dependence over variable pairs is apparent. Table (ref) reports the true rank correlations and asymmetric dependence metrics for DGP2 across the three variable pairs. The posterior mean and posterior standard deviations for these dependence metrics from the fitted copulas are also given. Copula 2 (the TrUST copula with $q=2$) captures these well, whereas Copula 1 (the skew-t copula with $q=1$) is unable to capture the variability in these metrics across variable pairs. Overall, the results suggest that a TrUST copula with $q\geq 2$ is preferable whenever dependence is heterogeneous over variable pairs, as the next section shows is the case for intraday equity returns.
Tail dependence between the returns on equity pairs is often asymmetric patton2004out,harvey2010 and recent studies suggest it also exhibits high levels of heterogeneity over asset pairs, including during the recent COVID-19 pandemic le2021covid,ando2022quantile. deng2024large found evidence of this in 15-minute equity returns using an AC skew-t copula model, and we extend their study here. We show that the TrUST copula with $q>1$ is more effective than the skew-t copula at capturing the dependence between the market volatility VIX index published by the Chicago Board of Exchange (CBOE) and returns of leading US stocks.
Observations on 15-minute equity returns and the VIX index were obtained from the LSEG DataScope Select Database. The intraday GARCH(1,1) model of Engle_Sokalska_2012 with student t innovations is used as a marginal for the equity returns. For the VIX index, the marginal is a first-order autoregression with a nonparametric disturbance given by a kernel density estimator. The copula data are computed from these marginal models. The same filtering approach was also employed by deng2024large, to whom we refer for further details.
The first study is of dependence during the COVID-19 crash between two key banking stocks, Bank of America (BAC) and JP Morgan (JPM), and the VIX. TrUST copulas with \(q=1\) and \(q=2\) of dimension $d=3$ are fit to the $n=1040$ intraday observations from February 11 to April 15, 2020. The posterior estimates of the copula parameters $\text{\boldmath$\theta$}$, pairwise rank correlations and asymmetric dependencies are given in Table (ref) in the Online Appendix. They show that increasing the number of latent variables in the TrUST copula from $q=1$ to $q=2$ has a strong impact on the estimated dependence structure. In particular, the DIC when $q=2$ is 5742, while for the skew-t copula with $q=1$ it is 6653, suggesting an improved fit even though the example is low dimensional.
To visualize the asymmetric dependence between each pair of variables we use a “quantile dependence plot” as in patton2012review. This is where \(\lambda_{\mathrm{LL}}(\kappa)\) (blue line) and \(\lambda_{\mathrm{UL}}(\kappa)\) (purple line) are plotted against quantile $\kappa$, and \(\lambda_{\mathrm{UR}}(\kappa)\) (red line), \(\lambda_{\mathrm{LR}}(\kappa)\) (yellow line) are plotted against $1-\kappa$. Asymmetric dependence in either the minor or major diagonal results in visual asymmetry around $\kappa=0.5$ in the plot. Moreover, greater consistency with the noisy empirical quantiles (black line) suggest an improved fit. Figure (ref) gives the quantile dependence plots for the pairs (BAC, VIX), (JPM, VIX), and (JPM, BAC) for the fitted TrUST copulas with $q=1$ and $q=2$. The quantile estimates with $q=2$ presented in panels (d)-(f) more closely align with the empirical quantiles for all three pairs, reinforcing the improvement also indicated by the DIC values.
Cross industry dependence in U.S. equity returns is estimated using $d=6$ dimensional TrUST copulas. The filtered copula data was computed for the VIX and leading stocks from five key sectors of the economy: Apple (AAPL) from Electronics, JPMorgan (JPM) from Banking, Visa (V) from Business Services, Exxon Mobil (XOM) from Energy, and United Parcel Service (UPS) from Freight/Transport. We consider two periods each with $n=1040$ observations: a pre-COVID period from 18 December 2018 to 15 February 2019, and the COVID period from 11 February 2020 to 15 April 2020. Market volatility was high during both periods, allowing investigation of how extreme market conditions influence dependence between the variables.
The TrUST copula model with \(q=1\) and \(q=2\) is estimated for both periods. Table (ref) reports the posterior means of the copula parameters for the pre-COVID period. The DIC for both copulas is also reported, indicating that the fit with $q=2$ provides a substantial improvement over the AC skew-t copula with $q=1$. This is also the case for the COVID period data, for which the estimates are given in Table (ref) in the Online Appendix.
Estimates of the asymmetric dependence measures $\Lambda_{\scalebox{0.6}{Major}}(0.05)$ and $\Lambda_{\scalebox{0.6}{Minor}}(0.05)$ for all variable pairs are presented as heatmaps in Figure (ref). The level of asymmetry varies across variable pairs in both periods and for both fitted copulas, particularly for $\Lambda_{\scalebox{0.6}{Major}}(0.05)$ which is the more important measure. However, there are substantive differences in direction and degree between the two copulas for some variable pairs due to the improved fit of the TrUST copula with $q=2$. Equivalent results are given for quantile $\tau=0.01$ in the Online Appendix.
To measure whether these differences are meaningful, we consider their impact on density forecasts. To do so, the out-of-sample one-step-ahead predictive distributions were computed from both copula models over 40 trading day evaluation periods (equivalent to 1040 intraday observations) following the two estimation periods. Figure (ref) presents the cumulative log-score (LS) differences between the TrUST copula models and a Student-t copula model with the same marginals. Panel (a) gives the results for the earlier pre-COVID estimation period, and panel (b) for the later COVID estimation period. Positive values of the cumulative LS difference indicate improved performance, and the TrUST copula with $q=2$ outperforms both the AC skew-t copula and the benchmark Student-t copula in both evaluation periods.
The UST distribution can capture richer asymmetry than the popular AC skew-t distribution which it nests. However, this potential has not been exploited in data analysis to date because inference is difficult. The current paper addresses this by specifying a novel tractable subclass for which likelihood-based inference can be computed using data augmentation methods. The flexibility of this TrUST distribution is illustrated empirically using both simulated and regional Australian electricity price data.
The implicit copulas of skew-elliptical distributions demarta_mcneil_2005,smith_gan_kohn_2010,lucas2017,Yoshiba_2018,oh_patton_2023,deng2024large are popular because they are both scalable and invariant to permutation of the random vector. A major contribution of the current paper is to show how the implicit copula of the TrUST distribution is also tractable, and how Bayesian data augmentation method can also be used to compute the posterior of its parameters. The TrUST copula with $q\geq 2$ allows for greater heterogeneity in asymmetric dependence across variable pairs than the AC skew-t copula. An important application is the modeling of dependence between equity returns, where an improvement is demonstrated using a TrUST copula model for the intraday returns of large U.S. equities.
We finish by discussing three possible future extensions of our work. First, in this paper we compute exact posterior inference for both the TrUST distribution and its copula using MCMC methods. However, approximate Bayesian inference methods---notably variational inference---are faster and have the potential to allow estimation for higher dimensions $d$. Second, the prior for $\Omega$ is uniform on its $d(d-1)/2$ hyper-spherical angles. For high values of $d$ these angles can be regularized, such as through Bayesian global-local priors. Alternatively, a factor decomposition can be used for $\Omega$, as in Murray_Dunson_Carin_Lucas_2013 or deng2024large, so that parameterization of $\Omega$ only increases linearly with $d$. Interestingly, this is very similar to the approach used in USE distributions to parameterize skew using $\Delta$. Here, $\Delta$ is a $(d\times q)$ matrix where $q$ is fixed, so that the number of skew parameters only increases linearly with $d$. Last, adopting a dynamic specification of the TrUST copula, as in the models of creal2015high,oh2017 and oh_patton_2023, would introduce time variation in the dependence structure.
\oldappendix {\appendixname Part \Alph{section}\quad}
The term “extended” follows from Azzalini_Capitanio_2003 that has hidden truncation on $ L + \tau > 0$.
The pairwise quantile dependencies of the TrUST copula are computed from the bivariate marginal copula function of \( (U_1, U_2) \), which is \[ C(u_1, u_2) = F_{\mathrm{UST}, q}{\left((z_1, z_2)^\top; \Omega, \Delta, \Sigma, \nu\right)} = \frac{T_{2+q}{\left((z_1, z_2, \mathbf{0}_q^\top)^\top; R^*, \nu \right)} }{ T_q{\left(\mathbf{0}^\top; \Sigma, \nu \right)} } \] with \[ R^* = \left(
\right), \] which can be derived from Equations (ref) and (ref) that \[ \text{$\mathds{P}$}{\left(\text{\boldmath$Z$} \leq \text{\boldmath$z$}\right)} = \text{$\mathds{P}$}{\left(\text{\boldmath$X$} \leq \text{\boldmath$z$} | \text{\boldmath$L$} > \mathbf{0}\right)} = \frac{\text{$\mathds{P}$}{\left(\text{\boldmath$X$} \leq \text{\boldmath$z$}, -\text{\boldmath$L$} \leq \mathbf{0}\right)} }{\text{$\mathds{P}$}{\left( -\text{\boldmath$L$} \leq \mathbf{0}\right)} }. \] Each pseudo observation \( z_{ij}=F_{\mathrm{UST},q}^{-1}(u_{ij};1, \Delta_j, \Sigma, \nu) \) is obtained via numerical optimization.
\singlespacing
\spacing{1.5} \setcounter{page}{1} \setcounter{figure}{0} \setcounter{table}{0} \setcounter{section}{0} \setcounter{equation}{0} \setcounter{algorithm}{1} \makeatletter \makeatother
This Online Appendix has four parts:
The augmentation leverages the hidden truncation mechanism, as described in Eq. (ref), within the framework of the joint Student-t distribution without scale mixture of normals. Specifically, the joint distribution \( {\left( \text{\boldmath$X$}^\top, \text{\boldmath$L$}^\top \right)}^\top \sim \mathrm{T}_{d+q}{\left(\mathbf{0}, R, \nu\right)} \). Then from this expression, we can derive the \(\text{\boldmath$Z$}\) in the conditional distribution
The augmented density of \({\left\{ \text{\boldmath$Z$}, \text{\boldmath$L$} \right\}} \) is
Still, the latent variable $\text{\boldmath$L$}$ can be sampled as a block by:
for \(k = 1,\ldots,q\), where $m_k= \text{\boldmath$\delta$}_k^\top \Omega^{-1}\text{\boldmath$Z$}$, $ r_k = 1 - \text{\boldmath$\delta$}_k^\top \Omega^{-1} \text{\boldmath$\delta$}_k$.
The generative representation in the Section (ref) using the scaling then hidden truncation ordering. The other method follows a sequence where hidden truncation is applied first, then scaling. First, we have the joint distribution \( (\text{$\widetilde{\text{\boldmath$X$}}$}^\top, \text{$\widetilde{\text{\boldmath$L$}}$}^\top )^\top \sim \mathrm{N}_{d+q}{\left(\mathbf{0}, R\right)}, \) then expressed the scaling as:
where $\odot$ is the element-wise multiplication.
Then the augmented likelihood is expressed as:
with the posterior distribution for latent variable $\text{\boldmath$L$}$ given by:
and the posterior of $W$ is not in closed form.
Both (GR1) and (GR2) provide alternative augmentations for the model, offering flexibility in computational implementation while adhering consistency with the underlying hidden truncation framework. The choice between representations, including the (GR1) without scaling, depends on the specific modeling objectives and computational requirements, ensuring adaptability to diverse data contexts.