The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
91,139 characters
myblue Tractable Unified Skew-t Distribution and Copula for Heterogeneous Asymmetries
\setlength{\abovedisplayskip}{0.15cm}
\setlength{\belowdisplayskip}{0.15cm}
\pagestyle{empty}
\begin{titlepage}
\title{\LARGE\bfseries\color{myblue}
Tractable Unified Skew-t Distribution and Copula for Heterogeneous Asymmetries}
\author{Lin Deng, Michael Stanley Smith and Worapree Maneesoonthorn}
\date{\today}
\maketitle
\noindent
{\small Lin Deng is a PhD student and Michael Smith is Professor of Management (Econometrics), both at the Melbourne Business School, University of Melbourne, Australia.
Worapree Maneesoothorn is Associate Professor at the Department of Econometrics and Statistics at Monash University, Australia. Correspondence should be directed to Michael Smith at {\tt [email removed]}.
\\
\noindent \textbf{Acknowledgments:} This research was supported by The University of Melbourne’s Research Computing Services and the Petascale Campus Initiative. Michael Stanley Smith research
has been partially supported by the Australian Research Council (ARC) Discovery Project grant DP250101069, and that
of Worapree Maneesoonthorn by ARC Discovery Project Grant DP200101414.}\\
\newpage
\begin{center}
\mbox{}\vspace{2cm}\\
{\title{\LARGE\bfseries\color{myblue}
Tractable Unified Skew-t Distribution and Copula for Heterogeneous Asymmetries}}\\
\vspace{1cm}
{\Large Abstract \\}
\end{center}
\vspace{-1pt}
\onehalfspacing
\noindent
{
Multivariate distributions that allow for asymmetry and heavy tails are important building blocks
in many econometric and statistical models. The Unified Skew-t (UST) is a promising
choice because it is both scalable and allows for a high level of flexibility in the asymmetry in the distribution.
However, it suffers from parameter identification and computational hurdles that have to date inhibited its use for modeling data.
In this paper we propose a new tractable variant of the unified skew-t (TrUST) distribution
that addresses both challenges.
Moreover, the copula of this distribution is shown to also
be tractable, while allowing for greater heterogeneity in asymmetric dependence over variable pairs than the popular skew-t copula.
We show how Bayesian posterior inference for both the distribution and its copula can be computed using an extended likelihood derived from a generative representation of the distribution. The efficacy of this Bayesian method, and the enhanced flexibility of both the TrUST distribution and its implicit copula, is first demonstrated using simulated data. Applications of the TrUST distribution to highly skewed regional Australian electricity prices, and the TrUST copula to intraday U.S. equity returns, demonstrate how our proposed distribution and its copula can provide substantial increases in accuracy over the popular skew-t and its copula in practice.
}
\vspace{20pt}
\noindent
{\bf Keywords}: Asymmetric Dependence, Bayesian Data Augmentation, Electricity Prices, Implicit Copulas, Intraday Equity Returns, Markov chain Monte Carlo, Skew-Elliptical Distributions.
\end{titlepage}
\newpage
\pagestyle{plain}
\setcounter{equation}{0}
\section{Introduction}\label{sec:intro}
Skew-elliptical distributions are models of multivariate asymmetry that scale well.
While there are different variants~(for some of these see~\citealp{genton2004skew})
those formed by conditioning on a latent truncated variable as suggested by~\cite{branco2001general} are prominent. This approach is often called ``hidden truncation'' and is used to form the skew-t distribution of~\cite{Azzalini_Capitanio_2003} (AC hereafter).
The AC skew-t distribution is a particularly popular choice in applied data analysis, with diverse applications in
renewable energy~\citep{Hering01032010}, actuarial studies~\citep{eling2012}, macroeconomic modeling~\citep{adrian2019vulnerable}, atmospheric science~\citep{morris2017space},
and in flow cytometric analysis~\citep{PyneetalPNAS2009,fruhwirth2010bayesian} among others. However, in $d$ dimensions skew-elliptical distributions typically have only $d$ parameters to control asymmetry across
$d(d-1)/2$ variable pairs, which can limit their effectiveness for modeling asymmetry. Unified skew-elliptical (USE)
\footnote{\cite{arellano2006unification} use the abbreviation SUN for the Unified Skew-Normal distribution, later adopting SUE for the Unified Skew-Elliptical family and SUT for the Unified Skew-t. In this paper, we break with this convention and adopt the word-ordered acronyms USN, USE and UST for these three distributions.}
distributions~\citep{arellano2010multivariate} are an extension that introduces more latent variables and parameters to allow for a much more flexible form of asymmetry.
In particular, the unified skew-t (UST) discussed by~\cite{wang2024multivariate} is a promising
generalization of the impactful AC skew-t distribution.
However, despite the strong potential of USE distributions, they have not been adopted for data analysis because (i)~the parameters are unidentified~\citep{wang2023non}, and (ii)~likelihood evaluation and optimization is computationally difficult in even low dimensions.
The current paper addresses these problems by
specifying a novel subclass of the USE that is tractable.
For this proposed subclass, it is shown that the USE parameters are identified
under an ordering of the eigenvalues of the conditional scale matrix of the latent variables.
The special case of a tractable UST distribution, which we label a TrUST distribution, is considered in detail.
A generative representation is given for the proposed TrUST distribution that
is used to define an extended likelihood that is easier to evaluate than the regular likelihood. From this, a Bayesian augmented posterior can be defined and computed using a
proposed Markov chain Monte Carlo (MCMC) sampler that is both fast and efficient, and
for which
the identifying eigenvalue constraint is easy to impose.
The proposed TrUST distribution nests the popular AC skew-t as a special case, but generalizes it
to allow for richer multivariate asymmetry. An application of the TrUST distribution to both simulated data and highly skewed regional daily Australian electricity prices shows it captures asymmetry more
accurately than the popular AC skew-t, and that this improves the quality of fit considerably.
However, the main advantage of the proposed TrUST distribution is that its implicit copula is also tractable and has parameters that are identified. Every multivariate
continuous distribution has a unique implicit copula; see~\cite{smith2021implicit} for
an introduction to this class of copulas. Currently, implicit copulas of differing skew-t distributions are used to capture the strong
asymmetric dependence often found in financial data; see~\cite{smith_gan_kohn_2010}, \cite{Christoffersen_Errunza_Jacobs_Langlois_2012}, \cite{creal2015high}, \cite{lucas2017}, \cite{opschoor2021}, \cite{oh_patton_2023} and \cite{deng2024large} for examples. Of these,~\cite{deng2024large} show that the implicit
copula of the AC skew-t distribution allows for the greatest level of asymmetric dependence. But the implicit copula of the TrUST distribution (hereafter the TrUST copula) allows for a much greater level of
heterogeneity in asymmetric dependencies over variable pairs than the AC skew-t copula. We show here that by doing so, the TrUST copula can better
capture the dependence between financial returns data.
To estimate the TrUST copula parameters a Bayesian augmented posterior is specified using
a modification of the extended likelihood of the TrUST distribution, which is then evaluated using an MCMC sampler. Expressions for the Kendall and Spearman
rank correlations for this TrUST copula are also derived.
As far as we are aware, ours is the first paper to construct the implicit copula of any USE
distribution and show how to compute statistical inference for its parameters.
\cite{deng2024large} use the implicit copula of the AC skew-t distribution to capture dependence between intraday equity returns.
In our main application, we extend their analysis to show that the TrUST copula is more effective at capturing such
dependence than the AC skew-t copula. This is done for five large equities and an
equity market volatility index (VIX) using data collected both before and during the COVID-19 pandemic. The performance of the TrUST copula is compared with that of both the symmetric Student-t copula and the AC skew-t copula using the Deviance Information Criterion (DIC) and cumulative log-score to assess accuracy. The results indicate that the TrUST copula better captures heterogeneity in asymmetric dependence across variables. High levels of heterogeneity is also found
by~\cite{le2021covid} and \cite{ando2022quantile} using network models of tail dependencies.
The rest of this paper is organized as follows. Section~\ref{sec:02} briefly introduces the USE distribution and the special case of the UST, after which our proposed tractable USE variant is outlined.
Section~\ref{sec:03} considers the case of the TrUST distribution in detail. Section~\ref{sec:04} outlines estimation methodology for the TrUST distribution,
and illustrates its efficacy through a simulated example and application to electricity prices. Section~\ref{sec:05} presents the TrUST copula, derivation of its rank correlations, Bayesian inference and a simulation example. Section~\ref{sec:06} shows how our TrUST copula captures heterogeneity in asymmetric dependence between high-frequency financial variables; Section~\ref{sec:07} concludes. Some extra results are given in the Appendix, while an Online Appendix contains additional information, figures, tables and all proofs.
\section{Unified Skew-Elliptical Distributions}\label{sec:02}
In this section, the unified skew-elliptical (USE) distribution is introduced, including the special case of the unified skew-t (UST). Our tractable variant of
the USE
distribution is then outlined in detail. We consider the standardized case (i.e. with zero mean and
unit scale) because this is used in forming the implicit copula in Section~\ref{sec:05}. Generalization of the distribution by
multiplying the random vector by a diagonal scaling matrix and adding a location parameter is straightforward.
\subsection{Unified Skew-Elliptical Distribution}
\cite{arellano2006unification} extend the popular skew-elliptical distribution
to allow for additional skew parameters.
It can be specified for the standardized case as follows.
Let $\text{\boldmath$X$}\in \text{$\mathds{R}$}^d$ and $\text{\boldmath$L$} \in \text{$\mathds{R}$}^q$, for $q\geq 1$, be jointly (radially symmetric) elliptically
distributed as
\begin{equation}\label{eq:sue_rep}
{\left( \begin{array}{c} \text{\boldmath$X$} \\ \text{\boldmath$L$} \end{array} \right)} \sim \mathrm{EL} {\left(\mathbf{0}, R, g\right)},
\quad R = {\left( \begin{array}{ll} \Omega & \Delta^\top \\ \Delta & \Sigma \end{array} \right)},
\end{equation}
where $R$ is a $(d+q) \times (d+q)$ correlation matrix partitioned to be consistent with $(\text{\boldmath$X$}^\top,\text{\boldmath$L$}^\top)^\top$, $\mathbf{0}$ is a zero mean vector\footnote{Throughout this paper we use $\mathbf{0}$ to denote a conformable vector or
matrix of zeros without explicitly denoting its dimension.},
and $g: [0, \infty) \to [0, \infty)$ is the density generator function; see~\citet[Chp. 2]{Fang_Kotz_Ng_1990} for specification of an elliptical distribution.
Then the USE distribution can be defined as
\begin{equation}\label{eq:cond_trunc}
\text{\boldmath$Z$} \overset{\mathrm{d}}{=} {\left(\text{\boldmath$X$} | \text{\boldmath$L$} > \mathbf{0}\right)}\,,
\end{equation}
where ``$\overset{\mathrm{d}}{=}$'' denotes equality in distribution, and the inequality $\text{\boldmath$L$}>\mathbf{0}$ holds element-wise. The vector $\text{\boldmath$L$}$ are latent variables and the
specification at~\eqref{eq:cond_trunc}
is often called ``hidden truncation'' \citep{arnold2000hidden,arnold2004elliptical}.
The $(q\times d)$ matrix $\Delta$ contains parameters that control the level and direction
of skew, where $q$ is typically low, such as $q\in\{1,2,3,4\}$.
Let
$f_{\mathrm{EL}}(\text{\boldmath$y$};V,g)$ and $F_{\mathrm{EL}}(\text{\boldmath$y$};V,g)$ denote the density and distribution
functions, respectively, of a zero-mean elliptically distributed random variable $\text{\boldmath$Y$}\sim \mathrm{EL} {\left(\mathbf{0}, V, g\right)}$. Also, for the partition $\text{\boldmath$Y$}=(\text{\boldmath$Y$}_1^\top,\text{\boldmath$Y$}_2^\top)^\top$, the conditional $\text{\boldmath$Y$}_1 \mid \text{\boldmath$Y$}_2$ is elliptically distributed~\citep{Fang_Kotz_Ng_1990} with density generator
denoted here as \( g_{\text{\boldmath$Y$}_2} \).
Then, the density of $\text{\boldmath$Z$}$ is obtained by using Bayes theorem to evaluate the
conditional density on the righthand side of~\eqref{eq:cond_trunc}, giving
\begin{equation}\label{eq:sue_pdf}
f_{\mathrm{USE}}{\left(\text{\boldmath$z$}; \Omega, \Delta, \Sigma, g \right)} = f_{\mathrm{EL}}{\left(\text{\boldmath$z$}; \Omega, g\right)} \frac{ F_{\mathrm{EL}}{\left(\Delta \Omega^{-1} \text{\boldmath$z$}; \Sigma - \Delta \Omega^{-1} \Delta^\top, g_{\text{\boldmath$X$}} \right)} }{ F_{\mathrm{EL}} {\left(\mathbf{0}; \Sigma, g \right)} }\,.
\end{equation}
If a random variable $\text{\boldmath$Z$}$ has this density, then we write
$\text{\boldmath$Z$} \sim \mathrm{USE}_{q}{\left( \Omega, \Delta, \Sigma, g\right)} $.
The number of elements in $\Delta$ increases linearly with $q$. In this way, $q$ plays a role for multivariate asymmetry that is analogous to that of the number of factors in a traditional factor model for the covariance.
When \(q = 1\), the USE distribution reduces to the skew-elliptical distribution introduced by \cite{branco2001general} and \cite{Azzalini_Capitanio_2003}. In this case there is only a single skew parameter for each dimension, so that
\(\Delta \equiv \text{\boldmath$\delta$}^\top = {\left(\delta_1,\ldots,\delta_d\right)}^\top \).
Because $\text{\boldmath$\delta$}$ is constrained, for computing inference it is common to
transform it to an unconstrained parameterization via the one-to-one transformation
$\text{\boldmath$\alpha$} = (1 - \text{\boldmath$\delta$}^{\top} \Omega^{-1} \text{\boldmath$\delta$})^{-1 / 2} \Omega^{-1} \text{\boldmath$\delta$}$ with inverse $\text{\boldmath$\delta$} = (1+\text{\boldmath$\alpha$}^{\top} \Omega \text{\boldmath$\alpha$})^{-1 / 2} \Omega \text{\boldmath$\alpha$}$.
\subsection{Unified Skew-t Distribution}\label{sec:ust}
The elliptical distribution
with greatest applied potential is the t distribution, which
has density generator function $g(x)=(1+x/\nu)^{-(\nu+d)/2}$ that introduces the parameter $\nu>0$.
Assuming
$(\text{\boldmath$X$}^\top,\text{\boldmath$L$}^\top)^\top$ is multivariate t with location zero, scale matrix $R$
and $\nu$ degrees of freedom, the density at~\eqref{eq:sue_pdf} is given by
\begin{equation}\label{eq:pdf_SUT}
f_{\mathrm{UST},q}(\text{\boldmath$z$}; \Omega,\Delta,\Sigma,\nu)= t_d(\text{\boldmath$z$} ; \Omega, \nu) \frac{T_q{\left( \sqrt{\frac{\nu+d}{\nu+Q(\text{\boldmath$z$})}} \Delta \Omega^{-1} \text{\boldmath$z$}; \Sigma - \Delta \Omega^{-1} \Delta^\top,\nu+d \right)}}{T_q{\left(\mathbf{0}; \Sigma, \nu\right)}}.
\end{equation}
Here, $Q(\text{\boldmath$z$}) = \text{\boldmath$z$}^\top\Omega^{-1}\text{\boldmath$z$}$, while \( t_d{\left(\text{\boldmath$y$}; V, \upsilon\right)} \) and \( T_d(\text{\boldmath$y$}; V, \upsilon) \) denote the density and distribution functions, respectively, of a zero mean $d$-dimensional multivariate t random variable with scale matrix $V$ and degrees of freedom \(\upsilon\) evaluated at $\text{\boldmath$y$}$. If a random variable $\text{\boldmath$Z$}$ has density at~\eqref{eq:pdf_SUT}, then we write
$\text{\boldmath$Z$} \sim \mathrm{UST}_{q}{\left(\Omega, \Delta, \Sigma, \nu\right)} $.
Three special cases of the UST distribution are:
(i)~when $\nu\rightarrow \infty$, the UST converges to the unified skew-normal of~\cite{arellano2006unification};
(ii)~when $\Delta$
is a matrix with all elements equal to zero, then $\text{\boldmath$Z$}$ is distributed symmetric multivariate t with location zero, scale matrix $\Omega$
and $\nu$ degrees of freedom; and (iii)~when \(q=1\), the UST distribution is the AC skew-t distribution.
\cite{wang2024multivariate} provide a comprehensive overview of the UST distribution and its properties.
\subsection{Tractable Unified Skew-Elliptical Distribution}\label{sec:truse}
The novel subclass of the USE distribution is now outlined,
that we call here a Tractable Unified Skew-Elliptical (TrUSE) distribution. It employs
the following assumption of conditional independence in the latent variables.
\begin{assumption}[Conditional Independence]\label{asmp:CI} Let $(\text{\boldmath$X$}^\top,\text{\boldmath$L$}^\top)^\top$ follow
the joint elliptical distribution at~\eqref{eq:sue_rep}. Then assume each component of $\text{\boldmath$L$}$ is independent, conditional on $\text{\boldmath$X$}$, so that $L_i \perp\!\!\!\!\perp L_j | \text{\boldmath$X$} $, for all ${\left(i,j\right)} \subset {\left\{1,\ldots,q\right\}}$ and $q > 1$.
\end{assumption}
Under this assumption, the scale matrix $\Sigma$ is a deterministic function of
$\{\Omega,\Delta\}$, and the density at~\eqref{eq:pdf_SUT}
is simplified, as summarized by the following two lemmas.
\begin{lemma}\label{lem1}
If Assumption~\ref{asmp:CI} holds, then the scale matrix of the marginal elliptical distribution of $\text{\boldmath$L$}$ in~\eqref{eq:sue_rep} is given by
$\Sigma = I_q + \left(M-\operatorname{diag}(M)\right)$, with $M=\Delta \Omega^{-1}\Delta^\top$.
\end{lemma}
\begin{lemma}\label{lem2}
If Assumption~\ref{asmp:CI} holds, then
\[H=\Sigma-\Delta \Omega^{-1} \Delta^\top = \mbox{diag}\left((1 - \text{\boldmath$\delta$}_1^\top \Omega^{-1} \text{\boldmath$\delta$}_1),\ldots,(1 - \text{\boldmath$\delta$}_q^\top \Omega^{-1} \text{\boldmath$\delta$}_q)\right)\,,\]
is a diagonal matrix, and the distribution function of $\text{\boldmath$L$}|\text{\boldmath$X$}$ is given by
\[
F_{\mathrm{EL}}{\left(\Delta \Omega^{-1} \text{\boldmath$z$}; \Sigma - \Delta \Omega^{-1} \Delta^\top, g_{\text{\boldmath$X$}} \right)}=
\prod_{k=1}^{q} F_{\mathrm{EL}} {\left(\text{\boldmath$\alpha$}_k^\top \text{\boldmath$z$};1, g_\text{\boldmath$X$}\right)}
\]
where $\Delta^\top = [\text{\boldmath$\delta$}_1|\text{\boldmath$\delta$}_2| \cdots| \text{\boldmath$\delta$}_q]$ and
$\text{\boldmath$\alpha$}_k = {\left(1 - \text{\boldmath$\delta$}_k^\top \Omega^{-1} \text{\boldmath$\delta$}_k\right)}^{-1/2} \Omega^{-1} \text{\boldmath$\delta$}_k$ for $k=1,\ldots,q$.
\end{lemma}
Proofs of both lemmas are found in Appendix~\ref{app:proof}, and together they can be used to define our proposed tractable USE distribution as follows.
\begin{definition}[TrUSE Distribution]
\label{def:truse_dist}
If $\text{\boldmath$Z$} \sim \mathrm{USE}_{q}{\left( \Omega, \Delta, \Sigma, g \right)}$ and Assumption~\ref{asmp:CI} holds, then $\text{\boldmath$Z$}$ is said to be distributed Tractable Unified Skew-Elliptical (TrUSE) with density function
\begin{equation}\label{eq:fse_pdf}
f_{\mathrm{TrUSE}, q}{\left(\text{\boldmath$z$}; \Omega, A, g \right)}
= f_{\mathrm{EL}}{\left(\text{\boldmath$z$}; \Omega, g\right)}
\frac{ \prod_{k=1}^{q} F_{\mathrm{EL}} {\left(\text{\boldmath$\alpha$}_k^\top \text{\boldmath$z$};1, g_\text{\boldmath$X$} \right)} }{ F_{\mathrm{EL}} {\left(\mathbf{0}; \Sigma, g \right)} }.
\end{equation}
where $\Sigma$ is given in Lemma~\ref{lem1}, $\text{\boldmath$\alpha$}_k$ is given in Lemma~\ref{lem2}, $g$ is the density generator function at~\eqref{eq:sue_rep},
and $g_\text{\boldmath$X$}$ is the density generator function of the elliptical distribution of $\text{\boldmath$L$}|\text{\boldmath$X$}$.
We write
$\text{\boldmath$Z$} \sim \mathrm{TrUSE}_{q}{\left(\Omega, A, g \right)}$ where $A=[\text{\boldmath$\alpha$}_1|\text{\boldmath$\alpha$}_2|\cdots|\text{\boldmath$\alpha$}_q]$.
\end{definition}
Parameterization of the TrUSE distribution in terms of $A$ is attractive because each $\text{\boldmath$\alpha$}_k\in \mathbb{R}^d$ is unbounded given $\{\Omega,\text{\boldmath$\alpha$}_{j\neq k}\}$, whereas $\text{\boldmath$\delta$}_k$ has
complex nonlinear bounds given $\{\Omega,\text{\boldmath$\delta$}_{j\neq k}\}$. We show later that this property is useful for constructing MCMC schemes to generate $\text{\boldmath$\alpha$}_k$, rather than $\text{\boldmath$\delta$}_k$, from its posterior.
The TrUSE distribution is more tractable than the general USE for three reasons. First, it simplifies
the parameter space because $\Sigma$ is a deterministic function of $\{\Omega,A\}$,
whereas for a general USE the off-diagonal elements of $\Sigma$ have complex nonlinear bounds given $\{\Omega,\Delta\}$.
Second, it is more computationally tractable to compute inference because (i) the density at~\eqref{eq:sue_pdf} is simplified, and (ii) efficient MCMC sampling is possible from the augmented posterior constructed using an extended likelihood based on the joint of $(\text{\boldmath$X$},\text{\boldmath$L$})$. Third, a parameter identification constraint that is straightforward to impose exists as outlined below.
In contrast, it is hard to compute statistical inference
for the unconstrained parameters of the general USE distribution. The tractability of TrUST distribution is illustrated in Section~\ref{sec:04}.
Adopting
Assumption~\ref{asmp:CI} alone is insufficient to identify the USE distribution
under permutation of $\text{\boldmath$L$}$. \cite{wang2023non} discuss this problem in detail, which is characterized by the lemma below.
\begin{lemma}[Permutation Un-identification]
\label{theo:PI}
Let \(G(q) = \{\pi : \pi \text{ is a bijection from } \{1, \ldots, q\} \text{ to itself}\}\) represents the set of all possible orderings of the indices \(\{1, \ldots, q\}\). Then, the distribution of $ \text{\boldmath$Z$} \stackrel{\mathrm{d}}{=} {\left(\text{\boldmath$X$} | \text{\boldmath$L$} > \mathbf{0}\right)} \sim \mathrm{TrUSE}{\left( \Omega, A, g\right)}$ is unidentified under permutation of latent variable $\text{\boldmath$L$}$.
That is, for every permutation $\pi \in G(q) $, $\text{\boldmath$Z$} \overset{\mathrm{d}}{=} \text{\boldmath$Z$}_{\pi}$, where $ \text{\boldmath$Z$}_{\pi} \overset{\mathrm{d}}{=} \text{\boldmath$X$} | \text{\boldmath$L$}_\pi > \mathbf{0}$ and $\text{\boldmath$L$}_\pi = {\left(L_{\pi(1)},\ldots,L_{\pi(q)}\right)}^\top$.
\end{lemma}
To identify the parameters a constraint based on an ordering of
the eigenvalues of the scale matrix $H$ of the conditional
distribution of $\text{\boldmath$L$}|\text{\boldmath$X$}$ is used.
\begin{assumption}[Latent Permutation Constraint]\label{asmp:LP}
Let \(\text{\boldmath$Z$} \overset{\mathrm{d}}{=} {\left(\text{\boldmath$X$} | \text{\boldmath$L$} > \mathbf{0}\right)}\) follow a USE distribution, and
$H =\Sigma-\Delta \Omega^{-1}\Delta^\top$ be the scale matrix of the conditional
distribution of $\text{\boldmath$L$}|\text{\boldmath$X$}$ with eigenvalues $\lambda_1,\ldots,\lambda_q$. Then we assume the permutation \(\pi^*\in G(q)\) of the latent vector $\text{\boldmath$L$}$ and associated columns of $\Delta^\top$, solves the problem
\begin{equation*}
\pi^* = \operatorname*{{arg\,max}}_{\pi \in G(q)} \sum_{k=1}^q \pi(k) \lambda_{\pi(k)}.
\end{equation*}
\end{assumption}
This assumption, when applied to the TrUSE distribution under Assumption~\ref{asmp:CI}, results in a permutation of
the latent vector $\text{\boldmath$L$}$ that identifies the columns of \(A\) and \(\Delta^\top\), as below.
\begin{theorem}[Latent Permutation Identification]\label{thm1}
Under Assumption~\ref{asmp:CI} , $H=\Sigma-\Delta \Omega^{-1} \Delta^\top=\mbox{diag}(h_1,\ldots,h_q)$ is a diagonal matrix, so that $\lambda_k=h_k$ for $k=1,\ldots,q$.
Under the additional Assumption~\ref{asmp:LP}, the permutation of $\text{\boldmath$L$}$ (and therefore also the columns of $A$ and $\Delta^\top$) in the TrUSE distribution is based on the optimal ordering:
\begin{equation*}
\pi^*: 0 \leq h_{\pi^*(1)} \leq \cdots \leq h_{\pi^*(q)} \leq 1,
\end{equation*}
where \( h_{\pi^*(k)} = 1 - \text{\boldmath$\delta$}_{\pi^*(k)}^\top \Omega^{-1} \text{\boldmath$\delta$}_{\pi^*(k)}\), for \(k = 1, \ldots, q\). Moreover, when these inequalities are strict, the ordering is unique.
\end{theorem}
In practice, because the posterior distribution of $(h_1,\ldots,h_q)^\top$ is continuous, the inequalities in Theorem~\ref{thm1} are strict, so that the ordering is unique.
In the remainder of the paper Assumption~\ref{asmp:LP} is adopted to identify the parameters in
the TrUSE (and TrUST) distributions, so that $\Delta, A, \Sigma$ are all defined with respect
to the unique ordering $\pi^*$. Imposing this constraint is straightforward within
an MCMC scheme as discussed in Section~\ref{sec:bayesdist}.
We finish this section with some comments on the appropriateness of the TrUSE distribution.
We stress that Assumption~\ref{asmp:CI} is on the distribution of $\bm{L}$, not on that of $\bm{X}$, so that the TrUSE distribution can capture the same degree of rank correlation as the
USE distribution. Moreover, Assumptions~\ref{asmp:CI} and~\ref{asmp:LP} are identifying assumptions that are necessary
for the USE distribution to be well-defined for data analysis. In comparison, \cite{wang2024multivariate} suggest (but do not implement) some other ways to identify the parameters
in a UST distribution that involve much stronger restrictions on $\Delta$ and/or $\Omega$.
However, these
reduce substantially the ability to capture variation in pairwise correlations and/or asymmetry across variable pairs, which is the key advantage of the UST over skew-t distributions. \cite{wang2023non} suggest identification is secured by fixing the order of the eigenvalues of \(\Sigma\), but this condition is insufficient when \(q\geq 2\) because interchanging rows of \(\Delta\) does not alter the order of the eigenvalues of $\Sigma$.
\section{Tractable Unified Skew-t Distribution}\label{sec:03}
The special case of the tractable unified skew-t (TrUST) distribution that is the
focus of our paper is now outlined in detail.
\subsection{Joint Distribution}
The TrUST distribution results from selecting a t distribution with $\nu$ degrees of freedom as the choice of elliptical distribution in Section~\ref{sec:truse}. In this case, the density at Definition~\ref{def:truse_dist} is
\begin{equation}\label{eq:trust_joint_pdf}
f_{\mathrm{TrUST},q} {\left(\text{\boldmath$z$}; \Omega, A, \nu\right)} = t_d(\text{\boldmath$z$}; \Omega, \nu) \frac{\prod_{k=1}^{q} T_1{\left(\sqrt{\frac{\nu+d}{\nu+Q(\text{\boldmath$z$})}} \text{\boldmath$\alpha$}_k^\top \text{\boldmath$z$}; 1, \nu+d \right)}}{T_q{\left(\mathbf{0}; \Sigma, \nu\right)}}.
\end{equation}
If a random vector $\text{\boldmath$Z$}$ has the above density, then we write $\text{\boldmath$Z$} \sim \mathrm{TrUST}_{q}{\left(\Omega, A, \nu\right)}$.
If $\nu \rightarrow \infty$, then this reduces to the tractable unified skew-normal (TrUSN) distribution, with density
\begin{equation*}
f_{\mathrm{TrUSN},q} {\left(\text{\boldmath$z$}; \Omega, A\right)} = \phi(\text{\boldmath$z$}; \Omega) \frac{\prod_{k=1}^{q} \Phi {\left(\text{\boldmath$\alpha$}_k^\top \text{\boldmath$z$}; 1\right)}}{\Phi {\left(\mathbf{0}; \Sigma\right)}},
\end{equation*}
where $\phi_d{\left(\text{\boldmath$x$}; \Omega\right)}$ and $\Phi_d{\left(\text{\boldmath$x$}; \Omega\right)}$ are density and distribution functions, respectively, for a $N{\left(\mathbf{0}, \Omega\right)}$ distribution. If a random vector $\text{\boldmath$Z$}$ has the above density, then
we write $\text{\boldmath$Z$} \sim \mathrm{TrUSN}_{q}{\left(\Omega,A\right)}$. If $q=1$ then the TrUST, UST and
AC skew-t all coincide.
Figure~\ref{fig:fst_pdfs} presents the bivariate density contours of the TrUST distribution, when $d=2$, $q=2$, $\nu = 5$ and the first skew parameter vector is \(\text{\boldmath$\alpha$}_1 = (5, 5)^\top\). Each row corresponds to a different correlation parameter \(\omega \in {\left\{-0.5, 0, 0.5\right\}}\), ordered from top to bottom. The columns correspond to \(\text{\boldmath$\alpha$}_2 \in {\left\{ (5, 5)^\top, (0, 5)^\top, (-5, 5)^\top \right\}}\) and show the impact of varying the second skew parameter vector.
\begin{figure}[htbp]
\centering
\includegraphics[width=1\textwidth]{figs/FST_9_plot_2}
\caption{
Contour plots of the bivariate TrUST \((q = 2)\) density with \(\nu = 5\) and \(\text{\boldmath$\alpha$}_1 = (5, 5)^\top\) fixed. Rows vary by correlation \(\omega \in \{-0.5, 0, 0.5\}\) (top to bottom), and columns by second skewness vector \(\text{\boldmath$\alpha$}_2 \in \{(5, 5)^\top, (0, 5)^\top, (-5, 5)^\top\}\) (left to right) which vary by only the first element. The shadowed contours represent the corresponding Student-t distributions for the same values of \(\omega\) and \(\nu\).
}
\label{fig:fst_pdfs}
\end{figure}
\subsection{Marginal Distribution}\label{sec:fst_margin}
Because the UST distribution is closed under marginalization,
the marginals of a TrUST distribution are UST with parameters derived from those of the joint as below.
\begin{lemma}\label{lem:marginal}
Let $\text{\boldmath$Z$} \sim \mathrm{TrUST}_{q}{\left(\Omega, A, \nu\right)}$ with $\Delta$ computed from $A$, and $\Sigma$ computed from $\{\Delta,\Omega\}$ as in Lemma~\ref{lem1}. If \( J \subset {\left\{1,\ldots,d\right\}} \) is a subset of $d_J<d$ indices, then the marginal distribution of the random vector $\text{\boldmath$Z$}_J = (Z_{J(1)},Z_{J(2)},\ldots,Z_{J(d_J)})$ is \( \text{\boldmath$Z$}_J \sim \mathrm{UST}_{ q} {\left(\Omega_J, \Delta_J, \Sigma, \nu\right)} \),
where $\Delta_{J} = {\left(\Delta_{J(1)}, \ldots, \Delta_{J(d_J)}\right)}$ contains the
columns of $\Delta$ indexed by $J$, and $\Omega_J=\{\omega_{ij}\}_{i,j\in J}$ is the sub-matrix of $\Omega=\{\omega_{ij}\}$ comprising the elements indexed by set $J$.
\end{lemma}
The case where $d_J=1$ is of particular importance when constructing the implicit
copula of a TrUST distribution in Section~\ref{sec:05}. Here, the marginal for the $j$-th element is $Z_j \sim \mathrm{UST}_{q}( 1, \Delta_{j}, \Sigma, \nu)$ for $j \in {\left\{1,\ldots,d\right\}}$ with density function
\begin{equation}\label{eq:trust_unimargin_pdf}
f_{\mathrm{UST},q} {\left(z_j; 1,\Delta_j, \Sigma, \nu\right)} = t_1(z_j;1, \nu) \frac{T_q {\left( \sqrt{\frac{\nu+1}{\nu+Q(z_j)}} \Delta_{j} z_j ; \Sigma - \Delta_{j} \Delta_{j}^\top, \nu+1 \right)}}{T_q{\left(\mathbf{0}; \Sigma, \nu \right)}}\,,
\end{equation}
where $\Delta_j$ is the $j$th column of $\Delta$ and $Q(z_j)=z_j^2$.
The distribution
and quantile functions are computed numerically from \eqref{eq:trust_unimargin_pdf} using the
method given in~\cite{Yoshiba_2018} and~\cite{Smith_Maneesoonthorn_2018}.
\subsection{Generative Representation}\label{sec:gr}
For generating from the TrUST distribution, we use a scale mixture of normals representation of
a t distribution at~\eqref{eq:sue_rep}. If $W \sim \mathrm{Gamma} {\left(\nu/2, \nu/2\right)}$, then
$(\text{\boldmath$X$}^\top,\text{\boldmath$L$}^\top)^\top|W\sim N(\mathbf{0},W^{-1}\odot R)$, where
`$\odot$'
denotes the product of a scalar with each element of a vector or matrix.
From this representation, a draw of $\text{\boldmath$Z$}$ from the TrUST distribution can be obtained by completing the following three sequential
steps: (i) generate
$W \sim \mathrm{Gamma} {\left(\nu/2, \nu/2\right)}$, (ii) draw $\text{\boldmath$L$}|W \sim N(\mathbf{0},W^{-1}\odot\Sigma)$ constrained
so that $\text{\boldmath$L$}>\mathbf{0}$, and (iii)
draw from the conditional distribution
\begin{equation*}
{\left(\text{\boldmath$Z$} | \text{\boldmath$L$}, W\right)} \sim N {\left( \Delta^\top \Sigma^{-1} \text{\boldmath$L$} , W^{-1} \odot {\left(\Omega - \Delta^\top \Sigma^{-1} \Delta\right)}\right)}\right)}}\,.
\end{equation*}
Because the $(q\times 1)$ vector $\text{\boldmath$L$}$ is low dimensional (e.g. $1\leq q \leq 3$ in our empirical work) generating from the constrained distribution at step~(ii) is both easy and fast using the approach of \cite{botev2017normal} or another method.
From the generative algorithm, the joint density of $(\text{\boldmath$Z$},\text{\boldmath$L$},W)$ is
\begin{equation}\label{eq:jnt}
f_{Z,L,W}(\text{\boldmath$z$},\text{\boldmath$l$},w)=\phi_d\left(\text{\boldmath$z$};\Delta^\top \Sigma^{-1} \text{\boldmath$l$}, w^{-1} \odot (\Omega - \Delta^\top \Sigma^{-1} \Delta)\right) \phi_{\text{\boldmath$l$}>\mathbf{0}}(\text{\boldmath$l$} ;\mathbf{0},w^{-1}\odot\Sigma) f_{Gam}(w;\nu/2,\nu/2)\,,
\end{equation}
where $f_{Gam}(w;\alpha,\beta)$ denotes the density of a $\mbox{Gamma}(\alpha,\beta)$ distribution, and $\phi_{\text{\boldmath$x$}>\mathbf{0}}(\text{\boldmath$x$} ;\mathbf{0},V)$ is the density of a $N(\mathbf{0},V)$ distribution constrained so that $\text{\boldmath$x$}>\mathbf{0}$. Equation~\eqref{eq:jnt} is used to define an
extended likelihood when computing Bayesian inference in Section~\ref{sec:04} below.
Two other alternative generative representations of the TrUST distribution are given in Part~A of the Online Appendix. However, that above is preferred because it provides
for a numerically stable extended likelihood.
\subsection{Discussion of Distribution}
We make four additional comments on the proposed TrUST distribution. First, it is defined with zero mean and correlation matrix $R$ because it is convenient for formation of the implicit copula in
Section~\ref{sec:05}, where location and scale are unidentified in the copula. Second, generalization of the distribution is straightforward by adding a location parameter
and multiplying by scale parameters, as we do when modeling electricity prices in Section~\ref{sec:04}.
Third, Assumptions~\ref{asmp:CI} and~\ref{asmp:LP} can also be adopted in
the ``extended UST'', where an additional parameter $\text{\boldmath$\tau$} \in \text{$\mathds{R}$}^q$
is introduced and $\text{\boldmath$Z$}\overset{d}{=}\text{\boldmath$X$}|\text{\boldmath$L$}+\text{\boldmath$\tau$}>\mathbf{0}$ \citep{Azzalini_Capitanio_2003,arellano2006unification,wang2024multivariate}. Doing so further
generalizes the TrUST distribution; see Appendix~\ref{sec:etrust} for details. Fourth, the conditional distribution of a TrUST distribution is of this extended form, as discussed in Appendix~\ref{sec:condtrust}.
\section{Bayesian Inference and Application}\label{sec:04}
This section outlines how to compute Bayesian inference for the parameters of the TrUST distribution, and applies
the approach to both simulated data and highly skewed Australian regional electricity prices.
\subsection{Extended Likelihood}\label{sec:bayesdist}
The inferential problem is for the location-scaled version of the TrUST distribution, where \(\text{\boldmath$Y$} = {\left(\text{\boldmath$\mu$} + S \text{\boldmath$Z$}\right)}\), \( \text{\boldmath$Z$} \sim \mathrm{TrUST}_{q}{\left(\Omega, A, \nu\right)} \), $\text{\boldmath$\mu$} = {\left(\mu_1, \ldots, \mu_d\right)}^\top$ denotes the location vector, and $S = \operatorname{diag}{\left(\text{\boldmath$s$}\right)}$ is a diagonal scale matrix with leading diagonal $\text{\boldmath$s$} = {\left(s_{1}, \ldots, s_{d}\right)}^\top$. The vector \(\text{\boldmath$Y$}\) has density
\begin{equation}
\begin{aligned}
f{\left(\text{\boldmath$y$}; \text{\boldmath$\mu$}, S, \Omega, A, \nu\right)}
&= \det{\left(S\right)}^{-1} f_{\mathrm{TrUST},q}{\left(S^{-1}{\left(\text{\boldmath$y$}-\text{\boldmath$\mu$}\right)}; \Omega, A, \nu\right)}, A, \nu}\,.
\end{aligned} \label{eq:fy_trust}
\end{equation}
The likelihood based on~\eqref{eq:fy_trust} can exhibit a complex geometry, so that both its direct optimization and evaluation of the resulting Bayesian posterior can be difficult.
To simplify the problem the following extended likelihood is employed.
Let $\text{\boldmath$y$}_{\mbox{\tiny obs}}=\{\text{\boldmath$y$}_i\}_{i=1}^n$ be the observed data, with $\text{\boldmath$y$}_i=(y_{i1},\ldots,y_{id})^\top $ the $i$th observation of $\text{\boldmath$Y$}$. Also let $\text{\boldmath$w$} = \{w_i\}_{i=1}^n$ and $\text{\boldmath$l$} = \{\text{\boldmath$l$}_i\}_{i=1}^n$, with $\text{\boldmath$l$}_i = {\left(l_{i1}, \ldots, l_{iq}\right)}^\top$, be the corresponding $n$ values
of latent $W$ and $\text{\boldmath$L$}$. Then if $\text{\boldmath$\theta$}$ denotes the model parameters,
an extended likelihood based on the joint of $(\text{\boldmath$Y$},\text{\boldmath$L$},W)$ is
\[
{\cal L}(\text{\boldmath$\theta$},\text{\boldmath$l$},\text{\boldmath$w$};\text{\boldmath$y$}_{\mbox{\tiny obs}})=\mbox{det}(S)^{-n}\prod_{i=1}^n f_{Z,L,W}(\text{\boldmath$z$}_i,\text{\boldmath$l$}_i,w_i;\text{\boldmath$\theta$})\,.
\]
Here, $\text{\boldmath$z$}_i=S^{-1}(\text{\boldmath$y$}_i-\text{\boldmath$\mu$})$, while the density $f_{Z,L,W}$ is given at~\eqref{eq:jnt} and written here as a function of $\text{\boldmath$\theta$}$. Another advantage of the extended likelihood is that its evaluation does not involve computation of the
distribution functions in~\eqref{eq:trust_joint_pdf}, unlike with the conventional likelihood based on~\eqref{eq:fy_trust}.
\subsection{Parameterization, Prior and Augmented Posterior}\label{sec:augpost}
An effective parameterization of a correlation matrix is in terms of hyper-spherical
angles as in \citep{Rebonato_Jäckel_1999}.
We adopt this for \(\Omega = B B^\top\) by setting $B = \{b_{ij}\}$ to a lower triangular Choleskey factor with elements
\begin{equation*}
b_{ii} =
\begin{dcases*}
1 & for $i = 1$, \\
\prod_{k=1}^{i-1} \sin(\psi_{ik})& for $i > 1$,
\end{dcases*}
\quad \text{and} \quad
b_{ij} =
\begin{dcases*}
\cos(\psi_{i1}) & for $j = 1$, \\
\cos(\psi_{ij}) \prod_{k=1}^{j-1} \sin(\psi_{ik}) & for $j= 2,...,d-1$.
\end{dcases*} ,
\end{equation*}
that are functions of the angles $\psi_{i(i-1)}\in(0,2\pi]$ and $\psi_{ij}\in (0,\pi]$ for $j<i-1$.
The bounds on each angle $\psi_{ij}$ are unchanged when conditioning on the other angles, making
them an attractive parameterization for MCMC sampling as in~\cite{creal2011dynamic}. Denoting
the $d(d-1)/2$ unique angles as $\Psi=\{\psi_{ij}\}$, the parameters \(\text{\boldmath$\theta$}= \{\Psi, A, \nu,\text{\boldmath$\mu$},\text{\boldmath$s$}\}\).
The angles have uniform independent priors
\begin{equation}
\psi_{ij}
\sim \left\{\begin{array}{ll}
\mathrm{Unif}(\epsilon, 2 \pi - \epsilon) & \text { for } j = i-1
\\
\mathrm{Unif}(\epsilon, \pi - \epsilon) & \text { for } j < i-1\,,
\end{array}\right.
\label{eq:psi_prior}
\end{equation}
with \(\epsilon = 0.03\) chosen to bound $\Omega$ away from singular values.
Independent proper Gaussian priors are assigned to each component of \(A\), such that \(\alpha_{jk} \sim N_1(0, 25)\). The prior for \((\nu - 2) \sim \mathrm{Gamma}(3, 0.2)\), where the location shift ensures \(\nu > 2\), and we adopt the reference prior $p(\text{\boldmath$\mu$},\text{\boldmath$s$})\propto \prod_{j=1}^d s_j^{-1}$ as in \cite{liseo2006note} and others.
The prior
density is $p(\text{\boldmath$\theta$})\propto p_0(\text{\boldmath$\theta$})\mathds{1}(\Delta\in {\cal P}(\Omega))$, where $p_0(\text{\boldmath$\theta$})$
is the product of
the prior densities outlined above, and
the indicator function $\mathds{1}(X)=1$ if $X$ is true, and zero otherwise. This indicator function
constrains $\Delta$ to a set ${\cal P}(\Omega)$, which contains all values of $\Delta$ that satisfy
Assumption~\ref{asmp:LP}. This is where the elements of the diagonal matrix $H=\Sigma-\Delta \Omega^{-1}\Delta^\top$ (which is a deterministic function of $\{\Delta,\Omega\}$) are in ascending order so that ${\cal P}$ will depend on the value of $\Omega$ and we write it as such.
The Bayesian augmented posterior
\[
p(\text{\boldmath$\theta$},\text{\boldmath$l$},\text{\boldmath$w$}|\text{\boldmath$y$}_{\mbox{\tiny obs}})\propto {\cal L}(\text{\boldmath$\theta$},\text{\boldmath$l$},\text{\boldmath$w$};\text{\boldmath$y$}_{\mbox{\tiny obs}}) p_0(\text{\boldmath$\theta$})\mathds{1}(\Delta\in {\cal P}(\Omega))\,,
\]
is computed using the MCMC sampler below.
\noindent {\underline {\bf Algorithm~1}} {\em (MCMC Sampler for TrUST Distribution Parameters)}\\
\indent Step 1: Generate from $p(\text{\boldmath$l$}|\text{\boldmath$w$},\text{\boldmath$y$}_{\mbox{\tiny obs}})$ as a block.\\
\indent Step 2: Generate from $p(\text{\boldmath$w$}|\text{\boldmath$l$},\text{\boldmath$y$}_{\mbox{\tiny obs}})$ as a block.\\
\indent Step 3: Generate from $p(\text{\boldmath$\theta$}|\text{\boldmath$l$},\text{\boldmath$w$},\text{\boldmath$y$}_{\mbox{\tiny obs}})$.
At Step~1, $\text{\boldmath$l$}$ is sampled first as a block by exploiting the fact that each element is independently constrained Gaussian, where its conditional posterior density is
\[
p(\text{\boldmath$l$}|\text{\boldmath$w$},\text{\boldmath$y$}_{\mbox{\tiny obs}})=\prod_{i=1}^n \prod_{k=1}^q p(l_{ik}|\text{\boldmath$z$},w,\text{\boldmath$\theta$}) =\prod_{i=1}^n \prod_{k=1}^q\phi_{l_{ik}>0}(m_{ik},w_i^{-1}r_k)\,,
\]
$m_{ik}= \text{\boldmath$\delta$}_k^\top \Omega^{-1}\text{\boldmath$z$}_i$ and $r_k = 1 - \text{\boldmath$\delta$}_k^\top \Omega^{-1} \text{\boldmath$\delta$}_k$. The conditional independence of the elements of $\text{\boldmath$l$}$ is a key feature of
the TrUST distribution, and does not apply to the general UST distribution.
At Step~2, $\text{\boldmath$w$}$ is sampled next, also as a block from its posterior
$p(\text{\boldmath$w$}|\text{\boldmath$l$},\text{\boldmath$y$}_{\mbox{\tiny obs}})=\prod_{i=1}^n f_{Gam}(w_i;a,b_i)$,
with $a = (d+q+\nu)/2$ and
\[
b_i = \frac{1}{2}{\left[ {\left(\text{\boldmath$z$}_i - \Delta^\top \Sigma^{-1}\text{\boldmath$l$}_i\right)}^\top{\left(\Omega - \Delta^\top \Sigma^{-1} \Delta\right)}^{-1}{\left(\text{\boldmath$z$}_i - \Delta^\top \Sigma^{-1}\text{\boldmath$l$}_i\right)} + \text{\boldmath$l$}_i^\top \Sigma^{-1} \text{\boldmath$l$}_i + \nu \right]}.
\]
At Step~3, each element of \(\text{\boldmath$\theta$}\) is sampled conditional on $\text{\boldmath$l$}$ and $\text{\boldmath$w$}$. A range of
popular methods can be used here, and we simply sample each element
one-at-a-time\footnote{Each constrained or bounded element of $\text{\boldmath$\theta$}$ is first transformed to the real line using
a one-to-one transformation (e.g. logarithm for positive valued parameters) to
simplify sampling.} (in random order to improve mixing) using adaptive random walk Metropolis-Hastings. The parameter constraint in Assumption~\ref{asmp:LP} is imposed after generating each element of $\Psi$ and $A$ as follows.
Compute the diagonal matrix $H=\Sigma-\Delta \Omega^{-1} \Delta^\top$, where recall that $\Delta, \Omega$ and $\Sigma$ are all known functions of $\{\Psi,A\}$. Then simply reorder the columns of $\Delta$ to ensure the
leading diagonal elements of $H$ are ordered as permutation $\pi^*$; i.e. in ascending order. To see that the order of the leading diagonal elements of $H$ corresponds to the order of the columns of $\Delta$, recall that $\Sigma$ is a function of $\Delta$ as in Lemma~\ref{lem1}.
\subsection{Simulation Example}\label{sec:sim_dist}
The Bayesian method and increased flexibility of the TrUST distribution over the AC skew-t are first illustrated using a simulation, where we assume \(\text{\boldmath$\mu$}=\mathbf{0}\) and \(\text{\boldmath$s$}=(1,\ldots,1)^\top\)
for simplicity.
We generate \(n=1040\) observations from two data generating processes: DGP1 with $q=1$ (i.e. a skew-t), and DGP2 with \(q=2\). For both cases \(d=3\), \(\nu=10\), and elements of \(\Omega\) set to $\omega_{21}=0.5$, $\omega_{31}=0.3$, and \( \omega_{32}=0.811 \). The values of $A$ are:
\begin{itemize}
\item DGP1: \(A = (-5, 3, 5)^\top\), implying \(\Delta = {\left(-0.271, 0.618, 0.805\right)}^\top\).
\item DGP2: \(A = [\text{\boldmath$\alpha$}_1|\text{\boldmath$\alpha$}_2]\), where \(\text{\boldmath$\alpha$}_1 = (5, 3, 5)^\top\), \(\text{\boldmath$\alpha$}_2 = (-10, 0, 5)^\top\), implying $\Delta^\top = [\text{\boldmath$\delta$}_1| \text{\boldmath$\delta$}_2]$, where \(\text{\boldmath$\delta$}_1 = {\left(0.748, 0.894, 0.835\right)}^\top\), and \(\text{\boldmath$\delta$}_2 = {\left(-0.868, -0.097, 0.204\right)}^\top\).
\end{itemize}
\begin{table}[htbp]
\begin{center}
\caption{Posterior Mean Parameter Estimates for Simulation Example~1} \label{tab:sim_dist_case_1}
\resizebox{0.75\textwidth}{!}{
\begin{tabular}{ccccccccccrc}
\toprule
\toprule
Variable & & \multicolumn{3}{c}{$\Omega$} & & $\text{\boldmath$\delta$}_1$ & $\text{\boldmath$\delta$}_2$ & & $\nu$ & & DIC \\
\cmidrule{1-1}\cmidrule{3-5}\cmidrule{7-8}\cmidrule{10-10}\cmidrule{12-12} & & \multicolumn{10}{l}{Panel A: DGP1, fitted with $q=1$ (Correctly Specified)} \\
\cmidrule{3-12}$Z_1$ & & 1.000 & & & & -0.285 & --- & & \multirow{3}[1]{*}{10.658} & & \multirow{3}[1]{*}{\textbf{6680}} \\
$Z_2$ & & 0.504 & 1.000 & & & 0.606 & --- & & & & \\
$Z_3$ & & 0.294 & 0.815 & 1.000 & & 0.802 & --- & & & & \\
& & & & & & & & & & & \\
& & \multicolumn{10}{l}{Panel B: DGP1, fitted with $q=2$ (Misspecified)} \\
\cmidrule{3-12}$Z_1$ & & 1.000 & & & & -0.250 & -0.239 & & \multirow{3}[2]{*}{10.734} & & \multirow{3}[2]{*}{6702} \\
$Z_2$ & & 0.539 & 1.000 & & & 0.585 & 0.599 & & & & \\
$Z_3$ & & 0.341 & 0.814 & 1.000 & & 0.793 & 0.757 & & & & \\
\midrule
\midrule
& & \multicolumn{10}{l}{Panel C: DGP2, fitted with $q=1$ (Misspecified)} \\
\cmidrule{3-12}$Z_1$ & & 1.000 & & & & -0.270 & --- & & \multirow{3}[1]{*}{12.880} & & \multirow{3}[1]{*}{6413} \\
$Z_2$ & & 0.229 & 1.000 & & & 0.758 & --- & & & & \\
$Z_3$ & & -0.016 & 0.787 & 1.000 & & 0.951 & --- & & & & \\
& & & & & & & & & & & \\
& & \multicolumn{10}{l}{Panel D: DGP2, fitted with $q=2$ (Correctly Specified)} \\
\cmidrule{3-12}$Z_1$ & & 1.000 & & & & 0.761 & -0.864 & & \multirow{3}[2]{*}{10.548} & & \multirow{3}[2]{*}{\textbf{4997}} \\
$Z_2$ & & 0.530 & 1.000 & & & 0.901 & -0.133 & & & & \\
$Z_3$ & & 0.314 & 0.803 & 1.000 & & 0.829 & 0.190 & & & & \\
\bottomrule
\bottomrule
\end{tabular}
}
\end{center}
Note: Results are given the four combinations of DGP and TrUST distribution fit to the simulated data. Numbers reported are the posteriors means, except for final column which
reports the DIC values. Lower DIC indicates a better fit.
\end{table}
For each DGP, TrUST distributions with \(q =1 \) and \(q = 2\) are estimated and the quality of fit measured using
the Deviance Information Criterion (DIC)
\[ \operatorname{DIC} = -4\,\mathbb{E}_{\text{\boldmath$\theta$}}[\log p(\text{\boldmath$y$}_{\mbox{\tiny obs}} \mid \text{\boldmath$\theta$})] + 2\log p(\text{\boldmath$y$}_{\mbox{\tiny obs}} \mid \text{\boldmath$\hat \theta$}), \]
where \(p(\text{\boldmath$y$}_{\mbox{\tiny obs}}|\text{\boldmath$\theta$})=\prod_{i=1}^n f{\left(\text{\boldmath$y$}_i; \text{\boldmath$\mu$}, S, \Omega, A, \nu\right)} \) is evaluated as in~\eqref{eq:fy_trust}. The expectation \(\mathbb{E}_{\text{\boldmath$\theta$}}\) is computed from the MCMC samples after burn-in and \(\text{\boldmath$\hat \theta$}\) denotes the posterior mean.
Table~\ref{tab:sim_dist_case_1} presents the results. For DGP1 the misspecified TrUST distribution
with $q=2$ (Panel B) is almost as accurate at the correctly specified skew-t distribution with $q=1$ (Panel A).
In contrast, when the model is misspecified with too few latent variables, the impact is substantial with a much higher DIC value in Panel C compared to Panel D. Figure~\ref{fig:sim_dist_case_2} displays the major discrepancy in the two fitted densities for the case of DGP2, with the contours in the first row (Panels (a)–(c) for $q=1$) failing to align well with those in the second row (Panels (d)–(f) for $q=2$).
\begin{figure}[htbp]
\centering
\includegraphics[width=0.9\textwidth]{figs/Sim_TrUST_dist_Case_2_compare}
\caption{Contour plots of bivariate slices of the TrUST densities fitted to data from DGP2. Panels (a)-(c) give the contours when $q=1$, and panels (d)-(f) when $q=2$. The bivariate scatterplots are of the generated data and are the same in the first and second rows.}
\label{fig:sim_dist_case_2}
\end{figure}
\subsection{Daily Australian Electricity Prices}\label{sec:dist_app}
Over \$17.5bn of electricity is traded annually in the Australian National Electricity Market (NEM). This wholesale market comprises the five interconnected regions of New South Wales (NSW), Victoria (VIC), Queensland (QLD), South Australia (SA), and Tasmania (TAS). There are separate
prices in each region, which exhibit complex dependencies and extreme positive skew; see~\cite{panagiotelis2008bayesian} for a discussion of this market and the distribution of prices.
The data comprise the $n=2070$ daily prices from 1 December 2018 to 31 July 2024 for the \(d = 5\) regions. We consider TrUST distributions with values of $q\in\{0,1,2,3\}$ and for both Gaussian (i.e. $\nu \rightarrow \infty$) and unconstrained degrees of freedom $\nu$. Eight distributions are fit to the logarithm of prices using the Bayesian method, and Table~\ref{tab:nemdist} lists these, their dimension of $\text{\boldmath$\theta$}$ and DIC values.
The posteriors means of $\nu$ are also given, indicating the extreme kurtosis in this data, and the means of the other parameters can be found in
the Online Appendix. Overall,
the TrUST distributions with \(q=2\) and $q=3$ have the lowest DIC values, and clearly dominate the other benchmarks, including
the simpler skew-t distribution with $q=1$.
\begin{table}
\caption{Different TrUST Distributions fit to the Electricity Price Dataset}
\label{tab:nemdist}
\resizebox{1\textwidth}{!}{
\begin{tabular}{lcccccccccccccc}
\toprule
\toprule
Value of $q$ & & 0 & & 1 & & 2 & 3 & & 0 & & 1 & & 2 & 3 \\
\cmidrule{1-1}\cmidrule{3-3}\cmidrule{5-5}\cmidrule{7-8}\cmidrule{10-10}\cmidrule{12-12}\cmidrule{14-15}Dist. Name & & Gaussian & & \multicolumn{1}{l}{Skew-Normal} & & \multicolumn{2}{c}{TrUSN} & & Student-t & & Skew-t & & \multicolumn{2}{c}{TrUST} \\
$\mbox{dim}(\text{\boldmath$\theta$})$ & & 20 & & 25 & & 30 & 35 & & 21 & & 26 & & 31 & 36 \\
DIC & & -4943 & & -5863 & & -6280 & -6949 & & -12704 & & -13535 & & \textbf{-13756} & -13683 \\
$\mathbb{E}(\nu|\text{\boldmath$y$}_{\mbox{\tiny obs}} )$ & & $\infty$ & &$\infty$ & & $\infty$ & $\infty$ & & 2.189 & & 2.024 & & 2.029 & 2.031 \\
\bottomrule
\bottomrule
\end{tabular}
}
\caption*{Note: All distributions were fit using the Bayesian method outlined. The left-hand columns correspond to the Gaussian cases where $\nu\rightarrow \infty$. Lower values of DIC indicate better fit, with the lowest value indicated in bold. The full set of parameter estimates are given in the Online Appendix.}
\end{table}
\begin{figure}[htbp]
\centering
\includegraphics[width=1\textwidth]{figs/App_Dist_Pair_QLD_TAS}
\includegraphics[width=1\textwidth]{figs/App_Dist_Pair_SA_VIC}
\caption{
Contour plots of bivariate marginal densities from the fitted TrUST distributions.
These are given for the pairs (QLD, TAS) (top row) and (SA, VIC) (bottom row), while the columns correspond to TrUST distributions with \(q=0\) (Student-t), \(q=1\) (Skew-t), \(q=2\), and \(q=3\).}
\label{fig:app_NEM}
\end{figure}
Figure~\ref{fig:app_NEM} gives bivariate marginal density contour plots for the fitted TrUST distributions with \(q \in \{0,1,2,3\}\) and variable pairs (a--d) QLD, TAS, and (e--h) SA, VIC.
These are overlayed with scatterplots of the data, trimmed for extreme outliers for improved visualization only.
Increasing $q$ impacts the asymmetry in the distribution substantially. Further contour plots for all pairwise slices of the TrUST distribution with $q=2$, along with five univariate marginals, are given in the Online Appendix. Overall, this example highlights that the TrUST distribution fits this five-dimensional data more accurately than the skew-t.
\section{TrUST Copula}\label{sec:05}
This section outlines the copula of the TrUST distribution, which
extends the AC skew-t copula of~\cite{Yoshiba_2018} and~\cite{deng2024large} to allow for greater variability in asymmetric dependence across variable pairs.
\subsection{Copula Specification}
The copula of a continuous random vector $\text{\boldmath$Z$}$ is known as an implicit copula and is obtained by inversion of Sklar's theorem; see~\citet[p.51]{nelsen2006introduction} and~\cite{smith2021implicit}.
Let $\text{\boldmath$Z$} = {\left( Z_1, \ldots, Z_d\right)}^\top \sim \mathrm{TrUST}_q {\left(\Omega, A, \nu\right)}$,
then from~\eqref{eq:trust_joint_pdf} and~\eqref{eq:trust_unimargin_pdf}, the TrUST copula density is
\begin{equation}\label{eq:trust_copula_pdf}
c{\left(\text{\boldmath$u$}; \text{\boldmath$\theta$}\right)} = \frac{
f_{\mathrm{TrUST}, q}{\left(\text{\boldmath$z$}; \Omega, A, \nu\right)}
}{
\prod_{j=1}^{d} f_{\mathrm{UST}, q}{\left(z_j; 1, \Delta_j, \Sigma, \nu\right)}
}\,.
\end{equation}
Here, $\text{\boldmath$u$}=(u_1,\ldots,u_d)^\top \in [0,1]^d$, $z_j=F_{\mathrm{UST},q}^{-1}(u_j;1, \Delta_j, \Sigma, \nu)$ is the
marginal quantile function of $Z_j$ evaluated at $u_j$ for $j=1,\ldots,d$, $\Delta_j$ is the $j$th
column of $\Delta$,\footnote{Note that this is different than $\text{\boldmath$\delta$}_j$ which denotes the
$j$th row of $\Delta$.} and $\text{\boldmath$z$}=(z_1,\ldots,z_d)^\top$. The copula parameters
are
$\text{\boldmath$\theta$}=\{\Omega,A,\nu \}$, and $\Sigma, \Delta$ are both known functions of $\text{\boldmath$\theta$}$ as in Section~\ref{sec:03}.
To visualize the copula, Figure~\ref{fig:fstc_pdfs} in the Online Appendix plots contours of $c$ for the bivariate TrUST copulas of the bivariate distributions with $q=2$ in Figure~\ref{fig:fst_pdfs}.
This illustrates how the parameter $\text{\boldmath$\alpha$}_2$ (which is an additional copula parameter over that of the skew-t copula where $q=1$) affects the dependence asymmetries greatly.
\subsection{Rank Correlations}
\label{sec:copula_prop}
We now present expressions for both Kendall's correlation \(\rho_K\) and Spearman's correlation \(\rho_S\) for the UST distribution, which includes the TrUST subclass.
As far as we are aware, these general expressions are given for the first time in the literature.
\begin{lemma}[Kendall's Correlation of UST]\label{theo:kendall}
Let \( {\left(Z_1, Z_2\right)}^\top \sim \mathrm{UST}_{q}{\left(\Omega, \Delta, \Sigma, \nu\right)} \), then the Kendall's correlation for this pair is
\begin{equation}\label{eq:kendall}
\rho_{K, \mathrm{UST}}{\left(Z_1, Z_2\right)} = 4 \frac{ T_{2+2q} {\left(\mathbf{0}; R_K, \nu \right)}}{ {\left[T_{q} {\left(\mathbf{0}; \Sigma, \nu\right)}\right]}^2 } - 1,
\end{equation}
where
\begin{equation*}
R_K = \left(\begin{array}{lc}
2\Omega & \Delta_K^\top \\
\Delta_K & I_2 \otimes \Sigma \\
\end{array}\right)\,,
\quad \quad
\Delta_K = \left(\begin{array}{c}
\Delta \\
-\Delta \\
\end{array}\right)\,,
\end{equation*}
and \( \otimes \) denotes the Kronecker product.
\end{lemma}
For the specific case of $q=2$,
$\rho_{K, \mathrm{UST}}{\left(Z_1, Z_2\right)} = 4 \frac{T_{6}{\left(\mathbf{0}; R_K, \nu\right)}}{\left[1/4 + {\left(2\pi\right)}^{-1} \arcsin(\sigma)\right]^2} - 1$, where \(\sigma\) is the off-diagonal of \(\Sigma\). For the specific case of $q=1$ (i.e. the skew-t distribution)
$\rho_{K, \mathrm{UST}}{\left(Z_1, Z_2\right)} = 16\,T_{4}{\left(\mathbf{0}; R_K, \nu\right)} - 1$,
where the skew parameter is \(\Delta = \text{\boldmath$\delta$}^\top\).
\begin{lemma}[Spearman's Correlation of UST]
\label{theo:spearman}
Let ${\left(Z_1, Z_2\right)}^\top \sim \mathrm{UST}_{q}{\left(\Omega, \Delta, \Sigma, \nu\right)}$, then Spearman's correlation for this pair has is:
\begin{equation*}
\rho_{S}{\left(Z_1, Z_2\right)} = 12 \frac{T_{2+3q} {\left(\mathbf{0}; R_S, \nu \right)}}{{\left[T_q {\left( \mathbf{0}; \Sigma, \nu \right)}\right]}^3} - 3,
\end{equation*}
where
\begin{equation*}
R_S = \left(\begin{array}{cc}
\Omega + I_2 & \Delta_S^\top \\
\Delta_S & I_3 \otimes \Sigma \\
\end{array}\right) \,,
\quad \quad
\Delta_S = \left(\begin{array}{c}
\Delta \\
-{\left(\Delta_1 \oplus \Delta_2\right)}
\end{array}\right)\,,
\end{equation*}
and $\oplus$ denotes the direct sum operation that stacks each vector in \(\Delta\) diagonally to construct a block diagonal matrix.
\end{lemma}
For the specific case of $q=2$, $\rho_{S,\mathrm{UST}}{\left(Z_1, Z_2\right)} = 12\,\frac{T_{8}{\left(\mathbf{0}; R_S, \nu\right)}}{\left[1/4 + (2\pi)^{-1}\arcsin(\sigma)\right]^3} - 3$, where \(\sigma\) is the off-diagonal element of \(\Sigma\). For the specific case of $q=1$ (i.e. the skew-t distribution)
$\rho_{S,\mathrm{UST}}{\left(Z_1, Z_2\right)} = 96\,T_{5}{\left(\mathbf{0}; R_S, \nu\right)} - 3$,
where the skew parameter is \(\Delta = \text{\boldmath$\delta$}^\top\). When $q=0$, Lemmas~\ref{theo:kendall} and~\ref{theo:spearman} correspond to the usual expressions
for an elliptical distribution.
For the TrUST distribution, these expressions can be computed for a pair of elements in $\text{\boldmath$Z$}$ using their bivariate marginal. When \(q=1\) the expressions have been presented previously in \cite{heinen2022kendall} and \cite{lu2024kendall}. Table~\ref{tab:rankcorr} reports the the rank correlations for the nine bivariate distributions with $q=2$ shown in Figure~\ref{fig:fst_pdfs} and their implied copulas.
\begin{table}[htbp]
\centering
\caption{Rank Correlations for the TrUST Distributions in Figure~\ref{fig:fst_pdfs} and their Implicit Copulas }
\begin{tabular}{c
>{\centering\arraybackslash}p{1.5cm}
>{\centering\arraybackslash}p{1.5cm}
>{\centering\arraybackslash}p{1.5cm}
@{\hspace{1em}}
>{\centering\arraybackslash}p{1.5cm}
>{\centering\arraybackslash}p{1.5cm}
>{\centering\arraybackslash}p{1.5cm}}
\toprule
\toprule
& \multicolumn{3}{c}{Kendall Correlation} & \multicolumn{3}{c}{Spearman Correlation} \\
\cmidrule(lr){2-4} \cmidrule(lr){5-7}
$\omega$ & \multicolumn{3}{c}{Second Skew Parameter $\text{\boldmath$\alpha$}_2^\top$} & \multicolumn{3}{c}{Second Skew Parameter $\text{\boldmath$\alpha$}_2^\top$} \\
\cmidrule(lr){1-1} \cmidrule(lr){2-4} \cmidrule(lr){5-7}
& $[5,5]$ & $[0,5]$ & $[-5,5]$ & $[5,5]$ & $[0,5]$ & $[-5,5]$ \\
-0.5 & -0.589 & -0.417 & -0.298 & -0.776 & -0.588 & -0.434 \\
0 & -0.338 & -0.185 & 0.000 & -0.475 & -0.270 & 0.000 \\
0.5 & -0.017 & 0.083 & 0.298 & -0.010 & 0.130 & 0.434 \\
\bottomrule
\bottomrule
\end{tabular}
\label{tab:rankcorr}
\caption*{Note: The parameter values are given in the caption to Figure~\ref{fig:fst_pdfs} , where the first skew vector is $\text{\boldmath$\alpha$}_1 = [5,5]^\top$ in all nine panels.}
\end{table}
\subsection{Asymmetric Dependence Measurements}
To measure asymmetric dependence between any pair of variables $Y_1,Y_2$,
differences
between quantile dependence metrics are computed as follows. Let $U_1=F_1(Y_1)$ and $U_2=F_2(Y_2)$, then for quantile $\kappa\in(0,0.5]$ the quantile
dependence metrics in the four quadrants are:
\begin{align*}
& \text{Lower Left:}
\;\lambda_{\mathrm{LL}}(\kappa) \coloneq \mathbb{P}(U_2 < \kappa \mid U_1 < \kappa),
& \text{Upper Right:}
\; \lambda_{\mathrm{UR}}(\kappa) \coloneq \mathbb{P}(U_2 > 1-\kappa \mid U_1 > 1-\kappa), \\
& \text{Lower Right:}
\; \lambda_{\mathrm{LR}}(\kappa) \coloneq \mathbb{P}(U_2 > 1-\kappa \mid U_1 < \kappa),
& \text{Upper Left:}
\; \lambda_{\mathrm{UL}}(\kappa) \coloneq \mathbb{P}(U_2 < \kappa \mid U_1 > 1-\kappa).
\end{align*}
These probabilities are computed from the bivariate copula $C(U_1,U_2)$ of the joint
distribution of $(Y_1,Y_2)$ as in Appendix~\ref{app:quant_dep}. Then we measure
the level of asymmetry in the dependence along the major and minor
diagonals of the unit square support of $C$ as
\begin{equation}\label{eq:asymdep_quant}
\Lambda_{\scalebox{0.6}{Major}}(\kappa) \coloneq \lambda_{\scalebox{0.6}{UR}}(\kappa) - \lambda_{\scalebox{0.6}{LL}}(\kappa),
\quad \text{and} \quad
\Lambda_{\scalebox{0.6}{Minor}}(\kappa) \coloneq \lambda_{\scalebox{0.6}{UL}}(\kappa) - \lambda_{\scalebox{0.6}{LR}}(\kappa)\,.
\end{equation}
In most empirical applications, including when capturing dependence between equity returns,
$\Lambda_{\scalebox{0.6}{Major}}(\kappa)$ is the more important of these two measures.
\subsection{Bayesian Inference for the Copula Parameters}\label{sec:bayescop}
We now outline the Bayesian inference for the TrUST copula parameters $\text{\boldmath$\theta$}=\{\Omega,A,\nu\}$. For simplicity, the marginals of data $Y_{ij}\sim F_{Y_j}$ are assumed known, although joint estimation of these with $\text{\boldmath$\theta$}$ is possible by extending the MCMC scheme discussed here.
Let $U_j:=F_{\mathrm{UST},q}(Z_j;1, \Delta_j, \Sigma, \nu)$ for $j=1,\ldots,d$, where $\text{\boldmath$Z$}\sim \mbox{TrUST}_q(\Omega,A,\nu)$, and $f_{\mathrm{UST},q}=\frac{d}{dz_j} F_{\mathrm{UST},q}$ is the marginal density of $Z_j$ given at~\eqref{eq:trust_unimargin_pdf}. Then the joint density of $(\text{\boldmath$U$},\text{\boldmath$L$},W)$ can be evaluated using a change of variables from $\text{\boldmath$Z$}$ to $\text{\boldmath$U$}=(U_1,\ldots,U_d)^\top$ to obtain
\[
f_{U,L,W}(\text{\boldmath$u$},\text{\boldmath$l$},w;\text{\boldmath$\theta$})=f_{Z,L,W}(\text{\boldmath$z$},\text{\boldmath$l$},w)/\prod_{j=1}^{d} f_{\mathrm{UST}, q}{\left(z_j; 1, \Delta_j, \Sigma, \nu\right)}\,,
\]
where $f_{Z,L,W}$ is defined previously at~\eqref{eq:jnt}, and the denominator arises from the Jacobian of the transformation. This density can be used to define an extended likelihood and
augmented posterior for $\text{\boldmath$\theta$}$ and the latents $\text{\boldmath$l$},\text{\boldmath$w$}$.
For the observed copula data \(\text{\boldmath$u$}_{obs}= \{\text{\boldmath$u$}_i\}_{i=1}^n\), where \(\text{\boldmath$u$}_i=(u_{i1},\dots,u_{id})^\top\) and \(u_{ij}=F_{Y_j}(y_{ij})\), the extended likelihood is
simply $\mathcal{L}{\left(\text{\boldmath$\theta$},\text{\boldmath$l$},\text{\boldmath$w$};\text{\boldmath$u$}_{obs}\right)}
= \prod_{i=1}^n f_{U,L,W}(\text{\boldmath$u$}_i,\text{\boldmath$l$}_i,w_i;\text{\boldmath$\theta$})$.
When combined with the same prior for $\text{\boldmath$\theta$}$ in Section~\ref{sec:04}, the resulting Bayesian
augmented posterior is
\begin{equation} \label{eq:coppost}
p(\text{\boldmath$\theta$},\text{\boldmath$l$},\text{\boldmath$w$}|\text{\boldmath$u$}_{obs}) \propto \mathcal{L}{\left(\text{\boldmath$\theta$},\text{\boldmath$l$},\text{\boldmath$w$};\text{\boldmath$u$}_{obs}\right)} p_0(\text{\boldmath$\theta$})\mathds{1}(\Delta\in {\cal P}(\Omega))\,,
\end{equation}
where the identification constraint for $\Delta$ is imposed through the prior as before.
The MCMC scheme at Algorithm~1 with some minor adjustments can be used to evaluate this augmented posterior. Generating
$\text{\boldmath$l$}$ and $\text{\boldmath$w$}$ in Steps~1 and~2 is unchanged, except that $\text{\boldmath$z$}=\{\text{\boldmath$z$}_i\}_{i=1}^n$, with $\text{\boldmath$z$}_i=(z_{i1},\ldots,z_{id})^\top$, is computed
from the copula data using the $d$ quantile functions
\[
z_{ij}=F_{\mathrm{UST},q}^{-1}(u_{ij};1, \Delta_j, \Sigma, \nu)\,,\;\;j=1,\ldots,d\,;\quad i=1,\ldots,n\,.
\]
Evaluating $\text{\boldmath$z$}$ is the most demanding computation of the sampler.
The quantile functions above are unavailable in closed form, so that we use the efficient and accurate
numerical method in~\cite{Yoshiba_2018} and~\cite{Smith_Maneesoonthorn_2018}.
In Step~3 of Algorithm~1 each element of $\text{\boldmath$\theta$}$ is generated one element at a time using
an adaptive random walk Metropolis-Hastings scheme. We use the same approach here,
except that the target conditional posterior $p(\text{\boldmath$\theta$}|\text{\boldmath$l$},\text{\boldmath$w$},\text{\boldmath$u$}_{obs})$ is now given by~\eqref{eq:coppost} (up to proportionality).
\subsection{Simulation Example Continued} \label{sec:copsim}
The flexibility of the TrUST copula
and the efficacy of the Bayesian inference method
are first illustrated using the data simulated from the two DGPs in Section~\ref{sec:sim_dist}.
The copula data $\text{\boldmath$u$}_{obs}$ was computed using
the expressions for the univariate marginals $u_{ij}=F_{\mathrm{UST},q}(z_{ij};1, \Delta_j, \Sigma, \nu)$. The two DGPs have different levels of rank correlation and asymmetric dependence over variable pair. For the data from each DGP, we estimate two TrUST copulas: Copula 1 with \(q=1\) (i.e. skew-t copula) and Copula 2 with \(q=2\).
\begin{table}[bh]
\centering
\caption{Dependence Metrics for Data Generated using DGP2 (where \( q = 2 \))}
\resizebox{1.0\textwidth}{!}{
\begin{tabular}{ccrccrccrcc}
\toprule
\toprule
\multicolumn{2}{c}{Pair ${\left(U_i, U_j\right)}$} & & \multicolumn{2}{c}{True Values} & & \multicolumn{2}{c}{Copula 1 Estimates} & & \multicolumn{2}{c}{Copula 2 Estimates} \\
\cmidrule{1-2}\cmidrule{4-5}\cmidrule{7-8}\cmidrule{10-11} & & & \multicolumn{8}{l}{Panel A: Rank Correlations} \\
\cmidrule{4-11}$i$ & $j$ & & Kendall & Spearman & & Kendall & Spearman & & Kendall & Spearman \\
\multirow{2}[0]{*}{2} & \multirow{2}[0]{*}{1} & & 0.159 & 0.259 & & 0.133 & 0.206 & & 0.156 & 0.244 \\
& & & & & & (0.012) & (0.017) & & (0.013) & (0.019) \\
\multirow{2}[0]{*}{3} & \multirow{2}[0]{*}{1} & & 0.136 & 0.133 & & 0.108 & 0.166 & & 0.135 & 0.189 \\
& & & & & & (0.013) & (0.019) & & (0.013) & (0.019) \\
\multirow{2}[1]{*}{3} & \multirow{2}[1]{*}{2} & & 0.392 & 0.584 & & 0.391 & 0.556 & & 0.486 & 0.663 \\
& & & & & & (0.011) & (0.015) & & (0.008) & (0.010) \\
\midrule
\midrule
& & & \multicolumn{8}{l}{Panel B: Asymmetric Dependencies} \\
\cmidrule{4-11}$i$ & $j$ & & $\Lambda_{\scalebox{0.6}{Major}}(0.05)$ & $\Lambda_{\scalebox{0.6}{Minor}}(0.05)$ & & $\Lambda_{\scalebox{0.6}{Major}}(0.05)$ & $\Lambda_{\scalebox{0.6}{Minor}}(0.05)$ & & $\Lambda_{\scalebox{0.6}{Major}}(0.05)$ & $\Lambda_{\scalebox{0.6}{Minor}}(0.05)$ \\
\multirow{2}[0]{*}{2} & \multirow{2}[0]{*}{1} & & 0.338 & -0.052 & & 0.265 & -0.020 & & 0.368 & -0.035 \\
& & & & & & (0.037) & (0.016) & & (0.039) & (0.015) \\
\multirow{2}[0]{*}{3} & \multirow{2}[0]{*}{1} & & 0.362 & -0.105 & & 0.261 & -0.023 & & 0.363 & -0.067 \\
& & & & & & (0.031) & (0.019) & & (0.039) & (0.020) \\
\multirow{2}[1]{*}{3} & \multirow{2}[1]{*}{2} & & 0.459 & 0.002 & & 0.164 & 0.000 & & 0.454 & 0.000 \\
& & & & & & (0.051) & (0.004) & & (0.052) & (0.001) \\
\bottomrule
\bottomrule
\end{tabular}
}
\label{tab:sim_case_2}
\caption{Rank correlations and asymmetric dependence (at quantile $\kappa=0.05$) are reported for the three variable pairs. The true values from DGP2 are reported. The posterior mean estimates for the two fitted copulas are given, with the posterior
standard deviations in parentheses below.}
\end{table}
For DGP1, both Copula~1 and Copula~2 provide very similar and accurate estimates of the dependence metrics (see Table~\ref{tab:sim_case_1} in the Online Appendix). Therefore, we focus on DGP2 where greater heterogeneity in dependence over variable pairs is apparent.
Table~\ref{tab:sim_case_2} reports the true rank correlations and asymmetric dependence metrics for DGP2 across the three variable pairs. The posterior mean and posterior standard deviations for these
dependence metrics from the fitted copulas are also given.
Copula~2 (the TrUST copula with $q=2$) captures these well, whereas Copula~1 (the skew-t copula with $q=1$) is unable to capture the variability in these metrics across variable pairs.
Overall, the results suggest that a TrUST copula with $q\geq 2$ is preferable whenever
dependence is heterogeneous over variable pairs, as the next section shows is the case for
intraday equity returns.
\section{Asymmetric Dependence in Intraday Equity Returns}\label{sec:06}
Tail dependence between the returns on equity pairs is often asymmetric~\citep{patton2004out,harvey2010} and recent studies
suggest it also
exhibits high levels of heterogeneity over asset pairs, including during the recent COVID-19 pandemic~\citep{le2021covid,ando2022quantile}.
\cite{deng2024large} found evidence of this in
15-minute equity returns using an AC~skew-t copula model, and we extend their study here.
We show that the TrUST copula with $q>1$ is more effective than the skew-t copula
at capturing the dependence between the market volatility VIX index published by the Chicago Board of Exchange (CBOE) and returns of leading US stocks.
\subsection{Data and Marginal Models}
Observations on 15-minute equity returns and the VIX index were obtained from the LSEG DataScope Select Database. The intraday GARCH(1,1) model of~\cite{Engle_Sokalska_2012} with student t innovations is used as a marginal for the equity returns. For the VIX index,
the marginal is a first-order autoregression with a nonparametric disturbance given by a kernel density estimator.
The copula data are computed from these marginal models.
The same filtering approach was also employed by \cite{deng2024large}, to whom we refer for further details.
\subsection{Banking Stocks Example (3-dim)}
The first study is of dependence during the COVID-19 crash between two key banking stocks, Bank of America (BAC) and JP Morgan (JPM),
and the VIX. TrUST copulas with \(q=1\) and \(q=2\) of dimension $d=3$ are
fit to the $n=1040$ intraday observations from February 11 to April 15, 2020.
The posterior estimates of the copula parameters $\text{\boldmath$\theta$}$, pairwise rank correlations and asymmetric
dependencies are given in Table~\ref{tab:app_3d} in the Online Appendix. They
show that increasing the number of latent variables in the TrUST copula from $q=1$ to $q=2$ has
a strong impact on the estimated dependence structure.
In particular, the DIC when $q=2$ is 5742, while for the skew-t copula with $q=1$ it is 6653, suggesting an improved fit even though the example is low dimensional.
\begin{figure}[htbp]
\centering
\includegraphics[width=0.9\textwidth]{figs/App_3D_q_1}
\includegraphics[width=0.9\textwidth]{figs/App_3D_q_2}
\caption{
Quantile dependence plots for \(\lambda_{\mathrm{LL}}\) (blue), \(\lambda_{\mathrm{UR}}\) (red), \(\lambda_{\mathrm{LR}}\) (yellow), and \(\lambda_{\mathrm{UL}}\) (purple) are shown with empirical dependencies (black) for (BAC, VIX), (JPM, VIX), and (JPM, BAC) from left to right. The top and bottom rows depict TrUST copula estimates for \(q=1\) and \(q=2\) respectively.
}
\label{fig:app_3d}
\end{figure}
To visualize the asymmetric
dependence between each pair of variables we use a ``quantile dependence plot'' as in~\cite{patton2012review}. This is where \(\lambda_{\mathrm{LL}}(\kappa)\) (blue line) and
\(\lambda_{\mathrm{UL}}(\kappa)\) (purple line) are plotted against quantile $\kappa$, and
\(\lambda_{\mathrm{UR}}(\kappa)\) (red line), \(\lambda_{\mathrm{LR}}(\kappa)\) (yellow line) are plotted against $1-\kappa$. Asymmetric dependence in either the minor or major diagonal results in visual asymmetry around $\kappa=0.5$ in the plot. Moreover, greater consistency with the noisy empirical quantiles (black line) suggest an improved fit. Figure~\ref{fig:app_3d} gives the
quantile dependence plots for the pairs (BAC, VIX), (JPM, VIX), and (JPM, BAC) for the fitted TrUST copulas with $q=1$ and $q=2$.
The quantile estimates with $q=2$ presented in panels (d)-(f) more closely align with the empirical quantiles for all three pairs, reinforcing the improvement also indicated by the DIC values.
\subsection{Cross Industry Example (6-dim)}
Cross industry dependence in U.S. equity returns is estimated using $d=6$ dimensional TrUST copulas. The filtered copula data was computed for the VIX
and leading stocks from five key sectors of the economy: Apple (AAPL) from Electronics, JPMorgan (JPM) from Banking, Visa (V) from Business Services, Exxon Mobil (XOM) from Energy, and United Parcel Service (UPS) from Freight/Transport. We consider two periods each with $n=1040$ observations: a pre-COVID period from 18 December 2018 to 15 February 2019, and the COVID period from 11 February 2020 to 15 April 2020. Market volatility was high during both periods, allowing investigation of how extreme market conditions influence dependence between the variables.
\begin{table}[tbhp]
\centering
\caption{Estimates of the TrUST Copulas for the Cross-Industry Returns in the Pre-Covid Period}
\resizebox{1.0\textwidth}{!}{
\begin{tabular}{cccccccccccccrc}
\toprule
\toprule
Variable & & \multicolumn{6}{c}{$\Omega$} & & $\text{\boldmath$\delta$}_1$ & $\text{\boldmath$\delta$}_2$ & & $\nu$ & & DIC \\
\cmidrule{1-1}\cmidrule{3-8}\cmidrule{10-11}\cmidrule{13-13}\cmidrule{15-15}\rowcolor[rgb]{ .906, .902, .902} \multicolumn{15}{c}{Panel A: TruST Copula $q=1$} \\
VIX & & 1.000 & & & & & & & -0.199 & --- & & \multirow{6}[0]{*}{14.251} & & \multirow{6}[0]{*}{13271} \\
AAPL & & -0.537 & 1.000 & & & & & & 0.661 & --- & & & & \\
JPM & & -0.543 & 0.674 & 1.000 & & & & & 0.719 & --- & & & & \\
V & & -0.521 & 0.749 & 0.769 & 1.000 & & & & 0.830 & --- & & & & \\
XOM & & -0.576 & 0.594 & 0.674 & 0.657 & 1.000 & & & 0.536 & --- & & & & \\
UPS & & -0.459 & 0.658 & 0.715 & 0.745 & 0.613 & 1.000 & & 0.820 & --- & & & & \\
\rowcolor[rgb]{ .906, .902, .902} \multicolumn{15}{c}{Panel B: TruST Copula $q=2$} \\
VIX & & 1.000 & & & & & & & -0.319 & 0.307 & & \multirow{6}[1]{*}{27.266} & & \multirow{6}[1]{*}{\textbf{6613}} \\
AAPL & & -0.485 & 1.000 & & & & & & 0.942 & -0.848 & & & & \\
JPM & & -0.481 & 0.933 & 1.000 & & & & & 0.952 & -0.902 & & & & \\
V & & -0.435 & 0.952 & 0.963 & 1.000 & & & & 0.987 & -0.918 & & & & \\
XOM & & -0.597 & 0.795 & 0.803 & 0.788 & 1.000 & & & 0.745 & -0.628 & & & & \\
UPS & & -0.287 & 0.025 & -0.033 & -0.055 & 0.221 & 1.000 & & -0.110 & 0.394 & & & & \\
\bottomrule
\bottomrule
\end{tabular}
}
\label{tab:app_6d_estimates_1}
\caption*{Note: The posterior mean estimates are report for the TrUST copula with $q=1$ (Panel~A) and $q=2$ (Panel~B). Cells with ``---'' in $\text{\boldmath$\delta$}_2$ indicate that the parameter does not exist when $q=1$. The final column also reports the DIC values.}
\end{table}
The TrUST copula model with \(q=1\) and \(q=2\) is estimated for both periods.
Table~\ref{tab:app_6d_estimates_1} reports the
posterior means of the copula parameters for the pre-COVID period.
The DIC for both copulas is also reported, indicating that the fit with $q=2$ provides a substantial
improvement over the AC skew-t copula with $q=1$. This is also the case for the COVID period data, for which the estimates are given in Table~\ref{tab:app_6d_estimates_2} in the Online Appendix.
\begin{sidewaysfigure}[htbp]
\centering
\begin{minipage}[t]{1\linewidth}
\centering
\includegraphics[width=0.49\linewidth]{figs/Heatmap_Lambda_kappa5_Win25_q1} \hfill
\includegraphics[width=0.49\linewidth]{figs/Heatmap_Lambda_kappa5_Win39_q1} \\[3ex]
\includegraphics[width=0.49\linewidth]{figs/Heatmap_Lambda_kappa5_Win25_q2} \hfill
\includegraphics[width=0.49\linewidth]{figs/Heatmap_Lambda_kappa5_Win39_q2}
\end{minipage}
\caption{Pairwise asymmetric dependencies $\Lambda_{\scalebox{0.6}{Major}}(0.05)$ and $\Lambda_{\scalebox{0.6}{Minor}}(0.05)$ of the estimated TrUST copula models for 6-dimensional cross-industry example. Panels~(a,b,e,f) give results for the pre-COVID period data, and panels~(c,d,g,h) for the COVID period. The panels in the top row correspond to $q=1$ and panels in the
the bottom row when $q=2$.}
\label{fig:heatmap_6d_win}
\end{sidewaysfigure}
Estimates of the asymmetric dependence measures $\Lambda_{\scalebox{0.6}{Major}}(0.05)$ and $\Lambda_{\scalebox{0.6}{Minor}}(0.05)$ for all variable pairs are presented as heatmaps in Figure~\ref{fig:heatmap_6d_win}. The level of asymmetry varies across
variable pairs in both periods and for both fitted copulas, particularly for $\Lambda_{\scalebox{0.6}{Major}}(0.05)$ which is the more important measure.
However, there are substantive differences in direction and degree between the two copulas for some variable pairs due to the improved
fit of the TrUST copula with $q=2$. Equivalent results are given for quantile $\tau=0.01$
in the Online Appendix.
To measure whether these differences are meaningful, we
consider their impact on density forecasts. To do so, the out-of-sample one-step-ahead predictive distributions were computed from both copula models over 40 trading day evaluation periods (equivalent to 1040 intraday observations) following the two estimation periods. Figure~\ref{fig:logscore_6d} presents the cumulative log-score (LS) differences between the TrUST copula models and a Student-t copula model with the same marginals.
Panel~(a) gives the results for the earlier pre-COVID estimation period, and panel~(b) for the later COVID estimation period. Positive values of the cumulative LS difference indicate improved performance, and the TrUST copula with $q=2$ outperforms both the AC skew-t copula and the benchmark Student-t copula in both evaluation periods.
\begin{figure}[htbp]
\centering
\includegraphics[width=1\textwidth]{figs/fig_LogScore}
\caption{
Cumulative difference in log-scores between (1) TrUST copula ($q=1$), (2) TrUST copula ($q=2$) and a benchmark student-t copula model. All copula models have the same marginals.
}
\label{fig:logscore_6d}
\end{figure}
\section{Discussion}\label{sec:07}
The UST distribution can capture richer asymmetry than the popular AC skew-t distribution which it nests. However, this potential has not been exploited in data analysis to date because inference is difficult.
The current paper addresses this by specifying a novel tractable subclass for which likelihood-based
inference can be computed
using data augmentation methods. The flexibility of this TrUST distribution is illustrated empirically using both simulated and regional Australian electricity price data.
The implicit copulas of skew-elliptical distributions~\citep{demarta_mcneil_2005,smith_gan_kohn_2010,lucas2017,Yoshiba_2018,oh_patton_2023,deng2024large}
are popular because they are both scalable and invariant to permutation of the
random vector.
A major contribution of the current paper is to show how the implicit copula of the
TrUST distribution is also tractable, and how Bayesian data augmentation method can also be used to compute the posterior of its parameters. The TrUST copula with $q\geq 2$ allows for greater
heterogeneity in asymmetric
dependence across variable pairs than the AC skew-t copula. An important application
is the modeling of dependence between equity returns, where an improvement is demonstrated using a TrUST copula model for the intraday returns of large U.S. equities.
We finish by discussing three possible future extensions of our work. First, in this paper we compute exact posterior inference for both the TrUST distribution and its copula using MCMC methods. However, approximate Bayesian inference methods---notably variational inference---are
faster and have the potential to allow estimation for higher dimensions $d$. Second,
the prior for $\Omega$ is uniform on its $d(d-1)/2$ hyper-spherical angles. For high values of $d$ these angles can be regularized, such as through Bayesian global-local priors. Alternatively, a factor decomposition can be used for $\Omega$, as in~\cite{Murray_Dunson_Carin_Lucas_2013} or~\cite{deng2024large}, so that parameterization of
$\Omega$ only increases linearly with $d$. Interestingly, this is very similar to the approach used in USE distributions to parameterize skew using $\Delta$. Here, $\Delta$ is a $(d\times q)$ matrix where $q$ is fixed, so that the number of skew parameters only increases linearly with $d$.
Last, adopting a dynamic specification of the TrUST copula, as in the models of \cite{creal2015high,oh2017} and~\cite{oh_patton_2023}, would introduce time variation
in the dependence structure.
\clearpage
\noindent
\oldappendix
\newcommand{\appendixname~Part~\Alph{section}\quad}{\appendixname~Part~\Alph{section}\quad}
\section{Extended Expression}\label{app:extended}
\subsection{Extended TrUST Distribution}\label{sec:etrust}
The term ``extended'' follows from \cite{Azzalini_Capitanio_2003} that has hidden truncation on $ L + \tau > 0$.
\begin{corollary}[Extended TrUST Distribution]\label{theo:etrust}
For joint random vector \( {\left(\text{\boldmath$X$}^\top, \text{\boldmath$L$}^\top\right)}^\top \sim \mathrm{T}{\left(\mathbf{0}, R, \nu\right)}\). We denote the Extended-TrUST (ETrUST) distribution for random vector \(\text{\boldmath$Z$} \overset{\mathrm{d}}{=} \text{\boldmath$X$} | \text{\boldmath$L$} + \text{\boldmath$\tau$} > \mathbf{0}\), where with $\text{\boldmath$\tau$} = {\left(\tau_1,\ldots,\tau_q\right)}^\top $. Then $\text{\boldmath$Z$} \sim \mathrm{ETrUST}_q {\left(\mathbf{0}, \Omega, A, \nu, \text{\boldmath$\tau$}\right)}$ has density function
\begin{equation}\label{eq:etrust}
f_{\mathrm{ETrUST}, q}{\left(\text{\boldmath$z$}; \Omega, A, \nu, \text{\boldmath$\tau$}\right)} = t {\left(\text{\boldmath$z$}; \Omega, \nu\right)} \frac{\prod_{k=1}^{q} T {\left( \sqrt{\frac{\nu+d}{\nu+Q {\left(\text{\boldmath$z$}\right)}}} {\left(\text{$\tilde \tau$}_k + \text{\boldmath$\alpha$}_k^\top \text{\boldmath$z$}\right)}; 1, \nu +d \right)}ght)}; 1, \nu +d } }{ T {\left(\text{\boldmath$\tau$}; \Sigma, \nu\right)} },
\end{equation}
where $\text{$\tilde \tau$}_k = \tau_k {\left(1-\text{\boldmath$\delta$}_k^\top \Omega^{-1}_k \text{\boldmath$\delta$}_k\right)}^{-1/2}$, for $k=1,\ldots,q$.
\end{corollary}
\subsection{Conditional Distribution}\label{sec:condtrust}
\begin{corollary}[Conditional TrUST Distribution]\label{theo:condtrust}
Let \(\text{\boldmath$Z$} \sim \mathrm{TrUST}_{q}(\Omega, A, \nu)\). Partition
\[
\text{\boldmath$Z$} \;=\; (\text{\boldmath$Z$}_1,\text{\boldmath$Z$}_2)^\top
\;\overset{d}{=}\; (\text{\boldmath$X$}_1,\text{\boldmath$X$}_2 \mid \text{\boldmath$L$}>\mathbf{0}),
\]
where \(\text{\boldmath$Z$}_1\in\mathbb{R}^{d_1}\) and \(\text{\boldmath$Z$}_2\in\mathbb{R}^{d_2}\), and
\[
\Omega =
\begin{pmatrix}
\Omega_1 & \Omega_{12} \\
\Omega_{21} & \Omega_2
\end{pmatrix},
A =
\begin{pmatrix}
A_1 \\ A_2
\end{pmatrix}
\]
with \( A_1 = {\left(\text{\boldmath$\alpha$}_{k(1)}\right)}_{d_1 \times q},\; A_2 = {\left(\text{\boldmath$\alpha$}_{k(2)}\right)}_{d_2 \times q} \).
Define the conditional random vector
\[
\text{\boldmath$Z$}_{1\mid 2}
\;:=\;
(\text{\boldmath$Z$}_1 \mid \text{\boldmath$Z$}_2 = \text{\boldmath$z$}_2)\,\overset{d}{=}\,
(\text{\boldmath$X$}_1 \mid \text{\boldmath$X$}_2 = \text{\boldmath$x$}_2,\;\text{\boldmath$L$}>\mathbf{0})\,.
\]
Then \(\text{\boldmath$Z$}_{1\mid 2}\) admits an “extended” TrUST distribution, in which the latent location is shifted by conditioning \(\text{\boldmath$Z$}_2\). Its density is given by:
\[
p_{\text{\boldmath$Z$}_{1|2}}(\text{\boldmath$z$}_1) = t_{d_1}{\left(\text{\boldmath$z$}_1 - \text{\boldmath$\mu$}_*; \frac{\nu+Q(\text{\boldmath$z$}_2)}{\nu+d_2} \Omega_*, \nu+d_2\right)}
\frac{ \prod_{k=1}^{q}
T_1{\left(\sqrt{\frac{\nu+d_2+d_1}{\nu + Q(\text{\boldmath$z$})} } (\text{\boldmath$\alpha$}_{k(1)}^\top \text{\boldmath$z$}_1 + \text{\boldmath$\alpha$}_{k(2)}^\top \text{\boldmath$z$}_2); 1, \nu+d_2+d_1\right)} }{ T_q{\left(\Delta_2 \Omega_{2}^{-1} \text{\boldmath$z$}_2; \frac{\nu+Q(\text{\boldmath$z$}_2)}{\nu+d_2} (\Sigma - \Delta_2 \Omega_{2}^{-1} \Delta_2^\top), \nu+d_2 \right)} },
\]
where \( \text{\boldmath$\mu$}_* = \Omega_{12}\Omega_2^{-1} \text{\boldmath$z$}_2\), \( \Omega_{*} = \Omega_1 - \Omega_{12} \Omega_2^{-1} \Omega_{21}\), \(Q(\text{\boldmath$z$}_2) = \text{\boldmath$z$}_2^\top \Omega_{2}^{-1} \text{\boldmath$z$}_2 \).
\end{corollary}
\section{Quantile Dependence}\label{app:quant_dep}
The pairwise quantile dependencies of the TrUST copula are computed from the bivariate marginal copula function of \( (U_1, U_2) \), which is
\[
C(u_1, u_2) = F_{\mathrm{UST}, q}{\left((z_1, z_2)^\top; \Omega, \Delta, \Sigma, \nu\right)} = \frac{T_{2+q}{\left((z_1, z_2, \mathbf{0}_q^\top)^\top; R^*, \nu \right)} }{ T_q{\left(\mathbf{0}^\top; \Sigma, \nu \right)} }
\]
with
\[
R^* = \left(\begin{array}{cc}
\Omega & -\Delta^\top \\
-\Delta & \Sigma \\
\end{array}\right),
\]
which can be derived from Equations~\eqref{eq:sue_rep} and~\eqref{eq:cond_trunc} that
\[
\text{$\mathds{P}$}{\left(\text{\boldmath$Z$} \leq \text{\boldmath$z$}\right)} = \text{$\mathds{P}$}{\left(\text{\boldmath$X$} \leq \text{\boldmath$z$} | \text{\boldmath$L$} > \mathbf{0}\right)} = \frac{\text{$\mathds{P}$}{\left(\text{\boldmath$X$} \leq \text{\boldmath$z$}, -\text{\boldmath$L$} \leq \mathbf{0}\right)} }{\text{$\mathds{P}$}{\left( -\text{\boldmath$L$} \leq \mathbf{0}\right)} }.
\]
Each pseudo observation
\(
z_{ij}=F_{\mathrm{UST},q}^{-1}(u_{ij};1, \Delta_j, \Sigma, \nu)
\)
is obtained via numerical optimization.
\newpage
\singlespacing
\bibliography{references}
\newpage
\noindent
\begin{center}
{\bf \Large{Online Appendix for ``Tractable Unified Skew-t Distribution and Copula for Heterogeneous Asymmetries''}}
\end{center}
\spacing{1.5}
\vspace{10pt}
\setcounter{page}{1}
\setcounter{figure}{0}
\setcounter{table}{0}
\setcounter{section}{0}
\setcounter{equation}{0}
\setcounter{algorithm}{1}
\makeatletter
\makeatother
\noindent
This Online Appendix has four parts:
\begin{itemize}
\item[] {\bf Part~A}: Alternative Generative Representations
\begin{itemize}
\item[] A.1: Generative Representation 1
\item[] A.2: Generative Representation 2
\end{itemize}
\item[] {\bf Part~B}: Supplementary Figures
\item[] {\bf Part~C}: Supplementary Tables
\item[] {\bf Part~D}: Proofs
\end{itemize}
\newpage
\section{Alternative Generative Representations}\label{sm:grs}
\subsection{Generative Representation 1}\label{sm:gr1}
The augmentation leverages the hidden truncation mechanism, as described in Eq.~\eqref{eq:sue_rep}, within the framework of the joint Student-t distribution without scale mixture of normals. Specifically, the joint distribution
\(
{\left( \text{\boldmath$X$}^\top, \text{\boldmath$L$}^\top \right)}^\top \sim \mathrm{T}_{d+q}{\left(\mathbf{0}, R, \nu\right)}
\). Then from this expression, we can derive the \(\text{\boldmath$Z$}\) in the conditional distribution
\begin{equation*}
{\left(\text{\boldmath$Z$} \mid \text{\boldmath$L$}\right)} \sim \mathrm{T} {\left(\Delta^\top \Sigma^{-1} \text{\boldmath$L$}, \frac{\nu+Q{\left(\text{\boldmath$L$}\right)}}{\nu+q} {\left(\Omega - \Delta^\top \Sigma^{-1} \Delta\right)}, \nu+q\right)}ta\right)}, \nu+q}.
\end{equation*}
The augmented density of \({\left\{ \text{\boldmath$Z$}, \text{\boldmath$L$} \right\}} \) is
\begin{equation}
p{\left(\text{\boldmath$z$}, \text{\boldmath$l$}\right)} = t_d{\left(\text{\boldmath$z$}; \Delta^\top \Sigma^{-1} \text{\boldmath$l$}, \frac{\nu+Q{\left(\text{\boldmath$l$}\right)}}{\nu+q} {\left(\Omega - \Delta^\top \Sigma^{-1} \Delta\right)}, \nu+q\right)}ta\right)}, \nu+q} t_{\text{\boldmath$l$} > \mathbf{0}}{\left(\text{\boldmath$l$}; \mathbf{0}, \Sigma, \nu\right)}.
\end{equation}
Still, the latent variable $\text{\boldmath$L$}$ can be sampled as a block by:
\begin{equation*}
{\left(L_k | \text{\boldmath$Z$}, W, \text{\boldmath$\theta$}\right)} \sim \mathrm{T}^+_1 {\left(m_{k}, \frac{\nu + Q{\left(\text{\boldmath$Z$}\right)} }{ \nu + d} r_k, \nu+d\right)}k, \nu+d},
\end{equation*}
for \(k = 1,\ldots,q\), where $m_k= \text{\boldmath$\delta$}_k^\top \Omega^{-1}\text{\boldmath$Z$}$, $ r_k = 1 - \text{\boldmath$\delta$}_k^\top \Omega^{-1} \text{\boldmath$\delta$}_k$.
\subsection{Generative Representation 2}\label{sm:gr2}
The generative representation in the Section~\ref{sec:gr} using the scaling then hidden truncation ordering. The other method follows a sequence where hidden truncation is applied first, then scaling. First, we have the joint distribution
\(
(\text{$\widetilde{\text{\boldmath$X$}}$}^\top, \text{$\widetilde{\text{\boldmath$L$}}$}^\top )^\top \sim \mathrm{N}_{d+q}{\left(\mathbf{0}, R\right)},
\)
then expressed the scaling as:
\begin{equation*}\label{eq:GR3}
\text{\boldmath$Z$} \overset{\mathrm{d}}{=} W^{-1/2} \odot {\left( \text{$\widetilde{\text{\boldmath$X$}}$} \mid \text{$\widetilde{\text{\boldmath$L$}}$} > \mathbf{0}\right)},
\end{equation*}
where $\odot$ is the element-wise multiplication.
Then the augmented likelihood is expressed as:
\begin{equation*}
{\left(\text{\boldmath$Z$} | \text{$\widetilde{\text{\boldmath$L$}}$}, W\right)} \sim \mathrm{N}_d {\left( W^{-1/2}\odot \Delta^\top \Sigma^{-1} \text{$\widetilde{\text{\boldmath$L$}}$}, W^{-1}\odot {\left(\Omega - \Delta^\top \Sigma^{-1} \Delta \right)}\right)}\right)}},
\end{equation*}
with the posterior distribution for latent variable $\text{\boldmath$L$}$ given by:
\begin{equation*}
L_k | \text{\boldmath$Z$}, W, \text{\boldmath$\theta$} \sim \mathrm{N}^+ {\left(W^{1/2} m_k, r_k\right)},
\end{equation*}
and the posterior of $W$ is not in closed form.
Both (GR1) and (GR2) provide alternative augmentations for the model, offering flexibility in computational implementation while adhering consistency with the underlying hidden truncation framework. The choice between representations, including the (GR1) without scaling, depends on the specific modeling objectives and computational requirements, ensuring adaptability to diverse data contexts.