EconBase
← Back to paper

Bayesian Modular Inference for Copula Models with Potentially Misspecified Marginals

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

69,490 characters · 21 sections · 68 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Bayesian Modular Inference for Copula Models with Potentially Misspecified Marginals

titlepage\thispagestyle{empty} \begin{center} { Abstract} \end{center} Copula models of multivariate data are popular because they allow separate specification of marginal distributions and the copula function. These components can be treated as inter-related modules in a modified Bayesian inference approach called “cutting feedback” that is robust to their misspecification. Recent work uses a two module approach, where all $d$ marginals form a single module, to robustify inference for the marginals against copula function misspecification, or vice versa. However, marginals can exhibit differing levels of misspecification, making it attractive to assign each its own module with an individual influence parameter controlling its contribution to a joint semi-modular inference (SMI) posterior. This generalizes existing two module SMI methods, which interpolate between cut and conventional posteriors using a single influence parameter. We develop a novel copula SMI method and select the influence parameters using Bayesian optimization. It provides an efficient continuous relaxation of the discrete optimization problem over $2^d$ cut/uncut configurations. We establish theoretical properties of the resulting semi-modular posterior and demonstrate the approach on simulated and real data. The real data application uses a skew-normal copula model of asymmetric dependence between equity volatility and bond yields, where robustifying copula estimation against marginal misspecification is strongly motivated. {\bf Keywords}: Bayesian optimization, Cutting feedback, Modular inference, Semi-modular inference, Variational approximation. {\smallAcknowledgments: David Nott's research was supported by the Ministry of Education, Singapore, under the Academic Research Fund Tier 2 (MOE-T2EP20123-0009), and he is affiliated with the Institute of Operations Research and Analytics at the National University of Singapore. Michael S. Smith and David T. Frazier gratefully acknowledge support by the Australian Research Council through a Discovery Project grant DP250101069.} { $^1$ Department of Statistics and Data Science, National University of Singapore\\ $^2$ Department of Econometrics and Business Statistics, Monash University\\ $^3$ Melbourne Business School, University of Melbourne\\ $^\ast$ Correspondence should be directed to Michael Smith at {\tt [email removed]}\,. }

\spacingset{1.5}

Introduction

Copula models of multivariate continuous data are very popular because they allow the marginal distributions and copula function characterizing the dependence structure to be specified separately. However, in many problems specification of the copula function or the marginals can be difficult. Here we consider modified Bayesian inference for copula models, in settings where we wish to robustify inference about the copula parameters to misspecification of marginal distributions. Our work contributes to the literature on modular Bayesian inference, which decomposes a joint model into simpler inter-related sub-models called modules, and attempts to limit the influence that misspecified modules can have on parameters in well-specified ones. We consider each of the $d$ marginals as a separate module, and develop a “semi-modular inference” CarNic2020,CarNic2022,FraNot2025 method that robustifies inference for the copula parameters from potential misspecification of one or more of the marginals. We establish the theoretical properties of the resulting semi-modular posterior and demonstrate the approach empirically.

A common modular Bayesian inference approach is to calculate the so-called “cut posterior”, which is where feedback from misspecified modules to the parameters of other well-specified modules is completely cut Plu2015,liu+g25. In recent work, SmiYuNotFra2025 develop such cut posteriors for copula models, but treat all marginals as a single module. However, in practice marginals can exhibit differing levels of misspecification, making it attractive to treat them as $d$ separate modules, and to modulate the influence of each according to the degree of misspecification.

SMI CarNic2020,CarNic2022,FraNot2025 extends cutting feedback approaches to the case where the effects of a misspecified module on a joint SMI posterior are modulated continuously by an influence parameter, which interpolates between a cut and the conventional posterior. Here, we specify a completely novel SMI approach for copulas, with a separate influence parameter for each marginal distribution. These influence parameters are introduced through a novel extended pseudo likelihood, which is a significant departure from the SMI approach of CarNic2020,CarNic2022. We develop a Bayesian optimization garnett23 approach to learn the influence parameters of this new semi-modular posterior, which allows some marginals to be fully cut, some fully uncut, and some partially cut. Even when searching for a modular structure where each marginal is fully cut or uncut, the SMI optimization is an effective continuous relaxation of the discrete search problem that is not tractable even in modest dimensions.

We study the properties of the SMI posterior both theoretically and empirically. Re-interpreting our SMI posterior as a type of generalized Bayesian posterior (bissiri2016general), we can view the influence parameters in the SMI posterior as akin to the “learning rates” required to perform generalized Bayes. However, in contrast to standard generalized Bayes, our theoretical results demonstrate that our SMI posterior depends critically on the value of the influence parameters, which determine the extent of posterior modularity. In general, the learning rates in generalized Bayes posteriors only impacts the posterior scale, while the influence parameters in the SMI posterior impact both posterior scale and location. This makes existing results for generalized posteriors inapplicable, so that we must tailor our theoretical analysis to this novel setting. The strong dependence on the influence parameters also suggests that, in contrast to the case of learning rates in generalized Bayes, values can be learned that deliver accurate results, where accuracy is measured by an external criterion. A simulation study illustrates the impact of the influence parameters empirically. Considering a bivariate example, where only one of the marginals is misspecified, we show that tuning the influence parameter to down-weight the impact of the misspecified module improves inference for the copula and the misspecified marginal at the cost of a mild deterioration for the well-specified marginal. In this example, an SMI posterior with a partial cut of the misspecified module is preferred over both fully cut and uncut posteriors.

In a substantive application to U.S. financial data, we estimate a copula model of equity market volatility and the yields on domestic debt graded AAA and BBB by Standard & Poors. A skew-normal copula captures strongly asymmetric contemporaneous and serial dependence between the variables. Parametric skewed distributions with heavy tails are used for the marginals, which are preferred over nonparametric ones because the extreme tail behavior is important. However, one or more of these marginals may be misspecified, so that computing inference robust to marginal misspecification is attractive. We show that the SMI posterior differs substantially from the conventional and fully cut posteriors, with results that are more economically intuitive and consistent with the data.

The rest of the paper is organized as follows. Section (ref) gives the necessary background on copula models, variational inference and Bayesian modular inference methods. Section (ref) briefly outlines the cutting feedback approach for copula models of SmiYuNotFra2025, how to generalize it to allow each marginal to be a separate module, and outlines the proposed SMI approach. Efficient variational inference methods to compute the SMI posterior are then given, where Bayesian optimization is used to learn the influence parameters. Section (ref) explores the properties of our SMI posterior, Section (ref) presents the simulation study, Section (ref) contains the financial application, and Section (ref) concludes.

Background

Copula models

Let $Y=(Y_1,\dots,Y_d)^\top$ be a random vector with joint distribution $F_Y$ and marginal distributions $F_j$, $j=1,\dots,d$. Sklar's theorem Skl1959 shows that the joint distribution function is

equation[equation omitted — 71 chars of source]

where $C:[0,1]^d\to[0,1]$ is a copula function determining the dependence structure of $Y$. The copula function is unique if all marginals are continuous, and differentiating (ref) gives the joint density of $Y$,

align[align omitted — 93 chars of source]

where $c(u)=\frac{\partial}{\partial u}C(u)$ is the “copula density”, and $f_j(y_j)=\frac{\partial}{\partial y_j}F_j(y_j)$, $j=1,\dots,d$, are the marginal densities.

A major advantage of copula models is that the marginals and copula function can be specified separately. In many applications parametric families are employed for both, and we write $C(u;\psi)$ for the copula function with parameters $\psi$, and $F_j(y_j;\eta_j)$ for the $j$-th marginal with parameters $\eta_j$. The model parameters are thus $\theta=(\psi^\top,\eta^\top)^\top$, where $\eta=(\eta_1^\top,\dots,\eta_d^\top)^\top$ are the marginal parameters. We write $\mathcal{D}=\{y_1,\dots,y_n\}$ for a set of $n$ independent observations $y_i=(y_{i1},\dots,y_{id})^\top \sim F_Y$. Under the prior $p(\theta)=p(\psi)p(\eta_1)\cdots p(\eta_d)$, the joint posterior is then

align*[align* omitted — 156 chars of source]

Variational inference

Variational inference BleKucMca2017 is a computationally efficient alternative to exact Bayesian inference computed via Markov chain Monte Carlo (MCMC). The main idea is to approximate the posterior $p(\theta|\mathcal{D}) \propto p(y|\theta)p(\theta)$ by a distribution $q_\lambda(\theta)$ chosen from a tractable variational family $\mathcal{Q}=\{q_\lambda(\theta)\mid\lambda\in\Lambda\}$ parameterized by $\lambda$. An optimal value $\lambda^*\in\Lambda$ is commonly obtained by minimizing the reverse Kullback-Leibler (KL) divergence,

equation[equation omitted — 184 chars of source]

where $\mathbb{E}_{q_\lambda(\theta)}$ denotes expectation with respect to $q_\lambda(\theta)$. Minimizing (ref) is equivalent to maximizing the evidence lower bound OrmWan2010

equation*[equation* omitted — 142 chars of source]

A general approach to optimizing the ELBO is to employ stochastic gradient methods. These update an initial value $\lambda^{(0)}$ using an iterative scheme of the form $\lambda^{(m+1)}=\lambda^{(m)}+\rho^{(m)}\circ\widehat{\nabla_\lambda\mathcal{L}\left(\lambda^{(m)}\right)}$ at each step $m$, where $\rho^{(m)}$ is an adaptive vector-valued step size, $\circ$ denotes element-wise multiplication, and $\widehat{\nabla_\lambda\mathcal{L}\left(\lambda^{(m)}\right)}$ is an unbiased estimate of the gradient of $\mathcal{L}(\lambda)$ evaluated at $\lambda^{(m)}$.

While $\nabla_\lambda\mathcal{L}\left(\lambda^{(m)}\right)$ can be estimated by Monte Carlo estimation of the expectation, variance reduction techniques such as the reparameterization trick KinWel2014,RezMohWie2014 are often employed to ensure fast convergence. The reparameterization trick assumes that samples $\theta\sim q_\lambda(\theta)$ can be generated as $\theta=f(\varepsilon,\lambda)$ for a deterministic differentiable function $f$ and a random variable $\varepsilon\sim\pi(\varepsilon)$ from a base distribution not depending on $\lambda$. Then,

align*[align* omitted — 232 chars of source]

A single sample from $\pi(\varepsilon)$ in each iteration $m$ is often enough to approximate the ELBO gradient. The choice of step sizes $\rho^{(m)}$ is usually made adaptively, with popular approaches including ADAM KinBa2014 and ADADELTA Zei2012. VI has been used to estimate copula models by LoaizaMayaSmith2019 and NguyenAusinGaleano2020.

Modular inference under misspecification

In this paper we consider Bayesian modular inference methods for misspecified copula models. Modular inference is used for statistical models which are formed by interacting sub-models that are called “modules”. Cutting feedback Plu2015,liu+g25 is a modular Bayesian inference method that aims to cut feedback from misspecified modules to well specified ones. It has been considered previously by SmiYuNotFra2025 for copula models. The current work extends cutting feedback for copulas to semi-modular inference, and we first briefly outline both cutting feedback and semi-modular inference. We discuss these for two modules for clarity, but our later semi-modular extensions treat each marginal distribution in a copula model as its own module, and hence involve many modules. A general approach to cutting feedback with multiple modules is given in liu+g25.

Consider a model for data $\mathcal{D}$ with density $p(\mathcal{D}\mid\theta)$ and parameter vector $\theta=(\psi^\top,\eta^\top)^\top$ that can be factorized as

equation[equation omitted — 133 chars of source]

In many cutting feedback applications, the terms $p_1$ and $p_2$ are likelihoods for different data sources for the two modules, but in general they only need to represent different factors in a decomposition of the joint likelihood. We show later that this is the case with copula models. Module 1 comprises $p_1(\mathcal{D}\mid \psi)$ and prior $p(\psi)$, and module 2 of $p_2(\mathcal{D}\mid \eta,\psi)$ and prior $p(\eta\mid\psi)$. The main idea of cutting feedback is that module 1 is trusted, but module 2 is not, and we wish to avoid the misspecification in module 2 corrupting inferences about the parameters in module 1.

The joint posterior is $p(\psi,\eta\mid\mathcal{D})=p(\psi\mid\mathcal{D})p(\eta\mid\psi,\mathcal{D})$, where

equation*[equation* omitted — 219 chars of source]

The term $\Bar{p}_2(\mathcal{D}\mid \psi)=\int p(\eta\mid\psi)p_2(\mathcal{D}\mid \eta,\psi)d\eta$ is called the feedback term because it represents the feedback of module 2 on the marginal for $\psi$. Cutting feedback is a way to stop misspecification in the second module from corrupting inference for parameters in the trusted module 1. The feedback term is deleted from $p(\eta\mid\psi,\mathcal{D})$ to define the cut posterior

align*[align* omitted — 205 chars of source]

where $p_\text{cut}(\psi\mid\mathcal{D})\propto p(\psi)p_1(\mathcal{D}\mid \psi)$. CarNic2020,CarNic2022 extend this idea to “semi-modular inference” (SMI), which introduces an influence parameter $\gamma$ to smoothly interpolate between the conventional and cut posteriors. This approach offers several practical advantages. First, even under model misspecification, allowing some information flow from the misspecified to the trusted module can be beneficial; see FraNot2025 for a formalization of the bias-variance trade-off involved. Second, in a model with $m>2$ modules, deciding which to cut can be challenging as it involves a search over $2^m$ possible configurations. Later we show that relaxing this choice into a continuous search problem over a multivariate influence parameter $\gamma$ simplifies it greatly.

CarNic2022 define a SMI posterior using the power posterior

equation[equation omitted — 205 chars of source]

which introduces auxiliary parameters $\tilde \eta$ to define the joint

equation*[equation* omitted — 192 chars of source]

Integrating out $\tilde \eta$ gives the SMI posterior $p_{\text{SMI},\gamma}(\psi,\eta\mid\mathcal{D})=\int p_{\text{SMI},\gamma}(\psi,\Tilde{\eta},\eta\mid\mathcal{D})d\Tilde{\eta}$. Setting $\gamma=1$ results in the conventional posterior, $\gamma=0$ gives the cut posterior $p_\text{cut}(\psi,\eta\mid\mathcal{D})$, whereas $0<\gamma<1$ interpolates between the two. In this paper, we later propose a variation of SMI posteriors for copula models that is not based on the power posterior.

The original motivation of CarNic2020 for introducing SMI was that the conventional posterior can be preferable to the cut posterior when misspecification is only moderate. Their crucial insight was to frame the problem in terms of a bias-variance trade-off, with the influence parameter managing this trade-off by interpolating between the cut and conventional posteriors. FraNot2025 formalized this intuition for a variant of SMI introduced in ChaNotDroFraSis2023, and Nicholls+lwc22 develop connections between SMI posteriors and generalized Bayesian approaches; related generalized Bayesian treatments of cutting feedback appear in frazier2025cutting and Tan+nf25. Selecting influence parameters has been studied in CarNic2022 and battaglia+chlbn25 using amortized variational inference, and we also adopt variational methods here to compute SMI posteriors, following YuNotSmi2023.

Generalized posteriors for copula models with misspecified marginals

Cutting feedback in copula models

SmiYuNotFra2025 consider copula models as a two module system, where the copula function $C(u;\psi)$ forms one module, and the marginals $F_1(y_1;\eta_1)$, \dots, $F_d(y_d;\eta_d)$ form a second module. They consider a scenario where a researcher is confident in the copula function, but the marginals may be misspecified, and therefore wishes to cut the influence of $\eta$ on $\psi$. They call this a “Type 2” cut posterior and define it using a pseudo likelihood of the rank data to bring the likelihood into the form (ref). Denote the rank data as $r(\mathcal{D})=\{r(y_{ij}); i=1,\dots,n, j =1,\dots,d\}$, where $r(y_{ij})=\sum_{k=1}^n\mathds{1}\left\{ y_{kj}\leq y_{ij} \right\}$ is the rank of $y_{ij}$ within marginal $j$. Setting

align*[align* omitted — 134 chars of source]

SmiYuNotFra2025 define the pseudo rank likelihood as

equation[equation omitted — 203 chars of source]

where for $v=(v_1,v_2,\ldots,v_d)^\top$,

equation*[equation* omitted — 148 chars of source]

is a differencing operator over the $j$th element. Equation (ref) is similar to replacing each marginal in the copula model with its empirical distribution function to remove dependence on the marginals, but accounting for the discrete nature of the ranks. By setting $p_1(\mathcal{D}|\psi)=p_\text{pl}(r(\mathcal{D})\mid\psi)$ and $p_2(\mathcal{D}\mid \eta,\psi)=p(\mathcal{D}\mid \eta,\psi)/p_\text{pl}(r(\mathcal{D})\mid\psi)$ at (ref), then $p_\text{cut}(\psi\mid r(\mathcal{D}))\propto p_\text{pl}(r(\mathcal{D})\mid\psi)p(\psi)$ and the cut posterior is

equation[equation omitted — 147 chars of source]

Evaluating (ref) is difficult computationally. To avoid this the authors introduce auxiliary variables $u=(u_1^\top,\ldots,u_n^\top)^\top$, with $u_i=(u_{i1},\ldots,u_{id})^\top$, and use the more tractable extended pseudo likelihood

equation[equation omitted — 196 chars of source]

Integrating (ref) over $u$ returns the pseudo rank likelihood. They then compute the cut posterior augmented with $u$ given by $p_\text{cut}(\psi,u,\eta\mid\mathcal{D})=p_\text{cut}(\psi,u\mid r(\mathcal{D}))p(\eta\mid \psi, \mathcal{D})$, where $p_\text{cut}(\psi,u\mid r(\mathcal{D}))\propto p_\text{epl}(r(\mathcal{D}),u\mid\psi)p(\psi)$. The marginal in $\psi,\eta$ is the desired cut posterior at (ref).

Multiple cuts of marginals

The approach of SmiYuNotFra2025 allows for only a single cut from one module for all marginals to the copula. However, in many applications some marginal distributions might be more trusted than others, and a researcher may wish to separately specify whether or not to cut each marginal. To do this, we extend their framework as follows.

Let $\delta=(\delta_1,\dots,\delta_d)^\top\in\{0,1\}^d$ be a vector of binary indicators such that $\delta_j=0$ indicates that the $j$th marginal is cut and $\delta_j=1$ that the $j$th marginal is uncut. The set $D(\delta)=\{j; \delta_j=0\}=\{j_1,\ldots,j_r\}$ includes the indices of the marginal distributions that are cut, and $D^c(\delta)=\{j; \delta_j=1\}=\{j_{r+1},\ldots,j_d\}$ those that remain uncut. Then we can partition the marginal parameters $\eta$ into cut $\eta_{D(\delta)}=(\eta_{j_{1}}^\top,\ldots,\eta_{j_r}^\top)^\top$ and uncut $\eta_{D^c(\delta)}=(\eta_{j_{r+1}}^\top,\ldots,\eta_{j_d}^\top)^\top$ components.

To define the cut posterior we replace the extended pseudo likelihood at (ref) with the following extension:

align[align omitted — 337 chars of source]

Integrating over $u$ gives a pseudo likelihood where the marginals for $D(\delta)$ only use their rank data; see Supporting Information A. This resulting pseudo likelihood is a mixed density, so that we call it a mixed pseudo likelihood. When $\delta=(0,\ldots,0)^\top$, $D^c(\delta)=\emptyset$ and (ref) is exactly equal to (ref), whereas when $\delta=(1,\ldots,1)^\top$, $D(\delta)=\emptyset$ and (ref) is the copula density evaluated at points $u_{ij}=F_j(y_{ij};\eta_j)$ for all $i,j$.

For a fixed vector $\delta$, the joint cut posterior (augmented with $u$) is given by setting

align[align omitted — 325 chars of source]

Here, the conventional conditional posterior $p(\eta_{D(\delta)}\mid \psi,\eta_{D^c(\delta)},u,\mathcal{D})=p(\eta_{D(\delta)}\mid \psi,\eta_{D^c(\delta)},\mathcal{D})$ is unaffected by $u$. The marginal of (ref) in $(\psi,\eta)$ is the required cut posterior.

This approach is a natural extension to Section (ref), and it agrees with the approach of SmiYuNotFra2025 when $\delta=(0,\dots,0)^\top$. For $\delta=(1,\dots,1)^\top$, $D(\delta)=\emptyset$, so that from (ref),

equation*[equation* omitted — 146 chars of source]

is the conventional posterior.

However, conducting inference using this framework has two challenges. First, selecting the optimal cut $\delta$ can be difficult as this leads to a discrete search problem over $2^d$ potential cuts. Second, the degree of misspecification in each marginal can vary, and a full cut of each misspecified marginal may be disadvantageous depending on the objective of the analysis. To overcome both challenges, we instead propose a novel way of controlling the information flow between marginals and copula parameters in a continuous manner which we describe below. It allows each marginal to be either fully cut, fully uncut, or partially cut. By doing so, it produces a continuous relaxation of the hard cut, where the search is no longer over a discrete space of dimension $2^d$, but over the hypercube $[0,1]^d$.

Our SMI approach

Let $\gamma=(\gamma_1,\dots,\gamma_d)^\top$ denote a vector of parameters, $0\leq\gamma_j\leq1$, where $\gamma_j$ controls the influence of the $j$th marginal parameter $\eta_j$ on $\psi$. When $\gamma_j=0$ the influence of the $j$th marginal is fully cut from the copula module. Then, we propose using the following SMI extended likelihood

equation[equation omitted — 242 chars of source]

where the bounds interpolate between the ranks and $F_j(y_{ij},\eta_j)$:

align*[align* omitted — 227 chars of source]

Equation (ref) is a continuous relaxation of (ref) and exactly equals $p_{\text{mpl},\gamma}$ when $\gamma=\delta\in\{0,1\}^d$. In particular, it recovers the extended pseudo likelihood at (ref) when $\gamma=(0,\dots,0)^\top$, and when $\gamma=(1,\dots,1)^\top$

align*[align* omitted — 155 chars of source]

we recover the copula density term from the parametric likelihood. The full likelihood (ref) can therefore be written as

align*[align* omitted — 265 chars of source]

Based on this observation we now define a generalized Bayes posterior for fixed $\gamma\in[0,1]^d$ as follows. Similar to the SMI framework of CarNic2022 we introduce an additional auxiliary parameter $\widetilde{\eta}=(\widetilde{\eta}_1^\top,\ldots,\widetilde{\eta}_d^\top)^\top$ and set

equation[equation omitted — 253 chars of source]

This is a posterior derived from a tempered likelihood, where $\gamma_j$ controls the influence of $\widetilde{\eta}_j$ on $\psi$. Combining (ref) with the conditional posterior $p(\eta\mid\psi,\mathcal{D})$ defines an SMI posterior distribution augmented with both $u$ and $\widetilde{\eta}$:

align[align omitted — 193 chars of source]

Information from the marginals can influence $\psi$ through the axillary parameter $\widetilde\eta$, which can influence the bounds $a_{ij}(\gamma_j,\widetilde{\eta}_j,\mathcal{D})$ and $b_{ij}(\gamma_j,\widetilde{\eta}_j,\mathcal{D})$ in $p_{\text{SMI},\gamma}(\psi,u,\widetilde{\eta}\mid\mathcal{D})$, to an extent controlled by $\gamma$. Integration over the auxiliary parameters $u$ and $\widetilde{\eta}$ in (ref) gives the SMI posterior of the copula and marginal parameters:

equation[equation omitted — 185 chars of source]

The pseudo posterior at (ref) is similar to the power posterior at (ref) in the sense that it uses a $\gamma$-modified likelihood controlling the influence of auxiliary parameters $\widetilde{\eta}$ on the parameter $\psi$ of the trusted module. However, this novel approach to SMI in copula models differs substantively from the SMI framework in CarNic2022. Under the power posterior approach of CarNic2022, the untrusted module for $\eta_j$ would reduce to the prior for $\gamma_j=0$, whereas in our case the likelihood term for the $j$th marginal, $f_j(y\mid\widetilde{\eta}_j)$, is unaffected by $\gamma_j$. That is, even if $\widetilde{\eta}_j$ is fully cut from the copula module, $\widetilde{\eta}_j$ is still informed by the data. This difference is crucial. Our structure ensures that, for each $j$, $F_j(y_{ij};\widetilde{\eta}_j)$ is fully informed by the marginal likelihood for all $\gamma$, and the bounds $a_{ij}(\gamma_j,\widetilde{\eta}_j,\mathcal{D})$ and $b_{ij}(\gamma_j,\widetilde{\eta}_j,\mathcal{D})$ in (ref) are a mixture between the (parameter independent) rank data and the parametric marginals, with mixture weight given by $\gamma_j$.

Efficient variational inference

Exact sampling from $p_{\text{SMI},\gamma}(\psi,u,\eta,\widetilde{\eta}\mid \mathcal{D})$ lends itself to a two step approach, where in a first step a sample from $p_{\text{SMI},\gamma}(\psi,u,\widetilde{\eta}\mid \mathcal{D})$ is generated and then in a second step $\eta\sim p(\eta\mid\psi,\mathcal{D})$ is drawn. However, this is computationally prohibitive as it would not only involve two nested MCMC samplers, but also the need to derive an efficient sampler for the high-dimensional density (ref). Therefore, we propose to use an efficient variational approximation instead. VI has been previously employed to derive cut and semi-modular posteriors YuNotSmi2023,CarNic2022, and here we combine a structured Gaussian variational family with “stop gradient” operators CarNic2022 that allow updating of all variational parameters jointly.

All model parameters $\theta=(\psi^\top,\eta^\top)^\top$ are either unconstrained or transformed monotonically to be so. Similarly, let $\Phi$ denote the standard normal distribution function, then each $u_{ij}$ is transformed to an unconstrained value $z_{ij}=\Phi^{-1}\left((u_{ij}-a_{ij})/(b_{ij}-a_{ij})\right)$ for all $i,j$. A priori, $u_{ij}\sim \operatorname{U}(a_{ij}(\gamma_i,\eta_j,\mathcal{D}),b_{ij}(\gamma_i,\eta_j,\mathcal{D}))$, so that $u$ depends on $\eta$, whereas each $z_{ij}\sim \operatorname{\cal N}(0,1)$ is independent of $\eta$ a priori. Thus $z=(z_1^\top,\ldots,z_n^\top)^\top$, with $z_i=(z_{i1},\ldots,z_{id})^\top$, is an attractive re-parameterization that simplifies the geometry of the prior and posterior.

The variational approximation of $p_{\text{SMI},\gamma}(\psi,u,\eta,\widetilde{\eta}\mid \mathcal{D})$ is assumed to be of the form

align*[align* omitted — 136 chars of source]

where $q_{\lambda}(\psi), q_\lambda(\eta\mid\psi)$, and $q_\lambda(\tilde{\eta}\mid\psi)$ are multivariate Gaussian distributions parameterized as follows:

align*[align* omitted — 410 chars of source]

where $T_\psi$, $T_\eta$, and $T_{\tilde\eta}$ are lower triangular matrices with positive diagonals, and $T_{\psi\tilde{\eta}}$, $T_{\psi\eta}$ are unrestricted matrices. This parameterization is chosen so that $q_\lambda(\psi,\eta,\tilde{\eta})$ is jointly Gaussian with $\eta$ and $\tilde{\eta}$ being independent conditional on $\psi$ TanNot2018. Following LoaizaMayaSmith2019, $q_\lambda(z)=\prod q_\lambda(z_{ij})$ are independent Gaussian approximations with $q_\lambda(z_{ij})=\varphi(z_{ij},\zeta_{ij},\omega_{ij}^2)$. The full vector of variational parameters to be optimized is thus $$\lambda=\left(\zeta^\top,\omega^\top,\mu_\psi^\top,\mu_\eta^\top,\mu_{\tilde{\eta}}^\top,\text{vec}(T_\psi),\text{vec}(T_\eta),\text{vec}(T_{\tilde{\eta}}),\text{vech}(T_{\psi\eta}),\text{vech}(T_{\psi\tilde{\eta}})\right)^\top,$$ where $\text{vec}(A)$ and $\text{vech}(A)$ denote the vectorization and half-vectorization of a matrix $A$, respectively.

Following YuNotSmi2023, the variational parameters $\lambda$ can be learned in two VI steps. First, the sub-vector parameterizing $q_\lambda(\psi,z,\tilde{\eta})$ approximating $p_{\text{SMI},\gamma}(\psi,u,\widetilde{\eta}\mid \mathcal{D})$ can be learned. Then, in a second step the variational parameters $(\mu_{\eta}^\top,\text{vech}(T_{\eta})^\top,\text{vec}(T_{\psi,\eta})^\top)^\top$ can be trained while keeping the remaining parameters fixed. However, this two step approach involves two separate stochastic gradient optimizations. Instead, we update the full vector of variational parameters $\lambda$ in a single joint step enabling end-to-end training applying stop gradient operators as now described. From a computational point of view, a stop gradient operator prevents gradients from flowing backwards during automatic differentiation.

Denote a copy of $\lambda$ which takes the same value as $\lambda$ but is treated as a constant when deriving the gradient $\nabla_\lambda$ as “$\cancel{\lambda}$”. To employ the reparameterization trick, a draw from $q_\lambda(\psi,z,\eta,\tilde{\eta})$ is obtained by transforming independent standard normal draws as follows:

itemize[label=] • Draw $\varepsilon_{ij}^{(z)}\sim\operatorname{\cal N}(0,1)$ and set $z_{ij}=\zeta_{ij}+\omega_{ij}\varepsilon_{ij}^{(z)}$ for $i=1,\dots,n$, $j=1,\dots,d$. • Draw $\varepsilon^{(\psi)}\sim\operatorname{\cal N}(0,I)$ and set $\psi=\mu_\psi+T_\psi^{-\top}\varepsilon^{(\psi)}$. • Draw $\varepsilon^{(\eta)}\sim\operatorname{\cal N}(0,I)$ and set $\eta=\mu_\eta-T_\eta^{-\top}T_{\psi\eta}^\top(\cancel{\psi}-\cancel{\mu_\psi})+T_\eta^{-\top}\varepsilon^{(\eta)}$. • Draw $\varepsilon^{(\tilde{\eta})}\sim\operatorname{\cal N}(0,I)$ and set $\tilde{\eta}=\mu_{\tilde{\eta}}-T_{\tilde{\eta}}^{-\top}T_{\psi\tilde{\eta}}^\top(\psi-\mu_\psi)+T_{\tilde{\eta}}^{-\top}\varepsilon^{(\tilde{\eta})}$.

Writing $h_\gamma(\psi,u,\widetilde{\eta})$ for the right hand side in (ref), we note that $h_\gamma(\psi,u,\widetilde{\eta})$ is the kernel of $p_{\text{SMI},\gamma}(\psi,u,\widetilde{\eta}\mid\mathcal{D})$, and $h_1(\cancel{\psi},\cancel{u},\eta)$ is the kernel of the conventional conditional posterior $p(\eta\mid\psi,\mathcal{D})$. Hence, the ELBO for the variational optimization can be expressed as

equation[equation omitted — 221 chars of source]

An unbiased estimate of the gradient of $\mathcal{L}(\lambda)$ is thus given as $\widehat{\nabla_\lambda\mathcal{L}\left(\lambda\right)}=\nabla_\lambda \left(\log h_\gamma(\psi,u,\tilde{\eta})+\log h_1(\cancel{\psi},\cancel{u},\eta)-\log q_\lambda(\psi,z,\eta,\tilde{\eta})\right)$. Due to the way the stop-gradient operator is employed in generating draws from $q_\lambda$ and in the evaluation of (ref) this updating scheme is equivalent to the two-step approach described above.

Selecting the influence parameter

So far, we have considered $\gamma$ as a fixed vector. However, selecting $\gamma$ is challenging, with similar issues also being encountered for the power posterior in the original SMI of CarNic2022. The theoretical results presented in Section (ref) demonstrate that the (asymptotic) behaviour of the SMI posterior depends directly on $\gamma$. Section (ref) gives an empirical example, in which selecting $\gamma$ to correctly match the misspecification of the marginals in the data generating process does not necessarily lead to an optimal joint SMI posterior.

Each influence parameter $\gamma_j$ only directly controls the influence of the $j$th marginal on the copula function module, and its influence on the other marginals is indirect. For this reason, we suggest choosing the optimal vector $\gamma^*$ using an external utility function $u(\gamma)$. The utility function may depend on additional data sources not considered for training and is largely determined on the objectives of the empirical application. For example, if the goal is prediction, $u(\gamma)$ can be a predictive metric evaluating the sharpness and precision of the copula-based forecasting model. Alternatively, if the main interest is estimation of the dependence structure, $u(\gamma)$ can be the expected log-copula likelihood.

Even for moderate $d$, optimizing $\gamma$ directly is infeasible, and so we propose to use BO garnett23. BO is a commonly employed strategy for efficiently optimizing black-box functions lacking gradients. BO builds a probabilistic surrogate of the objective $u(\cdot)$ by training a Gaussian Process on observed values and uses that surrogate to choose the most informative next point to evaluate balancing both prediction and learned uncertainty. Standard software for BO can be readily combined with our variational framework, and we use BoTorch botorch in our empirical work.

Theoretical behavior of SMI copulas

This section discusses the theoretical behavior of the SMI posterior $p_\mathrm{SMI}(\eta,\psi\mid\mathcal{D})$ developed in Section (ref). While the discussion pertains to copula models as presented here, many of the points raised apply more generally to SMI posteriors constructed in the fashion of CarNic2020.

$\gamma$-dependent concentration results

Our SMPs dependence on $\gamma$ ultimately implies that the point onto which the posterior concentrates must be viewed as being $\gamma$-dependent. This result suggests that, unlike the usual learning rate encountered in generalized Bayesian inference, the value of $\gamma$ directly influences not just posterior width but location (i.e., concentration) as well. This suggests that one should optimize over $\gamma$ to determine which values deliver accurate inferences and predictions.

Given that the SMI posterior contains the fully cut posterior at (ref) as a special case, $\gamma_j=0$ for all $j$, and given that the theoretical behavior of standard posteriors, $\gamma_j=1$ for all $j$, are well-known, we restrict our theoretical analysis to values where $\gamma_j\in[0,1]$ and where $\gamma_j\not\in\{0,1\}$ for at least some $j$.

Define the joint vector of parameters $\varphi=(\psi^\top,\widetilde{\eta}^\top)^\top\in\Psi\times\mathcal{E}$, and the following negative SMI “log-pseudo-likelihood” and its corresponding limit counterpart $$ M_n(\varphi;\gamma):=-\log \int p_{\mathrm{SMI},\gamma}(\mathcal{D},u\mid \psi,\widetilde\eta)\mathrm{d} u,\quad \mathcal{M}(\varphi;\gamma):=\lim_{n\rightarrow\infty}\mathbb{E} [M_n(\varphi;\gamma)/n]. $$Further, define the $\gamma$-dependent pseudo-true value as $$ \varphi_0(\gamma):=\operatornamewithlimits{argmin}_{\varphi\in\Psi\times\mathcal{E}}\mathcal{M}(\varphi;\gamma). $$Likewise, for $ \ell_n(\eta,\psi)=\prod_ic(u_{i1},\dots,u_{id};\psi)\prod_jf(y_{ij}\mid\eta)$, and $\mathcal{L}_{\infty}(\eta,\psi)=\lim_n\mathbb{E}[\ell_n(\eta,\psi)]/n $ define $$ \eta_0(\gamma):=\operatornamewithlimits{argmin}_{\eta\in\mathcal{E}}\mathcal{L}_\infty\{\eta,\psi_0(\gamma)\}.$$

The following result clarifies the SMI posteriors dependence on $\gamma$: posterior concentration occurs at the standard rate but towards a $\gamma$-dependent quantity.

theoremUnder Assumptions 1-4 in Supporting Information B, for a positive sequence $r_n\rightarrow0$, $M_n$ large enough, some $C>0$ and $\alpha>0$, with probability converging to one, $$ \int \mathds{1}\{\|\psi-\psi_0(\gamma)\|>CM_n^\alpha r_n^\alpha\}p_{\mathrm{SMI},\gamma}(\psi\mid\mathcal{D})\mathrm{d}\psi\rightarrow0 $$and $$ \int \int \mathds{1}\{\|\eta-\eta_0(\gamma)\|>CM_n^\alpha r_n^\alpha\}p_{\mathrm{SMI},\gamma}(\eta,\psi\mid\mathcal{D})\mathrm{d}\psi\mathrm{d}\eta\rightarrow0. $$

The nature of the SMP construction means that results beyond posterior concentration derived in Theorem (ref) are unlikely to be enlightening due to the SMPs nonlinear dependence on $\gamma$. In particular, it is not obvious to us that the existing theoretical results on the asymptotic shape of cut posteriors, frazier2025cutting, would be satisfied for the SMP, and additional regularity conditions would be required to obtain results on the uncertainty quantification for this class of SMPs. Due to the technical nature of such constructions, we propose to study these issues in a more general setting separately.

Interpretation of results

The proposed SMP behaves in a distinctly different fashion to power-posterior SMP approaches: since we have multiple values of $\gamma$, so long as at least one $\gamma_j\in(0,1)$, the likelihood for the $j$-th marginal parameter $\widetilde{\eta}_j$ is still informed by data. Consequently, our theoretical results should be interpreted as showing that the SMP defines a posterior process $\{\gamma\in[0,1]^d: p_{\mathrm{SMI},\gamma}(\psi,\eta\mid\mathcal{D})\}$, where the points onto which $p_{\mathrm{SMI},\gamma}(\psi,\eta\mid\mathcal{D})$ concentrate explicitly depend on the values of $\gamma$.

This behavior means that, even though the value of $\gamma$ is treated as a hyper-parameter in $p_{\mathrm{SMI},\gamma}$, its impact on the resulting inferences is substantially different from other hyper-parameters in generalized Bayesian inference such as the learning rate: it is well-known that the choice of learning rate does not impact the point onto which the posterior concentrates, at least in regular models; see syring2023gibbs and mclatchie2025predictive for discussion of this point.

SMPs dependence on $\gamma$ suggests that it may be possible to learn an optimal value of $\gamma$ according to some measure of “accuracy”. In particular, since each individual value of $\gamma\in[0,1]^d$ indexes a generalized posterior, and thus some loss function, we could in-principle select $\gamma$ via the loss function selection method proposed in jewson2022general where we place uniform priors (on 0 to 1) over each component $\gamma_j$. However, in contrast to the analysis in jewson2022general, there is no reason to suspect that the value of $\gamma$, and hence $\varphi_0(\gamma)$, should be uniquely identified from the data.

To demonstrate this fact, consider that the marginal models $F_j\left(y_{i j} ; \eta_j\right)$ are well-specified. Then, writing $r(y_{ij})/(n+1)=\frac{n}{n+1}F_n(y_{ij})$ we see that, for $n$ large,

flalign*a_{i j}\left(\gamma_j, \eta^0_j, \mathcal{D}\right)&=\gamma_j F_j\left(y_{i j} ; \eta^0_j\right)+\left(1-\gamma_j\right) \frac{r\left(y_{i j}\right)-1}{n+1}=F(y_{ij};{\eta}_j^0)-\frac{(1-\gamma_j)}{n+1}+o(n),\\ b_{i j}\left(\gamma_j, \eta^0_j, \mathcal{D}\right)&=\gamma_j F_j\left(y_{i j} ; \eta_j\right)+\left(1-\gamma_j\right) \frac{r\left(y_{i j}\right)}{n+1}= \left(\frac{n}{n+1}\right)F(y_{ij};{\eta}_j^0)+o(n),

since $\frac{n}{n+1}F_n(y_{ij})\approx F(y_{ij};\widetilde{\eta}_j^0)$. Hence, if the marginals are well-specified, the value of $\gamma_j$ is irrelevant, and unidentifiable asymptotically. Conversely, if the model is misspecified, then the value of $\gamma_j$ may have a non-negligible impact: define $F_0$ to be the true distribution and let $\Delta(\gamma_j)=\gamma_j\{F(y_{ij};\eta_j)-F_0(y_{ij})\}$, then we see that

flalign*a_{i j}\left(\gamma_j, \eta_j, \mathcal{D}\right)&=\Delta(\gamma_j)+F_0(y_{ij})\left(\frac{n}{n+1}\right)-\frac{(1-\gamma_j)}{n+1}+o(n),\\ b_{i j}\left(\gamma_j, \eta_j, \mathcal{D}\right)&=\Delta(\gamma_j)+F_0(y_{ij})\left(\frac{n}{n+1}\right)+o(n),

which will clearly depend on $\gamma_j$, and will therefore influence the values of $\psi,\eta$ onto which the posterior concentrates.

This does not mean that one cannot attempt to target some value in the equivalence class of optima. Herein, we have attempted to select such a value of $\gamma$, and ultimately the value of $\varphi_\star(\gamma)$, through BO. We note here that the theoretical analysis of such a procedure is complicated by the lack of identification and so we leave a detailed theoretical analysis of the optimal value of $\gamma$, and thus the point onto which the SMP concentrates to further research.

Before finishing this section, we note that many of the points raised by our theory are not unique to our SMP but are more generally true, under different conditions, for the SMP of CarNic2020. That is, many of the techniques we have used to derive results for our SMP are directly applicable to the SMP of CarNic2020, and in this way minor modifications of our results will deliver concentration results for the SMP approach of CarNic2020. Consequently, similar to our SMP, the SMP of CarNic2020 will also concentrate onto a $\gamma$-dependent quantity. We do note, however, that this will not be the case for the SMP of FraNot2025, which builds a linear pool between two different posteriors, and ultimately has different behavior to the SMI considered here. That being said, the SMP of FraNot2025 does not allow for dimension-by-dimension cutting as the modularity is applied at the level of the posterior, which is much less flexible than the approach proposed herein.

Simulations

To illustrate how the SMI posterior mediates information flow between modules we consider a bivariate copula model where the dependence is correctly specified and only one marginal is misspecified. We generate repeated datasets from the data generating process (DGP), and vary $\gamma=(\gamma_1,\gamma_2)^\top$ over a grid with the following three goals: (i) to illustrate how reducing the influence of the misspecified marginal can improve estimation of the dependence structure, (ii) to characterize the trade-offs the SMI posterior introduces for inference on the correctly specified marginal that is coupled to the misspecified marginal through the copula, and (iii) to show how $\gamma$ can be selected using a simple utility function in practice.

\paragraph{Data generating process} We generate a total of $m=100$ datasets, each with $n=1,000$ observations, from a bivariate Gumbel copula model. The copula is parameterized through Kendall's tau correlation coefficient, with true value $\tau^\ast=0.7$. The first marginal with density $f_1^*$ is a log-normal distribution with location $\mu^\ast=-1$, and scale $\sigma^\ast=1$. The second marginal $f_2^*$ is gamma distributed with shape $\alpha^\ast=1$ and rate $\beta^\ast=1$.

\paragraph{Statistical model} To each dataset we fit a Gumbel copula model with log-normal marginals with densities $f_1, f_2$, so that the copula and the first marginal are correctly specified, while the second marginal is misspecified. The model parameters are $\tau$ for the Gumbel copula, as well as $(\mu_j,\sigma^2_j)$, $j=1,2$, for the marginals. Writing $\psi=\log\frac{\tau}{1-\tau}$, $\eta_j=\left(\mu_j,\log\sigma_j\right)^\top$, $j=1,2$, $\eta=(\eta_1^\top,\eta_2^\top)^\top$ gives the vector of unconstrained parameters $\theta=\left(\psi,\eta^\top\right)^\top$. We adopt the following weakly informative priors $\tau\sim\operatorname{U}(0,1)$, $\mu_1,\mu_2\sim\operatorname{\cal N}(0,100^2)$, and $\sigma^2_1,\sigma^2_2\sim\operatorname{HN}(0,100^2)$, where $\operatorname{HN}$ is a half-normal distribution.

\paragraph{Performance metrics} To evaluate the fit to the copula, we consider the mean squared error for $\tau$, integrating over the SMI posterior for $\psi$,

equation*[equation* omitted — 113 chars of source]

Since the marginals are potentially misspecified, we evaluate the model fit for each of the marginals by considering their predictive Kullback-Leibler divergences, integrating over the SMI posterior for $\eta_j$,

equation*[equation* omitted — 140 chars of source]

The integrals are numerically approximated using Monte Carlo samples from $p_{\text{SMI},\gamma}(\theta\mid\mathcal{D})$.

\paragraph{Results} Figures (ref)(A-C) show average values for $\mathcal{L}_1$, $\mathcal{L}_2$, and $\mathcal{L}_\text{cop}$ across the $100$ replicates. In each panel, $\gamma=(\gamma_1,\gamma_2)^\top$ is varied over a mesh on $[0,1]^2$ with a grid size of $0.1$. Both $\mathcal{L}_2$ and $\mathcal{L}_\text{cop}$ exhibit a similar structure: small values for $\gamma_2$ and values close to $1$ for $\gamma_1$ improve the estimation of the copula and the second marginal. That is, inference is improved when influence from the second but not from the first marginal is reduced, which matches the misspecification of the model. In contrast to this $\mathcal{L}_1$ is lowest when $\gamma_2=1$ (i.e. the misspecified marginal is uncut). Notably, $\gamma$ controls the influence of $\eta$ on $\psi$ while the information flow between $\eta_1$ and $\eta_2$ through the copula is only indirectly controlled by $\gamma$.

Under the conventional posterior, $\gamma=(1,1)^\top$, $\tau$ is underestimated in this example. The maximum a-posteriori estimator for $\tau$ derived from the conventional posterior has an average bias of $-0.0113$ over all repetitions, so that the parameters of the misspecified marginal $\eta_2$ have less influence on $\eta_1$. Because $\tau$ is underestimated under the conventional posterior, dependence between the two marginals is weakened. Counteracting this, by choosing $\gamma$ so that the estimation of the copula is improved, increases the information flow from the misspecified second marginal to the well-specified first marginal resulting in poorer inference for the first marginal. Since the first marginal is well-specified the influence of $\gamma_1$ will vanish asymptotically as shown in Section (ref). Figure (ref) illustrates this behaviour already for finite $n$. If $\gamma_2$ is small, the choice of $\gamma_1$ has only very limited influence on the three performance metrics selected. In particular, $\mathcal{L}_\text{cop}$ is almost completely flat on $[0,1]\times[0,0.5]$. This also indicates that an optimal $\gamma$ might not be uniquely identified.

figure[figure omitted — 619 chars of source]

\paragraph{Selecting $\gamma$} The results presented above illustrate the difficulties in selecting $\gamma$. Even if prior knowledge on the misspecification of the marginals is available, choosing $\gamma$ to match the misspecification can lead to unforeseen and undesired consequences in the estimation. In this example, the first marginal is correctly specified while the second marginal is misspecified, indicating a choice of $\gamma=(1,0)^\top$, where only influence from the second marginal is cut. While this choice improves the estimation of the dependence structure and the misspecified marginal, the estimation of the first marginal deteriorates.

As shown in Section (ref) the behaviour of the SMI posterior and in particular the point on which the posterior concentrates asymptotically, depends directly on $\gamma$. Here, we consider the utility function

align[align omitted — 151 chars of source]

as an external metric to select $\gamma^*$. (ref) balances the trade-off between the marginal fits described above. Figure (ref)(D) shows average values for $-u(\gamma)$ across the $100$ repetitions. Optimal values are reached for $\gamma_2$ close to $0$ and $\gamma_1$ close to 1 matching the misspecification in the statistical model. Truly SMI posteriors with $\gamma\in(0,1)^2$ outperform possible cut-posteriors with multiple cuts, $\gamma\in\{0,1\}^2$, in this metric.

Stock market volatility and bond yields

Problem description and data

There is growing evidence that volatility in the stock market and bond yields are dependent; see Campbelletal2003,Jubinski2012 and MoeSorJFE2022. However, the form of this dependence has not been explored previously. We do so here using a copula model with a skew-normal (SN) copula that allows for asymmetric dependence between variable pairs. Parametric marginal distributions are preferred over nonparametric ones because focus is on the tail behavior of the variables, and we use a “Sinh–Arcsinh” (SAS) distribution jones2009sinh for each marginal. This distribution is a four-parameter family constructed from a transformation of a Gaussian base distribution that models skewness and kurtosis flexibly. However, one or more of these marginals may be misspecified, so that computing the SMI posterior is attractive here.

For bond yields we consider the effective yield on domestic U.S. publicly issued debt graded AAA ($Y_{\text{AAA},t}$) and BBB ($Y_{\text{BBB},t}$) by Standard & Poors. The former grade is the least risky, while the latter is more risky and the most liquid. Daily values from 1 January 2022 to 30 June 2025 were sourced from the Federal Reserve Bank of St. Louis, which corresponds to the post-pandemic period when U.S. interest rates rose quickly from historical lows. We consider the dependence of the yields with U.S. equity market volatility measured using the average daily values of the VIX ($Y_{\text{VIX},t}$) sourced from Yahoo Finance. All variables are transformed to the real line using their logarithms, and we fit the copula model to the $d=6$ dimensional vector $Y_t=(Y_{\text{VIX},t},Y_{\text{AAA},t},Y_{\text{BBB},t},Y_{\text{VIX},t-1},Y_{\text{AAA},t-1},Y_{\text{BBB},t-1})^\top$. Our goal is to gain insights into the temporal and cross-sectional dependence structure.

Copula model

The SN copula is the implicit copula of the SN distribution of AzzVal1996; see smith2023implicit for an overview of implicit copulas. Let $X=(X_1,\ldots,X_d)^\top\sim F_X$ be distributed SN with location zero, scale matrix $\Omega$, and skewness parameter $\alpha\in\mathbb{R}^d$. Then it has density

equation*[equation* omitted — 84 chars of source]

where $x=(x_1,\ldots,x_d)^\top$, $\varphi(x;\mu,\Sigma)$ denotes the density of an $N(\mu,\Sigma)$ distribution, and $\Phi$ is the standard normal distribution function. The SN distribution is closed under marginalization, with $F_{X_j}$ a zero mean univariate SN distribution with skewness parameter $\bar{\alpha}_j$ that is a function of $\alpha,\Omega$. Then, $X$ has copula function $C(u)=F_X(F^{-1}_{X_1}(u_1),\dots,F^{-1}_{X_d}(u_d))$ with density

equation*[equation* omitted — 115 chars of source]

where $x_j=F^{-1}_{X_j}(u_j)$ and $f_{X_j}=\frac{\partial}{\partial x_j}F_{X_j}$. Setting $\Omega$ to a correlation matrix identifies the copula parameters $\{\Omega,\alpha\}$ and enforces $F_{X_j}$ to have unit scale for $j=1,\ldots,d$.

We assume that the three series $(Y_{\text{VIX},t},Y_{\text{AAA},t},Y_{\text{BBB},t})$ follow a 3-dimensional stationary stochastic process. Because $Y_t$ contains lagged values of these three series, this assumption further restricts the parameter space, with $\Omega$ a block symmetric matrix, $\alpha$ constrained, and the marginals for $Y_{j,t}$ and $Y_{j,t-1}$ being the same for each $j\in\{\text{VIX},\text{AAA},\text{BBB}\}$. Full specification of the copula parameterization under this stationary assumption is given in Supporting Information D.1, where the dimensions of $\psi$, $\eta$ and $\gamma$ are 15, 12 and 3, respectively. Finally, when evaluating the copula density and function, calculating $F_{X_j}^{-1}$ is a key computation, and we do this using the spline interpolation approach outlined in SmiMan2018.

Empirical results

\paragraph{Selecting $\gamma$} We consider the utility function

equation[equation omitted — 206 chars of source]

which we optimize using BO. This utility function is similar to that discussed in Section (ref). The fully cut model proposed by SmiYuNotFra2025 corresponds to $\gamma=(0,0,0)^\top$ and has a higher utility value than the conventional posterior where $\gamma=(1,1,1)^\top$. The optimal vector of influence parameters maximizing (ref) is $\gamma^\ast=(1.00,0.61,0.00)^\top$.

\paragraph{Marginal fit} Figure (ref) shows histograms for the three marginals as well as estimated densities under the posterior means for the conventional, fully cut, and optimal SMI posteriors. Differences between these three posteriors are clearly visible, with a SAS distribution capturing the VIX marginal well, but not those of the AAA and BBB yields. The estimated marginal density under the conventional posterior does not fully capture the modes of these two marginals, which are better captured by the cut and SMI posteriors. The marginal of the BBB yield variable has the strongest misspecification, followed by that of the AAA yield variable, which is consistent with the optimal $\gamma^\ast$ value.

figure[figure omitted — 613 chars of source]
figure[figure omitted — 499 chars of source]

\paragraph{Dependencies} The main objective of the analysis is estimate the dependence structure between the variables, which we summarize using pairwise metrics. Let $U=(U_1,\dots,U_d)\sim C(u;\psi)$, we define for a pair $(U_j,U_l)$, $j\not=l$, and a given quantile $\zeta\in(0,0.5)$ dependence in each of the quadrants as:

itemize[itemsep=0pt] • Lower Left: $\rho_\text{LL}(\zeta;U_j,U_l)=\mathbb{P}(U_l\leq\zeta\mid U_j\leq \zeta)$, • Upper Right: $\rho_\text{UR}(\zeta;U_j,U_l)=\mathbb{P}(U_l>1-\zeta\mid U_j>1-\zeta)$, • Lower Right: $\rho_\text{LR}(\zeta;U_j,U_l)=\mathbb{P}(U_l\leq\zeta\mid U_j>1-\zeta)$, • Upper Left: $\rho_\text{UL}(\zeta;U_j,U_l)=\mathbb{P}(U_l>1-\zeta\mid U_j\leq\zeta)$,

which are computed from the bivariate marginal copula of $(U_j,U_l)$. Note that the variable pairs are strictly ordered, and switching their order gives the identities $\rho_\text{LL}(\zeta;U_j,U_l)=\rho_\text{UR}(\zeta;U_l,U_j)$ and $\rho_\text{LR}(\zeta;U_j,U_l)=\rho_\text{UL}(\zeta;U_l,U_j)$. Following OhPat2013, Figure (ref) presents “quantile dependence plots” for the five ordered pairs VIX-AAA, VIX-BBB, AAA-BBB, VIXlag-AAA and VIXlag-BBB. These visualize asymmetric dependence along the major diagonal by plotting $\rho_\text{LL}(\zeta)$ and $\rho_\text{UR}(1-\zeta)$ (blue and orange lines), and along the minor diagonal by plotting $\rho_\text{LR}(\zeta)$ and $\rho_\text{UL}(1-\zeta)$ against $\zeta\in(0,0.5)$ (green and red lines). Estimates are given for the conventional, cut and SMI posterior means of $\psi$ in the three columns. In each panel, the empirical (i.e. nonparametric) equivalent is also given in gray for comparison. Major diagonal dependence between AAA and BBB are strong for all estimators, reflecting the fact that bond yields move together very tightly in response market events. However, the estimated dependence structure between the yields and the VIX (contemporaneous or lagged) differs greatly between conventional and cut/SMI posteriors. The conventional posterior suggests dependence is symmetric, whereas the cut and SMI posteriors suggest strong asymmetry which is more consistent with the empirical equivalents. The asymmetry is most apparent in the minor diagonal, as overall dependence between the VIX and yields is negative.

Following DenSmiMan2025, pairwise asymmetry in the major and minor diagonals for the bivariate copula of pair $(U_j,U_l)$ can be measured at quantile $\zeta$ by

eqnarray*[eqnarray* omitted — 234 chars of source]

Switching the order of the variables gives the identities $\Delta_\text{Major}(\zeta;U_j,U_l)=-\Delta_\text{Major}(\zeta;U_l,U_j)$ and $\Delta_\text{Minor}(\zeta;U_j,U_l)=-\Delta_\text{Minor}(\zeta;U_l,U_j)$. Figure (ref) plots the posterior means of these at quantile $\zeta=0.1$ for all variable pairs in the form of heatmaps. Results are not given for the conventional posterior because they are either exactly or very close to zero for all variable pairs. However, there is a clear difference between those for the cut and SMI posteriors, with the latter indicating strong positive asymmetric dependence between the VIX (contemporanous or lagged) and AAA and BBB yields. This is more economically intuitive, as volatility spikes typically coincide with risk-off episodes that induce flight-to-quality and nonlinear repricing of credit risk. It is also consistent with existing empirical evidence that equity market volatility is more strongly linked to bond yields and credit spreads in stress states than in tranquil periods Jubinski2012.

figure[figure omitted — 746 chars of source]

Further empirical results, including copula parameter estimates and Kendall's tau values, are given for all three posteriors in Supporting Information D.2.

Conclusion

We have developed a novel semi-modular inference framework for copula models, designed for settings where inferences about copula parameters must be protected from corruption by misspecified marginal distributions, and where the degree of misspecification varies across marginals. Our approach treats each marginal as its own module with a marginal-specific influence parameter, allowing the framework to adapt to the degree of misspecification by controlling how strongly each marginal shapes the semi-modular joint posterior. We provide efficient variational computational methods, a principled strategy for selecting influence parameters via Bayesian optimization of a utility function, and theoretical validation in the form of a posterior concentration result. The practical value of the methodology is demonstrated through simulations and a substantive real application examining the relationship between bond yields and a volatility index, where the SMI posterior most clearly reveals the asymmetric dependence structure compared to the fully cut and conventional posterior.

There are many directions for future work and we mention one of them. Our paper has focused on treating each marginal as its own module, rather than treating all marginals as a single module. It would also be possible to consider decomposing the copula function in some way, perhaps using a vine copula, and decomposing the copula part of the model into multiple distinct modules. When different copula components are misspecified to varying degrees, the semi-modular framework could then be used to limit the influence of misspecified copula modules on inference for parameters elsewhere in the model, offering a more flexible approach to robust dependence modelling.

Supplementary material

Supporting Information An online Supporting Information contains A) further discussion on the mixed pseudo likelihood (ref), B) the assumptions and proofs for the theoretical analysis, C) additional information for the simulation study, and D) further results for the financial application (.pdf).

\noindentPython code Code is publicly available at \href{https://github.com/kocklucx/CopulaSMI}{github.com/kocklucx/CopulaSMI}.

\FloatBarrier {0pt plus 0.3ex} \FloatBarrier