Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
85,992 characters · 16 sections · 86 citation commands
Model selection in a high-dimensional setting is a common challenge in statistical and econometric inference. The introduction of Bayesian model averaging (BMA) techniques in the statistical literature raf-etal:bay_mod,bro-etal:bayJRSSB,cot-etal:var has led to many interesting applications, see, among others, koo-pot:for,sal-etal:det,kle-van:baymod,fru-tue:bay for early references in econometrics.
Predictor selection for possibly very high-dimensional regression problems though shrinkage priors is an attractive alternative to BMA which relies on discrete mixture priors, see bha-etal:las for an excellent review. There is a vast and growing literature on shrinkage priors for regression problems that focuses on the following aspects. First, how to choose sensible priors for high-dimensional model selection problems in a Bayesian framework, second, how to design efficient algorithms to cope with the associated computational challenges and third, to investigate, both from a theoretical and a practical viewpoint, how such priors perform in high-dimensional problems.
A striking duality exists in this very active area between Bayesian and traditional approaches. For many shrinkage priors, the mode of the posterior distribution obtained in a Bayesian analysis can be regarded as a point estimate from a regularization approach, see fah-etal:bay and pol-sco:loc. One such example is the popular Lasso tib:reg which is equivalent to a double-exponential shrinkage prior in a Bayesian context par-cas:bay. However, the two approaches differ when it comes to selecting penalty parameters that impact the sparsity of the solution. One advantage of the Bayesian framework in this context is that the penalty parameters are considered to be unknown hyperparameters which can be learned from the data. Such \lq\lq global-local\rq\rq\ shrinkage priors pol-sco:shr adjust to the overall degree of sparsity that is required in a specific application through a global shrinkage parameter and separates signal from noise through local, individual shrinkage parameters.
While predictor selection though shrinkage priors in regression models is addressed in a vast literature, the use of shrinkage priors for more general econometric models for time series analysis, such as state space models and time-varying parameter (TVP) models is, in comparison, less well-studied. Sparsity in the context of such models refers to the presence of a few large variances among many (nearly) zero variances in the latent state processes that drive the observed time series data. A common goal in this setting is to recover a few dynamic states, driven by such a state space model, among many (nearly) constant coefficients. As shown by fru-wag:sto, this variance selection problem can be cast into a variable selection problem in the non-centered parametrization of a state space model. Once this link has been established, shrinkage priors that are known to perform well in high-dimensional regression problems can be applied to variance selection in state space models, as demonstrated for the Lasso bel-etal:hie_tv and the normal-gamma gri-bro:hie,bit-fru:ach.
Despite this already existing variety, we introduce a new shrinkage prior for variance selection in sparse state space and TVP models in the present paper, called triple gamma prior as it has a representation involving three gamma distributions. This prior can be related to various shrinkage priors that were found to be useful for high-dimensional regression problems such as the generalized beta mixture prior arm-etal:gen_bet and contains the popular Horseshoe prior car-etal:han,car-etal:hor as a special case. Furthermore, the half-$t$ and the half Cauchy gel:pri,pol-sco:hal, suggested as robust alternatives to the inverse gamma distribution for variance parameters in hierarchical models, as well as the Lasso and the double gamma, are special cases of the triple gamma. In this context, the triple gamma can also be regarded as an extension of the scaled beta2 distribution per-etal:sca.
Among Bayesian shrinkage priors, usually a clear distinction is made between two-group mixture or spike-and-slab priors and continuous shrinkage priors, of which the triple gamma is a special case. An important contribution of the present paper is to show that the triple gamma provides a bridge between these two approaches and has the following property which is favourable both in sparse and dense situations. One of the hyperparameters allows high concentration over the region in the shrinkage profile that is relevant for shrinking noise, while the other hyperparameter allows high concentration over the region that prevents overshrinking of signals. This leads to a behaviour of the triple gamma prior that very much resembles Bayesian model averaging based on discrete spike-and-slab priors, with a strong prior concentration at the corner solutions where some of the variances are nearly close to zero. While this is reminiscent of the Horseshoe prior, the shrinkage profile induced by the triple gamma is more flexible than that of a Horseshoe. Thanks to the estimation of the hyperparemters, it is not constrained to be symmetric around one half, enabling adaption to varying degrees of sparsity in the data.
The triple gamma prior also scores well from a computational perspective. While exploring the full posterior distribution for spike-and-slab priors leads to computational challenges due to the combinatorial complexity of the model space, Bayesian inference based on Markov chain Monte Carlo (MCMC) methods is straightforward for continuous shrinkage priors, exploiting their Gaussian-scale mixture representation mak-sch:sim,bit-fru:ach. An extension of these schemes to the triple gamma prior is fairly straightforward.
We will study the empirical performance of the triple gamma for a challenging setting in econometric time series analysis, namely for time-varying parameter vector autoregressive models with stochastic volatility (TVP-VAR-SV models). Since the influential paper of pri:tim (see del-pro:tim for a corrigendum), this model has become a benchmark for analyzing relationships between macroeconomic variables that evolve over time, see nak:tim, koo-kor:lar, eis-etal:sto, cha-eis:bay, fel-etal:sop and car-etal:lar, among many others. Due to the high dimensionality of the time-varying parameters, even for moderately sized systems, shrinkage priors such as the triple gamma prior are instrumental for efficient inference.
The rest of the paper is organized as follows. In Section (ref), we define the triple gamma prior and discuss some of its properties. The close relationship between the triple gamma and spike-and-slab priors applied in a BMA context is investigated in Section (ref). Section (ref) introduced an efficient MCMC scheme and Section (ref) provides applications to TVP-VAR-SV models. Section (ref) concludes the paper.
Let us recall the state space form of a TVP model. For $t = 1, \ldots, T$, we have that
where ${\mathbf{Q}}=\mbox{\rm Diag}\left(\theta_1, \ldots, \theta_d\right)$, $y_t$ is a univariate response variable, $ { \mathbf x}_t = (x_{t 1}, \ldots, x_{t d})$ is a $d$-dimensional row vector containing the regressors at time $t$, with $x_{t 1}$ corresponding to the intercept, and the initial value follows a normal distribution, $\boldsymbol{\beta}_{0} \sim \mathcal{N} _{d}\left(\boldsymbol{\beta}, {\mathbf{Q}}\right)$, with initial mean $\boldsymbol{\beta} = (\beta_1, \ldots, \beta_d)^\top$. Model ((ref)) can be rewritten equivalently in the non-centered parametrization introduced in fru-wag:sto as
with $ \tilde{\boldsymbol{\beta}}_{0} \sim \mathcal{N} _{d}\left({\mathbf{0}}, {{\mathbf I}}_{d}\right) $, where ${{\mathbf I}}_{d}$ is the $d$-dimensional identity matrix. The error variance in the observation equation is either homoscedastic ($\sigma^2_t \equiv \sigma^2$ for all $t=1,\dots,T$) or follows a stochastic volatility (SV) specification jac-etal:bayJBES, where the log volatility $h_t = \log \sigma^2_t $ follows an AR(1) process. Specifically,
To motivate the triple gamma prior, let us recall that, in TVP models, shrinkage priors are placed on each scale parameter $\sqrt{\theta_j}$, $j=1, \ldots,d$, in order to shrink dynamic coefficients to static ones, hence avoiding overfitting. One of such priors is the double gamma prior, employed recently by bit-fru:ach for shrinkage of variances. The double gamma prior can be expressed as a scale-mixture of gamma distributions, with the following hierarchical representation:
In the double gamma prior, each innovation variance $\theta_j$ is mixed over its own $\xi_j^2$ , each of which has an independent gamma distribution, with a common hyperparameter $\kappa_B^2$. Moreover, the parameters $\xi_j^2$ play the role of local (component specific) shrinkage parameters, while the parameter $\kappa_B^2$ is a (common) global shrinkage parameter.
We propose an extension of the double gamma prior to a triple gamma prior, where another layer is added to the hierarchy:
The main difference with the double gamma prior is that the $\xi_j$ are not identically distributed, but each one depends on its component specific parameter $\kappa_j^2$. Prior ((ref)) contains many well-known shrinkage priors as a special case, as will be discussed in Section (ref).
To make the shrinkage behaviour of the triple gamma prior more apparent, we will work with representations that involve the scale parameter $\sqrt{\theta_j}$, rather than the variance $\theta_j$, using the fact that $\theta_j | \xi_j^2 \sim \xi_j^2 \chi^2 _{1}$ follows a re-scaled $\chi^2 _{1}$-distribution. If we consider both the positive and the negative root of $\theta_j$, then we obtain
Hence, prior ((ref)) corresponds to $\sqrt{\theta_j}$ following the so-called normal-gamma-gamma prior consider by gri-bro:hie in the context of defining hierarchical shrinkage priors for regression models.
To allow shrinkage of dynamic coefficients toward fixed, but significant ones, we extend bit-fru:ach further by assuming such a normal-gamma-gamma prior on the fixed parameter $\beta_1, \ldots, \beta_d$:
In Section (ref), we will discuss hierarchical versions of both priors, by putting a hyperprior on the parameters $\kappa^2_B$, $\lambda^2_B$, $a^\xi$, $a^\tau$, $c^\xi$, and $c^\tau$.
It will be shown in Theorem (ref) that the triple gamma prior is a global-local shrinkage prior in the sense of pol-sco:loc where the local shrinkage parameters arise from the $\mbox{\rm F}\left(2a^\xi, 2 c^\xi\right)$ distribution. This representation allows to relate the triple gamma to the well-known Horseshoe prior, see Section (ref). Furthermore, a closed form of the marginal shrinkage prior $p(\sqrt{\theta_j}|\phi^{\xi},a^\xi, c^\xi)$ is given in Theorem (ref), which is proven in Appendix (ref).
In Figure (ref) we can see the marginal prior distribution of $\sqrt{\theta_j}$ under the triple gamma prior $a^\xi=c^\xi=0.1$ and under other well-known shrinkage priors which are special cases of the triple gamma, see Table (ref). Using bit-fru:ach, Theorem (ref) also allows to give a closed form for the prior $p(\theta_j|\phi^{\xi}, a^\xi, c^\xi)= p(\sqrt{\theta_j}|\phi^{\xi}, a^\xi, c^\xi)/\sqrt{\theta_j}$.
Global-local shrinkage priors are typically compared in terms of the concentration around the origin and the tail behaviour. For the triple gamma prior $p(\sqrt{\theta_j}|\phi^{\xi},a^\xi, c^\xi)$, the two shape parameters $a^\xi$ and $c^\xi$ play a crucial role in this respect, see Theorem (ref) which is proven in Appendix (ref).
From Theorem (ref), Part (a) and (b), we find that the triple gamma prior $p(\sqrt{\theta_j}|\phi^{\xi},a^\xi, c^\xi)$ has a pole at the origin, if $a^\xi \leq 0.5$. According to Part (a), the pole is more pronounced, the closer $a^\xi $ gets to 0. For $a^\xi > 0.5$, we find from Part (c) that $p(\sqrt{\theta_j}|\phi^{\xi},a^\xi, c^\xi)$ is bounded at zero by a positive upper bound which is finite, as long as $0< c^\xi< \infty$. Part (d) shows that the triple gamma prior $p(\sqrt{\theta_j}|\phi^{\xi},a^\xi, c^\xi)$ has polynomial tails, with the shape parameter $c^\xi$ controlling the tail index. Prior moments $\mbox{\rm E}((\sqrt{\theta_j})^k|\phi^{\xi},a^\xi, c^\xi)$ exist up to $k< 2 c^\xi$. Hence, the triple gamma prior has no finite moments for $c^\xi< 1/2$.
Finally, additional useful representations of the triple gamma prior as a global-local shrinkage prior are summarized in Lemma (ref) which is proven in Appendix (ref). Representations (a) shows that the triple gamma is an extension of the double gamma prior where the Gaussian prior $\sqrt{\theta_j} | \xi_j^2 \sim \mathcal{N}(0,\xi_j^2)$ is substituted by a heavier-tailed Student-$t$ prior, making the prior more robust to large values of $\sqrt{\theta_j}$. Representation (b) and (c) will be useful for MCMC inference in Section (ref). Representations (c) and (d) show that for a triple gamma prior with finite $a^\xi$ and $c^\xi$, $\phi^{\xi}$ acts as a global shrinkage parameter, in addition to $2/\kappa_B^2$.
The triple gamma prior can be related to the very active research on shrinkage priors in a Bayesian framework in various ways. On the one hand, popular priors for variance parameters introduced as robust alternatives to the inverse gamma prior are special cases of the triple gamma, see Table (ref). For instance, in ((ref)), $\psi^2 _j$ converges a.s. to 1, as $a^\xi \to \infty$ and $c^\xi \to \infty$, and the triple gamma reduces to a normal distribution for $\sqrt \theta _j$, applied for univariate TVP models fru:eff and unobserved component state space model fru-wag:sto. For $c^\xi \to \infty$, $\mbox{\rm F}\left(2a^\xi, 2 c^\xi\right) $ converges to the $ \mathcal{G}\left(a^\xi,a^\xi\right)$ distribution and the triple gamma reduces to the Bayesian Lasso for $a^\xi=1$ bel-etal:hie_tv and otherwise to the double gamma bit-fru:ach applied in sparse TVP models.
gel:pri introduced the half-$t$ and the half-Cauchy prior for variance parameters in hierarchical models, by assuming that $\sqrt \theta _j$ follows a \lq\lq folded\rq\rq\ $t$-distribution, i.e. a $t$-distribution truncated to $[0, \infty)$, see also pol-sco:hal. In ((ref)), $\tilde{\xi}_j ^2$ converges a.s. to 1 as $a^\xi \to \infty$ and the triple gamma reduces to a $t _{2 c^\xi}$- distribution and to the Cauchy distribution for $ c^\xi =1/2$, however without being \lq\lq folded\rq\rq , since we allow $\sqrt \theta _j$ to take on negative values. A half-$t_{\nu}$ with $\nu= 2 c^\xi$ and a triple gamma with $a^\xi=\infty$ obviously imply the same prior for $\theta _j$, so does the negative half matter? It matters, whenever inference is performed in a parametrization involving $\sqrt \theta _j$ such as the non-centered parametrization ((ref)). Restricting the prior to the positive half will lead to automatic truncation of the full conditional posterior $p(\sqrt \theta_j| \tilde{\boldsymbol{\beta}}_{0}, \ldots, \tilde{\boldsymbol{\beta}}_{T}, \bm y, \cdot)$, e.g. during MCMC sampling (see Section (ref)). If the positive and the negative mode of the marginal posterior $p(\sqrt \theta_j| \bm y)$ are well-separated, then this will not matter. However, if the true value of $\theta _j$ is close to or equal to zero, then $p(\sqrt \theta_j| \bm y)$ is concentrated at zero and truncation at 0 will introduce a bias, because the negative half is missing.
On the other hand, the triple gamma prior is related to popular shrinkage priors in regression models. It extends the generalized beta mixture prior introduced by arm-etal:gen_bet for variable selection in regression models,
to variance selection in state space and TVP models. This is evident from rewriting ((ref)) as $ \xi_j^2 \sim \mathcal{G}\left(a^\xi, \lambda_j\right), \lambda_j \sim \mathcal{G}\left(c^\xi, \phi^{\xi}\right)$. We exploit this relationship in Section (ref) to investigate the shrinkage profile of a triple gamma prior. Using arm-etal:gen_bet, the triple gamma prior can be written as
where $\mathcal{TPB}\left(a^\xi, c^\xi, \phi^{\xi}\right)$ is the three-parameter beta (TPB) distribution with density:
From ((ref)) and ((ref)), it becomes evident that the Strawderman-Berger prior $\sqrt \theta_j \sim \mathcal{N}\left(0, 1/\rho_j-1\right)$, $\rho_j \sim \mathcal{B}\left(1/2,1\right) $ str:pro,ber:rob is that special case of the triple gamma prior where $\phi^{\xi}=1$, $a^\xi=1/2$, and $c^\xi=1$.
The special case of a triple gamma, where $a^\xi=c^\xi =1/2$, corresponds to a Horseshoe prior car-etal:han,car-etal:hor on $\sqrt \theta_j $ with global shrinkage parameter $\tau^2=2/\kappa_B^2$, since $\psi^2 _j \sim \mbox{\rm F}\left(1,1\right)$ implies that $ \psi _j \sim t _{1}$. The Horseshoe prior has been introduced for variable selection in regression models and has been shown to have excellent theoretical properties in this context for the \lq\lq nearly black\rq\rq\ case van-etal:hor. The triple gamma is a generalization of the Horseshoe prior, with a similar shrinkage profile, however with much more mass close to the corner solutions. Most importantly, as will be discussed in Section (ref), this leads to a BMA-type behaviour of the triple gamma prior for small values of $a^\xi$ and $c^\xi$.
The vast literature on shrinkage priors contains many more related priors. Rescaling $ \xi_j^2= 2/(\kappa_B^2) \psi^2 _j$ in ((ref)), for instance, yields a representation involving a scaled beta2 distribution,\footnote{The pdf of a $\mbox{SBeta2} \left(a, c, \phi\right)$-distribution reads:
}
as is easily derived from ((ref)). The scaled beta2 was introduced by per-etal:sca in hierarchical models as a robust prior for scale parameters, $\sqrt \theta_j $, and variance parameters, $\theta_j$, alike. Based on ((ref)), the triple gamma can be seen as a hierarchical extension of this prior which puts a scaled beta2 distribution on the scaling parameter $\xi_j^2$ of a Gaussian prior for $\sqrt \theta_j $, see Table (ref). gri-bro:hie termed prior ((ref)) gamma-gamma distribution, denoted by $\mbox{GG}\left(a^\xi,c^\xi,\phi\right)$.
For $a^\xi=1$, the triple gamma reduces to the normal-exponential-gamma which has a representation as a scale-mixture of double exponential $ \mathcal{DE}\left(0, \sqrt{2} \psi _j\right)$-distributions, see Table (ref). It has been considered for variable selection in regression models gri-bro:bay and locally adaptive B-spline models sch-kne:loc. The R2-D2 prior suggested by zha-etal:hig for high-dimensional regression models is another special case of the triple gamma. It reads
where $a=d a^\tau$ and $\sigma^2$ is the residual error variance of the regression model. As shown by zha-etal:hig, this implies following prior for the coefficient of determination: $ R^2 \sim \mathcal{B}\left(a,b\right)$ which motivates holding $a$ fixed, while $a^\tau$ decrease as $d$ increases. Using that $ \phi _j \omega \sim \mathcal{G}\left(a^\tau, \tau\right)$, we can show that the R2-D2 prior is equivalent to following hierarchical normal gamma prior applied in bit-fru:ach for TVP models:
The popular Dirichlet-Laplace prior, $\sqrt \theta_j | \psi _j \sim \mathcal{DE}\left(0, \psi _j\right)$, however, is not related to the triple gamma as the {\em prior scale} $\psi _j$ rather than $\psi^2 _j$ follows a gamma distribution, see again Table (ref).
A challenging question is how to choose the parameters $a^\xi$, $ c^\xi$ and $\kappa_B^2$ or $\phi^{\xi}$ of the triple gamma prior in the context of variance selection for TVP models. In addition, in a TVP context, the shrinkage parameters $a^\tau$, $ c^\tau$ and $\lambda_B^2$ or $\phi^{\tau}=2 c^\tau /( a^\tau \lambda_B^2) $ for the prior ((ref)) of the initial values have to be selected.
In high-dimensional settings it is appealing to have a prior that addresses two major issues: first, high concentration around the origin to favor strong shrinkage of small variances toward zero; second, heavy tails to introduce robustness to large variances and to avoid over-shrinkage. For the triple gamma prior, both issues are addressed through the choice of $a^\xi$ and $ c^\xi$, see Theorem (ref). First of all, we need values $ 0 < a^\xi \leq 0.5$ to induce a pole at 0. Second, values of $0 < c^\xi < 0.5 $ will lead to very heavy tails. For very small values of $a^\xi$ and $ c^\xi$, the triple Gamma is a proper prior that behaves nearly as the improper normal-Jeffrey's prior fig:ada, where $p(\sqrt \theta_j ) \propto 1/\sqrt \theta_j $ and $p(\rho_j ) \propto \rho_j^{-1} (1- \rho_j)^{-1}$.
Ideally, we would place a hyper prior distribution on all shrinkage parameters which would allow us to learn the global and the local degree of sparsity, both for the variances and the initial values. Such a hierarchical triple gamma prior introduces dependence among the local shrinkage parameters $\xi^2_1, \ldots, \xi^2_d$ in ((ref)) and, consequently, among $\theta_1, \ldots, \theta_d$ in the joint (marginal) prior $p(\theta_1, \ldots, \theta_d)$. Introducing such dependence is desirable in that it allows to learn the degree of variance sparsity in TVP models, meaning that how much a variance is shrunken toward zero depends on how close the other variances are to zero. However, first na\"{i}ve approaches with rather uninformative, independent priors on $\kappa_B^2$, $a^\xi$, $ c^\xi$ and $\lambda_B^2$, $a^\tau$, $ c^\tau$ were not met with much success and we found it necessary to carefully design appropriate hyper priors.
Hierarchical versions of the Bayesian Lasso bel-etal:hie_tv and the double gamma prior bit-fru:ach in TVP models are based on the gamma prior $ \kappa_B^2 \sim \mathcal{G}\left(d_1, d_2\right)$. Interestingly, this choice can be seen as a heavy-tailed extension of both priors, where each marginal density $p(\sqrt \theta_j |d_1, d_2)$ follows a triple gamma prior with the same parameter $a^\xi$ (being equal to one for the Bayesian Lasso) and tail index $c^\xi =d_1$. In light of this relationship, it is not surprising that very small values of $d_1$ were applied in these papers to ensure heavy tails of $p(\sqrt \theta_j |d_1, d_2)$. Since a triple gamma prior has already heavy tails, we choose a different hyperprior in the present paper.
For the case $a^\xi=c^\xi =1/2$, the global shrinkage parameter $\tau$ of the Horseshoe prior typically follows a Cauchy prior, $\tau \sim t _{1}$ car-etal:han,bha-etal:hor_reg, see also bha-etal:las. The relationship $\phi^{\xi} = 2/\kappa_B^2= \tau^2 $ between the various global shrinkage parameters (see Table (ref)) implies in this case $\phi^{\xi} \sim \mbox{\rm F}\left(1,1\right)$ or, equivalently, $\kappa^2_B/2 \sim \mbox{\rm F}\left(1,1\right)$.
For a triple gamma prior with arbitrary $a^\xi$ and $c^\xi$, this is a special case of the following prior:
which will be motivated in Section (ref). Under this prior, the triple gamma prior exhibits a BMA-like behavior with a uniform prior on an appropriately defined model size (see Theorem (ref)). Prior ((ref)) is equivalent with following representations:
Concerning $a^\xi $ and $ c^\xi$, we choose the following priors:
Hence, we are restricting the support of $a^\xi $ and $ c^\xi$ to $(0, 0.5)$, following the insights brought to us by Theorem (ref).
We follow a similar strategy for the parameters $a^\tau$, $ c^\tau$ and $\lambda_B^2$ ($\phi^{\tau}$) of the prior ((ref)):
which is equivalent with $\lambda_B^2 | a^\tau \sim \mathcal{G}\left(a^\tau, e_2\right)$, $ e_2 | c^\tau \sim \mathcal{G}\left(c^\tau, 2c^\tau/a^\tau\right)$, and $\phi^{\tau} | a^\tau , c^\tau \sim \mathcal{BP}\left(c^\tau,a^\tau \right)$.
An interesting special case is the \lq\lq symmetric\rq\rq\ triple gamma, where $a^\xi = c^\xi$. Despite this constraint, the favourable shrinkage behaviour is preserved and decreasing $a^\xi=c^\xi$ toward zero simultaneously leads to a high concentration around the origin and a heavy-tailed behaviour. For a symmetric triple gamma prior, $\phi^{\xi}=2/\kappa_B^2$ is independent of $a^\xi$ and $ c^\xi$ and the two global shrinkage parameters are related through $\phi^{\xi}=2/\kappa_B^2$. This induces shrinkage profiles that are symmetric around 1/2, see Section (ref). Interestingly, a symmetric triple gamma resolves the question whether to choose a gamma or an inverse gamma prior for a variance parameter $ \psi^2 _j$. It implies the same symmetric beta-prime distribution on the variance, $ \psi^2 _j \sim \mbox{\rm F}\left(2a^\xi, 2 a^\xi\right)=\mathcal{BP}\left(a^\xi,a^\xi\right)$, and the information, $( \psi^2 _j)^{-1} \sim \mathcal{BP}\left(a^\xi,a^\xi\right) $, and can be represented as a gamma prior with the scale arising from an inverse gamma prior or, equivalently, as an inverse gamma prior with the scale arising from a gamma prior:
In the sparse normal-means problem where $ \bm y|\boldsymbol{\beta} \sim \mathcal{N} _{d}\left(\boldsymbol{\beta}, \sigma^2 {{\mathbf I}}_{d}\right)$ and $\sigma^2=1$, the parameter $\rho_j=1/(1+ \psi^2 _j)$ appearing in ((ref)) is known as shrinkage factor and plays a fundamental role for comparing different shrinkage priors, as $\rho_j$ determines shrinkage toward 0.
Also in a variance selection context, it is evident from ((ref)) that values of $\rho_j\approx 0 $ will introduce no shrinkage on $\theta_j$, whereas values of $\rho_j\approx 1 $ will introduce strong shrinkage of $\theta_j$ toward 0. Hence, the prior $p(\rho_j)$, also called shrinkage profile, will play an instrumental role in the behaviour of different shrinkage priors. Following car-etal:hor, shrinkage priors are often compared in terms of the prior they imply on $\rho_j$, i.e. how they handle shrinkage for small \lq\lq observations\rq\rq\ (in our case innovations) and how robust they are to large \lq\lq observations\rq\rq . Note that we ideally want a shrinkage profile that has a pole in zero (heavy tails to avoid over-shrinking signals) and a pole in one (spikiness to shrink noise). The Horseshoe prior, e.g., implies $\rho_j \sim \mathcal{B}\left(1/2,1/2\right)$ which is a shrinkage profile that takes this much desired form of a \lq\lq Horseshoe\rq\rq , see Figure (ref).
For the triple gamma prior, the shrinkage profile is given by the three-parameter beta prior $p(\rho_j)$ provided in ((ref)). For {$ \phi^{\xi}=1$}, $\rho_j \sim \mathcal{B}\left(c^\xi,a^\xi\right)$ and {$\kappa_B^2=2 c^\xi/ a^\xi$}. Choosing small values $a^\xi<<1$ will put prior mass close to 1, choosing small values $c^\xi<<1$ will put prior mass close to 0, whereas values for both $a^\xi$ and $c^\xi$ smaller than one will induce the form of a Horseshoe prior for $\rho_j$. Evidently, for $ \phi^{\xi}=1$, a symmetric triple gamma prior with $a^\xi=c^\xi $ implies a Horseshoe prior for $\rho_j$ that is symmetric around 0.5. This is illustrated in Figure (ref) for a symmetric triple gamma with $a^\xi=c^\xi=0.1$.
In Figure (ref) we can also see the shrinkage profile for the Bayesian Lasso and the double gamma, which correspond to a triple gamma where $c^\xi \rightarrow \infty$. \footnote{Using ((ref)), we obtain the following prior for $\rho_j=1/(1+ \psi^2 _j)$ by the law of transformation of densities:
} For the Bayesian Lasso with $a^\xi=1$ it is clear that the shrinkage profile $p(\rho_j)$ converges to a constant for $\rho_j \to 1$, while there is no mass around $\rho_j=0$. This means that this prior tends to over-shrink signals, while not shrinking the noise completely to zero. A double gamma prior with $a^\xi < 1$ has the potential to shrink the noise completely to zero, as $p(\rho_j)$ has a pole at $\rho_j = 1$, but $p(\rho_j)$ has also zero mass around $\rho_j=0$, meaning the prior encourages over-shrinking of signals.
When we make $\kappa_B^2$ random, we obtain a \lq\lq prior density\rq\rq\ of shrinkage profiles, see Figure (ref). We can see that such hierarchical versions of the Lasso and the double gamma have shrinkage profiles that resemble the ones of the Horseshoe and the triple gamma. We have used $\kappa_B^2 \sim \mathcal{G}\left(0.01, 0.01\right)$ for the Lasso and the double gamma, $2/\kappa_B^2 \sim \mbox{\rm F}\left(1,1\right)$ for the Horseshoe and $2/\kappa_B^2 \sim \mbox{\rm F}\left(0.2, 0.2\right)$ for the triple gamma, see {Section (ref)}.
From the perspective of Bayesian model averaging (BMA), an ideal approach for handling sparsity in TVP models would be the use of discrete mixture priors as suggested in fru-wag:sto,
with $ \delta_0$ being a Dirac measure at 0, while $p_{\mbox{\tiny slab}}(\sqrt{\theta}_j)$ is the prior for non-zero variances. In terms of shrinkage profiles, the discrete mixture prior ((ref)) has a spike at $\rho_j=1$, with probability $1-\pi$, and a lot of prior mass at $\rho_j=0$, provided that the tails of $p_{\mbox{\tiny slab}}(\sqrt{\theta}_j)$ are heavy enough. The mixture prior (ref) is considered the \lq\lq gold standard\rq\rq\ in BMA, both theoretically and empirically, see e.g. joh-sil:nee. However, MCMC inference under this prior is extremely challenging. As opposed to this, MCMC inference for the triple gamma prior is straightforward, see Section (ref).
In this section, we relate the triple gamma prior to BMA based on the discrete mixture prior ((ref)). An interesting insight is that the triple gamma prior shows a very similar behaviour as a discrete mixture prior, if both $a^\xi$ and $c^\xi$ approach zero. This induces a BMA-type behaviour on the joint shrinkage profile $p(\rho_1, \ldots, \rho_d)$, with a spike at all corner solutions, where some $\rho_j$ are very close to one, whereas the remaining ones very close to zero.
The bivariate shrinkage profiles shown in Figure (ref) give us some intuition about the convergence of a symmetric triple gamma prior with $a^\xi=c^\xi \rightarrow 0$ toward a discrete spike and slab mixture. As opposed to the Lasso and the double gamma prior, the Horseshoe and the triple gamma prior put nearly all prior mass on the \lq\lq corner solutions\rq\rq , which correspond to the four possibilities (a) $\rho_1= \rho_2 =0$, i.e. no shrinkage on $\theta_1$ and $\theta_2$, (b) $\rho_1=1, \rho_2 =0$, i.e. shrinkage of $\theta_1$ toward 0 and no shrinkage on $\theta_2$, (c) $\rho_1=0, \rho_2 =1$, i.e. shrinkage of $\theta_1$ toward 0 and no shrinkage on $\theta_2$, and (d) $\rho_1= \rho_2 =1$, i.e. shrinkage of both $\theta_1$ and $\theta_2$ toward 0.
A very important aspect of BMA is the one of choosing a prior for the model dimension, $K$, see e.g. fer-etal:ben and ley-ste:eff. In the discrete mixture prior ((ref)), the distribution of $K$ depends on the choice of $\pi$. Fixing $\pi$ corresponds to a very informative prior on the model dimension, for example $\pi=0.5$ assigns more prior probability to models of dimension $d/2$ and lower prior probability to empty or full models. In fact, let $\delta_j$ be the indicator that tells us if the $ j$-th coefficient is included in the model, then we have that $K=\sum_{j=1}^d{\delta_j} \sim \mbox{\rm BiNom}\left(d, \pi\right)$. Placing a uniform prior for $\pi$ has been shown to be a good choice, since it corresponds to placing a prior on $K$ which is uniform on $\{0, \ldots, d\}$. Note that $\pi$ will be learned using information from all the variables, that it is a global parameter and will adapt to the degree of sparsity.
Following ideas in car-etal:han, we believe that a natural way to perform variable selection in the continuous shrinkage prior framework is though thresh-holding. Specifically, we say that when $ (1-\rho_j) > 0.5$, or $\rho_j< 0.5$, the variable is included, otherwise it is not. Notice that this classification via thresh-holding makes perfectly sense in the case of a triple gamma of which the Horseshoe is a special case, but less so for a Lasso or double gamma prior, even if the shrinkage profile shows a Horseshoe-like behaviour for hierarchical versions of these priors (see again Figure (ref)). Notice that this implies a prior on the model dimension $K$. Specifically,
where $\rho_j| a^\xi, \phi^\xi \sim \mathcal{TPB}\left(a^\xi, b^\xi, \phi^\xi\right)$, see ((ref)). The choice of $\phi^\xi$ (or $\kappa^2_B$) will strongly impact the prior on $K$. For a symmetric triple gamma with $a^\xi=c^\xi$, for instance, and fixed $\phi^{\xi}=1$, that is $\kappa^2_B=2$, we obtain $K \sim \mbox{\rm BiNom}\left(d, 0.5\right)$, since $\pi^\xi=0.5$ regardless of $a^\xi$. Hence, we have to face similar problems as with fixing $\pi=0.5$ for the discrete mixture prior ((ref)).
Placing a hyper prior on $\phi^{\tau}$ and $\phi^{\xi}$ or, equivalently on, $\lambda_B^2$ and $\kappa^2_B$, as we did in Section (ref), is as instrumental for BMA-type variable and variance selection for the triple gamma prior, as is making $\pi$ random for the discrete mixture prior ((ref)). Ideally, we would like to have a uniform distribution on the model size $K$. We show in Theorem (ref) that the hyperprior for $\kappa_B^2 $ defined in ((ref)) achieves exactly this goal, since $ \pi^\xi$ is uniformly distributed, see Appendix (ref) for a proof.
Let $\bm y=(y_1, \ldots, y_T)$ be the vector of time series observations and let $\bm z$ be the set of all latent variables and unknown model parameters in a TVP model. Moreover, let $\bm z _{-x}$ denote the set of all unknowns but $x$. Bayesian inference based on MCMC sampling from the posterior $p(\bm z|\bm y)$ is summarized in Algorithm (ref). The hierarchical priors introduced in Section (ref) are employed, where $(a^\tau, c^\tau,\lambda_B^2)$ follow ((ref)), $( a^\xi , c^\xi)$ follow ((ref)), and $\kappa_B^2$ follows ((ref)). For certain sampling steps, the hierarchical representation ((ref)) is used for $\kappa_B^2$, and similarly for $\lambda_B^2$.
Algorithm (ref) extends several existing algorithms such as the MCMC schemes introduced for the Horseshoe prior by mak-sch:sim and for the double gamma prior by bit-fru:ach. We exploit various representations of the triple gamma prior given in Lemma (ref) and choose representation ((ref)) as the baseline representation of our MCMC algorithm:
where $\phi^{\tau}=2 c^\tau/(\lambda_B^2 a^\tau)$ and $\phi^{\xi}=2 c^\xi/(\kappa_B^2 a^\xi)$. All conditional distributions in our MCMC scheme are available in closed form, expect the ones for $a^\xi$, $c^\xi$, $a^\tau$ and $c^\tau$, for which we will resort to a MH step within Gibbs. Several conditional distributions are the same as for the double gamma prior and we apply Algorithm 1 of bit-fru:ach. We provide more details on the derivation of the various densities in Appendix (ref).
The MCMC scheme in Algorithm (ref) is not a full conditional scheme, as several steps are based on partially marginalized distributions. That means that the sampling order matters. For instance, in Step (b), we marginalize w.r.t. $\check \xi_1^2, \ldots, \check \xi_d^2$, hence we need to update $\check \xi_1^2, \ldots, \check \xi_d^2$ after sampling $a^\xi$, before we update $c^\xi$ in Step (d) conditional on $\check \xi_1^2, \ldots, \check \xi_d^2$. Similarly, due to marginalization in Step (d), we need to update $\check \kappa_1^2, \ldots, \check \kappa_d^2$, before we update $d_2$ in Step (f). Furthermore, both Step (b) and Step (d) are based on the marginal prior of $\kappa_B^2$, given in ((ref)). Hence, in Step (f), $d_2$ has to be updated from $d_2 | a^\xi, c^\xi, \kappa_B^2 $, before $\kappa_B^2$ is updated conditional on $d_2$.
For a symmetric triple gamma prior, where $a^\xi= c^\xi$, the MCMC scheme in Algorithm (ref) has to be modified only slightly. Either $q_a ( a^\xi)$ in Step (b) is adjusted and Step (d) is skipped, setting $c^\xi=a^\xi$, or $q_c ( c^\xi)$ in Step (d) is adjusted and Step (b) is skipped, setting $a^\xi= c^\xi$. In Appendix (ref), we provide details in ((ref)) for the first case and in ((ref)) for the second case. Similar modifications are needed, if $a^\tau= c^\tau$. All other steps in Algorithm (ref) remain the same for $a^\xi= c^\xi$ and/or $a^\tau= c^\tau$.
Consider an ${{m}}$-dimensional time series, $\bm Y_1, \ldots, \bm Y_T$. The joint dynamics of such time series can be modeled through a time-varying parameter vector autoregressive model with stochastic volatility (TVP-VAR-SV). Since the influential paper of pri:tim (see del-pro:tim for a corrigendum), this model has become a benchmark for analyzing relationships between macroeconomic variables that evolve over time, see nak:tim, koo-kor:lar, eis-etal:sto, cha-eis:bay, fel-etal:sop and car-etal:lar, among many others. A TVP-VAR-SV model of order $p$ can be expressed as follows:
Our algorithm exploits the aforementioned unitriangular decomposition to estimate the model parameters equation-by-equation. {Due to the prior structure introduced in ((ref)),} the estimation of the {$\boldsymbol{\beta}_{t}^i$} and the $a_{ij, t}$s is separated into two blocks, with the algorithm cycling through the equations, alternating between {sampling $\boldsymbol{\beta}_{t}^i$ conditional on $\boldsymbol{\Sigma}_t$ and sampling the $a_{ij, t}$s and $d_{i,t}$s conditional on the VAR coefficients $\boldsymbol{\beta}_{t}^i$}. Given a set of initial values, the algorithm repeats the following steps:
In the following applications, we run our algorithm for $M = 200000$ iterations, discarding the first $100000$ iterations as burn-in, and then keeping the output of one every $100$ iterations.
To illustrate the merit of our methodology in the context of TVP-VAR-SVs, we simulate data from two TVP-VAR-SVs with $T = 200$ points in time, $p = 1$ lags and ${{m}} = 7$ equations, with varying degrees of sparsity. In the dense regime, approximately 30% of the values of $\bm \beta$ and $\bm \theta$ (here referring to the means of the initial states and the variances of the innovations as defined in Section (ref), respectively) are truly zero, while in the sparse regime approximately 90% are truly zero. We show results for the triple gamma prior, the Horseshoe prior, the double gamma and the Lasso.
Regarding the priors on the hyperparameters, {we use prior ((ref)) with $\alpha_{a^\tau}=\alpha_{c^\tau} = \alpha_{a^\xi}= \alpha_{c^\xi}=1$ and $\beta_{a^\tau}=\beta_{c^\tau}=\beta_{a^\xi}=\beta_{c^\xi}=6$ for the triple gamma.} The probability density function of the corresponding beta prior is monotonically increasing, with a maximum at $0.5$. This prior places positive mass in a neighborhood of the Horseshoe, but allows for more flexibility. In practice, placing a prior on the spike and slab parameters of the triple gamma, instead of fixing them to 0.5 as in the Horseshoe, allows us to learn the shrinkage profile from the data. Moreover, since the spike and the slab parameters are allowed to be different, the shrinkage profile can be asymmetric and adapt to the sparseness of the data.
{We assume that the global shrinkage parameters $\lambda ^ {{\beta,2}}_{B,i},\kappa ^ {{\beta,2}}_{B,i}, \lambda ^ {{a,2}}_{B,i}$, and $\kappa ^ {{a,2}}_{B,i}$ follow a $\mbox{\rm F}\left(1,1\right)$ distribution for the Horseshoe prior which corresponds to the prior in car-etal:han and a $ \mathcal{G}\left(0.001, 0.001\right)$ distribution for the Lasso and the double gamma prior, as suggested in bel-etal:hie_tv and bit-fru:ach.} {Concerning the spike parameters $a ^ {\tau ,{a}}_{i}, a ^ {\xi ,{a}}_{i}, a ^ {\tau ,{\beta}}_{i}$, and $a ^ {\xi ,{\beta}}_{i}$ of the double gamma, we employ a rescaled beta prior to force them to be smaller than $0.5$. Specifically, we use a $B (4 , 6)$ prior which places most of its mass between $0.05$ and $0.4$}, a range that bit-fru:ach have found to induce desirable shrinkage characteristics.
Figure (ref) shows the posterior path against time for a constant non-significant parameter, that is one for which $\theta_{ij}^\beta = 0$ and $\beta_{ij}^\beta = 0$ for all times, in the sparse regime. The entire set of states for the triple gamma prior can be found in Appendix (ref). Note that, while the zero line is contained in the 95% posterior credible interval for all priors, said interval is thinner under the triple gamma prior and the double gamma prior than under the Lasso and the Horseshoe prior. However, the light tails of the double gamma prior, as the ones of the Lasso, can over-shrink weak signals.
The above statement becomes clearer when looking at the posterior inclusion probabilities. We calculate the posterior inclusion probabilities based on the thresholding approach introduced in Section (ref), comparing the fully unimpeded triple gamma prior to widely used special cases. Figure (ref) show the posterior inclusion probability for the variance of the innovations ($\theta^\beta_{ij}$'s) under four different shrinkage priors, for the sparse and the dense scenario, respectively. The cells are shaded in gray when the corresponding true state parameter is time-varying ($\theta^\beta_{ij} \neq 0 $), while the background is white when the corresponding true state parameter is not time-varying ($\theta^\beta_{ij} = 0$). The posterior inclusion probabilities under the triple gamma prior are consistently higher for the states for which the parameters are actually time-varying. In some cases, that heavier tails of our prior pick up signals that the other shrinkage priors are not able to capture. On the other hand, the triple gamma identifies constant parameters, effectively pulling the appropriate $\theta^\beta_{ij}$'s to zero.
Our application investigates a subset of the area wide model of the European Union of RePEc:eee:ecmode:v:22:y:2005:i:1:p:39-59, which comprises quarterly macroeconomic data spanning from 1970 to 2017. We include 7 of the variables present in the dataset, namely real output (YER), prices (YED) , short-term interest rate (STN), investment (ITR), consumption (PCR), exchange rate (EEN) and unemployment (URX). A more detailed description of the data and the transformations performed to make the time series stationary can be found in Table (ref) in Appendix (ref). To stay in line with the literature, e.g. feldkircher2017sophisticated, we estimate a TVP-VAR-SV model with $p = 2$ lags on all endogenous variables. The hyperparameter choices are the same as in Section (ref). As in the example with simulated data, we run the algorithm for $M = 200000$ iterations, discarding the first $100000$ iterations as burn-in, and then keeping the output of one every $100$ iterations.
Figures (ref) and (ref) display the posterior inclusion probabilities for the means of the initial states and the innovation variances of the VAR coefficients, respectively. A few things about Figure (ref) are striking. First, the posterior inclusion probabilities on the diagonal, meaning those belonging to the parameter of each equation's own autoregressive term, appear to be those that are the highest, while off diagonal elements are more likely to be excluded. Second, the equation for the short-term interest rate is characterized by a large amount of parameters with a high inclusion probability, across all priors. Third, the first lag tends to have higher posterior inclusion probabilities than the second lag, which is in line with the literature. Finally, the triple gamma prior can be seen to often have either the largest or the smallest posterior inclusion probability compared to the other priors, which can be seen as a reflection of the high amount of prior mass placed near {the shrinkage factors $\rho^\beta_{ij} = 1$ and $\rho^\beta_{ij} = 0$ of $\beta_{ij}^\beta$}, as illustrated in Section (ref). This BMA-like behavior yields a prior that is prone to be more absolute when it comes to inclusion decisions.
Now, we shift our focus to the posterior inclusion probabilities for the $\theta^\beta_{ij}$'s plotted in Figure (ref). Compared to the means of the inital states, almost all inclusion probabilities are essentially zero, with virtually only the triple gamma picking up (faint) signals, in particular with respect to the equations for the financial variables in the model, namely interest rate and nominal exchange rate. This lack of variability is unsurprising, as it is well known (see, e.g., feldkircher2017sophisticated) that stochastic volatility in a TVP-VAR model for macroeconomic variables can explain a large part of the variability in the data. Despite this, the triple gamma, thanks to its heavy tails, is still capable of picking up weak signals in the data that the other shrinkage priors we considered are not able to discern from noise.
Given that the triple gamma tends to include more time variation than the other priors, overfitting might be considered a concern. However, Figures {(ref) and (ref)} put these fears to rest. They display the posterior median of {$\beta ^\beta_{ij}$ and $\left\lvert \sqrt \theta ^\beta_{ij}\right\rvert$}, respectively. Here the triple gamma can be seen to be quite conservative, both in terms of which parameters to include, as well as their magnitude. In particular the medians of the $\left\lvert \sqrt \theta ^\beta_{ij}\right\rvert$ are interesting, as they are closest to zero under the triple gamma prior, despite having the highest posterior inclusion probabilities among all considered priors, pointing towards the triple gamma's ability to pick up even small signals with a higher degree of confidence than other priors.
In Figures (ref) and (ref) in Appendix (ref), all the posterior paths of {$\boldsymbol{\Phi}_{1,t}$ and $\boldsymbol{\Phi}_{2,t}$} under the triple gamma prior are shown.
In the present paper, shrinkage for time-varying parameter (TVP) models was investigated within a Bayesian framework with the goal to automatically reduce time-varying parameters to static ones, if the model is overfitting. This goal was achieved by suggesting the triple gamma prior as a new shrinkage priors for the process variances of varying coefficients, extending previous work using spike-and-slab priors, the Bayesian Lasso, or the double gamma prior. The triple gamma prior is related to the normal-gamma-gamma prior applied for variable selection in highly structured regression models gri-bro:hie. It contains the well-known Horseshoe prior as a special case, however it is more flexible, with two shape parameters that control concentration at zero and the tail behaviour. This leads to a BMA-type behaviour which allows not only variance shrinkage, but also variance selection.
In our application, we considered time-varying parameter VAR models with stochastic volatility. Overall, our findings suggest that the family of triple gamma priors introduced in this paper for sparse TVP models is successful in avoiding overfitting, if coefficients are, indeed, static or even insignificant. The framework developed in this paper is very general and holds the promise to be useful for introducing sparsity in other TVP and state space models in many different settings. Nevertheless, a number of extensions seem to be worth pursuing.
In particular, in ultra-sparse settings, modifications seem sensible. Currently, the hyperprior for the global shrinkage parameter of the triple gamma prior is selected in a way that it implies a uniform prior on \lq\lq model size\rq\rq . A generalization of Theorem (ref) would allow to choose hyper priors that induce higher sparsity. Furthermore, in the variable selection literature, special priors such as the Horseshoe+ bha-etal:hor_est were suggested for very sparse, ultra-high dimensional settings. Exploiting once more the non-centered parametrization of a state space model, it is straightforward to extend this prior to variance selection using following hierarchical representation:
We leave both extensions for future research.
An important limitation of our approach is that shrinking a variance toward zero implies that a coefficient is fixed over the entire observation period of the time-series. In future research we will investigate dynamic shrinkage priors kal-gri:tim,kow-etal:dyn,roc-mca:dyn where coefficients can be both fixed and dynamic.
\authorcontributions{The authors contributed equally to the work.}
\conflictsofinterest{The authors declare no conflict of interest.}
\appendixtitles{no}