EconBase
← Back to paper

Sparse time-varying parameter VECMs with an application to modeling electricity prices

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

84,082 characters · 17 sections · 24 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Sparse time-varying parameter VECMs with an application to modeling electricity prices

\onehalfspacing\thispagestyle{empty}

center[center omitted — 1,123 chars of source]

\doublespacing

Introduction

This paper discusses econometric tools to achieve dynamic model specification for vector error correction models (VECMs) automatically. Our main idea is to start with a suitably flexible and sophisticated specification amongst this model class, and to impose data-driven shrinkage on the parameter space to obtain the simplest adequate nested version. This approach is motivated by our applications, where no clear theoretical guidance is available about how to choose crucial modeling aspects deterministically. Using a very general model guards against underfitting and misspecification, while pushing its parameters towards a simpler sparsified solution avoids overfitting and poor out-of-sample (OOS) predictive performance.

Specifically, we propose a time-varying parameter (TVP) VECM with heteroskedastic errors and apply it to model and forecast European electricity prices. Deregulation and increasingly competitive markets in the power sector have led to a surge of interest in statistical methods for modeling and forecasting electricity demand and price dynamics. Competing approaches include both univariate and multivariate time series models in linear and nonlinear settings (e.g., structural breaks in the conditional means and variances). For an overview of the related literature, see \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{WERON20141030}. Most directly related to our approach, \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{DeVany1999} show that electricity prices in US states are cointegrated, with long-run relationships driven by no-arbitrage conditions. The more recent literature also finds evidence in favor of common dynamics and cointegration between electricity and gas prices for major power exchanges in the European Union \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep[see, among others,][]{Bosco2010,Houllier2012,Bello2013,Marcos2016,gianfreda2019Energy}.

We aim at capturing these empirical features and regularities in energy markets and model them explicitly. This motivates our proposed TVP-VECM. However, estimating VECMs, particularly with TVPs, poses several econometric challenges \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep[see, e.g.,][]{koop2011bayesian}. First, even when it is agreed upon that cointegration should be taken into account, there is often no compelling argument about how a set of (economic) variables are cointegrated (especially in higher dimensions). This complicates introducing reasonable restrictions to identify the long-run behavior of such time series. For electricity prices, this aspect is even more apparent, since there is no clear intuition from economic theory about how to a priori restrict the cointegration space. Second, and relatedly, the cointegration rank is unknown and may be subject to change over time depending on the application. The previous literature often conditions on the rank and then compares measures of model fit ex post \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep[see][]{geweke1996bayesian}. This may be impractical, either due to computational limits, or due to varying the rank requiring additional (possibly ad hoc) identifying restrictions.\footnote{Notable exceptions are \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{jochmann2015regime} and \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{chua2018bayesian}, who use regime-switching models to estimate a time-varying cointegration rank.} Third, due to their flexibility, large TVP models are prone to overfitting. Many papers thus propose to restrict the parameter space relying on hierarchical prior distributions or approximations.\footnote{These three aspects are discussed to a varying extent in, e.g., \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{bunea2012joint,jochmann2013stochastic,eisenstat2016stochastic,huber2019threshold,chakraborty2016bayesian,huber2020bayesian, hauzenberger2020dynamic, huber2018stochastic}.}

Our approach to solving these interrelated issues combines several recent econometric techniques used for large-scale TVP models and reduced rank regressions. We employ continuous global-local priors for pushing the parameter space towards sparsity. \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{pruser2023data} is a related paper with a similar approach to deal with shrinkage in (constant parameter) VECMs. However, as noted by \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{chakraborty2016bayesian}, such priors solely achieve approximate zeroes; the probability of observing exact zeroes is zero. In simple terms, this implies that shrinkage may guide a general model specification towards a simpler one, but it cannot achieve the nested version exactly (in terms of model selection).

As a remedy, we postprocess our posterior via minimizing distinct least absolute shrinkage and selection operator (Lasso)-type loss functions to obtain truly sparse estimates that may feature exact zeroes. The choices of loss functions are due to different implications of varying sparsity patterns across partitions of the parameter space. In particular, we propose to use distinct loss functions for the cointegration matrix (grouped Lasso), the autoregressive parameters (element-wise Lasso), and the covariances (graphical Lasso). A key feature of our framework is that it selects the number of cointegration relationships for each period, limiting the need for imposing ad hoc restrictions a priori. Moreover, sparsifying the coefficients \textit{ex post} alleviates overfitting concerns and reduces parameter estimation uncertainty \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep[see also][]{huber2020inducing}.

We apply our model framework in two different energy-related contexts. The first uses daily data from several European markets jointly. Our approach detects distinct patterns in both dynamic and static interdependencies across markets (i.e., the procedure selects relevant relationships in a multicountry context, see, e.g., feldkircher2022approximate), and the results corroborate previous empirical evidence about the importance of addressing cointegration, nonlinearities and heteroskedasticity. In our second application, we conduct an extensive OOS forecast exercise for hourly electricity prices for Germany. In this case, our approach uncovers and/or excludes intra-day relationships between energy prices within a single country. We consider forecasts for each hour of the following day, and find that multivariate cointegration models with TVPs and heteroskedastic errors provide improvements relative to various simpler benchmarks. In particular, the forecast exercise indicates that our proposed sparsified TVP-VECM yields competitive and in many cases superior forecasts for German hourly electricity prices.

Summarizing, the VECM allows for discriminating between long-run equilibria and short-run adjustment dynamics, which we find to be important for modeling electricity prices. TVPs capture structural breaks in the dynamic relationships of the underlying data. This is particularly useful when addressing complex latent pricing mechanisms and varying importance of variables such as fuel prices that may be subject to change over time. Moreover, our shrink-then-sparsify approach allows for specification search and variable selection in high-dimensional data by imposing exact zeroes in the coefficient matrices. Heteroskedastic errors capture large unanticipated shocks in prices, a crucial feature when interest centers on producing accurate density forecasts \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep[see, e.g.,][]{gianfreda2020large}.

The paper proceeds as follows. Section (ref) presents a flexible TVP-VECM model equipped with global-local shrinkage priors and heteroskedastic errors, while Section (ref) discusses our proposed dynamic sparsification techniques. Section (ref) applies the TVP-VECM to modeling European electricity prices. Section (ref) summarizes and concludes the paper.

Time-varying parameter vector error correction models

We begin by introducing our baseline econometric framework using rather general notation. This reflects the notion that while these methods are developed in light of our applications, they are also applicable in other contexts. Let $\bm{y}_t$ be an $M\times1$ vector of endogenous variables for $t=1,\hdots,T$, and denote the first difference operator by $\Delta$ such that $\Delta \bm{x}_t = \bm{x}_t - \bm{x}_{t-1}$. A general specification of the TVP-VECM is:

equation[equation omitted — 229 chars of source]

Here, $\bm{w}_t = (\bm{y}_{t-1}',\bm{f}_{t-1}')'$, where $\bm{y}_{t-1}$ is the first lag of the endogenous variables and $\bm{f}_{t-1}$ are a set of $q_f$ exogenous factors such that $\bm{w}_t$ is of size $q\times1$ with $q=M+q_f$.

The left-hand side of Eq. ((ref)) is stationary, i.e., integrated of order zero or $I(0)$; this in turn requires the product $\bm{\Pi}_{t}\bm{w}_t$ to be $I(0)$. Assuming unit roots of the endogenous variables in levels, this implies that $\bm{\Pi}_{t}$ is an $M\times q$ matrix of reduced rank. The rank ${r}_t$ reflects the number of linearly independent cointegrating relationships, with ${r}_t<M$. $\bm{A}_{pt}$ refers to an $M\times M$ time-varying coefficient matrix related to the $p$th lag $\Delta \bm{y}_{t-p}$, and $\bm{\gamma}_t$ of size $M\times N$ relates an $N\times1$ vector $\bm{c}_t$ of deterministic terms (such as seasonal dummies, trends, or intercepts) to $\Delta \bm{y}_t$. In our baseline version of the model, we assume a zero mean Gaussian error term $\bm{\epsilon}_t$ with an $M\times M$ time-varying covariance matrix $\bm{\Sigma}_t$.

The cointegration matrix

A more thorough discussion of the matrix $\bm{\Pi}_{t}$ of reduced rank ${r}_t$, governing the cointegration relationships, is in order. In most applications using VECMs, ${r}_t = \bar{r}$ is some fixed (time-invariant) integer with $1\leq\bar{r}\leq(M-1)$. The rank order is commonly motivated either based on economic theory \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep[see, e.g.,][]{giannone2019priors}, determined by calculating marginal likelihoods for a set of possible choices \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep[see, e.g.,][]{geweke1996bayesian}. For large-scale models, these approaches are often computationally prohibitive and rather restrictive.\footnote{Notable exceptions are huber2019threshold and pruser2023data, who use global-local shrinkage priors to estimate cointegration relations in a data-driven manner.} As a solution, we adapt the approach of \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{chakraborty2016bayesian} for the TVP-VECM and estimate ${r}_t$ for each period.

It is convenient to consider a reparameterized version of Eq. ((ref)), where $\bm{\Pi}_t = \bm{\alpha}_t\bm{\beta}'$, see also \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{liu1999parameter}. Here, the short-run adjustment coefficients are collected in $\bm{\alpha}_t$ of dimension $M\times q$, and the long-run relationships are captured by $\bm{\beta}$ which is $q\times q$. Note that we follow \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{yang2018state} and assume the long-run relations to be constant over time.\footnote{Assuming both $\bm{\alpha}_t$ and $\bm{\beta}_t$ to vary over time further complicates achieving identification. First, note that $\bm{\Pi}_t=\bm{\alpha}_t\bm{\beta}_t' = \bm{\alpha}_t\bm{Q}\bm{Q}^{-1}\bm{\beta}_t'$ for any non-singular matrix $\bm{Q}$ which results in the so-called global identification problem. It is common in the literature to use linear normalization schemes such as $\bm{\beta}_t = (\bm{I}_{r_t},\bm{\beta}_t')'$, see also \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{villani2001bayesian} or \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{strachan2003valid}. Second, the local identification problem appears for $\bm{\alpha}_t=\bm{0}$ which implies that $\text{rank}(\bm{\Pi}_t) = 0$ \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep[see][]{kleibergen1994shape,kleibergen1998bayesian,paap2003bayes}.} This amounts to the assumption that long-run fundamental relations do not change over time, since our interest centers on the combined matrix $\bm{\Pi}_t$, where nonlinearities appear through $\bm{\alpha}_t$, which is sometimes also referred to as a loadings matrix.

We refrain from restricting the cointegration space by imposing a deterministic structure on $\bm{\beta}$ \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep[see also][]{strachan2003valid,strachan2004bayesian,villani2006bayesian}. Instead, we follow \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{koop2009efficient} and \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{koop2011bayesian} and use the transformations ${\bm{\tilde\alpha}}_t = \bm{\alpha}_t\bm{\zeta}^{-1}$, $\bm{\tilde\beta} = \bm{\beta}\bm{\zeta}$, with $\bm{\zeta} = (\bm{\tilde\beta}'\bm{\tilde\beta})^{-0.5}$. This allows for employing a linear state-space modeling approach assuming conditional Gaussianity with a cointegration space prior.

Time-varying parameters and shrinkage

Using $\bm{\tilde w}_t=\bm{\tilde\beta}'\bm{w}_t$ and $\bm{z}_t=(\bm{\tilde w}_t',\bm{x}_t')'$, a more compact version of Eq. ((ref)), providing notational simplicity, is given by:

equation[equation omitted — 148 chars of source]

with $\bm{A}_t=(\bm{A}_{1t},\hdots,\bm{A}_{Pt},\bm{\gamma}_t)$, $\bm{B}_t=(\bm{\tilde\alpha}_t,\bm{A}_t)$ and $\bm{x}_t=(\Delta\bm{y}'_{t-1},\hdots,\Delta\bm{y}'_{t-P},\bm{c}_t')'$, where $\bm{A}_t$ is of size $M\times J$ and $\bm{x}_t$ of size $J\times1$ with $J=(MP+N)$. Furthermore, it is convenient to factor $\bm{\Sigma}_t = \bm{L}_t \bm{H}_t \bm{L}_t'$, with a diagonal matrix $\bm{H}_t =$ diag$(\exp(h_{1t}),\hdots,\exp(h_{Mt}))$ and $\bm{L}_t$ denoting the normalized lower Cholesky factor (i.e., a lower triangular matrix with ones on its diagonal). Note that Var$(\bm{L}_t\bm{\eta}_t) =$ Var$(\bm{\epsilon}_t)$; and this triangularization of the multivariate system allows for equation-by-equation estimation \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep[see, e.g.,][]{primiceri2005,carriero2019large, carriero2022corrigendum}.

We select the $i$th row of $\bm{B}_t$, and define $\bm{b}_{it} = \bm{B}_{i\bullet,t}'$ which refers to the parameters of the $i$th equation of the VECM. In addition, we stack all free elements of the matrix $\bm{L}_t$ in a vector $\bm{l}_t$. We then assume a random walk law of motion for these TVPs:

align[align omitted — 284 chars of source]

The state innovation variances, which govern the amount of time variation, are collected in diagonal covariance matrices $\bm{\Theta}_{(b)i}$ and $\bm{\Theta}_{(l)}$. To impose shrinkage, we use the non-centered parameterization of the TVP model \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep[see][for details]{fs_wagner}, which splits the parameters into a constant and time-varying part:

equation*[equation* omitted — 271 chars of source]

for $i = 1,\hdots,M,$ with an analogously transformed state equation for the parameters $\bm{l}_t$. This allows to impose shrinkage on the constant part of the coefficients $\bm{b}_{i0}$, with $j$th element $b_{ij,0}$ and the amount of time variation determined by $\sqrt{\theta}_{(b)ij}$, the $j$th diagonal element of $\sqrt{\bm{\Theta}}_{(b)i}$.

The sparsification methods proposed in this paper (see Section (ref)) may be combined with any desired setup from the class of global-local shrinkage priors, see \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{cadonna2020triple} for a review. The respective parameters, in our case, the constant part of the coefficients and the square root of the state innovation variances, are assumed to follow a Gaussian distribution with zero mean, and a global variance parameter pushes all coefficients strongly towards zero. These global parameters are multiplied with local scalings, which allow to pull prior mass away from zero for specific parameters even in cases where the underlying parameter vector is very sparse. Both of these shrinkage factors are equipped with another prior hierarchy, and shrinkage properties arise from choices about these mixing distributions. From this class of priors, we choose the horseshoe of \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{carvalho2010horseshoe} for its lack of prior tuning parameters and excellent shrinkage properties.

Heteroskedastic errors

To address possible heteroskedasticity, we use stochastic volatility (SV) models for the structural errors. The logarithms of the diagonal elements of $\bm{H}_t$ follow independent AR(1) processes,

equation[equation omitted — 133 chars of source]

where $\mu_{i}$ is the unconditional mean, $\phi_i$ is the persistence parameter and $\varsigma_i$ is the error variance of the log-volatility process. Note that $\bm{\eta}_t = (\eta_{1t},\hdots,\eta_{Mt})'$ in Eq. ((ref)) features Gaussian errors. We also consider specifications where we replace this Gaussian with a t-distribution,

equation*[equation* omitted — 58 chars of source]

which renders the model even more flexible. Note that as the degrees of freedom $\nu_i\rightarrow\infty$, we obtain a Gaussian as the limiting case. In our applied work, we place a prior on the degrees of freedom and estimate them alongside all other parameters.

This completes our baseline model specification. Additional details about priors and the sampling algorithm are provided in Appendix (ref). It is worth noting that our prior choices are mostly standard and result in a fairly straightforward Markov chain Monte Carlo (MCMC) sampling algorithm which provides draws from the joint posterior of all model parameters.

Dynamic sparsification

The model proposed above combined with continuous global-local priors yields posterior draws that are pushed towards approximate sparsity.\footnote{As highlighted by hahncarvalho2015dss, the success of the two-step shrink-then-sparsify approach depends on the shrinkage properties of the prior. They note that the horseshoe prior is well suited for such procedures, which is another reason why we illustrate our proposed framework with this specific choice.} We rely on ex post sparsification of these draws for each point in time to obtain exact sparsity. In cointegration models, particularly with TVPs, it is beneficial to adjust loss functions for specific parts of the parameter space. This is due to subtle implications of these blocks of parameters for the overall model structure.

Designing suitable loss functions

To perform variable selection and obtain sparse coefficient matrices we postprocess $\bm \Pi_t$, $\bm A_t$ and $\bm \Sigma_t$ by minimizing three coefficient-type specific Lasso loss functions. In particular, we rely on methods proposed in friedman2008sparse, hahncarvalho2015dss, chakraborty2016bayesian, bhattacharya2018signal, and bashir2019post, which have been successfully used in a range of multivariate and univariate macroeconomic and finance applications \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep[see][]{puelz2017variable, puelz2019portfolio, huber2020inducing, hauzenberger2020combining}. Given the properties of the reduced rank matrix $\bm \Pi_t$, we propose modifications when compared to sparsification of the autoregressive coefficients, $\bm A_t$, and the covariance matrix $\bm \Sigma_t$.

As a general remark on notation, draws from the non-sparsified posteriors are indicated by a hat (e.g., $\hat{\bm{\Pi}}_t$) and sparse estimates are marked with an asterisk (e.g., $\bm{\Pi}^{\ast}_t$). The key difference between the two is that the non-sparsified posterior draws may have many entries close to but not exactly zero, while the sparsified estimate has exact zeros which are imposed using auxiliary loss functions. The general idea is that these loss functions are designed to reward model fit by minimizing a distance measure between the non-sparse and sparse solutions, while an additional tuning parameter penalizes non-zero parameters. Choosing this tuning parameter --- which, loosely speaking, governs the number of zeroes --- is generally not a trivial task. To avoid extensive pre-estimation tuning procedures, we use the signal adaptive variable selection (SAVS) estimator proposed in \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{bhattacharya2018signal}, which selects an appropriate amount of sparsity automatically.

In a TVP context, dynamic sparsification poses some additional challenges with respect to how sparsity is imposed with respect to $t$:

enumerate[align=left] • Ex post sparsification is commonly applied to point estimators, such as the posterior median or mean. We will deviate from this procedure and solve the respective optimization problem for each draw from the posterior distribution. The procedure of woody2019model is closely related to this approach. They provide a theoretical motivation of conducting uncertainty quantification of sparse posterior estimates. • Our loss functions will be defined in terms of full-data matrices instead of $t$-specific covariates. The latter might be considered as the natural candidate when transforming a TVP model to its static representation \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep[for details, see][]{chan2009efficient, hhko2019}. Commonly, sparsification is applied in standard regression frameworks with constant coefficients, implying that all information over time is considered. Using $t$-by-$t$ draws would result in dependence on a single observation in time $t$.

We illustrate these issues and our solutions in more detail below in the context of sparsifying the cointegration matrix. It is worth noting that these concerns apply to all three blocks of the parameters that we intend to dynamically sparsify.

Sparsifying the cointegration matrix

Our basic approach to sparsifying the cointegration relationships follows \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{chakraborty2016bayesian}. Ex post sparsfying $\bm{\Pi}_t$ is of crucial importance to obtain an estimate for the rank. Using solely a global-local prior, some estimates in $\bm{\Pi}_t$ are pushed towards zero, but they are never exactly zero. In the case of the cointegration matrix, this implies that $\bm{\Pi}_t$ would always be of full rank (i.e., $r = M$ for all $t$). The main goal is thus to minimize the predictive loss between a non-sparsified draw $\hat{\bm{\Pi}}_t$ and a column-sparse solution $\bm{\Pi}_t^{\ast}$.

Our interest centers on dynamic sparsification to obtain a sparse cointegration matrix at each point in time. The loss function is specified in terms of the full-data matrix $\bm{W}$, a $T\times q$ matrix with $\bm{w}'_t$ in the $t$th row:

equation[equation omitted — 220 chars of source]

with $\lVert \bm C \rVert_F$ denoting the Frobenius norm of a matrix $\bm C$, $\lVert \bm c \rVert_2$ the Euclidean norm of a vector $\bm c$ and $\bm \Pi_{\bullet j,t}$ referring to the $j$th column of $\bm \Pi_t$ (i.e., the $j$th row of its transpose). Equation (ref) denotes a grouped Lasso problem with a row (column)-specific penalty $\kappa_{jt}$ and aims at finding a row (column)-sparse solution of $\bm{\Pi}'_t$ ($\bm{\Pi}_t$).\footnote{The Frobenius norm is a common distance measure between subspaces and is given by $\lVert \bm C \rVert_F = \sqrt{\text{tr}(\bm C' \bm C)}$ with $\text{tr}(\bm C)$ denoting the trace of a matrix $\bm C$. For a detailed discussion on properties of the grouped Lasso, see yuan2006model and wang2008note.}

The first part controls the distance between an estimate and its sparse solution (measured by the Frobenius norm), while the second part penalizes non-zero elements in $\bm \Pi_t$ (in terms of column-specific Euclidean norms). We use the grouped Lasso as opposed to an element-wise Lasso to establish a loss function that penalizes the cointegration matrix towards a lower rank structure. It is worth noting that using an element-wise Lasso could yield situations where the penalty introduces spurious cointegration relationships. Summarizing in simple terms, this loss function is designed to obtain an adequate reduced rank estimate of $\bm{\Pi}_t$ that avoids sacrificing model fit relative to the full rank case.

Notice that we rely on the full data matrix $\bm W$ instead of the $t$-specific covariates $\bm w_t'$. This is in line with our state equations for the TVPs, where all information over time is used for filtering and smoothing. In constant parameter regressions, losses would be based on variation explained by $\bm W \bm \Pi'$. Using the static regression framework to perform dynamic sparsification and solely relying the $t$th observation $\bm w_t'$, instead of the full data matrix $\bm W$, would make the penalty highly sensitive to individual observations over time. To illustrate this, we focus on the $j$th columns in both $\bm W_{\bullet j}$ and $\bm \Pi_{\bullet j}$. The norm of $\bm W_{\bullet j}$ is defined by $\lVert \bm W_{\bullet j} \rVert_2 = \sqrt{\sum_{t=1}^T w_{tj}^2}$. When using $t$-by-$t$ estimates independently, the norm of the $t$th observation is given by $\lVert w_{tj} \rVert_2 = \sqrt{w_{tj}^2}$. A simple solution for this issue would be to down-weight the penalty in Eq. ((ref)) by a factor $T$. However, although theoretically in line with the sparsification techniques proposed in hahncarvalho2015dss, this has the disadvantage of being exposed to idiosyncrasies of the $t$th observation.\footnote{This poses the risk of unstable penalties in Eq. ((ref)), which may be particularly problematic with seasonal patterns in the data such as is the case for electricity prices.} Therefore, using $\bm W$ is arguably the most practicable solution, where each $t$-specific estimate is used for the entire sample.

Note that Eq. ((ref)) can be interpreted as minimizing the expected loss \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep[see][]{hahncarvalho2015dss}. Setting $\hat{\bm \Pi}_t = \bm \Pi_t^{(s)}$ with $(s)$ indicating the $s$th MCMC draw (rather than a posterior point estimate), has the attractive feature of allowing for uncertainty quantification about the rank of $\bm \Pi_t$ \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep[see also][]{huber2020inducing,hauzenberger2020combining}.

In fact, the resulting reduced rank cointegration matrix can be used to extract a model-based estimate of the number of cointegration relationships. An estimate of this time-varying rank $r_t$ is obtained using:

equation*[equation* omitted — 84 chars of source]

with $\mathbb{I}(\bullet)$ denoting the usual indicator function and $\mathfrak{s}_{it}$ for $i=1,\hdots,M,$ are the singular values of $\bm W \bm{\Pi}_t^{\ast}{'}$. We follow \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{chakraborty2016bayesian} and define the rank as the number of non-thresholded singular values, with $\varphi$ defined as the largest singular value of the residuals of the full data specification of (ref), which corresponds to the maximum noise level.

To avoid cross-validation for the penalty term \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep[as in][]{hahncarvalho2015dss}, we rely on the SAVS estimator proposed in \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{bhattacharya2018signal} and set the penalty term to $\kappa_{jt} = 1/\lVert \bm{\hat\Pi}_{\bullet j,t} \rVert_2^2$. This yields the following soft threshold estimate:\footnote{To solve the optimization problem in (ref), the SAVS estimator can be interpreted as special case of the coordinate descent algorithm \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep{friedman2007pathwise} by relying on a single iteration to obtain a closed-form solution. bhattacharya2018signal and hauzenberger2020combining both provide evidence that the coordinate descent algorithm already converges after the first pass through.}

equation[equation omitted — 410 chars of source]

As shown by bhattacharya2018signal this choice of $\kappa_{jt}$ has properties similar to the adaptive Lasso proposed by zou2006adaptive. To extend our discussion about specifying the penalty in terms of full data matrices with respect to SAVS, it is worth noting that the only part that changes is how the norm of the data is calculated ($\lVert \bm W_{\bullet j} \rVert_2^2$ instead of $w_{jt}^2$). In this case, the penalty specified in chakraborty2016bayesian, $\kappa_{jt} = 1/(\lVert \hat{\bm \Pi}_{\bullet j,t} \rVert_2^2$), can be used (without correcting for $T$). Note that this is a practicable solution from an applied perspective, since the norm over the full data matrix is more robust than considering a single observation in period $t$.

Sparsifying the autoregressive coefficients

For sparsifying the time-varying autoregressive coefficients, we define a full-data matrix $\bm{X}$ of dimension $T\times J$ with $\bm{x}_t'$ in the $t$th row, an $MJ\times1$ vector $\bm{a}_t = \text{vec}(\bm{A}_t')$ and the corresponding $TM \times MJ$ regressor matrix $\tilde{\bm X} = (\tilde{\bm x}_1', \dots, \tilde{\bm x}_T')'$ with $\tilde{\bm x}_t = (\bm I_M \otimes \bm x_t')$.

We assume a loss function of the form:

equation[equation omitted — 202 chars of source]

with $a_{jt}$ denoting the $j$th element of $\bm{a}_t$ and $\delta_{jt}$ a covariate-specific penalty. It is worth noting that we minimize a standard Lasso-type predictive loss where the first part controls the distance between an estimate and a sparse solution and the second part penalizes non-zero elements in $\bm a_t$, different to the sparsification of the cointegration relationships.

An optimal choice for the penalty is $\delta_{jt} = 1/(|\hat a_{jt}|^2)$. We again rely on the soft threshold estimate implied by SAVS to obtain a sparse draw of $\bm{a}_t$:

equation[equation omitted — 315 chars of source]

As shown by bhattacharya2018signal this choice of $\delta_{jt}$ again has properties similar to the adaptive Lasso proposed by zou2006adaptive.

Sparsifying the covariance matrix

A sparse draw of the covariance matrix can be obtained by relying on methods proposed in friedman2007pathwise and \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{bashir2019post}. This involves using the precision rather than the covariance matrix, as the precision defines the conditional independence structure of the variables. It is worth noting that directly postprocessing estimates of the covariance matrix can result in a rather dense precision matrix and, hence, induce spurious contemporaneous relationships. Sparse estimates of the precision matrix $\bm{\Sigma}_t^{-1}$ are based on the graphical Lasso penalty:

equation[equation omitted — 257 chars of source]

with $\text{tr}(\bm C)$ and $\log \det(\bm C)$ denoting the trace and the log-determinant of a square matrix $\bm C$, {$\lambda_{ij,t}$ an element-specific Lasso penalty and $\sigma_{ij,t}^{-1}$ the $j$th element in the $i$th row of $\bm{\Sigma}_t^{-1}$.} Similar to (ref) and (ref) the parts $\text{tr}\left(\bm{\Sigma}_t^{-1} \bm{\hat\Sigma}_t \right) - \log\det \left(\bm{\Sigma}_t^{-1}\right)$ are measures of fit, while the third term penalizes non-zero elements in the precision matrix. Here, it is worth noting that if $\sigma_{ij,t}^{-1\ast}$ is set to zero, the $i$th and the $j$th endogenous variable in the system do not feature a contemporaneous relationship. {Thus, postprocessing estimates of the precision matrix capture the notion of obtaining a truly sparse set of relationships between elements in $\Delta \bm y_t$.}

Following bashir2019post, the penalty in (ref) is chosen as $\lambda_{ij,t} = 1/|\sigma_{ij,t}^{-1}|^{0.5}$, which constitutes a semi-automatic procedure to circumvent cross-validation. Here, we refrain from showing the exact form of the soft threshold estimates for each element in $\bm \Sigma_t^{-1 \ast}$ and refer to friedman2008sparse instead, who define a set of soft threshold problems, similar to (ref), to solve for a optimal solution for each element in $\bm \Sigma_t^{-1}$ (and thus $\bm \Sigma_t$). We use the coordinate descent algorithm provided by glasso and only iterate once in line with the SAVS estimator.\footnote{Alternatively, one could also regularize the precision matrix by writing $\bm \Sigma^{-1}_t$ as $M$-dimensional set of nodewise regressions by using the triangularization decomposition outlined by \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{meinshausen2006high}. \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{friedman2008sparse} and \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{banerjee2008model} show that this approach constitutes a special case of (ref) with $M$ independent Lasso problems.}

Estimating the model proposed in Section (ref) yields MCMC draws from the posterior distributions of all relevant parameters. These are subsequently postprocessed using the procedures described above in Section (ref). The baseline framework may be applied to any dataset where one suspects the presence cointegration relationships. In the next section, we discuss further specification details in the context of applying this framework to a study of electricity prices.

Modeling European electricity prices

We use our proposed model in two different applications related to European electricity prices. First, we use daily data from several European electricity markets jointly. This serves to illustrate our approach in terms of detecting suitable cointegration relationships and nonlinearities in these relationships between interconnected energy markets in different European countries. Second, we produce forecasts of hourly electricity prices one-day-ahead. Here, we perform an extensive OOS forecast exercise, selecting German hourly electricity prices as the market of interest. We then evaluate our approach against a large set of competing models. This exercise serves to demonstrate that flexibly controlling for cointegration patterns in data (in this case, between hours during the day), in the absence of prior knowledge of such relationships, is beneficial for forecast accuracy.

Dataset

For the first application, we use daily prices (in levels, averaged over the hours of the day) to estimate our model jointly for nine different regional markets: Baltics (BALT), Denmark (DK), Finland (FI), France (FR), Germany (DE), Italy (IT), Norway (NO), Sweden (SE) and Switzerland (CH); i.e., $M=9$. The data are available for the period from January $1$st, $2017$ to December $31$st, $2019$ in EUR per megawatt-hour (MWh). We follow the literature and choose day-ahead prices determined on a specific day for delivery in a certain hour on the following day.

Prices for BALT and the Nordic countries (DK, FI, NO, and SE) are obtained from Nord Pool; the German, Swiss and French hourly auction prices are from the power spot market of the European Energy Exchange (EEX); for the Italian prices, we use the single national prices (PUN) from the Italian system operator Gestore dei Mercati Energetici (GEM). We preprocess the data for daylight saving time changes to exclude the $25$th hour in October and to interpolate the $24$th hour in March. As additional exogenous factors, we consider daily prices for coal and fuel and interpolate missing values for weekends and holidays. In particular, we use the closing settlement prices for coal (LMCYSPT) and one month forward ICE UK natural gas prices (NATBGAS) due to their importance for the dynamic evolution of electricity prices and potential cointegration relationships (i.e., $q = M+2 = 11$). The model also features deterministic seasonal terms (encoded in $\bm{c}_t$) based on the respective day of the week.

In our second application, the forecast comparison, we choose hourly day ahead prices (in levels) for Germany as our main country of interest, and focus on daylight hours ($8$ a.m. until $6$ p.m.) and an average of the night hours \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep[see also][]{raviv2015forecasting}. We use a hold-out period of approximately a year and a half ranging from July $3$rd, 2018 to December $31$st, 2019 (in total $550$ observations). We estimate TVP-VECMs with the individual hours per day being treated as dependent variables such that $M=12$. Our exercise is based on a pseudo OOS simulation using a rolling window of $T=400$ observations at a time.\footnote{{The main reason for the use of a rolling rather than expanding window is to limit the computational burden. A rolling window implies quicker parameter change over the holdout than an expanding one. However, the daily frequency provides enough observations for making more abrupt parameter change --- to possibly be captured by the TVPs --- likely also for the rolling window.}} We consider one-step ahead predictions, which implies that we forecast each individual hour for the following day. All models (both daily and hourly) feature $P=2$ lags.

Nonlinearities and cointegration in European electricity prices

One major feature of the proposed approach is that it remains relatively agnostic on the precise form of the cointegration and parameter space spanned by short-run coefficients. We do not rule out complex patterns a priori, but our flexible modeling approach is also capable of supporting a fairly parsimonious specification when the data suggests so.

In this sub-section, we examine the sparsification patterns and nonlinearities when daily electricity price data from multiple electricity markets are modeled jointly with a sparse TVP-VECM. We first assess whether these electricity prices are indeed cointegrated and whether these cointegration relationships change over time. We roughly gauge the nonlinear features of the cointegration space by assessing estimates for a time-varying cointegration rank. Second, we assess the sparsification patterns and nonlinearities of static and dynamic interdependencies across the European electricity market. To this end, we examine the sparsified estimates of both the short-run adjustment coefficients (dynamic interdependencies) and the contemporaneous relationships (static interdependencies). Third and finally, we investigate heteroskedastic data features and the role of large variance shocks.

Figure (ref) shows the posterior probability of the rank (PPR) over time based on our MCMC output. The probabilities are indicated in various shades of red. Several findings are worth noting. A rank {larger than six is hardly ever supported}, and while most posterior mass is concentrated on $r_t = 4$ for all $t$, we detect subtle differences over time. At the beginning of the sample in $2017$, our estimates are more dispersed, with non-negligible probabilities for no cointegration. The precision of our rank estimate increases over time, with a much narrower corridor of probabilities starting around $2018$. { After a brief period in late $2018$ and early $2019$ with probabilities shifting towards lower ranks, we find increases in cointegration relationships towards the end of the sample. We conjecture that this pattern of the cointegration rank results from the fact that some of the countries (or regions) experience a large (idiosyncratic) positive shock in electricity prices in early 2018, while others do not. For example, electricity prices in the Baltics and in two Nordic countries (Finland and Sweden) increase from around 30 euros to 90 euros. In all other countries, these jumps are either less pronounced (e.g., Denmark and Norway) or almost nonexistent (e.g., Italy) during this episode.}

figure[figure omitted — 193 chars of source]
figure[figure omitted — 699 chars of source]

Next, we turn to sparsified estimates of the autoregressive coefficients and the error covariance matrix in panels (a) and (b) of Figure (ref). Rather than showing magnitudes of the estimated coefficients, we use this exercise to illustrate the sparsification approach. As noted earlier, conventional shrinkage approaches push coefficients towards zero, but they are never exactly zero. The sparsification approach, on the other hand, introduces exact zeros in these matrices. We use this fact to compute the posterior inclusion probabilities (PIPs) of all coefficients by calculating the relative share of zeroes over the iterations of the algorithm. In other words, these figures showcase the sparsification approach as a variable selection tool \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep[see][]{hahncarvalho2015dss}.

We start with the coefficients linked to $\Delta\bm{y}_{t-p}$, that is, the dynamic interdependencies of the multivariate system in panel (a). A few findings are worth noting. First, we detect different degrees of sparsity across countries. While the Nordic and Baltic countries look rather similar (with comparatively dense coefficient matrices), this is not the case for Switzerland, Germany, Italy and France (CH, DE, FR, and IT). For these countries, the model estimates rather sparse coefficient matrices. Second, the variables with the highest PIPs are typically the countries' own lagged series. Particularly France shows extremely sparse estimates, with non-zero inclusion probabilities only for its own lags. Third, we observe several interesting changes in PIPs over time. This implies changes in the importance of predictors over time, a feature which our model detects in a data-based fashion. Fourth, we observe some noteworthy patterns of dynamic interdependencies, namely between Nordic and Baltic countries on one side, and to a lesser degree between continental European economies.

A similar exercise of PIPs in the context of the sparse covariance matrix is displayed in panel (b). Here, we show the lower triangular part of the contemporaneous relationships over time. As in the case of the regression coefficients, we detect differences in the degree of sparsity over the cross-section and across time. Strong contemporaneous relationships are detected especially between the Nordic countries. The PIPs in this case are often exactly one, indicating that the respective coefficients feature non-zero draws across all iterations of the sampler. Similarly, albeit with lower inclusion probabilities, we find that covariances appear to be important between continental European electricity prices.

Interestingly, we observe a substantial degree of time-variation in the inclusion probabilities. Investigating these patterns in more detail, we find that the overall sparsity of the covariance matrix changes strongly over time, but does so in a specific way. In particular, there are periods where most covariance terms (apart from the always featured ones across Nordic countries) are sparsified strongly. Examples for such periods are in late $2017$ and early $2018$, or around the beginning of $2019$. Again, it is noteworthy that our model discovers these features in an automatic fashion.

We complete our discussion of the multi-country application by assessing whether heavy tailed errors are required to capture energy price fluctuations across the countries. For this purpose, we compare our estimates from our TVP-VECM benchmark model with SV to the same specification with t-distributed errors (see Appendix (ref)). We therefore assess the log-volatilities over time and across countries.

figure[figure omitted — 675 chars of source]

Figure (ref) shows our estimates for t-distributed errors in panel (a), while panel (b) indicates the standard SV specification. A few things are worth noting here. First, while the level of the volatilities varies substantially over the cross-section, the volatilities exhibit a substantial degree of co-movement. Second, even though we detect several differences between heavy tailed errors and conventional SVs, the first principal component of all volatility processes (marked by the red line) is almost identical for both error specifications. Third, t-distributed errors result in numerous high-frequency spikes for Denmark, Finland, and Sweden (i.e., the Nordic countries, apart from Norway). Table (ref) in the Appendix displays the parameters of the state equation for the SVs and gives a more rigorous assessment of the dynamics in volatility processes. A crucial parameter in the context of heavy tailed errors is $\nu_{i}$, the degrees of freedom in the t-distribution.

Overall, our results across countries are thus mixed. There is overwhelming empirical evidence for time-varying variances. On the other hand, we find that that t-distributed errors are not crucial for Baltic countries, Switzerland, Germany, France, Italy, and Norway, while the data supports heavier-than-normal tails for the remaining Nordic countries. Our model detects these features automatically, without the need of taking specific a priori decisions about model specification. Next, we discuss whether these features in fact pay off in terms of predictive performance.

Forecast results

In our forecasting comparison, we zoom in on a single country to conduct a thorough analysis featuring many competing models that have been shown to work well in the previous literature. In particular, we select hourly day ahead prices for Germany, and evaluate point and density forecasts for a set of models that are described in detail below.

Competing models

The competing specifications include a large set of univariate and multivariate models with constant parameters and TVPs. Moreover, we consider two variants for capturing heteroskedasticity. Our main interest centers on VECMs that are either sparsified or non-sparsified. For all these VECM specifications, we remain agnostic on the cointegration relations and estimate them from the data. As additional (non-sparsified) multivariate competitors, we consider VARs in levels and differences, while, as simple (non-sparsified) univariate competitors, we use AR($P$) models in levels and differences (with the latter model with constant coefficients serving as an overall benchmark). All these model variants are estimated with constant/time-invariant (TIV) and time-varying (TVP) parameters. With respect to the stochastic disturbances, we account for heteroskedasticity by relying either on a conventional SV specification with Gaussian errors (labeled n in the tables) as in Eq. ((ref)), or on an extension with t-distributed errors (labeled t in the tables) described in Appendix (ref). Recall that all models feature two lags, i.e., two full days worth of hourly lags ($P=2$) and are equipped with a horseshoe prior \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep{carvalho2010horseshoe}. To economize on space, we refrain from showing results for homoskedastic specifications, since these models are typically found to be inferior for forecasting \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep[see, for instance,][]{gianfreda2020large}.

Point and density forecasts

Root mean squared errors (RMSEs) serve to evaluate the point predictions across our models. As a density forecast measure we rely on the continuous ranked probability score \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep[CRPS,][]{gneiting2007strictly}. This measure captures not only the first moment but also higher-order moments of the predictive distribution. Table (ref) displays RMSEs and CRPSs across all model types (rows) and over the respective hours of the day (columns). Forecasts are produced hourly for a full day ahead ($8$ a.m. until $6$ p.m., and aggregate “Night” hours). Besides individual hours, we produce a summary figure for overall forecast performance across a full day (labeled “Total”). Model abbreviations and variants of features such as TVPs or the error variances specification are indicated in the previous sub-section. Values in the first row per model indicate RMSEs, those in parentheses in the second row are CRPSs. They are benchmarked relative (as ratios) to the AR($P$) model in differences with constant parameters and a standard SV specification (red shaded row, indicating raw values of RMSEs and CRPSs). In both cases, relative numbers below one mark superior forecast performance, with the best performing specification in bold.

The upper panel contains results for our main model variant (i.e., a VECM equipped with shrinkage priors) with both sparsified and non-sparsified estimates. The lower panel shows several non-sparsified benchmarks. It is worth mentioning that many of these benchmark specifications are nested in our proposed model. For example, in the case where our approach selects the cointegration matrix to be of full rank, we obtain a VAR in levels. Such specifications in essence differ in terms of the implied shrinkage prior on the reduced form coefficients \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep[see, for instance,][]{villani2009steady,eisenstat2016stochastic,giannone2019priors}.

Although our results indicate the lack of a single best performing specification, overall some interesting patterns emerge across the different model variants. While almost all multivariate models are superior compared to the univariate benchmarks, we obtain mixed evidence on whether one should choose VAR or VECM specifications to forecast electricity prices. Our results also indicate that taking into account cointegration relationships or directly modeling the data in levels is key to obtain accurate forecasts. Even a potentially too high number of cointegration relationships (by considering a VAR in levels) thus does not necessarily adversely affect predictive accuracy, whereas artificially removing all cointegration relationships (i.e., by considering a VAR in differences) results in a deteriorating forecast performance. By computing differences a priori, i.e., naively transforming towards stationarity, rich information useful to improve forecasts may be lost.

Moreover, a flexible modeling approach for the conditional mean (i.e., by considering TVPs) often yields forecast gains. It is worth mentioning that this is also the case for comparatively worse performing model specifications, at least relative to the corresponding model with constant parameters. No clear picture emerges about the necessity of heavy tailed error distributions. This finding is in line with our discussion of insample findings for Germany in Section (ref). Here, we found that when estimating a model with t-distributed errors, the large estimates for the degrees of freedom parameter suggests that a Gaussian error distribution is sufficient. It is worth mentioning, however, that t-distributed errors never substantially hurt predictive performance.

While sparsification techniques do not necessarily improve point forecast accuracy, there is compelling evidence that these techniques are particularly beneficial for density forecast performance. This could be mainly explained by an underlying bias-variance trade-off, as the multiple loss functions involved in our shrink-then-sparsify procedure strongly push coefficients of irrelevant predictors towards exact zeroes. Although this approach significantly reduces estimation uncertainty of the coefficients with the tighter predictive densities yielding improved density forecast performance, it may in turn introduce small biases resulting in a deterioration of point forecast performance.

Following this rather general discussion of our findings, we zoom into point and density forecast performance in more detail. Starting with point forecasts, Table (ref) indicates the non-sparsified TVP-VECM with Gaussian errors as the best performing specification. As suggested earlier, several other specifications, including our proposed model, display similar performance. These models show improvements upon the univariate benchmark model of about $15$ to $20$ percent lower RMSEs. This superior forecast performance is mainly driven by the excellent performance throughout the morning hours of the day. For night hours, on the contrary, improvements for point forecasts relative to the simple univariate benchmark are muted. We conjecture that electricity prices are more synchronized in the morning hours and late afternoon hours, where the demand for electricity (and thus prices) are typically high and thus more volatile, than during other hours of the day or at night. These considerations may provide an explanation why we observe the most pronounced gains in predictive accuracy with our proposed cointegrated approaches (or VAR models in levels) for (synchronised) peak hours, while the accuracy gains for (idiosyncratic) off-peak hours are rather muted.

Turning to density forecasts in terms of CRPSs, we observe several differences compared to point forecast performance orderings. While most models show substantial improvements around $20$ percent lower CRPSs compared to the constant parameter AR($P$) model with conventional SVs, this metric selects our proposed model with Gaussian errors as the best performing specification for the full day (see column “Total”). In particular, flexibly modeling the conditional mean and explicitly taking into account nonlinearities in the data pays off for density forecast performance. While allowing for nonlinear relationships in the conditional mean is beneficial, we observe that additional flexibility for the error variances (by considering t-distributed errors) is not necessarily needed.

landscape\begin{table*}[t] \caption{Forecast performance for point and density forecasts in parentheses relative to the benchmark.} \begin{tiny} \begin{center} \begin{threeparttable} \begin{tabular*}{\linewidth}{@{\extracolsep{\fill}} cclcccccccccccccc} \toprule \multicolumn{1}{l}&\multicolumn{1}{c}{ Coefficients }&\multicolumn{1}{c}{ SV}&\multicolumn{1}{c}&\multicolumn{13}{c}{ 1-day ahead}\tabularnewline \cmidrule{5-17} \multicolumn{1}{l}&\multicolumn{1}{c}&\multicolumn{1}{c}&\multicolumn{1}{c}&\multicolumn{1}{c}{Total}&\multicolumn{1}{c}{$8$ a.m.}&\multicolumn{1}{c}{$9$ a.m.}&\multicolumn{1}{c}{$10$ a.m.}&\multicolumn{1}{c}{$11$ a.m.}&\multicolumn{1}{c}{$12$ noon}&\multicolumn{1}{c}{$1$ p.m.}&\multicolumn{1}{c}{$2$ p.m.}&\multicolumn{1}{c}{$3$ p.m.}&\multicolumn{1}{c}{$4$ p.m.}&\multicolumn{1}{c}{$5$ p.m.}&\multicolumn{1}{c}{$6$ p.m.}&\multicolumn{1}{c}{Night}\tabularnewline \midrule \multicolumn{17}{c}\tabularnewline \multicolumn{17}{c}{ Sparsified VECMs}\tabularnewline \multicolumn{17}{c}\tabularnewline & TVP& t& & 0.83& 0.80& 0.79& 0.79& 0.79& 0.82& 0.83& 0.82& 0.82& 0.84& 0.86& 0.91& 0.91\tabularnewline & & & & (0.82)& (0.71)& (0.72)& (0.78)& (0.79)& (0.82)& (0.85)& (0.85)& (0.84)& (0.84)& (0.85)& (0.92)& (0.94)\tabularnewline \ifx\relax0pt\relax\else \kern-\the\dimexpr0pt\relax \fi \makebox[0pt][l]{ \fboxsep=0pt \colorbox{grey}{ \strut\kern\qrr@dimen@ } } \ifx\relax0pt\relax\else \kern\the\dimexpr0pt\relax \fi \ignorespaces & TVP& n& & 0.82& 0.80& 0.78& 0.78& 0.78& 0.81& 0.83& 0.83& 0.82& 0.84& 0.86& 0.91& 0.91\tabularnewline \ifx\relax0pt\relax\else \kern-\the\dimexpr0pt\relax \fi \makebox[0pt][l]{ \fboxsep=0pt \colorbox{grey}{ \strut\kern\qrr@dimen@ } } \ifx\relax0pt\relax\else \kern\the\dimexpr0pt\relax \fi \ignorespaces & & & & (0.77)& (0.68)& (0.69)& (0.73)& (0.74)& (0.78)& (0.80)& (0.79)& (0.78)& (0.80)& (0.82)& (0.88)& (0.87)\tabularnewline & TIV& t& & 0.82& 0.80& 0.78& 0.78& 0.78& 0.81& 0.83& 0.82& 0.82& 0.85& 0.87& 0.93& 0.91\tabularnewline & & & & (0.82)& (0.70)& (0.70)& (0.77)& (0.80)& (0.85)& (0.86)& (0.87)& (0.86)& (0.88)& (0.88)& (0.85)& (0.87)\tabularnewline \ifx\relax0pt\relax\else \kern-\the\dimexpr0pt\relax \fi \makebox[0pt][l]{ \fboxsep=0pt \colorbox{grey}{ \strut\kern\qrr@dimen@ } } \ifx\relax0pt\relax\else \kern\the\dimexpr0pt\relax \fi \ignorespaces & TIV& n& & 0.81& 0.79& 0.77& 0.76& 0.76& 0.79& 0.82& 0.81& 0.81& 0.84& 0.85& 0.91& 0.89\tabularnewline \ifx\relax0pt\relax\else \kern-\the\dimexpr0pt\relax \fi \makebox[0pt][l]{ \fboxsep=0pt \colorbox{grey}{ \strut\kern\qrr@dimen@ } } \ifx\relax0pt\relax\else \kern\the\dimexpr0pt\relax \fi \ignorespaces & & & & (0.81)& (0.70)& (0.71)& (0.76)& (0.79)& (0.84)& (0.85)& (0.86)& (0.85)& (0.86)& (0.87)& (0.84)& (\textbf{0.87})\tabularnewline \multicolumn{17}{c}\tabularnewline \multicolumn{17}{c}{ Non-sparsified VECMs}\tabularnewline \multicolumn{17}{c}\tabularnewline & TVP& t& & 0.80& 0.76& 0.75& 0.75& 0.76& 0.79& 0.82& \textbf{0.80}& 0.79& 0.82& \textbf{0.84}& 0.90& 0.90\tabularnewline & & & & (0.82)& (0.71)& (0.72)& (0.78)& (0.79)& (0.83)& (0.85)& (0.85)& (0.84)& (0.84)& (0.85)& (0.91)& (0.93)\tabularnewline \ifx\relax0pt\relax\else \kern-\the\dimexpr0pt\relax \fi \makebox[0pt][l]{ \fboxsep=0pt \colorbox{grey}{ \strut\kern\qrr@dimen@ } } \ifx\relax0pt\relax\else \kern\the\dimexpr0pt\relax \fi \ignorespaces & TVP& n& & 0.80& 0.76& 0.75& 0.75& 0.76& 0.79& 0.82& 0.81& 0.80& 0.83& 0.85& 0.90& 0.90\tabularnewline \ifx\relax0pt\relax\else \kern-\the\dimexpr0pt\relax \fi \makebox[0pt][l]{ \fboxsep=0pt \colorbox{grey}{ \strut\kern\qrr@dimen@ } } \ifx\relax0pt\relax\else \kern\the\dimexpr0pt\relax \fi \ignorespaces & & & & (0.82)& (0.75)& (0.75)& (0.78)& (0.78)& (0.82)& (0.82)& (0.83)& (0.83)& (0.84)& (0.86)& (0.90)& (0.92)\tabularnewline & TIV& t& & 0.80& 0.75& 0.75& 0.75& 0.76& 0.79& 0.82& 0.80& 0.79& 0.83& 0.85& 0.91& 0.90\tabularnewline & & & & (0.80)& (0.70)& (0.71)& (0.76)& (0.77)& (0.81)& (0.83)& (0.82)& (0.81)& (0.83)& (0.85)& (0.89)& (0.92)\tabularnewline \ifx\relax0pt\relax\else \kern-\the\dimexpr0pt\relax \fi \makebox[0pt][l]{ \fboxsep=0pt \colorbox{grey}{ \strut\kern\qrr@dimen@ } } \ifx\relax0pt\relax\else \kern\the\dimexpr0pt\relax \fi \ignorespaces & TIV& n& & \textbf{0.79}& 0.75& \textbf{0.74}& \textbf{0.74}& \textbf{0.75}& \textbf{0.78}& 0.82& 0.80& 0.80& 0.82& 0.85& 0.90& 0.88\tabularnewline \ifx\relax0pt\relax\else \kern-\the\dimexpr0pt\relax \fi \makebox[0pt][l]{ \fboxsep=0pt \colorbox{grey}{ \strut\kern\qrr@dimen@ } } \ifx\relax0pt\relax\else \kern\the\dimexpr0pt\relax \fi \ignorespaces & & & & (0.81)& (0.74)& (0.75)& (0.76)& (0.77)& (0.80)& (0.81)& (0.82)& (0.82)& (0.84)& (0.86)& (0.89)& (0.92)\tabularnewline \midrule \multicolumn{17}{c}\tabularnewline \multicolumn{17}{c}{ Non-sparsified VARs, estimated in levels}\tabularnewline \multicolumn{17}{c}\tabularnewline & TVP& t& & 0.81& 0.76& 0.76& 0.77& 0.79& 0.82& 0.83& 0.82& 0.81& 0.84& 0.87& 0.87& 0.90\tabularnewline & & & & (0.82)& (0.75)& (0.76)& (0.77)& (0.78)& (0.81)& (0.81)& (0.82)& (0.82)& (0.84)& (0.87)& (0.92)& (0.93)\tabularnewline \ifx\relax0pt\relax\else \kern-\the\dimexpr0pt\relax \fi \makebox[0pt][l]{ \fboxsep=0pt \colorbox{grey}{ \strut\kern\qrr@dimen@ } } \ifx\relax0pt\relax\else \kern\the\dimexpr0pt\relax \fi \ignorespaces & TVP& n& & 0.82& 0.77& 0.77& 0.78& 0.80& 0.83& 0.84& 0.83& 0.82& 0.85& 0.88& 0.88& 0.90\tabularnewline \ifx\relax0pt\relax\else \kern-\the\dimexpr0pt\relax \fi \makebox[0pt][l]{ \fboxsep=0pt \colorbox{grey}{ \strut\kern\qrr@dimen@ } } \ifx\relax0pt\relax\else \kern\the\dimexpr0pt\relax \fi \ignorespaces & & & & (0.78)& (0.71)& (0.72)& (\textbf{0.73})& (\textbf{0.73})& (\textbf{0.77})& (\textbf{0.77})& (\textbf{0.78})& (0.79)& (0.81)& (0.83)& (0.87)& (0.87)\tabularnewline & TIV& t& & 0.80& \textbf{0.75}& 0.74& 0.76& 0.78& 0.81& \textbf{0.81}& 0.81& \textbf{0.79}& \textbf{0.82}& 0.85& \textbf{0.85}& \textbf{0.87}\tabularnewline & & & & (0.86)& (0.72)& (0.73)& (0.80)& (0.83)& (0.89)& (0.90)& (0.91)& (0.90)& (0.92)& (0.93)& (0.90)& (0.93)\tabularnewline \ifx\relax0pt\relax\else \kern-\the\dimexpr0pt\relax \fi \makebox[0pt][l]{ \fboxsep=0pt \colorbox{grey}{ \strut\kern\qrr@dimen@ } } \ifx\relax0pt\relax\else \kern\the\dimexpr0pt\relax \fi \ignorespaces & TIV& n& & 0.81& 0.76& 0.75& 0.77& 0.79& 0.82& 0.82& 0.82& 0.80& 0.83& 0.86& 0.85& 0.88\tabularnewline \ifx\relax0pt\relax\else \kern-\the\dimexpr0pt\relax \fi \makebox[0pt][l]{ \fboxsep=0pt \colorbox{grey}{ \strut\kern\qrr@dimen@ } } \ifx\relax0pt\relax\else \kern\the\dimexpr0pt\relax \fi \ignorespaces & & & & (0.85)& (0.73)& (0.73)& (0.80)& (0.83)& (0.88)& (0.88)& (0.89)& (0.88)& (0.91)& (0.92)& (0.90)& (0.93)\tabularnewline \multicolumn{17}{c}\tabularnewline \multicolumn{17}{c}{ Non-sparsified VARs, estimated in first differences}\tabularnewline \multicolumn{17}{c}\tabularnewline & TVP& n& & 0.96& 0.94& 0.94& 0.94& 0.95& 0.96& 0.96& 0.97& 0.96& 0.99& 1.00& 0.98& 0.99\tabularnewline & & & & (0.99)& (0.89)& (0.89)& (0.97)& (0.98)& (1.02)& (1.01)& (1.04)& (1.03)& (1.06)& (1.05)& (1.00)& (1.00)\tabularnewline \ifx\relax0pt\relax\else \kern-\the\dimexpr0pt\relax \fi \makebox[0pt][l]{ \fboxsep=0pt \colorbox{grey}{ \strut\kern\qrr@dimen@ } } \ifx\relax0pt\relax\else \kern\the\dimexpr0pt\relax \fi \ignorespaces & TIV& n& & 0.96& 0.94& 0.94& 0.94& 0.95& 0.97& 0.97& 0.97& 0.96& 0.99& 1.00& 0.98& 1.00\tabularnewline \ifx\relax0pt\relax\else \kern-\the\dimexpr0pt\relax \fi \makebox[0pt][l]{ \fboxsep=0pt \colorbox{grey}{ \strut\kern\qrr@dimen@ } } \ifx\relax0pt\relax\else \kern\the\dimexpr0pt\relax \fi \ignorespaces & & & & (1.04)& (0.92)& (0.93)& (1.02)& (1.03)& (1.08)& (1.07)& (1.08)& (1.07)& (1.10)& (1.09)& (1.05)& (1.06)\tabularnewline \multicolumn{17}{c}\tabularnewline \multicolumn{17}{c}{ Non-sparsified AR($P$) models, estimated in levels}\tabularnewline \multicolumn{17}{c}\tabularnewline & TVP& n& & 0.86& 0.84& 0.82& 0.84& 0.86& 0.86& 0.86& 0.87& 0.86& 0.88& 0.90& 0.87& 0.92\tabularnewline & & & & (0.82)& (0.78)& (0.78)& (0.79)& (0.81)& (0.82)& (0.82)& (0.84)& (0.83)& (0.84)& (0.85)& (\textbf{0.83})& (0.89)\tabularnewline \ifx\relax0pt\relax\else \kern-\the\dimexpr0pt\relax \fi \makebox[0pt][l]{ \fboxsep=0pt \colorbox{grey}{ \strut\kern\qrr@dimen@ } } \ifx\relax0pt\relax\else \kern\the\dimexpr0pt\relax \fi \ignorespaces & TIV& n& & 0.91& 0.91& 0.89& 0.91& 0.90& 0.89& 0.91& 0.90& 0.89& 0.92& 0.94& 0.92& 0.97\tabularnewline \ifx\relax0pt\relax\else \kern-\the\dimexpr0pt\relax \fi \makebox[0pt][l]{ \fboxsep=0pt \colorbox{grey}{ \strut\kern\qrr@dimen@ } } \ifx\relax0pt\relax\else \kern\the\dimexpr0pt\relax \fi \ignorespaces & & & & (0.91)& (0.90)& (0.88)& (0.90)& (0.89)& (0.90)& (0.91)& (0.91)& (0.90)& (0.93)& (0.94)& (0.92)& (1.00)\tabularnewline \multicolumn{17}{c}\tabularnewline \multicolumn{17}{c}{ Non-sparsified AR($P$) models, estimated in differences}\tabularnewline \multicolumn{17}{c}\tabularnewline & TVP& n& & 1.00& 1.00& 1.01& 1.00& 1.00& 1.00& 0.99& 1.00& 1.00& 1.01& 0.99& 0.99& 1.00\tabularnewline & & & & (1.00)& (1.00)& (1.01)& (1.00)& (1.00)& (1.00)& (0.99)& (1.00)& (1.00)& (1.00)& (0.99)& (0.99)& (1.01)\tabularnewline \ifx\relax0pt\relax\else \kern-\the\dimexpr0pt\relax \fi \makebox[0pt][l]{ \fboxsep=0pt \colorbox{red}{ \strut\kern\qrr@dimen@ } } \ifx\relax0pt\relax\else \kern\the\dimexpr0pt\relax \fi \ignorespaces & TIV& n& & 0.94& 1.09& 1.03& 0.99& 0.98& 0.94& 0.95& 0.96& 0.96& 0.90& 0.88& 0.85& 0.69\tabularnewline \ifx\relax0pt\relax\else \kern-\the\dimexpr0pt\relax \fi \makebox[0pt][l]{ \fboxsep=0pt \colorbox{red}{ \strut\kern\qrr@dimen@ } } \ifx\relax0pt\relax\else \kern\the\dimexpr0pt\relax \fi \ignorespaces & & & & (0.54)& (0.68)& (0.62)& (0.58)& (0.57)& (0.54)& (0.55)& (0.54)& (0.55)& (0.52)& (0.50)& (0.49)& (0.37)\tabularnewline \bottomrule \end{tabular*} \begin{tablenotes}[para,flushleft] \scriptsize{\textit{Notes}: Point forecasts are evaluated using root mean squared errors (RMSEs), density forecasts are continuous ranked probability scores (CRPSs). The red shaded rows denote the benchmark and displays actual values for RMSEs and CRPSs. All other models are shown as ratios to the benchmark. Relative numbers below one mark superior forecast performance, with the best performing specification in bold.} \end{tablenotes} \end{threeparttable} \end{center} \end{tiny} \end{table*}

As is the case for point forecasts, we observe differences over the hours of the day. Our proposed sparsified TVP-VECMs show superior performance especially early in the morning and late in the afternoon. We conjecture that it is precisely at these hours that a single or a few cointgration relationships are the main driver of the synchronized price series. Since our sparsified TVP-VECM variant sufficiently accounts for nonlinear features and structural breaks in prices and, in addition, results in a relatively parsimonious and regularized cointegration space compared to a TVP-VAR in levels (i.e., the richest specification in terms of cointegration relationships) or non-sparsified TVP-VECMs, we suspect that this feature is the main reason for the excellent forecast performance.

Finally, to gauge the statistical significance of our results, we conduct a more thorough analysis of the robustness of our density forecast metrics over the holdout. Here, we employ the model confidence set (MCS) procedure of \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{hansen2011model}, with a loss function specified in terms of CRPSs (taking into account all the characteristics of predictive densities). Empty cells indicate that the corresponding model is eliminated from the MCS for the respective variable and loss function. The resulting MCS is displayed in Table (ref).

The MCS procedure yields a set of specifications which is a collection of models that contains the best ones with our pre-defined level of $75$ percent confidence. Given the informativeness of the data, this procedure may either select a single best performing specification, or a ranking of several comparable models on our chosen level of $25$ percent significance \bibpunct{(}{)}{;}{a}{,}{;}\natbibcitep[see][]{hansen2011model}. Considering the results of this exercise in Table (ref), we find that our proposed model performs well overall. The TVP-VECMs are always included in the MCS (i.e., among the statistically significant superior model set), and a single TVP-VECM variant appears among the top three ranked models for density forecasts in most hours. It is also worth mentioning that VECMs generally appear to yield superior forecasts to VARs. However, compared to our previous discussion of CRPSs, we observe some differences. The MCS procedure yields a different performance ranking in terms of significance when compared to evaluating solely CRPSs and producing a ranking in absolute terms. We conjecture that this is due to several periods in our holdout sample that affect the end-of-sample metrics shown in parentheses in Table (ref), whereas the MCS procedure is more robust to such idiosyncrasies. In particular, for the sparsified TVP-VECMs we observe high ranks particularly during the afternoon in terms of density forecasts.

landscape\begin{table*}[t] \caption{Model confidence set (MCS) for density forecasts.} \begin{tiny} \begin{center} \begin{threeparttable} \begin{tabular*}{\linewidth}{@{\extracolsep{\fill}} cclcrrrrrrrrrrrrr} \toprule \multicolumn{1}{l}&\multicolumn{1}{c}{ Coefficients }&\multicolumn{1}{c}{ SV}&\multicolumn{1}{c}&\multicolumn{13}{c}{ 1-day ahead}\tabularnewline \cmidrule{5-17} \multicolumn{1}{l}&\multicolumn{1}{c}&\multicolumn{1}{c}&\multicolumn{1}{c}&\multicolumn{1}{c}{Total}&\multicolumn{1}{c}{$8$ a.m.}&\multicolumn{1}{c}{$9$ a.m.}&\multicolumn{1}{c}{$10$ a.m.}&\multicolumn{1}{c}{$11$ a.m.}&\multicolumn{1}{c}{$12$ noon}&\multicolumn{1}{c}{$1$ p.m.}&\multicolumn{1}{c}{$2$ p.m.}&\multicolumn{1}{c}{$3$ p.m.}&\multicolumn{1}{c}{$4$ p.m.}&\multicolumn{1}{c}{$5$ p.m.}&\multicolumn{1}{c}{$6$ p.m.}&\multicolumn{1}{c}{Night}\tabularnewline \midrule \multicolumn{17}{c}\tabularnewline \multicolumn{17}{c}{ Sparsified VECMs}\tabularnewline \multicolumn{17}{c}\tabularnewline & TVP& t& & 6& 12& 11& 7& 6& 6& 4& 6& 4& 4 & 7& 8& 8\tabularnewline \ifx\relax0pt\relax\else \kern-\the\dimexpr0pt\relax \fi \makebox[0pt][l]{ \fboxsep=0pt \colorbox{grey}{ \strut\kern\qrr@dimen@ } } \ifx\relax0pt\relax\else \kern\the\dimexpr0pt\relax \fi \ignorespaces & TVP& n& & 4& 11& 12& 4& 2& 3& 2& 4& 5& 5& 8& 6& 6\tabularnewline & TIV& t& & 7& 10& 9& 5& 5& 5& 3& 5& 6& 6& 6& 7& 9\tabularnewline \ifx\relax0pt\relax\else \kern-\the\dimexpr0pt\relax \fi \makebox[0pt][l]{ \fboxsep=0pt \colorbox{grey}{ \strut\kern\qrr@dimen@ } } \ifx\relax0pt\relax\else \kern\the\dimexpr0pt\relax \fi \ignorespaces & TIV& n& & 2& 9& 10& 1& 1& \textbf{1}& \textbf{1}& \textbf{1}& \textbf{2}& \textbf{2}& \textbf{2}& 5& \textbf{2}\tabularnewline \multicolumn{17}{c}\tabularnewline \multicolumn{17}{c}{ Non-sparsified VECMs}\tabularnewline \multicolumn{17}{c}\tabularnewline & TVP & t& & 5& \textbf{3}& \textbf{3}& 6& 8& 7& 7& 7& 7& 7& 4& 11& 10\tabularnewline \ifx\relax0pt\relax\else \kern-\the\dimexpr0pt\relax \fi \makebox[0pt][l]{ \fboxsep=0pt \colorbox{grey}{ \strut\kern\qrr@dimen@ } } \ifx\relax0pt\relax\else \kern\the\dimexpr0pt\relax \fi \ignorespaces & TVP& n& & 3& \textbf{2}& \textbf{2}& \textbf{3}& \textbf{3}& 4& 5& \textbf{3}& \textbf{3}& \textbf{3}& \textbf{3}& 10& 5\tabularnewline & TIV & t& & 8& 4& 4& 8& 7& 8& 11& 8& 8& 8& 5& 13& 13\tabularnewline \ifx\relax0pt\relax\else \kern-\the\dimexpr0pt\relax \fi \makebox[0pt][l]{ \fboxsep=0pt \colorbox{grey}{ \strut\kern\qrr@dimen@ } } \ifx\relax0pt\relax\else \kern\the\dimexpr0pt\relax \fi \ignorespaces & TIV & n & & 1& \textbf{1}& \textbf{1}& \textbf{2}& 4& \textbf{2}& 6& \textbf{2}& \textbf{1}& \textbf{1}& \textbf{1}& 14& 4\tabularnewline \midrule \multicolumn{17}{c}\tabularnewline \multicolumn{17}{c}{ Non-sparsified VARs, estimated in levels}\tabularnewline \multicolumn{17}{c}\tabularnewline & TVP & t& & 10& 5& 6& 10& 11& 12& 12& 12& 12& 12& 12& 4& 11\tabularnewline \ifx\relax0pt\relax\else \kern-\the\dimexpr0pt\relax \fi \makebox[0pt][l]{ \fboxsep=0pt \colorbox{grey}{ \strut\kern\qrr@dimen@ } } \ifx\relax0pt\relax\else \kern\the\dimexpr0pt\relax \fi \ignorespaces & TVP& n& & 9& 6& 5& 9& 9& 11& 10& 10& 11& 11& 11& \textbf{3}& 12\tabularnewline & TIV& t& & 14& 7& 7& 13& 14& 14& 14& 14& 14& 14& 14& 12& \textbf{3}\tabularnewline \ifx\relax0pt\relax\else \kern-\the\dimexpr0pt\relax \fi \makebox[0pt][l]{ \fboxsep=0pt \colorbox{grey}{ \strut\kern\qrr@dimen@ } } \ifx\relax0pt\relax\else \kern\the\dimexpr0pt\relax \fi \ignorespaces & TIV& n& & 13& 8& 8& 12& 13& 13& 13& 13& 13& 13& 13& \textbf{2}& \textbf{1}\tabularnewline \multicolumn{17}{c}\tabularnewline \multicolumn{17}{c}{ Non-sparsified VARs, estimated in first differences}\tabularnewline \multicolumn{17}{c}\tabularnewline & TVP& n& & & & & & & & & & & & & & \tabularnewline \ifx\relax0pt\relax\else \kern-\the\dimexpr0pt\relax \fi \makebox[0pt][l]{ \fboxsep=0pt \colorbox{grey}{ \strut\kern\qrr@dimen@ } } \ifx\relax0pt\relax\else \kern\the\dimexpr0pt\relax \fi \ignorespaces & TIV& n& & & & & & & & & & & & & & \tabularnewline \multicolumn{17}{c}\tabularnewline \multicolumn{17}{c}{ Non-sparsified AR($P$) models, estimated in levels}\tabularnewline \multicolumn{17}{c}\tabularnewline & TVP& n& & 12& & & 11& 12& 9& 9& 11& 9& 9& 9& \textbf{1}& 7\tabularnewline \ifx\relax0pt\relax\else \kern-\the\dimexpr0pt\relax \fi \makebox[0pt][l]{ \fboxsep=0pt \colorbox{grey}{ \strut\kern\qrr@dimen@ } } \ifx\relax0pt\relax\else \kern\the\dimexpr0pt\relax \fi \ignorespaces & TIV& n& & 11& & & & 10& 10& 8& 9& 10& 10& 10& 9& 14\tabularnewline \multicolumn{17}{c}\tabularnewline \multicolumn{17}{c}{ Non-sparsified AR($P$) models, estimated in differences}\tabularnewline \multicolumn{17}{c}\tabularnewline & TVP& n& & & & & & & & & & & & & & 17\tabularnewline \ifx\relax0pt\relax\else \kern-\the\dimexpr0pt\relax \fi \makebox[0pt][l]{ \fboxsep=0pt \colorbox{grey}{ \strut\kern\qrr@dimen@ } } \ifx\relax0pt\relax\else \kern\the\dimexpr0pt\relax \fi \ignorespaces & TIV& n& & & & & & & & & & & & & & 15\tabularnewline \bottomrule \end{tabular*} \begin{tablenotes}[para,flushleft] \scriptsize{\textit{Notes}: Results for the model confidence set (MCS) procedure of \bibpunct{(}{)}{,}{a}{,}{,}\natbibcitet{hansen2011model} at a $25$ percent significance level. The loss function is specified in terms of continuous ranked probability scores (CRPSs) as a density forecast measure. Empty cells indicate that the model is \textbf{not} part of the MCS. The top three ranked models are marked in bold.} \end{tablenotes} \end{threeparttable} \end{center} \end{tiny} \end{table*}

Concluding remarks

In this paper, we propose a TVP-VECM equipped with shrinkage priors and heteroskedastic errors. We then discuss methods for inducing sparsity. This framework is capable of introducing exact zeroes in cointegration relationships, autoregressive coefficients and the covariance matrix. The main idea is to start with a suitably flexible and sophisticated specification, and to impose data-driven sparsity on the parameter space to obtain the simplest adequate nested version. Moreover, our procedure yields estimates for a time-varying cointegration rank, without the need for introducing prior information on the cointegration relationships.

We estimate our model using daily and hourly day-ahead prices for different European electricity markets. In the empirical section, we illustrate some features of our approach insample using a multicountry dataset, and conduct an extensive OOS forecast comparison for Germany. Regarding the insample analysis, we detect several interesting time-varying patters of sparsity in the autoregressive coefficients and the covariance matrix. Our proposed detects such features of the data automatically. In addition, we find that our approach is competitive when forecasting hourly one-day ahead electricity prices compared to a large set of univariate and multivariate benchmarks.

{\setstretch{0.85} \addcontentsline{toc}{section}{References} }

center[center omitted — 185 chars of source]