Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
44,380 characters · 8 sections · 40 citation commands
A mixture autoregressive model based on Student's $t$--distribution
Different types of mixture models are in widespread use in various fields. Overviews of mixture models can be found, for example, in the monographs of mclachlan2000finite and fruhwirth2006finite. In this paper, we are concerned with mixture autoregressive models that were introduced by le1996modeling and further developed by wong2000mixture,wong2001logistic,wong2001mixture (for further references, see kalliovirta2015gaussian).
In mixture autoregressive models the conditional distribution of the present observation given the past is a mixture distribution where the component distributions are obtained from linear autoregressive models. The specification of a mixture autoregressive model typically requires two choices: choosing a conditional distribution for the component models and choosing a functional form for the mixing weights. In a majority of existing models a Gaussian distribution is assumed whereas, in addition to constants, several different time-varying mixing weights (functions of past observations) have been considered in the literature.
Instead of a Gaussian distribution, wong2009student proposed using Student's $t$\textendash distribution. A major motivation for this comes from the heavier tails of the $t$\textendash distribution which allow the resulting model to better accommodate for the fat tails encountered in many observed time series, especially in economics and finance. In the model suggested by wong2009student, the conditional mean and conditional variance of each component model are the same as in the Gaussian case (a linear function of past observations and a constant, respectively), and what changes is the distribution of the independent and identically distributed error term: instead of a standard normal distribution, a Student's $t$\textendash distribution is used. This is a natural approach to formulate the component models and hence also a mixture autoregressive model based on the $t$\textendash distribution.
In this paper, we also consider a mixture autoregressive model based on Student's $t$\textendash distribution, but our specification differs from that used by wong2009student. Our starting point is the characteristic feature of linear Gaussian autoregressions that stationary distributions (of consecutive observations) as well as conditional distributions are Gaussian. We imitate this feature by using a (multivariate) Student's $t$\textendash distribution and, as a first step, construct a linear autoregression in which both conditional and (low-dimensional) stationary distributions have Student's $t$\textendash distributions. This leads to a model where the conditional mean is as in the Gaussian case (a linear function of past observations) whereas the conditional variance is no longer constant but depends on a quadratic form of past observations. These linear models are then used as component models in our new mixture autoregressive model which we call the StMAR model.
Our StMAR model has some very attractive features. Like the model of wong2009student, it can be useful for modelling time series with regime switching, multimodality, and conditional heteroskedasticity. As the conditional variances of the component models are time-varying, the StMAR model can potentially accommodate for stronger forms of conditional heteroskedasticity than the model of wong2009student. Our formulation also has the theoretical advantage that, for a $p$th order model, the stationary distribution of $p+1$ consecutive observations is fully known and is a mixture of particular Student's $t$\textendash distributions. Moreover, stationarity and ergodicity are simple consequences of the definition of the model and do not require complicated proofs.
Finally, a few notational conventions. All vectors are treated as column vectors and we write $\boldsymbol{x}=(x_{1},\ldots,x_{n})$ for the vector $\boldsymbol{x}$ where the components $x_{i}$ may be either scalars or vectors. The notation $\mathbf{X}\sim n_{d}(\boldsymbol{\mu},\mathbf{\Gamma})$ signifies that the random vector $\mathbf{X}$ has a $d$\textendash dimensional Gaussian distribution with mean $\boldsymbol{\mu}$ and (positive definite) covariance matrix $\mathbf{\Gamma}$. Similarly, by $\mathbf{X}\sim t_{d}(\boldsymbol{\mu},\mathbf{\Gamma},\nu)$ we mean that $\mathbf{X}$ has a $d$\textendash dimensional Student's $t$\textendash distribution with mean $\boldsymbol{\mu}$, (positive definite) covariance matrix $\mathbf{\Gamma}$, and degrees of freedom $\nu$ (assumed to satisfy $\nu>2$); the density function and some properties of the multivariate Student's $t$\textendash distribution employed are given in an Appendix. The notation $\mathbf{1}_{d}$ is used for a $d$\textendash dimensional vector of ones, $\imath_{d}$ signifies the vector $(1,0,\ldots,0)$ of dimension $d$, and the identity matrix of dimension $d$ is denoted by $I_{d}$. The Kronecker product is denoted by $\otimes$, and $vec(A)$ stacks the columns of matrix $A$ on top of one another.
In this section we briefly consider linear $p$th order autoregressions that have multivariate Student's $t$\textendash distributions as their stationary distributions. First, for motivation and to develop notation, consider a linear Gaussian autoregression $z_{t}$ ($t=1,2,\ldots$) generated by
where the error terms $e_{t}$ are independent and identically distributed with a standard normal distribution, and the parameters satisfy $\varphi_{0}\in\mathbb{R}$, $\boldsymbol{\varphi}=(\varphi_{1},\ldots,\varphi_{p})\in\mathbb{S}^{p}$, and $\sigma>0$, where
is the stationarity region of a linear $p$th order autoregression. Denoting $\boldsymbol{z}_{t}=(z_{t},\ldots,z_{t-p+1})$ and $\boldsymbol{z}_{t}^{+}=(z_{t},\boldsymbol{z}_{t-1})$, it is well known that the stationary solution $z_{t}$ to ((ref)) satisfies
where the last relation defines the conditional distribution of $z_{t}$ given $\boldsymbol{z}_{t-1}$ and the quantities $\mathbf{\Gamma}_{p}$, $\gamma_{0}$, $\boldsymbol{\gamma}_{p}$, $\mu$, and $\mathbf{\Gamma}_{p+1}$ are defined via
Two essential properties of linear Gaussian autoregressions are that they have the distributional features in ((ref)) and the representation in ((ref)).
It is not immediately obvious that linear autoregressions based on Student's $t$\textendash distribution with similar properties exist (such models have, however, appeared at least in spanos1994modeling and heracleous2006student). Suppose that for a random vector in $\mathbb{R}^{p+1}$ it holds that $(z,\boldsymbol{z})\sim t_{p+1}(\mu\mathbf{1}_{p+1},\mathbf{\Gamma}_{p+1},\nu)$ where $\nu>2$ (and other notation is as above in ((ref))). Then (for details, see the Appendix) the conditional distribution of $z$ given $\boldsymbol{z}$ is $z\mid\boldsymbol{z}\sim t_{1}(\mu(\boldsymbol{z}),\sigma^{2}(\boldsymbol{z}),\nu+p)$, where
We now state the following theorem (proofs of all theorems are in the Supplementary Material).
Results (i) and (ii) in Theorem 1 are comparable to properties ((ref)) and ((ref)) in the Gaussian case. Part (i) shows that both the stationary and conditional distributions of $z_{t}$ are $t$\textendash distributions, whereas part (ii) clarifies the connection to standard AR($p$) models. In contrast to linear Gaussian autoregressions, in this $t$\textendash distributed case $z_{t}$ is conditionally heteroskedastic and has an `AR($p$)\textendash ARCH($p$)' representation (here ARCH refers to autoregressive conditional heteroskedasticity).
Let $y_{t}$ ($t=1,2,\ldots$) be the real-valued time series of interest, and let $\mathcal{F}_{t-1}$ denote the $\sigma$\textendash algebra generated by $\{y_{t-j},\text{ }j>0\}$. We consider mixture autoregressive models for which the conditional density function of $y_{t}$ given its past, $f(\cdot\mid\mathcal{F}_{t-1})$, is of the form
where the (positive) mixing weights $\alpha_{m,t}$ are $\mathcal{F}_{t-1}$\textendash measurable and satisfy $\sum_{m=1}^{M}\alpha_{m,t}=1$ (for all $t$), and the $f_{m}(\cdot\mid\mathcal{F}_{t-1})$, $m=1,\ldots,M$, describe the conditional densities of $M$ autoregressive component models. Different mixture models are obtained with different specifications of the mixing weights $\alpha_{m,t}$ and the conditional densities $f_{m}(\cdot\mid\mathcal{F}_{t-1})$.
Starting with the specification of the conditional densities $f_{m}(\cdot\mid\mathcal{F}_{t-1})$, a common choice has been to assume the component models to be linear Gaussian autoregressions. For the $m$th component model ($m=1,\ldots,M$), denote the parameters of a $p$th order linear autoregression with $\varphi_{m,0}\in\mathbb{R}$, $\boldsymbol{\varphi}_{m}=(\varphi_{m,1},\ldots,\varphi_{m,p})\in\mathbb{S}^{p}$, and $\sigma_{m}>0$. Also set $\boldsymbol{y}_{t-1}=(y_{t-1},\ldots,y_{t-p})$. In the Gaussian case, the conditional densities in ((ref)) take the form ($m=1,\ldots,M$) \[ f_{m}(y_{t}\mid\mathcal{F}_{t-1})=\frac{1}{\sigma_{m}}\phi\Bigl(\frac{y_{t}-\mu_{m,t}}{\sigma_{m}}\Bigr), \] where $\phi(\cdot)$ signifies the density function of a standard normal random variable, $\mu_{m,t}=\varphi_{m,0}+\boldsymbol{\varphi}_{m}'\boldsymbol{y}_{t-1}$ is the conditional mean function (of component $m$), and $\sigma_{m}^{2}>0$ is the conditional variance (of component $m$), often assumed to be constant. Instead of a Gaussian density, wong2009student consider the case where $f_{m}(\cdot\mid\mathcal{F}_{t-1})$ is the density of Student's $t$\textendash distribution with conditional mean and variance as above, $\mu_{m,t}=\varphi_{m,0}+\boldsymbol{\varphi}_{m}'\boldsymbol{y}_{t-1}$ and a constant $\sigma_{m}^{2}$, respectively.
In this paper, we also consider a mixture autoregressive model based on Student's $t$\textendash distribution, but our formulation differs from that used by wong2009student. In Theorem 1 it was seen that linear autoregressions based on Student's $t$\textendash distribution naturally lead to the conditional distribution $t_{1}(\mu(\cdot),\sigma^{2}(\cdot),\nu+p)$ in ((ref)). Motivated by this, we consider a mixture autoregressive model in which the conditional densities $f_{m}(y_{t}\mid\mathcal{F}_{t-1})$ in ((ref)) are specified as
where the expressions for $\mu_{m,t}=\mu_{m}(\boldsymbol{y}_{t-1})$ and $\sigma_{m,t}^{2}=\sigma_{m}^{2}(\boldsymbol{y}_{t-1})$ are as in ((ref)) except that $\boldsymbol{z}$ is replaced with $\boldsymbol{y}_{t-1}$ and all the quantities therein are defined using the regime specific parameters $\varphi_{m,0}$, $\boldsymbol{\varphi}_{m}$, $\sigma_{m}$, and $\nu_{m}$ (whenever appropriate a subscript $m$ is added to previously defined notation, e.g., $\mu_{m}$ or $\mathbf{\Gamma}_{m,p}$). A key difference to the model of wong2009student is that the conditional variance of component $m$ is not constant but a function of $\boldsymbol{y}_{t-1}$. An explicit expression for the density in ((ref)) can be obtained from the Appendix and is
where $C(\nu)=\frac{\Gamma\left((1+\nu+p)/2\right)}{\left(\pi(\nu+p-2)\right)^{1/2}\Gamma\left((\nu+p)/2\right)}$ (and $\Gamma(\cdot)$ signifies the gamma function).
Now consider the choice of the mixing weights $\alpha_{m,t}$ in ((ref)). The most basic choice is to use constant mixing weights as in wong2000mixture and wong2009student. Several different time-varying mixing weights have also been suggested, see, e.g., wong2001logistic, glasbey2001non, lanne2003modeling, dueker2007contemporaneous, and kalliovirta2015gaussian,kalliovirta2016gaussian.
In this paper, we propose mixing weights that are similar to those used by glasbey2001non and kalliovirta2015gaussian. Specifically, we set
where the $\alpha_{m}\in(0,1)$, $m=1,\ldots,M$, are unknown parameters satisfying $\sum_{m=1}^{M}\alpha_{m}=1$. Note that the Student's $t$ density appearing in ((ref)) corresponds to the stationary distribution in Theorem 1(i): If the $y_{t}$'s were generated by a linear Student's $t$ autoregression described in Section 2 (with a subscript $m$ added to all the notation therein), the stationary distribution of $\boldsymbol{y}_{t-1}$ would be characterized by $t_{p}(\boldsymbol{y}_{t-1};\mu_{m}\mathbf{1}_{p},\mathbf{\Gamma}_{m,p},\nu_{m})$. Our definition of the mixing weights in ((ref)) is different from that used in glasbey2001non and kalliovirta2015gaussian in that these authors employed the $n_{p}(\boldsymbol{y}_{t-1};\mu_{m}\mathbf{1}_{p},\mathbf{\Gamma}_{m,p})$ density (corresponding to the stationary distribution of a linear Gaussian autoregression) instead of the Student's $t$ density $t_{p}(\boldsymbol{y}_{t-1};\mu_{m}\mathbf{1}_{p},\mathbf{\Gamma}_{m,p},\nu_{m})$ we use.
Equations ((ref)), ((ref)), and ((ref)) define a model we call the Student's $t$ mixture autoregressive, or StMAR, model. When the autoregressive order $p$ or the number of mixture components $M$ need to be emphasized we refer to an StMAR($p$,$M$) model. We collect the unknown parameters of an StMAR model in the vector $\boldsymbol{\theta}=(\boldsymbol{\vartheta}_{1},\ldots,\boldsymbol{\vartheta}_{M},\alpha_{1},\ldots,\alpha_{M-1})$ ($(M(p+4)-1)\times1$), where $\boldsymbol{\vartheta}_{m}=(\varphi_{m,0},\mathbf{\boldsymbol{\varphi}}_{m},\sigma_{m}^{2},\nu_{m})$ (with $\boldsymbol{\mathbf{\varphi}}_{m}\in\mathbb{S}^{p}$, $\sigma_{m}^{2}>0$, and $\nu_{m}>2$) contains the parameters of each component model ($m=1,\ldots,M$) and the $\alpha_{m}$'s are the parameters appearing in the mixing weights ((ref)); the parameter $\alpha_{M}$ is not included due to the restriction $\sum_{m=1}^{M}\alpha_{m}=1$.
The StMAR model can also be presented in an alternative (but equivalent) form. To this end, let $P_{t-1}\left(\cdot\right)$ signify the conditional probability of the indicated event given $\mathcal{F}_{t-1}$, and let $\varepsilon_{m,t}$ be a sequence of independent and identically distributed random variables with a $t_{1}(0,1,\nu_{m}+p)$ distribution such that $\varepsilon_{m,t}$ is independent of $\{y_{t-j},\ j>0\}$ ($m=1,\ldots,M$). Furthermore, let $\boldsymbol{s}_{t}=(s_{1,t},\ldots,s_{M,t})$ be a sequence of (unobserved) $M$\textendash dimensional random vectors such that, conditional on $\mathcal{F}_{t-1}$, $\boldsymbol{s}_{t}$ and $\varepsilon_{m,t}$ are independent (for all $m$). The components of $\boldsymbol{s}_{t}$ are such that, for each $t$, exactly one of them takes the value one and others are equal to zero, with conditional probabilities $P_{t-1}(s_{m,t}=1)=\alpha_{m,t}$, $m=1,\ldots,M$. Now $y_{t}$ can be expressed as
where $\sigma_{m,t}$ is as in ((ref)). This formulation suggests that the mixing weights $\alpha_{m,t}$ can be thought of as (conditional) probabilities that determine which one of the $M$ autoregressive components of the mixture generates the observation $y_{t}$.
It turns out that the StMAR model has some very attractive theoretical properties; the carefully chosen conditional densities in ((ref)) and the mixing weights in ((ref)) are crucial in obtaining these properties. The following theorem shows that there exists a choice of initial values $\boldsymbol{y}_{0}$ such that $\boldsymbol{y}_{t}$ is a stationary and ergodic Markov chain. Importantly, an explicit expression for the stationary distribution is also provided.
The stationary distribution of $\boldsymbol{y}_{t}$ is a mixture of $M$ $p$\textendash dimensional $t$\textendash distributions with constant mixing weights $\alpha_{m}$ . Hence, moments of the stationary distribution of order smaller than $\min\left(\nu_{1},\ldots,\nu_{M}\right)$ exist and are finite. As can be seen from the proof of Theorem 2 (in the Supplementary Material), the stationary distribution of the vector $(y_{t},\boldsymbol{y}_{t-1})$ is also a mixture of $M$ $t$\textendash distributions with density of the same form, $\sum_{m=1}^{M}\alpha_{m}t_{p+1}(\mu_{m}\mathbf{1}_{p+1},\mathbf{\Gamma}_{m,p+1},\nu_{m})$. Thus the mean, variance, and first $p$ autocovariances of $y_{t}$ are (here the connection between $\gamma_{m,j}$ and $\mathbf{\Gamma}_{m,p+1}$ is as in ((ref))) \[ \mu\overset{def}{=}E[y_{t}]=\sum_{m=1}^{M}\alpha_{m}\mu_{m},\quad \gamma_{j}\overset{def}{=}Cov[y_{t},y_{t-j}]=\sum_{m=1}^{M}\alpha_{m}\gamma_{m,j}+\sum_{m=1}^{M}\alpha_{m}(\mu_{m}-\mu)^{2},\ j=0,\ldots,p. \] Subvectors of $(y_{t},\boldsymbol{y}_{t-1})$ also have stationary distributions that belong to the same family (but this does not hold for higher dimensional vectors such as $(y_{t+1},y_{t},\boldsymbol{y}_{t-1})$).
The fact that an explicit expression for the stationary (marginal) distribution of the StMAR model is available is not only convenient but also quite exceptional among mixture autoregressive models or other related nonlinear autoregressive models (such as threshold or smooth transition models). Previously, similar results have been obtained by glasbey2001non and kalliovirta2015gaussian in the context of mixture autoregressive models that are of the same form but based on the Gaussian distribution (for a few rather simple first order examples involving other models, see tong2011threshold).
From the definition of the model, the conditional mean and variance of $y_{t}$ are obtained as
Except for the different definition of the mixing weights, the conditional mean is as in the Gaussian mixture autoregressive model of kalliovirta2015gaussian. This is due to the well-known fact that in the multivariate $t$\textendash distribution the conditional mean is of the same linear form as in the multivariate Gaussian distribution. However, unlike in the Gaussian case, the conditional variance of the multivariate $t$\textendash distribution is not constant. Therefore, in ((ref)) we have the time-varying variance component $\sigma_{m,t}^{2}$ which in the models of kalliovirta2015gaussian and wong2009student is constant (in the latter model the mixing weights are also constants). In ((ref)) both the mixing weights $\alpha_{m,t}$ and the variance components $\sigma_{m,t}^{2}$ are functions of $\boldsymbol{y}_{t-1}$, implying that the conditional variance exhibits nonlinear autoregressive conditional heteroskedasticity. Compared to the aforementioned previous models our model may therefore be useful in applications where the data exhibits rather strong conditional heteroskedasticity.
The parameters of an StMAR model can be estimated by the method of maximum likelihood (details of the numerical optimization methods employed and of simulation experiments are available in the Supplementary Material). As the stationary distribution of the StMAR process is known it is even possible to make use of initial values and construct the exact likelihood function and obtain exact maximum likelihood estimates. Assuming the observed data $y_{-p+1},\ldots,y_{0},y_{1},\ldots,y_{T}$ and stationary initial values, the log-likelihood function takes the form
where
An explicit expression for the density appearing in ((ref)) is given in ((ref)), and the notation for $\mu_{m,t}$ and $\sigma_{m,t}^{2}$ is explained after ((ref)). Although not made explicit, $\alpha_{m,t}$, $\mu_{m,t}$, and $\sigma_{m,t}^{2}$, as well as the quantities $\mu_{m}$, $\boldsymbol{\gamma}_{m,p}$, and $\boldsymbol{\Gamma}_{m,p}$, depend on the parameter vector $\boldsymbol{\theta}$.
In ((ref)) it has been assumed that the initial values $\boldsymbol{y}_{0}$ are generated by the stationary distribution. If this assumption seems inappropriate one can condition on initial values and drop the first term on the right hand side of ((ref)). In what follows we assume that estimation is based on this conditional log-likelihood, namely $L_{T}^{(c)}(\boldsymbol{\theta})=T^{-1}\sum_{t=1}^{T}l_{t}(\boldsymbol{\theta})$ which we, for convenience, have also scaled with the sample size. Maximizing $L_{T}^{(c)}(\boldsymbol{\theta})$ with respect to $\boldsymbol{\theta}$ yields the maximum likelihood estimator denoted by $\hat{\boldsymbol{\theta}}_{T}$.
The permissible parameter space of $\boldsymbol{\theta}$, denoted by $\boldsymbol{\Theta}$, needs to be constrained in various ways. The stationarity conditions $\boldsymbol{\varphi}_{m}\in\mathbb{S}^{p}$, the positivity of the variances $\sigma_{m}^{2}$, and the conditions $\nu_{m}>2$ ensuring existence of second moments are all assumed to hold (for $m=1,\ldots,M$). Throughout we assume that the number of mixture components $M$ is known, and this also entails the requirement that the parameters $\alpha_{m}$ ($m=1,\ldots,M$) are strictly positive (and strictly less than unity whenever $M>1$). Further restrictions are required to ensure identification. Denoting the true parameter value by $\boldsymbol{\theta}_{0}$ and assuming stationary initial values, the condition needed is that $l_{t}(\boldsymbol{\theta})=l_{t}(\boldsymbol{\theta}_{0})$ almost surely only if $\boldsymbol{\theta}=\boldsymbol{\theta}_{0}$. An additional assumption needed for this is
From a practical point of view this assumption is not restrictive because what it essentially requires is that the $M$ component models cannot be `relabeled' and the same StMAR model obtained. We summarize the restrictions imposed on the parameter space as follows.
Asymptotic properties of the maximum likelihood estimator can now be established under conventional high-level conditions. Denote $\mathcal{I}(\boldsymbol{\theta}) =E\bigl[\frac{\partial l_{t}(\boldsymbol{\theta})}{\partial\boldsymbol{\theta}}\frac{\partial l_{t}(\boldsymbol{\theta})}{\partial\boldsymbol{\theta}'}\bigr]$ and $\mathcal{J}(\boldsymbol{\theta})=E\bigl[\frac{\partial^{2}l_{t}(\boldsymbol{\theta})}{\partial\boldsymbol{\theta}\partial\boldsymbol{\theta}'}\bigr]$.
Of the conditions in this theorem, (i) states that a central limit theorem holds for the score vector (evaluated at $\boldsymbol{\theta}_{0}$) and that the information matrix is positive definite, (ii) is the information matrix equality, and (iii) ensures the uniform convergence of the Hessian matrix (in some neighbourhood of $\boldsymbol{\theta}_{0}$). These conditions are standard but their verification may be tedious.
Theorem 3 shows that the conventional limiting distribution applies to the maximum likelihood estimator $\hat{\boldsymbol{\theta}}_{T}$ which implies the applicability of standard likelihood-based tests. It is worth noting, however, that here a correct specification of the number of autoregressive components $M$ is required. In particular, if the number of component models is chosen too large then some parameters of the model are not identified and, consequently, the result of Theorem 3 and the validity of the related tests break down. This particularly happens when one tests for the number of component models. Such tests for mixture autoregressive models with Gaussian conditional densities (see ((ref))) are developed by meitz2017testing. The testing problem is highly nonstandard and extending their results to the present case is beyond the scope of this paper.
Instead of formal tests, in our empirical application we use information criteria to infer which model fits the data best. Similar approaches have also been used by wong2009student and others. Note that once the number of regimes is (correctly) chosen, standard likelihood-based inference can be used to choose regime-wise autoregressive orders and to test other hypotheses of interest.
Modeling and forecasting financial market volatility is key to manage risk. In this application we use the realized kernel of barndorff2008designing as a proxy for latent volatility. We obtained daily realized kernel data over the period 3 January 2000 through 20 May 2016 for the S&P 500 index from the Oxford-Man Institute's Realized Library v0.2 heber2009oxford. Figure (ref) shows the in-sample period (Jan 3, 2000--June 3, 2014; 3597 observations) for the S&P 500 realized kernel data ($\textup{RK}_t$), which is nonnegative with a distribution exhibiting substantial skewness and excess kurtosis (sample skewness 14.3, sample kurtosis 380.8). We follow the related literature which frequently use logarithmic realized kernel ($\log(\textup{RK}_t)$), to avoid imposing additional parameter constraints, and to obtain a more symmetric distribution, often taken to be approximately Gaussian. The $\log(\textup{RK}_t)$ data, also shown in Figure (ref), has a sample skewness of 0.5 and kurtosis of 3.5. Visual inspection of the time series plots of the $\textup{RK}_t$ and $\log(\textup{RK}_t)$ data suggests that the two series exhibit changes at least in levels and potentially also in variability. A kernel estimate of the density function of the $\log(\textup{RK}_t)$ series also suggest the potential presence of multiple regimes.
Table (ref) reports estimation results for three selected StMAR models (for further details, see the Supplementary Material). Following wong2001logistic, wong2009student, and li2015hysteretic, we use information criteria for model comparison. For the $\log(\textup{RK}_t)$ data in-sample period the Akaike information criterion (aic) favours the StMAR($4$,$3$) model, the Hannan-Quinn information criterion (hqc) the StMAR($4$,$2$) model, and the Bayesian information criterion (bic) the simpler StMAR($4$,$1$) model. In view of the approximate standard errors in Table (ref), the estimation accuracy appears quite reasonable except for the degrees of freedom parameters. Taking the sum of the autoregressive parameters as a measure of persistence, we find that the estimated persistence for the first regime of the StMAR($4$,$2$) is 0.909 and 0.489 for the second regime, suggesting that persistence is rather strong in the first regime and moderate in the second regime.
Numerous alternative models for volatility proxies have been proposed. We employ Corsi's (2009) \nocite{corsi2009simple} heterogeneous autoregressive (HAR) model as it is arguably the most popular reference model for forecasting proxies such as the realized kernel. We also consider a $p$th-order autoregression as the AR($p$) often performs well in volatility proxy forecasting. The StMAR models are estimated using maximum likelihood, and the reference AR and HAR models by ordinary least squares. We use a fixed scheme, where the parameters of our volatility models are estimated just once using data from Jan 3, 2000--June 3, 2014. These estimates are then used to generate all forecasts. The remaining 496 observations of our sample are used to compare the forecasts from the alternative models. As discussed in kalliovirta2016gaussian, computing multi-step-ahead forecasts for mixture models like the StMAR is rather complicated. For this reason we use computer driven forecasts to predict future volatility: For each out-of-sample date $T$, and for each alternative model, we simulate 500,000 sample paths. Each path is of length 22 (representing one trading month) and conditional on the information available at date $T$. In these simulations unknown parameters are replaced by their estimates. As the simulated paths are for $\log(\textup{RK}_t)$, and our object of interest is $\textup{RK}_t$, an exponential transformation is applied.
We examine daily, weekly (5 day), biweekly (10 day), and monthly (22 day) volatility forecasts generated by the alternative models; for instance, the weekly volatility forecast at date $T$ is the forecast for $\textup{RK}_{T+1}+\cdots+\textup{RK}_{T+5}$ (the 5-day-ahead {\em cumulative} realized kernel). Table (ref) reports the percentage shares of (1, 5, 10, and 22-day) cumulative $\textup{RK}_t$ out-of-sample observations that belong to the 99%, 95%, and 90% one-sided upper prediction intervals based on the distribution of the simulated sample paths; these upper prediction intervals for volatility are related to higher levels of risk in financial markets. Overall, it is seen that the empirical coverage rates of the StMAR based prediction intervals are closer to the nominal levels than the ones obtained with the reference models. By comparison, the accuracy of the prediction intervals obtained with the popular HAR model quickly degrade as the forecast period increases. The StMAR model performs well also when two-sided prediction intervals and point forecast accuracy are considered (for details, see the Supplementary Material).
The authors thank the Academy of Finland for financial support.