Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
88,114 characters · 20 sections · 62 citation commands
Option Pricing with State-dependent Pricing Kernel
Keywords:{ Option Pricing, Realized GARCH, Regime-switching, Variance Risk Premium, Edgeworth Expansion}
JEL Classification:{ G12, G13, C51, C52}
Risk aversion is a fundamental concept in economic theory with uncertainty. It is embedded in the pricing kernel that bridges the physical probability measure, $\mathbb{P}$, with the risk-neutral measures, $\mathbb{Q}$.\footnote{The risk-neutral measure (or equivalent martingale measure) is a probability measure such that each share price is exactly equal to the risk-free discounted expectation of the share price under this measure.} In the classical asset pricing model with a risk-averse investor, see Lucas1978, the pricing kernel is monotonically declining in aggregate wealth. In practice, the pricing kernel is unknown but can be estimated using a variety of econometric methods, see e.g. AitSahaliaLo2000, Jackwerth2000 and RosenbergEngle2002. When an empirical pricing kernel is plotted against the market return, it often has a upward sloping region. This is contrary to the classical model and has become known as a pricing kernel puzzle. The empirical pricing kernel is typically found to be U-shaped when estimated over a long sample period, see e.g. BakshiMadanPanayotov2010 and ChristoffersenHestonJacobs2013.
A popular explanation for the pricing kernel puzzle is the existence of additional risk factors beyond the equity risk premium. Compensation for additional risk factors can influence the projection onto the return dimension and induce the pricing kernel puzzle. A possible risk factor, which is supported by the empirical evidence, is the variance risk premium. The squared VIX is a measure of the squared variation under $\mathbb{Q}$ and it is, on average, larger than empirical measures of return variance under $\mathbb{P}$. When volatility is larger under the risk-neutral measure than under the physical measure it implies a negative variance risk premium, which can explain the observed U-shaped in empirical pricing kernels. An important contribution to this literature was made in ChristoffersenHestonJacobs2013, who proposed a variance-dependent pricing kernel with a variance risk premium in addition to the equity risk premium.
While pricing kernels are typically found to be U-shaped, there is also evidence that their shapes are unstable over time. In fact, there are periods when the U-shaped pricing kernel disappeared, and exhibited hump-shaped or even inverted U-shaped. Examples of this can be seen in ChristoffersenHestonJacobs2013 for their semi-parametric estimates during the year 2004 to 2007. For the same sample period, GrithHardleKratschmer2017 estimated a hump-shaped pricing kernel with DAX 30 index options. A related approach is to estimate the probability weighting function, as in PolkovnichenkoZhao2013. They estimate the probability weighting function to have a regular S-shaped in the same period (2004-2007) using S&P 500 index options, as opposed to an inverse S-shaped, which they estimate for most periods. The latter corresponds to a U-shaped pricing kernel.\footnote{An inverse S-shaped implies that investors tend to overweight low-probability events while they underweight the likelihood of events with high probability. The opposite is true with a regular S-shape.} Similarly, KieselRahe2017 found a slight positive variance risk premium from mid-2004 to mid-2007. So, these studies suggest that the variance risk aversion was relatively low during this period.
In classical asset pricing models, the risk aversion parameter is constant over time, and this greatly simplifies the implementation of these models. If risk aversion is time-varying, but incorrectly assumed to be constant, then this can induce large pricing errors because the misspecified model leads to an incorrect option pricing formula. The assumption of constant risk aversion is contradicted by empirical evidence, and several theoretical models with time-varying risk aversion have been proposed in the literature, see e.g. CampbellCochrane1999, Li2007, Chabi-YoGarciaRenault2008, and BekertEngstromXu2019.
In this paper, we propose a novel state-dependent pricing kernel with a variance risk premium and time-varying risk aversion. First, we introduce a new discrete-time volatility model within the Realized GARCH framework of HansenHuangShek:2012, in which a latent state follows a hidden Markov switching process. The same Markov switching process is used to introduce time-variation in the pricing kernel, and this leads to a pricing kernel with a state-dependent variance risk premium. The framework has a great deal of flexibility in terms of the statistical model for the observed variables under $\mathbb{P}$ and in terms of the plasticity of the pricing kernel. It is, nevertheless, possible to derive the corresponding pricing formula for European options by means of an analytical approximation method. The Markov-switching Realized GARCH model is relatively simple to estimate by quasi maximum likelihood, and the latent states can be inferred from the observed realized volatility measures and returns. The Realized GARCH framework is well suited for this problem because a key component in the model is a volatility-specific shock. This shock is inferred from the difference between the realized volatility measure and its conditional expectation, and it can be incorporated in the pricing kernel to include a variance risk premium.\footnote{A growing literature is exploring ways to utilize realized volatility measures for derivatives pricing, see e.g. CorsiFusariVecchia2013,ChristoffersenFeunouJacobsMeddahi2014,MajewskiBormettiCorsi2015,HuangWangHansen2017,HuangTongWang2019,TongHuang2021.} Estimation that includes the estimation of the parameters in the pricing kernel is more involved because it relies on the observed option prices.\footnote{HansenTong2021 also introduces an option pricing model with time-varying volatility risk aversion. Their framework is different in two important ways. First, they build on the Heston-Nandi GARCH model (HestonNandi2000). Second, they use a score-driven model, see Creal2013, to model time-variation in risk aversion.}
In this paper, we conduct an extensive empirical analysis with a large panel of S&P 500 index option prices based on 30 years of data (from 1990 to 2019). We find the volatility risk aversion to be time-varying. Investors tend to be more risk-averse during periods with high volatility, but there are also periods where investors have a slight appetite for the variance risk, such as during the low-volatility periods: 1993-1995, 2004-2007, and 2014-2017. This is consistent with other studies that report a positive volatility risk premium during these periods. On average, we find the volatility risk premium to be negative, which is consistent with existing literature. In terms of option pricing performance, our empirical results show that the proposed framework outperforms other benchmarks by reducing option pricing errors by 15% or more. These reductions in pricing errors are found both in-sample and out-of-sample.
The successful fusion of a GARCH model with a Markov switching structure is an accomplishment that deserves some commenting. Estimation of GARCH models with Markov switching is typically marred with complications. For instance, the path-dependence problem makes it nearly impossible to evaluate the sample likelihood function. The solution by Cai1994 and HamiltonSusmel1994 is only valid within an ARCH specification. Other approaches, such as those by Gray1996, and Klaassen2002, are not suitable for asset pricing due to the difficulties in risk neutralization.\footnote{Most option pricing models under the specification of Gray1996 are constructed without specifying risk premium, see e.g. SatoyoshiMitsui2011 and DaoukGuo2004.} ChenHung2010 directly assume a Markov-switching GARCH process for returns under the risk-neutral measure and obtain option prices by a lattice method, volatility discretization, and Monte Carlo simulation. However, their model is not amenable to estimation and the authors resort to calibration instead. ElliottSiuChan2006 develop a method for a Heston-Nandi GARCH model with Markov switching, but it requires the Markov states to be observable. For this reason, there is not much empirical literature on option pricing based on a GARCH model with a hidden Markov switching structure.\footnote{There are several Markov switching models for option pricing without GARCH structures, see e.g. Duan2002, DonaldDasMotwani2006, LiewSiu2010, and ShenFanSiu2014.}
The key to our successful coupling of a hidden Markov switching model with a GARCH-type model is the presence of the realized measures of volatility in the model. By providing accurate information about the contemporaneous volatility level, the realized measure adds valuable information about the latent state, and the Realized GARCH model maintains the structure that permits straightforward estimation based on returns and realized measures. The most complicated part of the framework relates to the analytical approximation for option pricing, because it requires many terms to be computed. The expressions for these terms are computed in the Appendix and are plugged into the generic expressions derived in DuanGauthierSimonato1999.
The rest of the paper is organized as follows. In Section 2, we propose the theoretical model under the physical measure, $\mathbb{P}$, the risk neutralization process under a state-dependent pricing kernel, and details about the estimation of the Markov switching Realized GARCH model. Section 3 introduces the analytical approximation formula for European call options. Section 4 presents the competing models used in our empirical comparisons. Section 5 describes the joint estimation method. All empirical results are presented in Section 6 including the summary statistics, parameters estimation, and in-sample and out-of-sample option pricing performance. All relevant proofs are presented in Appendix A, and Appendix B derives the terms needed for option pricing.
Let $\mathcal{F}_{t}=\ensuremath{\sigma}(\{R_{\tau},x_{\tau}\},\tau\leq t)$ denote the natural filtration, where $R_{t+1}=\log\left(S_{t+1}/S_{t}\right)$ is the daily log-return, and $x_{t}$ is the realized measure of volatility. The latter is, in our empirical analysis, computed from high-frequency data with the Realized Kernel by BNHLS:2008.
The Realized GARCH framework was introduced by HansenHuangShek:2012. In this paper, we build on the variant proposed in HansenHuang:2016 that was also used for option pricing in HuangWangHansen2017. The dynamic properties of returns, the conditional variance, $h_{t+1}=\mathrm{var}(R_{t+1}|\mathcal{F}_{t})$, and the realized volatility measure are given by the equations:
The intercept in the return equation, $r$, denotes the risk-free rate and $\lambda$ is the equity risk premium. The stochastic properties are driven by two i.i.d standard normally distributed “innovations”, $z_{t}$ and $u_{t}$, that represent return and volatility shocks, respectively.
Realized GARCH models are characterized by the measurement equation, ((ref)), that defines the relationship between the conditional variance, $h_{t}$, and the realized measure $x_{t}$. Unlike the conventional GARCH model, this model has two distinct innovations, $z_{t}$ and $u_{t}$, where the second is similar to the random innovation for volatility in stochastic volatility models. However, because the Realized GARCH model is an observation-driven model, it is simpler to estimate than stochastic volatility models. For the purpose of derivative pricing, the Realized GARCH structure is especially valuable because the two innovation terms can be used as risk factors, as we explore below.
Here, we introduce a new volatility model based on the Realized GARCH model above. Specifically, we will introduce time-variation in the intercept of the measurement equation, $\xi$, by assuming it is driven by a hidden Markov process, $\{s_{t}\}$. One may think of the states as representing different market conditions. We refer to this model as the Markov-switching Realized GARCH model. We assume that the state variable, $s_{t}\in\{e_{1},\ldots,e_{N}\}$, follows a stationary discrete-time hidden Markov chain process on $(\Omega,\mathcal{G},\mathbb{P})$, with transition probabilities, $\pi_{ij}\equiv\Pr(s_{t+1}=e_{j}|s_{t}=e_{i})$, $i,j=1,\ldots,N$. So, the Markov-switching Realized GARCH model is given by ((ref)), ((ref)), and a state-dependent measurement equation:
The conventional Realized GARCH model corresponds to the case with a single state ($N=1$). We can, without loss of generality, represent the states by the $N$ unit vectors, i.e. let $e_{j}$ be the $j$-th column of the $N\times N$ identity matrix $I_{N}$. With this convention we have $\xi_{s_{t}}=\xi^{\prime}s_{t}$, where $\xi\in\mathbb{R}^{N\times1}$ is a vector of parameters with $\xi_{j}\equiv\xi_{(s_{t}=e_{j})}$, $j=1,\ldots,N$. Note that the parameter $\xi_{j}$ controls the long-run volatility level within each state because the dynamic properties of $\log h_{t}$ are given by: \[ \log h_{t+1}=\omega+\gamma\xi^{\prime}s_{t}+(\beta+\gamma\phi)\log h_{t}+(\tau_{1}+\gamma\delta_{1})z_{t}+(\tau_{2}+\gamma\delta_{2})(z_{t}^{2}-1)+\gamma\sigma u_{t}. \] Thus, if $|\beta+\gamma\phi|<1$, then $\log h_{t}$ is mean-reverting towards $\left(\omega+\gamma\xi_{j}\right)/(1-\beta-\gamma\phi)$ in the $j$-th state, $j=1,\ldots,N$. Thus, the Markov switching model for $\xi$ is effectively a Markov switching model for the long-run level of volatility. This highlights the motivations for introducing time-variation in this parameter. Another advantage of mapping states to the measurement equation is that inference about states can be simply inferred from the realized measures and returns. Introducing Markov switching in other parameters would greatly complicate the analysis and, in some cases, make analytical option pricing formulae unobtainable.
Next, we turn to the risk neutralization as characterized by the pricing kernel. For this purpose, we generalized the exponentially affine pricing kernel \[ M_{t+1,t}=\frac{\exp[\psi z_{t+1}+\chi u_{t+1}]}{\mathbb{E}_{t}^{\mathbb{P}}[\exp(\psi z_{t+1}+\chi u_{t+1})]}, \] to be state-dependent. Here $\mathbb{E}_{t}^{\mathbb{P}}(\cdot)\equiv\mathbb{E}^{\mathbb{P}}(\cdot|\mathcal{F}_{t})$. This kernel was introduced within the Realized GARCH framework by HuangWangHansen2017, where the parameters, $\psi$ and $\chi$, govern the equity risk premium and the variance risk premium, respectively. We generalize this pricing kernel by substituting $\psi_{s_{t}}$ for $\psi$ and $\chi_{s_{t}}$ for $\chi$. However, it immediately follows that the former must be constant, because a no-arbitrage condition has the implication that $\psi_{s_{t}}=-\lambda$, which rules out time-variation in $\psi$, see Appendix (ref). Thus, the time-variation across states is confined to that in $\chi_{s_{t}}$, and our state-dependent pricing kernel is given by:
Note that the additional conditioning on $s_{t+1}$ is required due to the additional source of uncertainty induced by the hidden regime-switching part of the model. The appropriate martingale condition is, therefore, given by the enlarged filtration: $\mathcal{G}_{t+1}\vee\mathcal{F}_{t}$, where $\mathcal{G}_{t}=\ensuremath{\sigma}(\{s_{\tau}\},\tau\leq t).$ Naturally, by the law of iterated expectations it follows that the martingale condition holds with $\mathcal{F}_{t}$ alone, if it holds for the enlarged filtration. And $\mathcal{F}_{t}$ conveniently does not require knowledge about the latent state. The dynamic properties under the risk-neutral measure, $\mathbb{Q}$, are as stated in the following theorem.
We seek to characterize the implications of a state-dependent variance risk premium parameter, $\chi_{s_{t}}$. To this end, we consider the difference between conditional expectations of log-variance under $\mathbb{P}$ and under $\mathbb{Q}$,
which is a logarithmic variant of the variance risk premium. This quantity can be decomposed into compensation for equity risk and additional compensation for volatility risk. The former is given by the first two terms in ((ref)), which involve $\lambda$, and the latter is the last term in ((ref)), which is state-dependent. Thus, the parameter $\chi_{s_t}$ governs the variance risk aversion of investors in state $s_{t}$, where a large value of $\chi$ is associated with a high variance risk premium.
Note that the expectation $\mathbb{E}^{\mathbb{P}}_t(\chi_{s_{t+1}})$ can be also expressed as $\sum_{s_{t+1}}P_{t}(s_{t+1})\chi_{s_{t+1}}$, where $P_{t}(s_{t+1})$ is the notation for the conditional distribution $\Pr(s_{t+1}|\mathcal{F}_{t})$, over the possible states, $e_{1},\ldots,e_{N}$. Similarly, we use $P_{t}(s_{t})$ as compact notation for $\Pr(s_{t}|\mathcal{F}_{t})$. Next, we will derive model-based expressions for multi-step ahead forecasts of $h_{t}$, which defines the, so-called, variance term structure.
Corollary (ref) shows that the variance term structure, $\mathbb{E}_{t}^{\mathbb{Q}}(h_{t+n}|s_{t})$, for $n=1,2,\ldots$, depends on the current state $s_{t}$ and the current level of volatility, $h_{t+1}$, where the latter is $\mathcal{F}_{t}$-measurable. A major benefit of having a closed-from expression for $\mathbb{E}_{t}^{\mathbb{Q}}(h_{t+n}|s_{t})$ is that it leads to an analytical expression for the model-implied VIX price. In the present context, the VIX pricing formula is given by \[ {\rm{VIX}}_t = A\times\sqrt{\sum_{n=1}^{22}\mathbb{E}_{t}^{\mathbb{Q}}(h_{t+n})}=A\times\sqrt{\sum_{n=1}^{22}\sum_{j=1}^{N}\mathbb{E}_{t}^{\mathbb{Q}}(h_{t+n}|s_{t})P_{t}(s_{t}=e_{j})}, \] where $A=100\sqrt{252/22}$ is the annualized factor.
In this section, we discuss the estimation of the Markov-switching Realized GARCH model based on returns and realized measures alone. In Section (ref), we will discuss how option prices can be incorporated in the estimation of this model in conjunction with the parameters in the pricing kernel.
The model is relatively simple to estimate by maximum likelihood for $\{(R_{t},x_{t})\}_{t=1}^{T}$. The likelihood for $(R_{t+1},x_{t+1})$ is given by their density, conditional on $\mathcal{F}_{t}$, $L_{t}(R_{t+1},x_{t+1})=f(R_{t+1},x_{t+1}|\mathcal{F}_{t})$, which we can rewrite as
The last equality uses that the distribution of $R_{t+1}=r+\lambda\sqrt{h_{t+1}}-\tfrac{1}{2}h_{t+1}+\sqrt{h_{t+1}}z_{t+1},$ conditional on $\mathcal{F}_{t}$, does not depend on the state, $s_{t+1}$. The reason is that $h_{t+1}$ is $\mathcal{F}_{t}$-measurable and $z_{t+1}$ is independent of $s_{t+1}$, such that $L_{t}(R_{t+1}|s_{t+1})=L_{t}(R_{t+1})$. This term of the likelihood function is \[ L_{t}(R_{t+1})=\frac{1}{\sqrt{2\pi h_{t+1}}}\exp\left\{ -\frac{1}{2}\frac{[R_{t+1}-(r+\lambda\sqrt{h_{t+1}}-h_{t+1}/2)]^{2}}{h_{t+1}}\right\} . \] Next, the conditional distribution of the realized measure, $L_{t}(x_{t+1}|R_{t+1},s_{t+1})$, is log-normally distributed. From the measurement equation, ((ref)), and the Gaussian specification, we have \[ L_{t}(x_{t+1}|R_{t+1},s_{t+1})=\frac{1}{\sqrt{2\pi\sigma^{2}}}\exp\left\{ -\frac{1}{2}\frac{[x_{t+1}-(\xi_{s_{t+1}}+\phi\log h_{t+1}+\delta_{1}z_{t+1}+\delta_{2}(z_{t+1}^{2}-1))]^{2}}{\sigma^{2}}\right\} . \] Finally, the remaining term in ((ref)), $P_{t}(s_{t+1})$, is the conditional state distribution. The relation between $P_{t}(s_{t+1})$ and $P_{t}(s_{t})$ is given by
where the last equality follows by the transition probabilities being time-invariant and independent of $\{R_{t},x_{t}\}$. Before we can evaluate ((ref)), we need to compute $P_{t}(s_{t})$. From Bayes' Theorem we have, \[ \Pr(s_{t}|R_{t},x_{t},\mathcal{F}_{t-1})=\frac{\Pr(R_{t},x_{t},s_{t}|\mathcal{F}_{t-1})}{\Pr(R_{t},x_{t}|\mathcal{F}_{t-1})}=\frac{\Pr(R_{t},x_{t}|s_{t},\mathcal{F}_{t-1})\Pr(s_{t}|\mathcal{F}_{t-1})}{\Pr(R_{t},x_{t}|\mathcal{F}_{t-1})}, \] where the first term is identical to $P_{t}(s_{t})$ and the last is identical to $\frac{L_{t-1}(R_{t},x_{t}|s_{t})P_{t-1}(s_{t})}{L_{t-1}(R_{t},x_{t})}$. So,
In the third equality, we used that $L_{t-1}(R_{t}|s_{t})=L_{t-1}(R_{t})$. By combining ((ref)) and ((ref)), we arrive at the following recursive formula for $P_{t}(s_{t+1})$:
that defines the recursion from $P_{t-1}(s_{t})$ to $P_{t}(s_{t+1})$. Therefore, given an initial value for $P_{0}\left(s_{1}\right)$, we can evaluate the likelihood function for $\{(R_{t},x_{t})\}_{t=1}^{T}$. There are two obvious ways to obtain an initial value for $P_{0}\left(s_{1}\right)$. One can either model it as an unknown vector (of conditional state probabilities) to be estimated, or one can assign $P_{0}\left(s_{1}\right)$ to have its stationary distribution, which is given by the eigenvector, $\pi^{\prime}\Pi=\pi^{\prime}$. In our empirical implementation, we use the latter.
The recursion for state probabilities, ((ref)), depends on the conditional distribution of realized measures, but does not depend on the conditional distribution for returns. The reason is that the state, $s_{t}$, only impacts the intercept in the measurement equation. The recursive expression for state probabilities, ((ref)), also holds when option prices are included in the analysis and parameters are estimated by maximization of the total likelihood. However, the joint estimation will influence all parameter estimates, which will influence the likelihood for the realized measure, $L_{t-1}(x_{t}|R_{t},s_{t})$. Thus, the estimated state probabilities based on $\{(R_{t},x_{t})\}_{t=1}^{T}$, need not be identical to those obtained from joint estimation that also includes option prices.
The properties of the model under the $\mathbb{Q}$-measure were presented in Theorem (ref), and it is evident that there is no analytical formula for the moment generating function (MGF) of cumulative return. This complication is not a consequence of the Markov switching structure because the complication is also present without Markov switching ($N=1$). Without an analytical formula for the MGF, the conventional path to closed-form pricing expressions with Fourier inversion is not applicable. Instead, we follow the method developed in HuangWangHansen2017 and derive an analytical approximation based on an Edgeworth expansion of the density for the cumulative return.
The European call option price is given by \[ C_{t}=e^{-r(T-t)}\mathbb{E}_{t}^{\text{\ensuremath{\mathbb{Q}}}}\left(\max\left(S_{T}-K,0\right)\right)=\sum_{s_{t}}P_{t}(s_{t})C_{t}(s_{t}), \] where $C_{t}(s_{t})\equiv e^{-r(T-t)}\mathbb{E}_{t}^{\text{\ensuremath{\mathbb{Q}}}}\left(\max\left(S_{T}-K,0\right)|s_{t}\right)$, $T$ is the date of maturity, $S_{T}$ is the terminal stock price, and $K$ is the strike price. Let $R_{T}=\log(S_{T}/S_{t})$ be the future cumulated return, $\mu=\mathbb{E}_{t}^{\text{\ensuremath{\mathbb{Q}}}}(R_{T}|s_{t})$, and $\sigma^{2}={\rm Var}_{t}^{\mathbb{Q}}\left(R_{T}|s_{t}\right)$, then the expectation can be written as an integral of the standardized cumulated return $z_{T}=\left(R_{T}-\mu\right)/\sigma$: \[ C_{t}(s_{t})=e^{-r(T-t)}\int_{-\infty}^{k}\left[S_{t}\exp(\mu-\sigma z)-K\right]\tilde{g}(z)\mathrm{d}z, \] where $k=\left(\log\left(S_{t}/K\right)+\mu\right)/\sigma$ and $\tilde{g}(z)$ is the true conditional density function of $-z_{T}$ given $s_{t}$. Following JarrowRudd1982, we apply a second-order Edgeworth expansion by expanding the density of the standardized cumulated return, $z_{T}$, using the expression,
Here $\phi(z)=(2\pi)^{-\tfrac{1}{2}}\exp(-z^{2}/2$) is the density of the standard normal distribution, $\kappa_{3}=\mathbb{E}_{t}^{\mathbb{Q}}(z_{T}^{3}|s_{t})$, $\kappa_{4}=\mathbb{E}_{t}^{\mathbb{Q}}(z_{T}^{4}|s_{t})$, and $H_{n}(z)$ is the $n$-th order Hermite polynomial given by \[ H_{3}(z)=z^{3}-3z,\qquad H_{4}(z)=z^{4}-6z^{2}+3,\quad\text{and}\quad H_{6}(z)=z^{6}-15z^{4}+45z^{2}-15. \] From HuangWangHansen2017, the price of a European call option is approximated by
where
Before we can obtain options prices with ((ref)), we need to derive the third and fourth moments of the standardized cumulative return, $z_{T}$, \[ \kappa_{3}=\frac{1}{\sigma^{3}}\left[\mathbb{E}_{t}^{\mathbb{Q}}\left(R_{T}^{3}|s_{t}\right)-\mu^{3}\right]-3\frac{\mu}{\sigma},\quad\text{and}\quad\kappa_{4}=\frac{1}{\sigma^{4}}\left[\mathbb{E}_{t}^{\mathbb{Q}}\left(R_{T}^{4}|s_{t}\right)-\mu^{4}\right]-2\frac{\mu}{\sigma}\left(2\kappa_{3}+3\frac{\mu}{\sigma}\right). \] The remaining problem is to obtain expressions for $\mathbb{E}_{t}^{\mathbb{Q}}(R_{T}^{3}|s_{t})$ and $\mathbb{E}_{t}^{\mathbb{Q}}(R_{T}^{4}|s_{t})$. These expressions involve a large number of terms, which are derived in Appendix (ref).
In Figure (ref), we present the approximated density for $z_{T}$ (under $\mathbb{Q}$) and the true density. The latter is simulated from the model we estimated in our empirical application, see Table 2. The model has two states, and the densities are for the standardized cumulative return, $z_{T}$, over six months in the high-volatility state. The red solid line is the approximated risk-neutral distribution and the blue line is the “true” risk-neutral distribution for $z_{T}$. The latter was obtained with 100,000 simulations. The approximation method largely agrees with the true density.
We compare the newly proposed Markov-switching Realized GARCH model (denote MS-RG) to several existing models, including a range of models that have documented the benefit of incorporating realized measures into discrete-time option pricing models. We also include the variant of the Realized GARCH model by HuangWangHansen2017 (denote RG), which is nested in our framework as the case $N=1$.
The Generalized Affine Realized Volatility (GARV) model was proposed by ChristoffersenFeunouJacobsMeddahi2014. This model assumes that the conditional variance for returns, $\bar{h}_{t}$, has two components. The first, $h_{t}^{R}$, is driven by returns, and the second, $h_{t}^{{\rm RV}}$, is driven by the realized measure (${\rm RV}$). This model takes the form:
where $(z_{t},\epsilon_{t})$ follows a standard bivariate normal distribution with a correlation of $\rho$. The dynamic properties under $\mathbb{Q}$ can be obtained through $z_{t}^{*}=z_{t}+\lambda\sqrt{\bar{h}_{t}}$ and $\epsilon_{t}^{*}=\epsilon_{t}+\text{\ensuremath{\chi}}\sqrt{\bar{h}_{t}}$, where $\left(z_{t}^{*},\epsilon_{t}^{*}\right)$ also follows a standard bivariate normal distribution with correlation $\rho$. As in the Realized GARCH model, the parameter $\text{\ensuremath{\chi}}$ is associated with the variance risk premium, and a positive $\text{\ensuremath{\chi}}$ implies a negative variance risk premium.\footnote{We report $\gamma=\delta_{2}/\vartheta$ instead of $\vartheta$ because the former can be used to measure the contribution of the realized information to the volatility process.} Due to the affine structure of the GARV model, a closed-form option price is available.
In addition to GARCH-type models, the availability of high-frequency data and realized measures has boosted the development of reduced-form models such as the HAR model. In particular, the HAR model was adapted by using the leverage function of the Heston-Nandi GARCH as well as a noncentral gamma distribution (LHARG) to price European call options. MajewskiBormettiCorsi2015 provided a general framework for option pricing with an LHARG model. \[ R_{t+1}=r+(\lambda-\tfrac{1}{2}){\rm RV}_{t+1}+\sqrt{{\rm RV}_{t+1}}z_{t+1},\quad{\rm RV}_{t+1}|\mathcal{F}_{t}\sim\Gamma\left(\delta,\varTheta_{t},\theta\right), \] \[ \ensuremath{\varTheta_{t}\theta=d+\beta_{d}{\rm RV}_{t}^{(d)}+\beta_{w}{\rm RV}_{t}^{(w)}+\beta_{m}{\rm RV}_{t}^{(m)}}\ensuremath{+\alpha_{d}\bar{\ell}_{t}}^{(d)}\ensuremath{+\alpha_{w}\bar{\ell}_{t}}^{(w)}\ensuremath{+\alpha_{m}\bar{\ell}_{t}}^{(m)}, \] \[ \ensuremath{
} \] MajewskiBormettiCorsi2015 denoted this model by ZM-LHARG due to the zero-mean leverage function. They found it has the best option pricing performance because its less-constrained leverage allows the process to explain a larger fraction of the skewness and kurtosis observed in real data. The risk-neutral dynamics are given by: \[ R_{t+1}=r-\tfrac{1}{2}{\rm RV}_{t+1}+\sqrt{{\rm RV}_{t+1}}z_{t+1}^{*},\quad{\rm RV}_{t+1}|\mathcal{F}_{t}\sim\Gamma\left(\delta,\varTheta_{t}^{*},\theta^{*}\right), \] where $z_{t}^{*}=z_{t}+\lambda\sqrt{\mathrm{RV}_{t}}$ follows a standard normal distribution. The risk-neutral parameters are linked to the physical parameters as follows: \[ \varTheta_{t}^{*}=\text{\ensuremath{\text{\ensuremath{\chi}}}}\varTheta_{t},\ \theta^{*}=\text{\ensuremath{\chi}}\theta,\ d^{*}=\text{\ensuremath{\chi}}^{2}d,\ \beta_{j}^{*}=\text{\ensuremath{\chi}}^{2}(\beta_{j}+\alpha_{j}\left((\gamma+\lambda)^{2}-\gamma^{2}\right)),\ \alpha_{j}^{*}=\text{\ensuremath{\chi}}^{2}\alpha_{j},\ {\rm for}\ensuremath{\ j\in\{d,w,m\}}, \] where $\ensuremath{\chi=(1+\theta(\tfrac{1}{2}(\lambda-\tfrac{1}{2})^{2}-v_{0}-\tfrac{1}{8}))^{-\tfrac{1}{2}}}$. Here, a negative $v_{0}$ is associated with a negative variance risk premium. MajewskiBormettiCorsi2015 derive the option-pricing formula for this model.
The last competing model is the Heston--Nandi GARCH model (hereafter HNG) presented by HestonNandi2000. HNG is one of very few discrete-time volatility models that yield an analytical option pricing formula. We adopt the variance dependent pricing kernel introduced in ChristoffersenHestonJacobs2013 to risk neutralize the HNG model, which takes the variance premium into account. The dynamic properties under the physical measures are given by
The corresponding risk-neutral dynamics with the variance-augmented pricing kernel are given by:
where $z_{t+1}^{*}$ has a standard normal distribution and the risk-neutral parameters are \[ h_{t}^{*}=\chi h_{t},\quad\omega^{*}=\chi\omega,\quad\tau_{2}^{*}=\chi^{2}\tau_{1}\quad\tau_{1}^{*}=\left(\lambda+\tau_{1}-\frac{1}{2}\right)\chi^{-1}+\frac{1}{2}. \] Once again, the parameter $\chi$ is associated with the variance risk premium, where $\chi>1$ corresponds to negative variance risk premium.
In this section, we turn to model estimation including the pricing kernel. In Section (ref) we estimated the Markov-switching Realized GARCH model from $\{(R_{t},x_{t})\}_{t=1}^{T}$ alone. Now, we will include option prices in estimation, and simultaneously estimate parameters in the pricing kernel and the parameters in the Markov-switching Realized GARCH model. We adopt the quasi log-likelihood for the option prices from ChristoffersenHestonJacobs2013 and combine it with the log-likelihood for $\{(R_{t},x_{t})\}_{t=1}^{T}$. The quasi log-likelihood function for option prices by ChristoffersenHestonJacobs2013 assumes that option pricing errors, measured in units of the Black-Scholes Vega-units, \[ e_{i}=\frac{O_{i}^{\mathrm{Model}}-O_{i}^{\mathrm{Market}}}{\nu_{i}^{bs}},\qquad i=1,\ldots,N, \] are normally distributed, $e_{i}\sim iidN(0,\sigma_{e}^{2})$. Here, $O_{i}^{\mathrm{Model}}$ and $O_{i}^{\mathrm{Market}}$ represent the model-implied option price and the observed market-based option price, respectively, $\nu_{i}^{bs}$ is the corresponding Black-Scholes Vega, and $N$ is the total number of option prices in the sample period. The Vega-weighted pricing error mimics the difference in the implied volatilities, and $\sigma_{e}^{2}$ denotes the variance of the Vega-weighted pricing errors.
Parameters are estimated by maximizing the total quasi log-likelihood function, $\ell_{\mathrm{Total}}=\ell_{R,x}+\ell_{o}$, where the expression of $\ell_{R,x}$ is given in Section (ref), and $\ell_{o}$ is
In the inclusion of the quasi log-likelihood for option prices, we follow ChristoffersenJacobsOrnthanalai2012, Ornthanalai2014, and HuangTongWang2019, and scale $\ell_{o}$ by $\tfrac{T}{N}$. This amounts to an adjustment for the imbalance between the number of observed option prices and the number of observation of $(R_{t},x_{t})$, which serves to prevent that parameter estimation is largely dominated by the option data.
Our empirical analysis is based on close-to-close log-returns for the S&P 500 index and the panel of SPX option prices. We use the realized kernel estimator, by BNHLS:2008, as our realized measure of volatility, where the realized kernel is implemented with the Parzen kernel.\footnote{The S&P 500 index was collected from Yahoo Finance. The realized measure was obtained from the Realized Library of Oxford-Man institute for the years 2000--2019. Before 2000 we used high-frequency data on S&P 500 futures prices (front-month continuous) from TickData to construct the realized kernels.}
Our sample spans the period from January 1990 to December 2019, which has 7,559 trading days. The option prices are assembled from two different databases. Option price data for the first six years are based on the Optsum data from the CBOE DataShop, whereas option prices, 1996 or later, are based on the OptionMetrics data. We use out-of-the-money put and call options with positive trading volume and with maturity between two weeks and six months and we apply the filters proposed by BakshiCaoChen1997. We only use Wednesday option prices, as is common in this literature. For each maturity, we retain the three strike prices with the highest liquidity (as defined by daily trading volume).\footnote{The same type of inclusion criteria were used in ChristoffersenFeunouJacobsMeddahi2014 and HuangWangHansen2017, who retained the six most liquid strike prices with maturities between two weeks and six months.} This results in a total of 32,024 option prices.
Table (ref) reports descriptive statistics for S&P 500 returns, realized kernels, and CBOE VIX in Panel A. The S&P 500 returns exhibit a small negative skewness and a high level of kurtosis. The realized kernel and the VIX are both positively skewed and leptokurtic. The standard deviation of returns is, at 17.437%, substantially smaller than the average option-implied volatility, at 19.147%. Their difference reflects the (average) negative variance risk premium. Panel B of Table (ref) provides an extensive summary of the number of contracts, average prices, and the average implied volatility within each subcategory in our option data set. Following ChristoffersenFeunouJacobsMeddahi2014, the out-of-the-money put options are converted to in-the-money call options using put-call parity, and the Black-Scholes delta is used to measure the moneyness of options. The option prices included in our analysis are all out-of-the-money options. So, options with deltas larger than 0.5 are out-of-the-money put options, and options with deltas less than 0.5 are out-of-the-money call options. In the upper part of panel B, it can be seen that deep out-of-the-money put options (deltas above 0.7) are relatively expensive compared with out-of-the-money calls, which reflects the well-known volatility smirk when implied volatility is plotted against moneyness. The middle part of the panel B summarizes features of the option prices when sorted by maturity, and the implied volatility term structure is roughly flat on average. The bottom of Panel B sorts the data by the volatility state as measured by contemporary VIX level and, as expected, the implied volatility increases in VIX.
We present results for the Markov-switching Realized GARCH model with two states ($N=2$).\footnote{The number of states could be chosen by a suitable statistical criterion. Nevertheless, a model with two states will likely be preferred in many applications because additional states add many parameters to the model, and a two-state model already adds a high degree of flexibility to the model structure.} In Table (ref), we present the estimation results for each of the models. Parameters are estimated using the sample from January 1990 to December 2019. The estimated parameters are given along with their robust standard errors in brackets. We also report implied logarithmically transformed expectation of $h_{t}$, and the implied persistence of volatility dynamics, $\pi^{\ensuremath{\mathbb{P}}}$ and $\pi^{\ensuremath{\mathbb{Q}}}$, under the physical and risk-neutral measures, respectively. The values of the maximized log-likelihood, $\ell_{\mathrm{Total}}$, are provided. The maximized log-likelihoods are not comparable unless the underlying models describe the same data, which is not the case for all models. The first four models are GARCH-type models, which largely share a common notation for their parameters. To conserve space, we present the estimates for $\sigma$ (for MS-RG and RG) and $\rho$ (for GARV) on the same line. The last model is the LHARG model, which has a different set of parameters and these are labelled in a separate column.\footnote{The logarithm of unconditional variances are estimated instead of the intercepts in the variance equations, which are then implied from the unconditional variance formulas.}
The estimated models are consistent with several stylized facts in the related literature. First, the estimated models imply a highly persistent volatility process under both physical and risk-neutral measures. The only exception is the LHARG model which estimates the volatility to be less persistent, in particular under the physical measure. The likely explanation is that the LHARG model treats the noisy realized measure as the true underlying volatility. It is well known that the measurement errors in the realized measure will induce a downwards bias in autocorrelations, see Hansen2014a, and this may explain the lower estimate of persistence in this model. Second, all of the models estimate the equity premium parameter to be positive, $\lambda>0$, and significant. Third, the values of the variance risk parameter ($\chi$) all indicate a significant negative variance risk premium\footnote{Positive $\chi$ for MS-RG, RG, GARV models, and $\chi$ greater than one for HNG, LHARG models corresponds to a negative variance risk premium.} and higher risk-neutral volatility than their physical counterparts. Fourth, the models find the leverage effect, as indicated by $\tau_{1}$ (and $\gamma$ for LHARG), to be significant. Fifth, the two Realized GARCH models (columns 1 and 2), estimate the parameter, $\gamma$, to be positive and significant. This parameter measures the contribution of the realized volatility measure in describing the volatility dynamics, and the realized measures are found to be an important predictor. The estimate of $\phi$ is close to one, which implies that the realized measure is proportional to the conditional variance. This reinforces the label measurement equation for ((ref)) and ((ref)). Compared with other models in Table (ref), the two Realized GARCH models deliver the smallest Vega-weighted pricing error $\sigma_{e}$.
The estimation results for the Markov-switching Realized GARCH model (MS-RG), proposed in this paper, are quite interesting. First, the state-dependent intercept $\xi$ does improve the empirical fit of the model, as illustrated by the increased value of the log-likelihood for the “physical” variables $\ell^{\mathbb{P}}$. The parameter $\xi$ controls the long-run volatility level within each state. For this reason, we adopt labels “Low” and “High” to denote the low- and high-volatility states, respectively. Second, we also find the estimated variance risk premium parameter, $\chi$, to be state-dependent. In the high-volatility state, $\chi_{{\rm High}}$ is estimated to be 0.4586, and is highly significant. This estimate translates to a large negative variance risk premium (in absolute value). In the low-volatility state, we estimate $\chi_{{\rm Low}}$ to be a small negative value, which suggests that investors have a slight appetite for variance risk in this state, i.e. a positive variance risk premium. The difference between these two estimates reinforces the value of introducing a state-dependent pricing kernel. The estimate of $\chi$ in single-state RG model is 0.0519, which falls between $\chi_{{\rm High}}$ and $\chi_{{\rm Low}}$. Third, the estimated transition probabilities within the same state ($\pi_{{\rm High|High}}$ and $\pi_{{\rm Low|Low}}$) are both near one, which implies the time-varying process of the hidden states is highly persistent. Fourth, although we focus on the model with two states, the improvements relative to the benchmark model are impressive. The MS-RG model lowers the Vega-weighted pricing error $\sigma_{e}$ by 14.3% relative to the single-state RG model.
In Figure (ref), we present the time series of conditional state probabilities along with the VIX index. The blue solid line denotes the estimated value of $P_{t}(s_{t}=\text{\textquotedblleft}\text{High}\text{\textquotedblright})$, and the red dashed line is the logarithmically transformed VIX index. The \textquotedblleft High\textquotedblright volatility state is also the state with the highest variance risk premium (in absolute value), so a large value of $P_{t}(s_{t}=\text{\textquotedblleft}\text{High}\text{\textquotedblright})$ is a period where investors have relatively high variance risk aversion. There is a great deal of variation in $P_{t}(s_{t}=\text{\textquotedblleft}\text{High}\text{\textquotedblright})$ over time and the process is very persistent because $\pi_{ii}$ is estimated to be close to one. There are several period, where $P_{t}(s_{t}=\text{\textquotedblleft}\text{Low}\text{\textquotedblright})=1-P_{t}(s_{t}=\text{\textquotedblleft}\text{High}\text{\textquotedblright})\simeq1$, including most of the time during the year, 1993-1995, 2004-2007, and 2014-2017. These are times where investors have an appetite for variance risk (or very low volatility risk aversion). Investors demand additional compensation for taking on variance risk in the “High” volatility state. The onset of these periods often coincides with large jumps in the VIX index, such as those seen around the time of the Asian Crisis in 1997, the bursting of the dot-com bubble and corporate scandals in early 2000s, the global financial crises, and the Euro crisis. A large component of the VIX is, according to BekaertHoerovaDuca2013, driven by factors that relate to time-varying risk-aversion. The commonality between $P_{t}(s_{t}=\text{\textquotedblleft}\text{High}\text{\textquotedblright})$ and the VIX supports this view, and the clear pattern that large upwards jumps in these two series tend to coincide.
In this section, we turn to the models' ability to price options, first in-sample and then out-of-sample. We follow the existing literature and convert option prices to their corresponding implied volatilities, as defined by the Black--Scholes formula. We evaluated each of the models by comparing observed implied volatility, denoted by $\mathrm{IV}^{\mathrm{Market}}$, to the corresponding model-based implied volatility, denoted $\mathrm{IV}^{\mathrm{Model}}$. The criterion used in our evaluation is the root mean squared error (RMSE), \[ {\rm RMSE_{IV}}=\sqrt{\frac{1}{N}\text{\ensuremath{\sum_{i=1}^{N}\left(\mathrm{IV}_{i}^{\mathrm{Model}}-\mathrm{IV}_{i}^{\mathrm{Market}}\right)^{2}}}}\times100, \] where $i=1,\ldots,N$ indexes each of the option prices (converted to implied volatilities) that were included in the sample period.
In Table (ref) we report the in-sample performance for option pricing for each of the estimated models. The RMSEs are reported in the first row and we find the MR-RG to have the smallest RMSE followed by the RG. The HNG has the largest average RMSE. The percentage reduction in RMSE of the MS-RG model relative to each of the alternative models ranges from 15.3% to 30.8%.
In order to investigate if the improvements by the MS-RG model are seen for options with specific characteristics or are found across the board, we repeat the comparisons after sorting the options by moneyness, time to maturity, and the contemporaneous level of the VIX.
Sorting by moneyness can cast light on the models' ability to generate sufficient leverage effect. The MS-RG model has the smallest RMSE in all subcategories. Relative to the single-state RG model, the two-state MS-RG model performs particularly well at pricing deep out-of-the-money call options (Delta$<$0.3), with the reduction in RMSE up to 21%.
Maturity is related to the models' ability to explain volatility dynamics over longer time spans. The RMSEs of the MS-RG model are fairly uniform along this dimension and the MS-RG model also has the smallest RMSE across all subcategories. The RMSE of the MS-RG is a tad higher at the longest maturity compared to shorter maturities, but its RMSE is still substantially smaller than that of all competitors.
Finally, the comparisons across levels of the volatility index will illustrate the models' ability to generate a proper variance risk premium at different levels of volatility. Along this dimension, we also find that the MS-RG model has the smallest RMSE in all subcategories. The RMSEs are roughly proportional to the level of VIX, which is what one would expect if the distribution of pricing errors, when measured in percentage, is relatively homogeneous across subcategories.
In this section, we turn to out-of-sample comparisons of the pricing models. A Markov switching model is more heavily parameterized and this can lead to overfitting when the model is estimated and evaluated with the same data. Here we follow ChristoffersenJacobs2004ms and split the original sample into two subsamples. We use the first 15 years, 1990-2004, exclusively for model estimation (in-sample) and use the remaining 15 years, 2005-2019, for model evaluation (out-of-sample) of the estimated models. So, each of the models is estimated exactly once using the in-sample data and their abilities to price options are exclusively evaluated over the out-of-sample period.
We report the out-of-sample RMSEs in Table (ref). The relative ranking of the models is identical to the relative ranking we found in-sample. The MS-RG model continues to be the best-performing model. Not only is the MS-RG model the best on average, but it is also the best within all subcategories. Impressively, the MS-RG has nearly the same RMSE out-of-sample as it does in-sample. Given the great out-of-sample performance by the MS-RG, we conclude that its better option pricing performance is not a result of overfitting, but rather reflects a real and substantive improvement in model-based option pricing, which can be attributed to the time-varying risk aversion that the MS-RG model can capture.
We introduced Markov-switching to the Realized GARCH model and combined it with an exponentially affine pricing kernel with a time-varying aversion to volatility risk. A key feature of this framework is that time-variation in the pricing kernel is introduced with the same hidden Markov process that is used in the Realized GARCH model. In this way, the hidden Markov chain process brings time-variation in both the physical measure and the risk-neutral measure. We derived model-implied pricing formula for European options in this framework. This was achieved with an analytical approximation method that is based on an Edgeworth expansion of the density for cumulative return.
The MS-RG model is straightforward to estimate from returns and realized volatilities by quasi maximum likelihood estimation. Volatility models with a hidden Markov-switching are typically challenging to estimate. We circumvent the usual complications by introducing Markov switching in a manner where states probabilities can be inferred from the realized measures and returns. Estimating the parameters in the pricing kernel is more involved and requires that option prices are included in the empirical analysis. We estimated the model with a large panel of option prices using 30 years of data (from 1990 to 2019) and find that investors have a state-dependent tolerance for volatility-specific risk. Consistent with the existing literature, we find that the volatility risk premium is, on average, negative. However, there are periods when investors appear to have an appetite for variance risk. These periods tend to coincide with relatively low levels of the VIX. During other periods, we find that investors have a relatively high aversion to volatility risk, and these periods are often preceded by sudden upwards jumps in the VIX.
We compared the proposed model with several benchmarks and found it to outperform all competing models. The option pricing model that is based on the MS-RG framework leads to substantial improvements in option pricing. This reduction in the RMSE of option pricing errors range from 15.3% to 30.8% relative to the alternative models. The same magnitudes of improvements are seen across options with different characteristics. These improvements are also seen out-of-sample, where the RMSE reductions in option pricing errors range from 17% to 34.5%.
The S&P 500 index was collected from Yahoo Finance and CBOE VIX is from the CBOE website. The SPX options are available from the Optsum data (1990--1995) and OptionMetrics (1996--2019). The realized measure of volatility was obtained from the Realized Library of Oxford-Man institute (2000-2019) and TickData (1990-1999).