Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
123,275 characters · 21 sections · 123 citation commands
Option Pricing with Time-Varying Volatility Risk Aversion
The pricing kernel, which represents the ratio of risk-neutral to physical state probabilities, is a fundamental concept in economics with roots in Arrow-Debreu securities, see ArrowDebreu1954. A cornerstone of asset pricing theory, the Lucas1978 tree model, implies a monotonically decreasing pricing kernel with respect to aggregate wealth, see HansenRenault:2010. However, empirical studies consistently find the pricing kernel to be non-monotonic and time-varying, giving rise to the so-called pricing kernel puzzles, see e.g. AitSahaliaLo2000, Jackwerth2000, and RosenbergEngle2002.
In this paper, we propose a pricing kernel with time-varying volatility risk aversion to resolve these pricing kernel puzzles. Building on HestonNandi2000 and ChristoffersenHestonJacobs2013, we develop a novel framework for derivatives pricing, in which the variance risk ratio (VRR) emerges as the key quantity. The VRR, defined as the ratio of risk-neutral to physical conditional variance, characterizes both the shape and time variation of the pricing kernel and it is closely related to the variance risk premium by CarrWu2008. The dynamic pricing kernel introduces challenges for closed-form option pricing formulas, which we address through a novel analytical approximation method applicable to non-affine models. Our empirical implementation of the dynamic structure for VRR is based on the score-driven framework of CrealKoopmanLucas:2013, which intuitively updates time-varying quantities in response to first-order conditions. In an empirical application to the S&P 500 index, the CBOE VIX, and option prices, we show that the new model substantially reduces pricing errors, both in-sample and out-of-sample.
The discrete-time Heston-Nandi GARCH (HNG) model, developed by HestonNandi2000, conveniently yields closed-form expressions for option prices while allowing for time-varying volatility. However, it implies a monotonic pricing kernel. ChristoffersenHestonJacobs2013 extended the HNG model by introducing a variance-dependent pricing kernel, resulting in a non-monotonic shape. Their model remains tractable, retains closed-form option pricing expressions, and has become a benchmark in the option pricing literature. The shape of the pricing kernel in their framework is governed by the variance risk premium (VRP), where a negative VRP leads to a U-shaped pricing kernel, a pattern confirmed in their empirical analysis, see also CuesdeanuJackwerth2018.
While ChristoffersenHestonJacobs2013 explain non-monotonicity, their framework cannot account for the time variation observed in the pricing kernel. When estimated over long sample periods, the pricing kernel is typically U-shaped, yet the shape differs distinctly in certain periods, such as the years leading up to the global financial crisis. The shape of the pricing kernel can be expressed in terms of the probability weighting function. If individuals overweight low-probability events and underweight high-probability events, the probability weighting function will have an inverse S-shape, which corresponds to a U-shaped pricing kernel. PolkovnichenkoZhao2013 estimated the probability weighting function non-parametrically and found it to be time-varying. Their estimate has an inverse S-shape during most of their sample periods but a regular S-shape during the years 2004--2006. Similar results were obtained by ChabiSong2013, and an inverted U-shape was also documented with DAX 30 options during the same years, 2004--2006, as reported by GrithHardleKratschmer2017. Further evidence comes from KieselRahe2017 and BeareSchmidt2016. Interestingly, ChristoffersenHestonJacobs2013 also contains evidence of time variations in empirical pricing kernels. For most calendar years, their estimated empirical pricing kernel is U-shaped; however, in some calendar years, it takes an inverted U-shape.\footnote{The authors do not comment on this observation, but they estimate a more flexible structure, which is not based on a pricing kernel with a fixed shape, see ChristoffersenHestonJacobs2013.}
The derivative pricing model proposed in this paper combines a pricing kernel with time-varying volatility risk aversion and the GARCH model of HestonNandi2000. The VRR, defined as the ratio of risk-neutral to physical conditional variance, $\eta_{t}=h_{t+1}^{\ast}/h_{t+1}$, governs the time variation in the pricing kernel and is functionally linked to the volatility risk aversion parameter, which captures the curvature of the pricing kernel. We demonstrate that this structure can generate the observed time variation in the empirical pricing kernel, including its different shapes. Our new model nests ChristoffersenHestonJacobs2013 as the special case where $\eta_{t}$ is constant, which we denote as CHNG. The HNG option pricing model of HestonNandi2000 is also nested by imposing $\eta_{t}=1$. We refer to our new model as DHNG model, where “D” stands for dynamic, and highlight its key properties below.
We derive a closed-form CBOE VIX pricing formula for the DHNG model and develop an analytical option pricing formula, which constitutes a separate and important methodological contribution. When volatility risk aversion follows a stochastic process, the conditional moment-generating function (MGF) of future cumulative returns does not have an affine form, making it very challenging to derive option prices. However, we introduce a novel approximation method that is applicable to non-affine models. This approach constructs an auxiliary MGF for a simplified, related problem and then applies a Taylor expansion to obtain an analytical option pricing formula. The approximation method is shown to be highly accurate in an empirically relevant simulation design.
To implement the theoretical framework, we connect the model for $\eta_{t}$ with observed data using an observation-driven approach, which is a natural choice since the GARCH model under the physical measure also follows an observation-driven formulation. Specifically, we propose a score-driven model following CrealKoopmanLucas:2013, where the first-order conditions of the log-likelihood function define the innovations in the dynamic model for $\eta_{t}$. This design is intuitive because $\eta_{t}$ is updated to minimize pricing errors. Conveniently, this structure also enhances robustness to model misspecification. Another advantage of the observation-driven model is its simplicity in estimation, which is particularly useful in applications involving long sample periods and large panels of option prices.
An empirical application using 32 years of daily S&P 500 returns, the CBOE VIX, and a large panel of option prices provides strong support for the model. The estimation incorporates information from both the physical and risk-neutral measures, as advocated by ChernovGhysels2000. The results show that the pricing kernel with time-varying volatility risk aversion typically reduces pricing errors by 50% or more, as measured by the root mean square error. This reduction is achieved for both VIX and option prices, in both in-sample and out-of-sample comparisons. The estimated time variation in the shape of the pricing kernel aligns with previous studies, generally exhibiting a U-shape, but taking on an inverted U-shape during certain periods, including 2004--2007 and around 1993 and 2017.
Finally, having established the VRR as a fundamental quantity, it is natural to investigate its connection to key asset pricing variables, including those linked to pricing kernel puzzles. Theoretical and empirical studies suggest that heterogeneous beliefs and disagreements about the physical distribution can contribute to a U-shaped pricing kernel, see e.g. Shefrin2001,\nocite{Shefrin2008} BakshiMadan2008, BakshiMadanPanayotov2010. Similar arguments have been made about sentiment and uncertainty among investors, see e.g. Han2008, PolkovnichenkoZhao2013, BakerBloomDavis2016, and BollerslevLiXue2018. BaliZhou2016 argue that market uncertainty can be approximated by the variance risk premium, see CarrWu2008. The VRR is obviously related to the variance risk premium, which has been shown to predict stock returns, see BollerslevTauchenZhou2009. A straightforward explanation is that preferences are state-dependent, with market uncertainty being a possible state variable, see e.g. GrithHardleKratschmer2017. We connect our results to these studies by showing that the monthly average VRR is closely related to commonly used measures of sentiment, uncertainty, and disagreement, as well as other key economic indicators.
Our paper is related to and builds on a large body of literature, including studies that have sought to address pricing kernel puzzles by augmenting existing models with additional state variables, see e.g. Chabi-YoGarciaRenault2008, chabi2012, BrownJackwerth2012 and SongXiu2016. In our framework, the quantity VRR can be viewed as an additional state variable that emerges naturally from the model structure.
Our work is also related to Barone-AdesiEngleMancini2008, who estimated a GJR-GARCH model for returns under the physical measure. They did not specify a pricing kernel but simply assumed that the model under the risk-neutral measure is also a GJR-GARCH model, with time-varying parameters calibrated to minimize option pricing errors. While we also seek to minimize pricing errors, our approach is fundamentally different. Our starting point is a pricing kernel and a model for the physical probability measure, $\mathbb{P}$, and the two imply the model under the risk-neutral probability measure, $\mathbb{Q}$. An empirical comparison in ChristoffersenHestonJacobs2013 explores the same idea as in Barone-AdesiEngleMancini2008, as they estimate separate HNG models under both $\mathbb{P}$ and $\mathbb{Q}$ measures. They referred to this structure as an ad-hoc model, because it is not derived from their pricing kernel. However, we show that the pricing kernel with a time-varying VRR is incoherent with HNG models (with constant parameters) under both $\mathbb{P}$ and $\mathbb{Q}$ measures.
Additionally, our paper contributes to the literature on time-varying risk aversion, which relates to the local shape of the pricing kernel. A seminal paper by CampbellCochrane1999 addresses this concept, while Li2007 demonstrated how time-varying risk aversion links to a time-varying risk premium. Moreover, GonzalezNaveRubio2018 identified risk aversion dynamics as a key determinant of stock market betas. More recently, BekertEngstromXu2020 introduced a no-arbitrage asset pricing model with time-varying risk aversion for pricing equities and corporate bonds, where risk aversion is driven by uncertainty shocks.
The rest of the paper is organized as follows. We present the theoretical model in Section (ref) and derive closed-form pricing formula for the VIX and an option pricing formula, based on the new approximation method in Section (ref). Section (ref) introduces a dynamic model for the variance risk ratio using a score-driven approach. An empirical application with 32 years of S&P 500 returns, VIX, and option prices is presented in Section (ref). We show that the variance risk ratio relates to well-known measures of sentiment, disagreement, and uncertainty in Section (ref). A summary is presented in Section (ref). All proofs are provided in the Appendix (ref) and supplementary empirical results can be found in the Online Appendix.
\everymath{{11pt} {11pt}} \everydisplay{{11pt} {11pt}}
This section introduces the new model with dynamic variance risk aversion (DHNG), building on the models by HestonNandi2000 (HNG) and ChristoffersenHestonJacobs2013 (CHNG).
The observed variables include the daily returns, $R_{t}\equiv\log\left(S_{t}/S_{t-1}\right)$, where $S_{t}$ is the underlying asset price, and a vector of derivative prices, $X_{t}$. We will work with two filtrations, $\mathcal{F}_{t}=\ensuremath{\sigma}(\{R_{j},X_{j}\},j\leq t)$ and $\mathcal{G}_{t}=\ensuremath{\sigma}(\{R_{j},X_{j-1}\},j\leq t)$, where the latter arises naturally from a factorization of the joint likelihood function.\textcolor{red}{ }Clearly, $\mathcal{F}_{t-1}\subset\mathcal{G}_{t}\subset\mathcal{F}_{t}$ for all $t$. Throughout this paper, we will use the notation $\mathbb{E}_{t}^{\mathbb{P}}\left(\cdot\right)\equiv\mathbb{E}^{\mathbb{P}}\left(\cdot|\mathcal{G}_{t}\right)$ and $\mathbb{E}_{t}^{\mathbb{Q}}\left(\cdot\right)\equiv\mathbb{E}_{t}^{\mathbb{Q}}\left(\cdot|\mathcal{G}_{t}\right)$ to denote the conditional expectations with respect to $\mathcal{G}_{t}$ under $\mathbb{P}$ and $\mathbb{Q}$ measures, respectively.
We model returns in physical measure $\mathbb{P}$ using the classical Heston-Nandi GARCH model (HNG), which has a convenient structure for derivatives pricing. This model is given by
where the return shock, $z_{t+1}$, is independent and identically distributed (iid) with a standard normal distribution, $N(0,1)$; $r$ is the risk-free rate; $\lambda$ is the equity risk premium because the expected return is given by $\mathbb{E}^{\mathbb{P}}\left(\exp\left(R_{t+1}\right)|\mathcal{F}_{t}\right)=\exp\left(r+\lambda h_{t+1}\right)$; and $h_{t+1}=\mathrm{var}^{\mathbb{P}}(R_{t+1}|\mathcal{F}_{t})$ is the daily conditional variance, see HestonNandi2000. The HNG model captures time variation in the conditional variance, as is the case for ARCH and GARCH models, see Engle:1982 and bollerslev:86. However, the HNG model also allows for dependence between returns and volatility (leverage). It is straightforward to verify that $\mathrm{cov}^{\mathbb{P}}(R_{t},h_{t+1}|\mathcal{F}_{t-1})=-2\gamma\alpha h_{t}$, which shows that the magnitude of the leverage effect is defined by $\gamma$. A leverage effect is required to generate the empirically important volatility smirk in option prices. The dynamic structure of the HNG model is carefully crafted to yield a closed-form solution for option valuation, and HestonNandi2000 showed that the continuous limit of the HNG model (as the time interval between observations shrinks to zero) yields a variance process, $h_{t}$, that converges weakly to a continuous-time square-root variance process, see Feller1951, CoxIngersollRoss85, and Heston1993.
Option pricing with GARCH models can be achieved with a simple risk-neutralization by Duan1995, known as the locally risk-neutral valuation relationship (LRNVR). This approach is equivalent to imposing a pricing kernel with a single risk premium for equity risk, see HuangWangHansen2017. However, the LRNVR-based pricing kernel is inadequate for explaining many discrepancies between $\mathbb{P}$ and $\mathbb{Q}$, including the variance risk premium, see HaoZhang2013 and ChristoffersenHestonJacobs2013. We will review their model next and then proceed to introduce the new model with time-varying volatility risk aversion.
The model in ChristoffersenHestonJacobs2013 (CHNG) generalizes the HNG model by HestonNandi2000, to have a variance-dependent pricing kernel, which is given by, \[ \frac{M_{t+1}}{M_{t}}=\left(\frac{S_{t+1}}{S_{t}}\right)^{\phi}\exp\left[\delta+\pi h_{t+1}+\xi(h_{t+2}-h_{t+1})\right], \] and CorsiFusariVecchia2013 showed that it can be conveniently expressed as \[ M_{t+1,t}\equiv\frac{M_{t+1}}{\mathbb{E}^{\mathbb{P}}[M_{t+1}|\mathcal{F}_{t}^{R}]}=\frac{\exp\left(\phi R_{t+1}+\xi h_{t+2}\right)}{\mathbb{E}_{}^{\mathbb{P}}[\exp(\phi R_{t+1}+\xi h_{t+2})|\mathcal{F}_{t}^{R}]}, \] where $\mathcal{F}_{t}^{R}=\sigma\left(\{R_{j}\},j\leq t\right)$ is the natural filtration for returns only. The pricing kernel above depends on both equity risk and variance risk, where the latter is typically characterized by the parameter $\xi$. The equity risk is governed by $\phi$, as well as $\xi$, because $h_{t+2}$ depends on return $R_{t+1}$. This pricing kernel implies the following risk-neutral dynamics:
with $z_{t+1}^{*}|\mathcal{F}_{t}^{R}\overset{\mathbb{Q}}{\sim}iid\ N(0,1)$ and the following relations between $\mathbb{P}$ and $\mathbb{Q}$ parameters: \[ h_{t}^{*}=h_{t}\eta,\quad\omega^{*}=\omega\eta,\quad\beta^{\ast}=\beta,\quad\alpha^{*}=\alpha\eta^{2},\quad\gamma^{\ast}=\tfrac{1}{\eta}(\gamma+\lambda-\tfrac{1}{2})+\tfrac{1}{2},\quad\eta\equiv(1-2\alpha\xi)^{-1}. \]
The logarithm of the pricing kernel is a quadratic function of the daily market return,
where $\kappa_{0}=-\frac{1}{2}\log\eta$ and $\kappa_{1}=\frac{1}{2}(\lambda-\tfrac{1}{2})^{2}-\tfrac{1}{8}\eta$.\footnote{The expressions for $\kappa_{0}$ and $\kappa_{1}$ in ChristoffersenHestonJacobs2013 are $\kappa_{0}=\delta+\xi\omega+\phi r$ and $\kappa_{1}=\pi+\xi(\beta-1+\alpha(\lambda-\tfrac{1}{2}+\gamma)^{2})$. It follows from our results in Lemma (ref) that they are equivalent to those presented here.} Expression ((ref)) shows that $\lambda$ and $\xi$ define the linear and quadratic terms, respectively, with positive $\xi$ leading to a U-shaped and negative $\xi$ to an inverted U-shaped pricing kernel. Empirically, this relationship tends to be generally U-shaped, but it is unstable over time as seen in ChristoffersenHestonJacobs2013. Figure (ref) includes parts of figures 3 and 5 in ChristoffersenHestonJacobs2013; the four left panels are based on model-free market prices and the corresponding four right panels are based on model prices using their estimated CHNG model.
There are periods where the relationship has an inverted $U$-shape, such as the years 2004--2007 (of which 2005 and 2006 are shown in Figure (ref)), which is incompatible with a static quadratic coefficient $\xi$.
These observations, and the fact that $\xi$ defines the shape of the relationship between $\log M_{t}/M_{t-1}$ and returns, motivate us to develop a new model where $\xi$ can be time-varying. This parameter, $\xi$, is directly tied to volatility risk aversion, as it is the coefficient to the state variable, $h_{t+2}$, in the pricing kernel. Moreover, $\xi$ is also negatively related to the VRP, because $h_{t}-h_{t}^{\ast}=-\xi\times\tfrac{2\alpha h_{t}}{1-2\alpha\xi}\approx-\xi\times2\alpha h_{t}$, where $2\alpha h_{t}>0$.
We seek a more flexible pricing kernel that is coherent with the observed variation over time. Next, we introduce a new dynamic pricing kernel while maintaining the Heston-Nandi GARCH structure ((ref))-((ref)) under physical measure, $\mathbb{P}$.
Assumptions (ref) and (ref) induce the following dynamic properties under risk-neutral measure.
Theorem (ref) reveals the following six interesting properties of the risk-neutral probability measure and its dynamics.
First, $\eta_{t}=h_{t+1}^{\ast}/h_{t+1}$ emerges as a fundamental quantity\footnote{Parameterize the model with the inverse ratio, $\theta_{t}=h_{t+1}/h_{t+1}^{\ast}$, simplifies several expressions, such as $\xi_{t}=\tfrac{1}{2\alpha}(1-\theta_{t})$ and $\phi_{t}=(1-\theta_{t})(\gamma-\tfrac{1}{2})-\theta_{t}\lambda$. We adopt $\eta_{t}$, because it is more common in the literature.}, which is termed the variance risk ratio. It is $\mathcal{G}_{t}$-measurable since $\eta_{t}$ is a function of $\xi_{t}\in\mathcal{G}_{t}$ defined in Assumption (ref). This variable appears in all expressions relating $\mathbb{P}$ and $\mathbb{Q}$, and it has a straightforward interpretation. Its functional relationship with $\xi_{t}$ shows that $\eta_{t}$ is an indirect measure of volatility risk aversion. When $\eta_{t}$ is large, agents demand a large compensation for taking on variance risk, while a small value of $\eta_{t}$ corresponds to an appetite for variance risk. Moreover, $\eta_{t}$ is obviously related to the VRP, where the latter concerns the expectation of future values of $h_{t+1}-h_{t+1}^{\ast}$. This follows from $h_{t+1}-h_{t+1}^{\ast}=\left(1-\eta_{t}\right)h_{t+1}$. This relation could be used to construct a model-based measure of the VRP. Note that the VRP can be positive in the most general version of the model. A negative VRP can be guaranteed in the model design by restricting $\eta_{t}$ to be greater than one.\footnote{This can be be achieved by specifying a dynamic model for $\log(\eta_{t}-1)$. We do not impose this restriction in our analysis, as we specify a model for $\log\eta_{t}$.} For comparison, the ratio, $h_{t}^{\ast}/h_{t}$, is constant in CHNG model, and the relation $h_{t}-h_{t}^{\ast}=(1-\eta)h_{t}$ shows that CHNG model implicitly restricts the VRP to be proportional to the conditional variance under $\mathbb{P}$.
Second, the shape of the pricing kernel is time-varying, as shown by the following results.
The expression in Lemma (ref) is the generalization of ((ref)) to the case with time-varying preference parameter, $\phi_{t}$ and $\xi_{t}$. The quadratic coefficient can be time-varying, because it depends on $\xi_{t}$, which makes it clear that $\xi_{t}$ influences the shape of the pricing kernels. Or, equivalently, $\eta_{t}$ governs the shape. We have a U-shape if $\xi_{t}>0$ ($\Leftrightarrow\eta_{t}>1$) and an inverted U-shape if $\xi_{t}<0$ ($\Leftrightarrow\eta_{t}<1$).\footnote{We have $\alpha>0$ in the Heston-Nandi GARCH model.} Thus, this framework can generate a variety of shapes of the pricing kernel. This is illustrated in the upper left panel of Figure (ref), where the pricing kernel for cumulated returns over one month is presented for two levels of $\eta$. The design is based on the empirical estimates of the model in Section (ref), where $\eta_{\mathrm{low}}=0.70$ and $\eta_{\mathrm{high}}=1.70$ correspond to the $10\%$ and $90\%$ quantiles of the unconditional distribution of $\eta_{t}$, respectively. The low value of $\eta$ produce an inverted U-shape similar to that observed during the years 2004 to 2007, whereas the high value of $\eta$ results in a shape with a pronounced U-shape.
Third, the equity risk premium, $\lambda$, is constant, despite $\phi_{t}$ and $\xi_{t}$ being time-varying. This is an implication of the no-arbitrage condition, that ties the linear coefficient to $\lambda$, which is determined under the physical measure in the Heston-Nandi GARCH model.
Fourth, an important reason for permitting $\eta_{t}$ to be time-varying is that it enables $h_{t}^{\ast}$ to incorporate information from derivative prices. In conventional GARCH models, the conditional variance depends only on lagged returns, such that $h_{t+1}\in\mathcal{F}_{t}^{R}$. A restrictive implication of a constant VRR is that it requires the two conditional variances to be proportional to each other. Consequently, $h_{t+1}^{\ast}$ is driven solely by lagged returns, leaving no room for derivative prices to influence its dynamic properties. Allowing $\eta$ to be time-varying eliminates this rigidity and enables the model to incorporate information from derivative prices into the dynamic model of $h_{t+1}^{\ast}$. We take advantage of this feature in our empirical implementation.
Fifth, the structure in Theorem (ref) nests the CHNG model of ChristoffersenHestonJacobs2013 and the HNG model, which corresponds to $\eta_{t}=\eta$ and $\eta_{t}=1$, respectively. If $h_{t}$ is constant and $\eta_{t}=1$, then it leads to the pricing kernel given by the power utility function, see Rubinstein1976. So the famous Black-Scholes model is also nested as the special case. Interestingly, when $\eta_{t}$ is time-varying, there is time variation in the risk aversion parameters, which is similar to the habit formation models, see CampbellCochrane1999.
Sixth, the structure introduced in Theorem (ref) generates time-varying leverage under $\mathbb{Q}$, which manifests itself in a number of ways. A time-varying leverage effect will impact the distribution of cumulative returns, and the lower panels of Figure (ref) show the skewness and kurtosis of cumulative returns under $\mathbb{Q}$ as a function of the number of days that returns are cumulated over (the $x$-axis). Skewness and kurtosis are shown for the two levels of the VRR, $\eta_{\mathrm{low}}$ and $\eta_{\mathrm{high}}$, and these result in distinctively different levels of skewness and kurtosis under $\text{\ensuremath{\mathbb{Q}}}$. Both skewness and kurtosis are far more pronounced for the large value of $\eta$. Another way to illustrate the leverage effect is with the news impact curve by Engle_Ng_1993. The news impact curve under $\mathbb{Q}$ is defined by $\mathrm{NIC}(z^{\ast})=\mathbb{E}_{t}^{\mathbb{Q}}(h_{t+1}^{\ast}|z_{t}^{\ast}=z^{\ast})-\mathbb{E}_{t}^{\mathbb{Q}}(h_{t+1}^{\ast}|z_{t}^{\ast}=0)$, and from Theorem (ref) it follows that $\mathrm{NIC}(z^{\ast})=\alpha\eta_{t}\eta_{t-1}z^{\ast2}-2\alpha\gamma\eta_{t}\sqrt{h_{t}^{*}}z^{\ast}$, whose shape depends on the VRR. The upper right panel of Figure (ref) presents the news impact curve, with $h_{t}^{\ast}$ equal to its unconditional mean under $\mathbb{Q}$, as estimated in Section (ref), and with $\eta=\eta_{t}=\eta_{t-1}$ equal to either $\eta_{\mathrm{low}}$ (dashed line) or $\eta_{\mathrm{high}}$ (solid line). We observe the asymmetric shape of the news impact curve, see e.g. ChenGhysels:2011-news-impact-curve. However, the level of asymmetry depends on the level of $\eta$, such that volatility is more responsive to negative return shocks when $\eta$ is large, than when $\eta$ is small.
In this section, we consider cases where $\eta_{t}$ is time-varying and stochastic. We derive a closed-form expression for the VIX under the assumption that $\log\eta_{t}$ follows an autoregressive moving average process, ARMA$(p,q)$. We also obtain analytical expressions for option prices under the same assumptions. The option pricing formula is obtained with a novel approximation method based on a Taylor expansion of the MGF for cumulative returns.
It is relatively straightforward to obtain closed-form expressions for VIX pricing, once the dynamic properties of $\eta_{t}$ under $\mathbb{Q}$ are known. This only requires expressions for $\mathbb{E}_{t}^{\mathbb{Q}}(h_{t+k}^{*})$, $k=1,\ldots,M$, because the risk-neutral dynamics of returns under Theorem (ref) implies the following $M$-period ahead VIX pricing formula
where $A=100\sqrt{252}$ is the annualizing factor.
Assumption (ref) is explicit about the first two conditional moments of $\varepsilon_{t}$, but does not impose restrictive distributional assumptions. However, the exact distribution is important for VIX and option prices, as these depend on $\operatorname{mgf}_{\varepsilon}(s)$. That $\varepsilon_{t}$ is $\mathcal{G}_{t}$-measurable follows from Assumption (ref) and therefore not a new assumption.
In this section, we derive the European option pricing formula for the case where the variance risk aversion is time-varying.
The European call option price at time $t$ is given by the conditional risk-neutral expectation $C_{t}=\ensuremath{e^{-r\left(T-t\right)}\mathbb{E}_{t}^{\mathbb{Q}}\left[\max\left(S_{T}-K,0\right)\right]}$, where $T$ is the maturity date, $K$ is the strike price, and $S_{T}$ is the terminal price of underlying asset. Let $M=T-t$ be the number of periods to maturity, it is well-known that an affine structure of the MGF of future cumulative returns, \[ g_{t,M}(s)=\mathbb{E}_{t}^{\mathbb{Q}}\left[\exp\left(s\sum_{i=1}^{M}R_{t+i}\right)\right],\quad s\in\mathbb{R} \] is the key to closed-form option pricing formula, see HestonNandi2000. Unfortunately, $g_{t,M}(s)$ does not have an affine structure when the variance risk aversion follows a stochastic process. We propose a novel approximation method that extrapolates from the solution to a simpler auxiliary problem, leading to an approximation of $g_{t,M}(s)$.
The future values of $\log\eta_{t}$ that are relevant for pricing an option with $M$ periods to maturity are the elements of the vector, $\boldsymbol{\eta}_{t,M}=\left(\log\eta_{t+1},\ldots,\log\eta_{t+M}\right)^{\prime}$. Under Assumption (ref), this is a random vector. For later use, we use the notation, $\bar{\boldsymbol{\eta}}_{t,M}$, to represent a predetermined path (i.e. $\bar{\boldsymbol{\eta}}_{t,M}\in\mathcal{G}_{t}$), where the prime example is the conditionally expected trajectory, given by $\bar{\boldsymbol{\eta}}_{t,M}^{e}=\mathbb{E}_{t}^{\mathbb{Q}}(\boldsymbol{\eta}_{t,M})\in\mathbb{R}^{M\times1}$.
The assumption that $\log\eta_{t}$ will follow a deterministic trajectory, $\bar{\boldsymbol{\eta}}_{t,M}$, implies that $\phi_{t+j}$ and $\xi_{t+j}$ are also predetermined for $j=1,\ldots,M$, and similarly, $\omega_{t+j}^{*}$, $\alpha_{t+j}^{*}$, $\beta_{t+j}^{*}$, and $\gamma_{t+j}^{*}$ (the GARCH parameters under $\mathbb{Q}$) are also predetermined for $j=1,\ldots,M$. This is a direct consequence of Theorem (ref). We show that the affine structure is preserved in this case with the following closed-form expression for the MGF of cumulative returns.
The expression for the MGF in Lemma (ref) simplifies to that in ChristoffersenHestonJacobs2013 if $\bar{\boldsymbol{\eta}}_{t,M}=(\log\eta,\ldots,\log\eta)^{\prime}$, and it simplifies further to that in HestonNandi2000 when $\bar{\boldsymbol{\eta}}_{t,M}=0_{M\times1}$. Aside from these two special cases, this naive approach does not constitute a coherent model for option pricing. However, it serves as an important auxiliary “model” with the desired affine structure.
We now turn to the more challenging problem where $\boldsymbol{\eta}_{t,M}$ is stochastic. The lack of an affine structure thwarts the standard approach to obtaining closed-form option pricing formula, and it is common to resort to simulation methods in this case.
The idea behind our approximation method is simply to extrapolate from the simple case, $g_{t,M}(s|\bar{\boldsymbol{\eta}}_{t,M}^{e})$, and apply a second-order Taylor expansion about $\bar{\boldsymbol{\eta}}_{t,M}^{e}$. This leads to a quadratic expression in the $M$-dimensional vector, $\boldsymbol{\varepsilon}_{t,M}=\boldsymbol{\eta}_{t,M}-\bar{\boldsymbol{\eta}}_{t,M}^{e}$. After taking the conditional expectation, we arrive at the approximate MGF, which is a perturbation of $g_{t,M}(s|\bar{\boldsymbol{\eta}}_{t,M}^{e})$ to account for the randomness in $\boldsymbol{\eta}_{t,M}$. The adjustment of $g_{t,M}(s|\bar{\boldsymbol{\eta}}_{t,M}^{e})$ is simple, because Assumption (ref) implies that $\mathbb{E}_{t}^{\mathbb{Q}}[\boldsymbol{\varepsilon}_{t,M}]=0$ and makes it simple to evaluate $\mathbb{E}_{t}^{\mathbb{Q}}[\boldsymbol{\varepsilon}_{t,M}\boldsymbol{\varepsilon}_{t,M}^{\prime}]=\Sigma_{M}$.
Theorem (ref) clarifies how random variation in $\eta_{t}$ impacts the MGF.\footnote{The integrations in Theorem (ref) can be numerically evaluated through the method in Appendix (ref).} The core principle behind the proposed approximation method is akin to perturbation methods used to find approximate solutions to Schrödinger equations for which standard methods are not applicable. Perturbation methods also starts from an exact solution of a simpler, related problem, from which an approximate solution to the actual problem is deduced. Our approximation is based on a Taylor expansion about an $M$-dimensional vector, where the conditional moments of the terms in the expansion are deduced from Assumption (ref). Next, we evaluate the accuracy of the approximation.
The option pricing formula in Theorem (ref) involves a second-order approximation that accounts for the first two conditional moments of $\text{\ensuremath{\boldsymbol{\eta}_{t,M}}}$.\footnote{If the distribution of $\boldsymbol{\varepsilon}_{t,M}$ is symmetric, (e.g. Gaussian), then it is an approximation to third-order.} A first-order approximation is equivalent to assuming that $\eta_{t}$ will take the path of its conditional expectation, $\bar{\boldsymbol{\eta}}_{t,M}^{e}$. Thus, $g_{t,M}(\cdot|\bar{\boldsymbol{\eta}}_{t,M}^{e})$ is the MGF for the first-order approximation, which we previously labeled “Auxiliary”, because it was used as an intermediate step towards our preferred approximation, $\hat{g}_{t,M}(\cdot)$. Taking $\eta_{t}=\eta_{0}$ to be constant, $\bar{\boldsymbol{\eta}}_{t,M}^{c}=(\log\eta_{0},\ldots,\log\eta_{0})^{\prime}$, is a characteristic of CHNG, and this case can be interpreted as the zero-order “approximation”.
We first consider the risk-neutral density of cumulative returns. The blue solid lines in Figure (ref) present the true risk-neutral densities for cumulative returns over six months for three initial values, $\eta_{0}=0.70$ (left panel), $\eta_{0}=1.11$ (middle panel), and $\eta_{0}=1.70$ (right panel), based on 1 million simulations. The data generation process is based on the structure in Theorem (ref), using the parameter estimates from our empirical analysis of option prices with $\log\eta_{t}\sim\operatorname{AR}(1)$, normally distributed innovations, and $h_{1}$ initialized at its unconditional mean. In this design, the exponential of unconditional mean, $\exp\left(\mathbb{E}\log\eta_{t}\right)=\exp\left(\zeta\right)$, is used as the initial value $\eta_{0}$ in the middle panel.
The red dashed lines are the risk-neutral densities of DHNG implied by $\hat{g}_{t,M}(\cdot)$, the gray dot-dashed lines are based on the auxiliary, $g_{t,M}(\cdot|\bar{\boldsymbol{\eta}}_{t,M}^{e})$, and the yellow dotted line is the risk-neutral density for CHNG, i.e. $g_{t,M}(\cdot|\bar{\boldsymbol{\eta}}_{t,M}^{c})$. DHNG, which is based on the second-order approximation of Theorem (ref), is accurate, whereas the Auxiliary structure fails to match the upper tail of the densities. The CHNG is the worst approximation of the true density in all cases. It is tied with Auxiliary in the middle panel, because the two are identical when $\log\eta_{0}$ is set to have its unconditional mean $\zeta$.
The accuracy of the densities can also be assessed by the moments of cumulative returns under $\mathbb{Q}$. We plot the variance, skewness and kurtosis of multi-period cumulative returns in Figure (ref). The variance is presented in annualized units and scaled by 100, such that 4 corresponds to $\sqrt{0.04}=20\%$ annualized volatility. The $k$-th moment of cumulative returns can be computed by evaluating the $k$-th derivative of the MGF in Theorem (ref) at zero. However, Theorem (ref), presented below, provides a much simpler and more direct method for evaluating moments. Figure (ref) shows that neglecting the random variation in $\eta_{t}$ yields moments that are far from their true values, whereas the second-order approximation is quite accurate, especially for the second and third moment. Both CHNG and Auxiliary are far less accurate, with CHNG being the least accurate, except in the middle column of panels, where $\bar{\boldsymbol{\eta}}_{t,M}^{c}=\bar{\boldsymbol{\eta}}_{t,M}^{e}$.
The expression for the conditional moments in Theorem (ref) is, to the best of our knowledge, a new result, and it is applicable to any dynamic process with a well-defined MGF.
As a third way to assess the accuracy, we compare the (model) implied volatilities with the true implied volatilities for options with different levels of moneyness (Black-Scholes Delta) and days-to-maturity (DTM). Table (ref) reports the percentage errors, $e_{i}\equiv100\times\left[\log{\rm IV}_{i}^{{\rm model}}-\log{\rm IV}_{i}^{{\rm true}}\right]$, where the implied volatilities are deduced from option prices by the Black-Scholes formula. The data generating process is the same as that used in Figures (ref) and (ref). We consider three initial values of $\eta_{0}$ over a two-dimensional grid for Delta and DTM.
Table (ref) shows that the DHNG option pricing formula of Theorem (ref) is far more accurate than those of Auxiliary and CHNG. This is true across all types of options and initial values of $\eta_{0}$. The absolute approximation error (MAE) is just 0.08% on average over all designs in Table (ref), and the average error (ME) is -0.05%. Unsurprisingly, CHNG is the least accurate option pricing model. The assumptions that $\eta_{t}$ is constant leads to underpricing of options when $\eta_{0}$ is small and overpricing when $\eta_{0}$ is large. The option pricing formula for the Auxiliary structure systematically underprices options, and we observe that the pricing error is increasing in DTM. This is consistent with the results in Figure (ref), where the predicted values of $h_{t+1}^{\ast},\ldots,h_{t+M}^{\ast}$ for the Auxiliary structure are downward-biased and increasingly so as $M$ increases. This has implications for empirical estimations under the Auxiliary structure. To rationalize the observed option prices, the Auxiliary structure will need to exaggerate the predicted values of $\log\eta_{t}$, such that Auxiliary is expected to yield larger predictions of $\log\eta_{t}$ than those based on DHNG when estimated with option prices. This is indeed what we find in the supplementary results, see Table (ref) in Appendix (ref).
In this section, we develop an observation-driven model for $\eta_{t}$, which facilitates an implementation of the results from the previous section. In practice, this requires inferring the current value of $\eta_{t}$ and estimate its dynamic model. To this end, we adopt an intuitive observation-driven model for $\eta_{t}$, inspired by the score-driven framework of CrealKoopmanLucas:2013. We use the first-order conditions for minimizing pricing errors to define innovations to $\eta_{t}$, ensuring that $\eta_{t}$ is adjusted to reduce pricing errors on average.\footnote{Alternatively, one could adopt state space approach. This is computationally more complicated without necessarily providing any benefit. Even when the true model is a state space model, the score-driven models are typically found to be competitive, see KoopmanLucasScharth:2016.} The first-order condition of the log-likelihood function is presented in the empirical section. For now, it suffices to express the log-likelihood function in its generic form, $\sum_{t=1}^{T}\ell(R_{t},X_{t}|\mathcal{F}_{t-1})$, where $T$ is the sample size, and we should emphasize that it relies on information from both $\mathbb{P}$ and $\mathbb{Q}$. We factorize $\ell(R_{t},X_{t}|\mathcal{F}_{t-1})=\ell(R_{t}|\mathcal{F}_{t-1})+\ell(X_{t}|R_{t},\mathcal{F}_{t-1})$, where the latter can be expressed as $\ell(X_{t}|\mathcal{G}_{t})$. This likelihood term measures how well the observed derivative prices are explained by the statistical model for returns and the pricing kernel.
A score-driven model updates a parameter in the direction dictated by the first-order conditions of the log-likelihood function, known as the score.\footnote{The score-driven approach is locally optimal in the Kullback-Leibler sense, see BlasquesKoopmanLucas2015.} In our model, the relevant score is $\partial\ell(X_{t}|\mathcal{G}_{t})/\partial\log\eta_{t}$, because $\ell(R_{t}|\mathcal{F}_{t-1})$ does not depend on $\eta_{t}$. The required structure for $\varepsilon_{t}$, as specified in Assumption (ref), motivates the choice
where the normalized score is defined by
Since $s_{t}$ is $\mathcal{F}_{t}$-measurable, we have $\varepsilon_{t+1}\in\mathcal{F}_{t}\subset\mathcal{G}_{t+1}$ as required by Assumption (ref). Additionally, when the score is evaluated at the true parameters we have \[ \mathbb{E}_{t}^{\mathbb{P}}\left(s_{t}\right)=0,\quad\text{and}\qquad\mathrm{var}_{t}^{\mathbb{P}}\left(s_{t}\right)=1, \] such that $\mathbb{E}_{t}^{\mathbb{P}}(\varepsilon_{t+1})=0$ and $\mathrm{var}_{t}^{\mathbb{P}}(\varepsilon_{t+1})=\sigma^{2}$. Note that this definition of the score does not ensure that $\{s_{t}\}$ is a sequence of iid random variables with zero mean and unit variance under $\mathbb{Q}$, as needed by Assumption (ref). This requires additional distributional assumptions about the pricing errors. For instance, we will make assumptions about the pricing errors that imply $s_{t}|\mathcal{G}_{t}\sim iid\ N(0,1)$ under both in $\mathbb{P}$ and $\mathbb{Q}$ measures (see Theorem (ref) below), in which case the score satisfies all the requirements in Assumption (ref). From ((ref)) and the specifications ((ref))-((ref)), it follows that our constructed $\eta_{t}$ is $\mathcal{F}_{t-1}$-measurable, which ensures that the information about current derivative prices are not used to price themselves.
This approach to modeling $\log\eta_{t}$ is analogous to the way the conditional variance is modeled in GARCH models. For instance, the GARCH(1,1) model by bollerslev:86 implies that $h_{t}=c+bh_{t-1}+a(r_{t-1}^{2}-h_{t-1})$, such that $h_{t}\sim\operatorname{AR}(1)$ and changes in $h_{t}$ are driven by discrepancies between squared returns and the conditional variance. The score model invokes a similar self-adjusting property, where $\log\eta_{t}$ is updated in response to derivative prices that indicate that pricing errors can be reduced by revising the value of $\log\eta_{t}$.
In our empirical analysis, we will adopt a score-driven model with an $\operatorname{AR}(1)$ structure for $\log\eta_{t}$. Theorem (ref) can accommodate the case where $\log\eta_{t}\sim\operatorname{ARMA}(p,q)$, whereas more general models for $\log\eta_{t}$, such as exogenous/endogenous explanatory variables and long-memory specifications, would require analogous option pricing results to be established first.
Next, we turn to the log-likelihood function used to estimate model parameters. Our observed data consist of returns, $R_{t}$, and a vector of derivative prices, $X_{t}$. Without loss of generality, we can factorize the log-likelihood function as follows: \[ \ell(R_{1},\ldots,R_{T},X_{1},\ldots,X_{T},\mathcal{F}_{0})=\sum_{t=1}^{T}\ell(R_{t},X_{t}|\mathcal{F}_{t-1})=\sum_{t=1}^{T}\ell(R_{t}|\mathcal{F}_{t-1})+\ell(X_{t}|R_{t},\mathcal{F}_{t-1}). \] This decomposition was also used by ChristoffersenHestonJacobs2013. The log-likelihood function for returns, $\ell(R_{t}|\mathcal{F}_{t-1})$, is that of the HNG model in ((ref))-((ref)), \[ \ell(R_{t}|\mathcal{F}_{t-1})=-\tfrac{1}{2}\left[\log(2\pi)+\log h_{t}+(R_{t}-r-(\lambda-\tfrac{1}{2})h_{t})^{2}/h_{t}\right], \] which assumes that $z_{t}\overset{\mathbb{P}}{\sim}iid\ N(0,1)$. To complete the model, we need to specify the log-likelihood function for the vector of derivative prices, $\ell(X_{t}|R_{t},\mathcal{F}_{t-1})=\ell(X_{t}|\mathcal{G}_{t})$. To this end, we let $X_{t}^{m}\in\mathcal{G}_{t}$ denote the vector of model-based derivative prices and make the following assumption for pricing errors.
The vector of derivative “prices”, $X_{t}$, can be defined in several ways. WangShenJiangHuang2017 measured $X_{t}$ in the unit of volatility (the level of VIX), ChristoffersenHestonJacobs2013 used the Vega-weighted option price, and FeunouOkou2019 used the Black-Scholes implied volatility. In our empirical analysis, we use the logarithmically transformed VIX and the logarithmically transformed Black-Scholes implied volatilities for options, which are natural choices because we specific a model for $\log\eta_{t}$. Moreover, the logarithmic transformation reduces the risk of severe model misspecification, because the distribution of log-volatilities is often well-approximated by a Gaussian distribution, see Andersen2003.
The existing empirical studies that used derivative pricing errors in this manner, all assumed that the pricing errors were uncorrelated, i.e. $\Omega_{N_{t}}=I_{N_{t}}$ where $I_{N_{t}}$ is the $N_{t}$-dimensional identity matrix. This is unrealistic, because contemporaneous pricing errors tend to be positively correlated. We will therefore allow for a non-zero common correlation, which is important for the statistical properties of the score, see Theorem (ref) below. For instance, if we impose $\Omega_{N_{t}}=I_{N_{t}}$ in our empirical analysis with option prices, then it would result in a type of misspecification that induces an upward bias in the estimated variance of $s_{t}$.
The correlation matrix in Assumption (ref), $\Omega_{N_{t}}$, is known as an equicorrelation matrix. Having a common correlation for all pairs of pricing errors is particularly useful in applications where the panel of derivative prices is unbalanced over time. In this case, the dimension of $\Omega_{N_{t}}$ will be time-varying, but this is not problematic if an equicorrelation structure is assumed. A common correlation can be motivated by the assumption that pricing errors, $e_{i,t}=u_{t}+v_{i,t}$, have a common component, $u_{t}$, and uncorrelated idiosyncratic components, $v_{i,t}$, $i=1,\ldots,N_{t}$. This will bring about the structure in Assumption (ref) with $\rho=\sigma_{u}^{2}/\sigma_{e}^{2}$ and $\sigma_{e}^{2}=\sigma_{u}^{2}+\sigma_{v}^{2}$.
Under Assumption (ref), the part of log-likelihood function that relates to derivative prices is \[ \ell(X_{t}|\mathcal{G}_{t})=\ell(e_{t}|\mathcal{G}_{t})=-\tfrac{1}{2}\left(N_{t}\log(2\pi\sigma_{e}^{2})+\log\left|\Omega_{N_{t}}\right|+\sigma_{e}^{-2}e_{t}^{\prime}\Omega_{N_{t}}^{-1}e_{t}\right). \] The model for returns under $\mathbb{P}$ continues to be a time-homogeneous Heston-GARCH model, which does not depend on derivative prices. This is also true in the DHNG model, despite the time variations in the pricing kernel. For this reason, the model parameters can (if needed) be estimated by a two-stage estimation method, where the parameters in the GARCH model is estimated from returns in the first stage, followed by the remaining parameters being estimated in the second stage with derivative prices.\footnote{In our empirical application we estimated the model by maximizing the full log-likelihood function, which was straightforward and did not cause computational issues. If needed, two-stage estimation can be adopted to reduce the computational burden, and the approach to estimation is not uncommon in this setting, see e.g. BroadieChernovJohannes2007, CorsiFusariVecchia2013, ChristoffersenHestonJacobs2013, and MajewskiBormettiCorsi2015.}
To obtain the score, we need to derive $\partial X_{t}^{m}/\partial\log\eta_{t}$, and this must be done separately for the VIX and option prices. These terms are derived in Appendix (ref). Note that if $X_{t}^{m}$ is univariate (the case with a single derivative), then the scaled score simplifies to \[ s_{t}=\frac{1}{\sigma_{e}}{\rm sign}\left(\frac{\partial X_{t}^{m}}{\partial\log\eta_{t}}\right)e_{t}. \] From this expression, it is evident that the score will indicate the direction of change for $\eta_{t}$, which is expected to reduce pricing errors.
Our empirical analysis is based on daily returns for the S&P 500 index, the CBOE VIX, and the panel of SPX option prices based on the S&P 500 index. Our sample period spans 32 years from January 2nd, 1990, to December 31, 2021, with 8,064 trading days. Returns are defined from cum-dividend logarithmic transformed closing prices, which were downloaded from the CRSP of Wharton Research Data Services (WRDS). Daily VIX prices were obtained from the Chicago Board Options Exchange (CBOE) website. The SPX option prices were obtained from two sources. Prices for the first six years (1990--1995) are the so-called Optsum data, which were purchased from CBOE website. Option prices for the remaining 26 years are OptionMetrics data downloaded from the WRDS database.
Option prices were primarily preprocessed following ChristoffersenFeunouJacobsMeddahi2014 and BakshiCaoChen1997. Specifically, we include out-of-the-money put and call option prices with positive trading volume and maturities between two weeks and six months. Options with missing implied volatilities or prices below one dollar are excluded. Put option prices are converted to call options using the put-call parity. Very deep out-of-the-money options with deltas larger than 0.85 or less than 0.15 are discarded. Much of the existing literature uses weekly data, and a panel of option prices is typically sampled on Wednesdays, because liquidity tends to be highest on Wednesdays. However, in our analysis, we use daily option prices, because a daily score, $s_{t}$, is needed to update $\eta_{t}$. The number of available options has grown rapidly over the sample period, both in terms of available maturities and the range of moneyness at each maturity.\footnote{The number of available option prices has increased almost 50-fold over our sample period, initially from about 38 daily option prices to well over 1,700.} From the pool of available options, we select up to six options per trading day. The inclusion criteria are as follows: we first determine the most liquid option for each maturity on day $t$, as measured by daily trading volume. We sort these options by maturity in ascending order and index these by $j=1,\ldots,\tilde{N}_{t}$, where $\tilde{N}_{t}$ is the number of distinct maturities on day $t$. From this set, we include all options if $\tilde{N}_{t}\leq6$; the first six options if $7\leq\tilde{N}_{t}\leq10$; options $\{1,3,5,7,9,11\}$ if $11\leq\tilde{N}_{t}\leq15$; and options $\{1,4,7,10,13,16\}$ if $16\leq$ $\tilde{N}_{t}\leq20$; and so forth. This resulting set of options will be representative for the range of available maturities, and we have $N_{t}=6$ on most days. The total number of option prices in our full sample period is 37,152.
Table (ref) contains descriptive statistics for our S&P 500 returns and the CBOE VIX in Panel A, and option prices in Panel B. As expected, the sample average of the VIX (19.48%) is larger than the standard deviation for annualized returns (18.11%). This difference reflects the (average) negative volatility risk premium. The S&P 500 returns exhibit slight negative skewness and a high degree of kurtosis, whereas the VIX has positive skewness and slightly lower kurtosis. We also report summary statistics for option prices partitioned by moneyness, maturity, and the contemporaneous level of VIX. Options with deltas below 0.5 are out-of-the-money call options, and options with deltas above 0.5 are based on out-of-the-money put options. Deep out-of-the-money put options (i.e., deltas greater than 0.7) are expensive relative to out-of-the-money call options, reflecting the well-known volatility smirk. The implied volatility has a relatively flat term structure with respect to time to maturity. Unsurprisingly, the implied volatility increases with the VIX level, as shown at the bottom of Table (ref), where option prices are partitioned by the contemporaneous level of the VIX.
We estimate the model with both constant and time-varying variance risk aversion, CHNG and DHNG, respectively.\footnote{The corresponding results for the Auxiliary structure are presented in the Appendix (ref).} Both specifications are estimated using two types of derivative prices, VIX data and the panel of option prices. Parameters are estimated by maximizing the joint log-likelihood function for the full sample period from January 1990 to December 2021. Each column reports the parameter estimates for the specification listed in the first row of Table (ref). The type of derivatives, VIX or option prices, used in the estimation is indicated with {[}VIX{]} and {[}Opt{]}, respectively. Robust standard errors are give in parentheses below the estimates.\footnote{Following ChristoffersenHestonJacobs2013, we impose $\omega$ = 0 in estimation when the non-negativity constraint, which ensures positive variances, is binding.} We also report the implied persistence of volatility under both $\mathbb{P}$ and $\mathbb{Q}$. These are given by $\pi^{\ensuremath{\mathbb{P}}}=\beta+\alpha\gamma^{2}$ and $\pi^{\ensuremath{\mathbb{Q}}}=\mathbb{E}^{\mathbb{Q}}[\beta_{t}^{*}+\alpha_{t}^{*}\gamma_{t-1}^{*2}]$, respectively, where the latter simplifies to $\pi^{\ensuremath{\mathbb{Q}}}=\beta^{*}+\alpha^{*}\gamma^{*2}$ for the CHNG model with constant parameters. We also report the different terms of the log-likelihood (for returns and different sets of derivative prices). Some of these likelihood terms can be interpreted as pseudo out-of-sample log-likelihood values, indicated in italic. For instance, the CHNG{[}VIX{]} model is estimated using returns and the VIX, but the estimated model also yields model-based option prices that can be compared with actual option prices. The reported pseudo log-likelihood for option prices is evaluated with the resulting option pricing errors that are implicitly used to obtain estimates of $\rho$ and $\sigma_{e}$ for option prices. Similarly, we evaluate the log-likelihood for VIX pricing errors for the specifications estimated with option prices. In this case, we compute the implied estimate of $\sigma_{e}$ for VIX prices. The largest log-likelihood within each row is highlighted in bold.
There are several interesting observations to be made from Table (ref). First, the volatility process is found to be highly persistent across all specifications, and the persistence is larger under $\mathbb{Q}$ than under $\mathbb{P}$, which is consistent with the existing literature. Second, the estimate of the equity risk premium, $\lambda$, is positive and significant in all specification. Third, the VRR, $\eta$, is significantly larger than one in both CHNG specifications. Similarly, for the DHNG model the expected value of $\log\eta_{t}$, $\zeta$, is significantly positive in both DHNG specifications. Their corresponding expected $\eta_{t}$ (denoted $\mathbb{E}\eta$) inferred from the AR(1) model for $\log\eta_{t}$ are both notably larger than one. This implies that the risk-neutral volatility, $h^{\ast}$, is (on average) larger than the physical volatility, $h$. Fourth, the parameter that defines the leverage effects, $\gamma$, is also found to be significant across all specification. Fifth, for the DHNG specification we note that, $\eta_{t}$, is highly persistent, as $\varphi$ is estimated to be close to one. In fact, it is estimated to be more persistent than $h_{t}$ and $h_{t}^{\ast}$ across all DHNG specifications. Sixth, the coefficient, $\sigma$, which measures the impact of $s_{t-1}$ on $\eta_{t}$, is estimated to be positive and significant.\footnote{In principle, $\sigma$ could be negative, but $\operatorname{var}(\sigma s_{t})=\sigma^{2}$ is non-negative.} The implied unconditional variance for $\eta_{t}$ is similar for the two DHNG specifications, $\mathrm{var}(\eta_{t})=0.29$ when estimated with VIX and $\mathrm{var}(\eta_{t})=0.23$ when estimated with option prices.\footnote{The variance for $\eta_{t}$ is computed from AR(1) parameters for $\log\eta_{t}$ and the assumptions imply that the unconditional distribution for $\eta_{t}$ is log-normal.} Seventh, all specifications yield similar likelihood values for the returns. This is to be expected because they all rely on the same Heston-Nandi GARCH model for returns under physical measure.
Eighth, the key difference between the CHNG and DHNG models is the enhanced flexibility in DHNG's pricing kernel. This generalization leads to large improvements in the likelihood for derivative prices. The reason is that the adaptive pricing kernel with time-varying variance risk aversion yields model-implied derivative prices that are much closer to observed prices, and the substantial reduction in the pricing errors translates to much higher values of the log-likelihood for derivative prices. Ninth, estimating the models with VIX or option prices results in some differences across the estimated parameters. This is to be expected, as the vector of option prices contains more information about the distribution of future returns than the VIX. For instance, the leverage parameter, $\gamma$, is estimated to be larger with option prices than with the VIX. Tenth, the correlation among option pricing errors, $\rho$, is estimated to be positive and significant for both models. The estimate has the staggering large value of $82.2\%$ for the CHNG model and a more moderate value of $8.7\%$ for DHNG model. Ignoring this correlation (assuming it to be zero) would greatly underestimate the condition variance of $\text{\ensuremath{\nabla_{t}}}$, which makes the variance of $s_{t}$ larger than one.
Figure (ref) presents the estimated daily time series of the VRR, $\eta_{t}=h_{t+1}^{\ast}/h_{t+1}$, based on the VIX (blue line) and option prices (red line). The two have a high degree of comovement and, as we discussed earlier, the unconditional variance of the two series is very similar. Importantly, the VRR occasionally falls below one, as is seen during the years 1993--1995, 2004--2007, and around 2017. A value below one ($\eta_{t}<1$) indicates that investors have an increased appetite for variance risk, whereas a large value of $\eta$ implies that investors demand additional compensation for variance risk. The latter was particularly the case during financial crises, such as the immediate aftermath of the Lehman collapse and the recent COVID-19 pandemic. The observed conditional variance was unusually high during these episodes, and the large value of $\eta_{t}$ implies that $h_{t}^{\ast}$ increased to far higher levels and was, briefly, more than three times larger than $h_{t}$.
The level of $\eta_{t}$ based on the VIX tend to be slightly larger than that based on option prices. This discrepancy can be attributed to the fact that the VIX is based on options with 30 days to maturity, whereas the options have maturities up to 180 days. A possible explanation is that short-term investors demand larger compensation for variance risk than long-term investors, which would be consistent with the findings in EisenbachSchmalz2016 and AndriesEisenbachSchmalz2018.
We evaluate the performance of the models in terms of their ability to accurately price the VIX and options. We report the root mean squared errors (RMSE) for the logarithm of VIX and the logarithmically transformed implied volatilities.\footnote{As a robustness check, we also computed the RMSE for the level of VIX and implied volatilities. These results are reported in Appendix (ref).} The RMSE multiplied by 100 can be interpreted as the (absolute) relative pricing errors in percent. This loss function is coherent with our log-likelihood function, where the term that involves derivative prices is given in Assumption (ref). A similar loss function was adopted in Ornthanalai2014, who used the relative implied volatility to prevent days with high volatility from being given a disproportionately large weight in the comparisons. For each of the specifications listed in Table (ref), we report their (in-sample) RMSEs in Table (ref).
We present the in-sample RMSEs for the VIX in Panel A. These are defined by
where $\mathrm{VIX}_{t}^{m}$ is the model-based quantity and $\mathrm{VIX}_{t}$ is the observed market VIX value. Panel A of Table (ref) shows that a time-varying VRR, $\eta_{t}$, results in substantially smaller average pricing errors, relative to a constant $\eta$, which is a characteristic of CHNG. The improved VIX pricing is substantial and impressive. The most accurate VIX pricing is achieved with by the DHNG model, when it is estimated with VIX data. This reduces the RMSE by a factor of four relative to both CHNG specifications. Even the DHNG model that is estimated solely with option prices, has a much smaller RMSE for VIX prices than the CHNG specification that uses VIX prices as part of the objective in the estimation. This is impressive, because the DHNG model estimated with option prices is based on an objective that simultaneously seeks to price options with maturities ranging from one to six months. This entails a trade-off across maturities, unlike models estimated with solely with VIX, that specifically targets the one-month maturity.
Panel B of Table (ref) reports the in-sample performance for option pricing. We follow the literature and convert option prices to their corresponding Black-Scholes implied volatilities. We compare the model-based implied volatility, $\mathrm{IV}^{m}$, with the market-based implied volatility, $\mathrm{IV}$, where the latter is defined by the observed option price. The resulting RMSE,
is reported in the first row of Panel B. We also compute the ${\rm RMSE_{IV}}$ for options partitioned by moneyness, time to maturity, and the contemporaneous VIX level. The resulting RMSEs can be used to identify shortcomings in a model, such as its inability to generate sufficient leverage effect, capture the dynamic properties, and generate a proper variance risk premium.
The DHNG specifications clearly dominate the CHNG specifications in terms of option pricing. The average option pricing errors are substantially smaller for DHNG models than for CHNG models, and the smallest RMSE is obtained by the DHNG model estimated with option prices. Its RMSE is less than half that of any of the CHNG specifications. Impressively, the DHNG specifications uniformly dominate all CHNG specifications for both VIX and option pricing across all partitions. This further supports the advantage of having a dynamic variance risk aversion in the model. Note that the largest gains in pricing accuracy are found during periods with low volatility and for out-of-the-money call options, which tend to be options with low variance risk premia.
Figure (ref) displays the model-based VIX and implied volatility (red lines) alongside the corresponding market-based quantities (blue lines). The upper panels show the model-based VIX for CHNG and DHNG, with both models estimated using VIX data. The lower panels present the daily average of model-based implied volatilities for CHNG and DHNG, estimated with option prices. For CHNG, we see larger discrepancies between model-based and market-based quantities as evident in Figure (ref). Specifically, CHNG underprices when volatility is high and overprices when volatility is low, whereas no such patterns are observed for DHNG.
Interestingly, ChristoffersenHestonJacobs2013 also estimated an ad-hoc model where the ratio of volatility under $\mathbb{P}$ and $\mathbb{Q}$ is not held constant. This model was used as a benchmark for the CHNG model. The ad-hoc model consisted of two separate HNGs: one fitted to returns under $\mathbb{P}$, and one that was estimated with option prices under $\mathbb{Q}$. This model is referred to as ad-hoc, because it makes no attempt to uncover a pricing kernel that explains the differences between $\mathbb{P}$ and $\mathbb{Q}$. Their ad-hoc model did not improve option pricing accuracy relative to CHNG. Theorem (ref) provides a theoretical explanation for the poor performance of their ad-hoc model, as it shows that any time variation in $\xi_{t}$ (or, equivalently, in $\eta_{t})$ makes it impossible to have HNGs with time-invariant parameters under both $\mathbb{P}$ and $\mathbb{Q}$ measures.\footnote{The ad-hoc model is therefore internally inconsistent. Additionally, the ad-hoc model only leads to a minuscule improvement in the empirical fit, as the log-likelihood for option prices improves by only about 0.1 units.}
In the next section, we evaluate if the substantial in-sample improvements in derivatives pricing also holds out-of-sample.
In this section, we shift our focus to out-of-sample comparisons. Since the DHNG model nests the CHNG model as a special case, it naturally achieves a higher in-sample log-likelihood. Out-of-sample comparisons are crucial to determine whether the improved in-sample performance is due to overfitting or reflects the true quality of the DHNG model. We investigate whether the DHNG model also produces more accurate derivative prices in out-of-sample comparisons. The full dataset is divided into two subsamples: data from the years 1990--2007 (in-sample) are used to estimate each of the four specifications, while data from the period 2008--2021 (out-of-sample) are used to evaluate the estimated models. Similar to the in-sample comparisons, we assess and compare the models based on their root mean squared pricing errors.
We find that the out-of-sample performance of the DHNG model is just as impressive as its in-sample performance. The most accurate VIX pricing is once again achieved with the DHNG model estimated using VIX data, while the best option pricing is similarly achieved by the DHNG model estimated with option prices. Although these two specifications have similar point estimates and produce very similar paths for $\eta_{t}$, the small differences between the two estimated model do affect derivative pricing. Overall, the out-of-sample results are very encouraging and align closely with the in-sample findings.
Next, we turn our attention to the autocorrelations of the pricing errors.
The autocorrelations of pricing errors offer valuable insights into the improved derivative pricing achieved by the DHNG model. The estimated DHNG models reveal substantial time variation in $\eta_{t}$. In contrast, the CHNG model, which relies on a constant $\eta$, tends to produce positive pricing errors when $\eta_{t}$ is small and negative pricing errors when $\eta_{t}$ is large. Given the high persistence of $\eta_{t}$, this is expected to induce autocorrelation in the pricing errors for the CHNG model, which is indeed what we observe.
Figure (ref) displays the autocorrelation functions (ACF) for VIX pricing errors in the upper panels and ACFs for average option pricing errors in the lower panels, covering the full sample period. The very high and persistent ACFs for CHNG indicate that a constant $\eta$ induces a high degree of predictability in the pricing errors. In stark contrast, the results for DHNG, shown in the right panels, reveal autocorrelations that are substantially closer to zero. The horizontal lines in each plot represent two standard deviations from zero. For DHNG, the autocorrelations are largely insignificant, with the exception of the first few autocorrelations in the lower-right panel.
The remarkable reduction in these autocorrelations is a clear demonstration of the adaptive nature of score-driven models. The DHNG model continuously adjusts $\eta_{t}$ to minimize pricing errors, using the latest first-order conditions (and curvature) to determine the direction and magnitude of each adjustment. As a result, the first-order conditions are more consistently satisfied throughout the sample. In contrast, models with static parameters, such as CHNG, focus only on the average first-order condition. This approach allows for large pricing errors, as long as the positive errors are offset by negative errors. This explains why the CHNG model exhibits much larger RMSE and significant autocorrelations in its pricing errors.
The adaptive nature of the DHNG model, where parameters are adjusted according to the first-order conditions, enhances the robustness of the derivative pricing formulae against model misspecification. This robustness arises because $\eta_{t}$ is automatically adjusted to compensate for any misspecification that would otherwise lead to increased pricing errors. While this adaptive feature is a strength of the model, it also reveals a potential drawback: misspecification can undermine the interpretation of $\eta_{t}$ as the variance risk ratio.
Expected volatility under $\mathbb{Q}$ can be measured accurately using observed derivative prices. However, the situation differs under $\mathbb{P}$, where we must rely on a model to estimate expected volatility. If the model is misspecified, the resulting model-based expected volatility may be biased. One common form of misspecification is parameter instability. Sichert:2022 demonstrates that a GARCH model with structural breaks can resolve the pricing kernel puzzle related to the inverted U-shape. This explanation is further explored in the empirical analysis by TongHansenHuang2022, where they estimate a Markov switching Realized GARCH model with two states. Their findings show significant U-shaped pricing kernels in the high-volatility regime, which aligns with our results. However, the positive variance risk premium (associated with the inverted U-shaped pricing kernel) they observed in the low-volatility regime is close to zero and statistically insignificant.
If the model-based expected volatility is biased, this bias will affect $\eta_{t}$, meaning $\eta_{t}$ may not accurately reflect the true variance risk ratio. For example, a small $\eta_{t}$ might not indicate an investor's appetite for variance risk; instead, it could be an artifact of the Heston-Nandi model overestimating future volatility relative to actual expectations under $\mathbb{P}$.
As a robustness check, we computed an alternative VRR that does not rely on the Heston-Nandi GARCH model. This alternative measure is based on model-free realized variances and uses a simple AR(1) model to define expected variance under $\mathbb{P}$. The resulting time series of this alternative VRR, reported in Appendix (ref), closely resembles that of $\eta_{t}$ in Figure (ref). This similarity suggests that the observed fluctuations in $\eta_{t}$ are not driven by a flaw specific to the GARCH model ((ref))-((ref)). However, an oversimplified description of how expected volatility evolves can be the cause. The AR(1) model and standard GARCH models both produce volatility forecasts with mean-reverting characteristics.
Regarding misspecification as a potential source of time variation in $\eta_{t}$, we generated scatterplots of pricing errors against $\eta_{t}$. These scatterplots do not suggest a systematic relationship between pricing errors and $\eta_{t}$; see Figure (ref) in Appendix (ref).
Time variation in volatility risk aversion is the key characteristic of the new pricing kernel. In this section, we relate the time variation in $\eta_{t}$ to economic fundamentals and several other variables that are widely used in the asset pricing literature.
We focus on seven core measures of uncertainty, disagreement, and sentiment. The first set of variables is related to investor disagreements. Following BollerslevLiXue2018, we use two types of proxies for disagreement. The first concerns macroeconomic fundamentals, for which we use the forecast dispersion for both the unemployment rate and GDP growth from the Survey of Professional Forecasters.\footnote{BollerslevLiXue2018 only use the forecast dispersion for the one-quarter-ahead unemployment rate.} The second proxy is disagreement about economic policy, for which we adopt the Economic Policy Uncertainty index by BakerBloomDavis2016.
We also include the variance risk premium (VRP) by CarrWu2008, which GrithHardleKratschmer2017 suggest serves as a proxy for market uncertainty. Additionally, we incorporate three uncertainty measures from existing literature: the Economic Uncertainty Index by BaliBrownCaglayan2014, the Survey-based Uncertainty Index by OzturkSheng2018;\footnote{This index is constructed from the Consensus Forecasts publication by Consensus Economics Inc. The authors provide data up to October 2021, and we fill in the values for the last two months using a simple predictive regression on all our explanatory variables.} and the well-known Sentiment variable by BakerWurgler2006.
In addition to these seven core variables, we include the economic variables previously used in BakerWurgler2006 and WelchGoyal2007, which we label as “control variables”. Most variables are available at a monthly frequency. Therefore, we take the monthly average of $\eta_{t}$ and regress the logarithmically transformed average on the variables listed above, as well as subsets of these variables.
The full-sample monthly regression results are presented in Table (ref). The seven core variables are listed at the top of the table, followed by the six control variables that contributed the most to explaining variation in $\eta_{t}$. There are eleven additional control variables, labeled “Other Controls”, which were largely insignificant.\footnote{We follow {WelchGoyal2007 in the definitions of these variables. }For instance, the yield spread is the difference between BAA- and AAA-rated corporate bond yields, and the default return spread is the difference between the return on long-term corporate bonds and government bonds. See {WelchGoyal2007 for more details.}} To conserve space, the results for all control variables (along with their labels) are presented in the Appendix (ref).
The first column in Table (ref) corresponds to the specification that includes only the control variables, which explain 70.9% of the variation in the monthly average of $\eta$. The six most significant control variables are stock market volatility, the dividend-price ratio, and four variables related to interest rates and credit risk.
Columns two through eight present the results for specifications that include a single core variable in addition to the control variables. Every core variable is significant in these regressions, and the signs of their estimated coefficients align with expectations. The sentiment variable has a negative coefficient, indicating that low sentiment is associated with high volatility risk aversion. The five uncertainty-related variables all have positive coefficients, suggesting that high uncertainty is associated with high volatility aversion.
Finally, the variance risk premium is an empirical measure of the difference between volatility under $\mathbb{Q}$ and $\mathbb{P}$. As expected, this variable has a positive coefficient and contributes the most to the increase in $R^{2}$. The last column shows the results for a “kitchen sink regression” that includes all explanatory variables. Most of the variables in Table (ref) remain significant in this regression, and the $R^{2}$ increases to nearly 87%. However, three core variables--Forecast Dispersion of Unemployment, Forecast Dispersion of GDP Growth, and the Economic Uncertainty Index--become insignificant in the kitchen sink regression, possibly due to collinearity, as the first two and the last two variables exhibit the highest correlations among all core variables (88.6% and 54.8%), respectively.
We have introduced a novel option pricing model based on a flexible pricing kernel with time-varying risk aversion. This model, denoted DHNG, offers an elegant and tractable framework when combined with the Heston-Nandi GARCH model. The variance risk ratio, $\eta_{t}$, emerges as the fundamental variable that captures all time variation in the pricing kernel. This ratio is functionally linked to the variance risk aversion parameter, which defines the curvature of the pricing kernel, and the framework can generate the empirically observed shapes of the pricing kernel. Additionally, $\eta_{t}$ is closely related to a range of well-known measures of sentiment, disagreement, uncertainty, and other key economic variables.
Under the assumptions characterizing the DHNG model, we derived a closed-form expression for the VIX and an approximate option pricing formula. This formula is based on a novel approximation method that can yield affine-type pricing formulae for non-affine models. The method has proven to be highly accurate in simulation studies. Our comprehensive empirical analysis demonstrated that DHNG significantly reduces derivative pricing errors. Specifically, DHNG typically reduces the root mean square error of pricing errors by more than 50% compared to the CHNG models where $\eta_{t}$ is constant. This substantial reduction in pricing errors is observed both in-sample and out-of-sample.
The time-varying pricing kernel can also be combined with other GARCH models, such as EGARCH, GJR-GARCH, and NGARCH, see Nelson91, Glosten1993, and Engle_Ng_1993, as well as volatility models that utilize realized measures of volatility, such as Realized GARCH models by HansenHuangShek:2012 and HansenHuang:2016, the GARV model by ChristoffersenFeunouJacobsMeddahi2014, and the LHARG, see MajewskiBormettiCorsi2015. For some of these models, such as GARV and LHARG, it is possible to define an Auxiliary models with an affine MGF, and it may be possible to establish closed-form option pricing expressions with the approximation method we have proposed in this paper. Other approximation methods, such as that by DuanGauthierSimonato1999, may be applicable to the non-affine models.
Analytical results are generally easier to establish in continuous-time models, such as those by Heston1993 and CoxIngersollRoss85, which partly explains their popularity among practitioners. However, the discrete-time HNG model also offers an analytical option pricing formula, and we successfully derived an analytical option pricing formula for our discrete-time model using the new approximation method. In contrast, obtaining analytical results in a continuous-time framework with a time-varying VRR would be highly challenging due to the non-affine structure. Additionally, the estimation process would be further complicated by the presence of the latent variable $\eta_{t}$. The key advantage of discrete-time GARCH models over stochastic volatility models lies in their relative ease of estimation. This is particularly beneficial in empirical studies involving a large number of options. Our empirical implementation of the DHNG model retains a straightforward analytical log-likelihood function, which greatly simplifies the estimation process.