Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
47,117 characters · 9 sections · 46 citation commands
Method of Moments Estimation for Affine Stochastic Volatility Models
\affil[a]{School of Management, Fudan University, 670 Guoshun Road, Shanghai 200433, China.} \affil[b]{School of Management, Shandong University, 27 Shanda Nanlu, Jinan 250100, China.}
keywords: affine jump diffusion; stochastic volatility; Heston model; method of moments.\% MOS subject classification: 62F12; 62M05; 60J25.
Volatility plays a crucial role in the pricing of options and other derivative securities. Empirical evidence from financial markets, such as volatility clustering, the dependence between increments, and volatility smiles, shows that the assumption of constant volatility used in the Black-Scholes European options pricing model is inappropriate (see fama1965behavior). Therefore, several stochastic volatility (SV) models, such as heston1993closed, bates1996jumps and barndorff2001non, have been proposed to better capture the time-varying features of volatility. However, parameter estimation for these SV models poses a formidable computational challenge for practical implementation. Likelihood-based inference and moment-based inference are two major types of estimation methods. Our work in this paper offers a new and direct method of moments (MM) estimation for affine SV models.
There is a large body of literature on likelihood-based inference, in which Markov chain Monte Carlo (MCMC) and maximum likelihood estimation (MLE) are the two main approaches used. An MCMC model assumes a prior distribution of parameters, and then iteratively collects samples in a Bayesian framework until certain convergence is achieved eraker2001mcmc,jacquier2002bayesian,eraker2003impact,roberts2004bayesian,griffin2006inference,jasra2011inference. In some cases, this approach may become computationally prohibitive because of its slow convergence to a Markov chain equilibrium distribution. On the other hand, the MLE approach adopts a frequentist perspective and often requires a closed-form expression of the state transition density function; however, this is hard to obtain in SV models, let alone optimization. Hence, the MLE approach often involves some sort of approximation or simulation methods. Two prevail schemes are simulated MLE (SMLE) (see, e.g., durham2002numerical) and quasi-MLE (QMLE) (see, e.g., ruiz1994quasi, harvey1996estimation, and Feunou2018risk). hurn2014estimating design an MLE estimate using a particle filter to approximate the likelihood. ai2007maximum propose a closed-form approximation scheme based on the extracted volatility from observed option prices. andersen2015parametric employ optimization tools such as penalized nonlinear least squares to develop a new parametric estimation procedure. A gradient-based simulated MLE estimate is proposed in peng2014gradient and peng2016gradient. Nevertheless, approximations used in these MLE methods require enormous computational effort, and many assumptions may be difficult to satisfy in practice.
Modern financial markets give rise to numerous real-time decision-making instances (e.g. high-frequency trading) that request computationally fast and less restrictive estimators. The moment-based method is a case in point but has yet to be recognized fully. A few important contributions include the simulated MM (SMM); see duffie1993simulated), efficient MM (EMM); see Bansal1994 and gallant1996moments, and generalized MM (GMM); see singleton2001estimation, bollerslev2002estimating, jiang2002estimation, and chacko2003spectral. SMM simulates sequences from the target models and estimates parameters by matching simulated data moments with actual data moments numerically. GMM can be applied when we have more sample moments than the number of parameters. EMM combines the efficiency of MLE with the flexibility of GMM, which is a variant of SMM. In general, the major issue of the moment-based method is the so-called statistical inefficiency in the sense that the higher order moments are used, the greater likelihood of estimation bias occurs. Fortunately, we can mitigate this problem by using better computational techniques (e.g., see wu2019moment and wu2022moment). For instance, yang2020method recently develop the first closed-form MM parameter estimators for a class of L\'{e}vy-driven Ornstein-Uhlenbeck stochastic volatility models.
In this paper, we consider affine SV models, and develop explicit MM estimators for the parameters of these models. To this end, we first derive the analytical formulas for the required moments and covariances of the affine SV models and then use them to develop our MM estimators. The MM estimators can be expressed explicitly in terms of the sample moments and sample covariances of the observed asset prices. In fact, we have established a recursive equation for deriving closed-form expressions for moments of any order, while our MM estimators only require those moments of relatively low order. We also prove that the large-sample behaviors of our MM estimators satisfy the central limit theorem and derive the explicit formulas for the asymptotic covariance matrix. In addition, numerical experiments are provided to test the efficiency of our method. The advantage of our estimators is that they are simple and easy to implement in practice and do not require the availability of high-frequency data or option price data.
The rest of this paper is organized as follows. In Section 2, we present the affine SV model and its baseline SV model, then calculate the moments and covariances of the returns of the asset price. In Section 3, we derive our MM estimators and establish the central limit theorem. Some numerical experiments are provided in Section 4. Section 5 concludes this paper. The appendices contain some of the detailed calculations which are omitted in the main body of the paper and also some extensions of the baseline affine SV model.
We first delineate the general framework of affine SV models. Subsequently, we employ the Heston model, a prototypical example of an affine SV model, to illustrate the derivation of closed-form expressions for moments and covariances of any order. We then extend this moment derivation methodology to a broader array of affine SV models, encompassing those with jump components in either the price or variance processes, as well as models incorporating multiple latent volatility factors. These extensions are detailed in the appendices.
Consider $s(t)$ as the price of a certain asset at time $t$. The general affine SV model explored in this study is governed by the following stochastic differential equations (SDEs):
where $\boldsymbol{v}(t)$ denotes a vector of latent variance factors, $w^s(t)$ and $\boldsymbol{w^v}(t)$ are potentially correlated Wiener processes, $z^s(t)$ and $\boldsymbol{z^v}(t)$ are possibly correlated compound Poisson processes. The drift functions \textemdash $\mu(\cdot), \boldsymbol{\mu_v}(\cdot)$ and the variance $\sigma^2(\cdot)$ and the covariance matrix $\boldsymbol{\sigma_v}^T(\cdot)\boldsymbol{\sigma_v}(\cdot)$ \textemdash all have affine dependence on the hidden state vector $\boldsymbol{v}(t)$ duffie2000transform. It is assumed that the drifts, diffusions and compound Poisson processes are well-behaved enough to ensure the unique strong solution of the SDEs.
This general model includes several common SV models as special cases, such as the Heston model, the stochastic volatility with jumps (SVJ) model, the stochastic volatility with independent jumps (SVIJ) model, the stochastic volatility with contemporaneous jumps (SVCJ) model, and even the BNS model. Within the framework of the Heston model heston1993closed, the vector $\boldsymbol{v}(t)$ simplifies to a single variance factor $v(t)$, with drift and diffusion functions defined as:
and without any jump components $z^s(t)$ and $z^v(t)$ in both the price and variance equations. The SVJ model, as described by bates1996jumps, extends the Heston framework by including jumps in the price process. The SVIJ and SVCJ models, detailed in eraker2003impact, further incorporate jumps into the variance process, either independently or contemporaneously with price jumps. An additional variant is the multi-factor Heston model, which characterizes the latent variance as a vector of square-root diffusions, i.e., $\boldsymbol{v}(t) = (v_1(t),v_2(t))^T$, and
In contrast, the BNS model, proposed by barndorff2001non posits a one-dimensional non-Gaussian Ornstein-Uhlenbeck (OU) process for the latent variance, driven solely by a compound Poisson process $z^v(t)$, devoid of a diffusion component. The drift and diffusion functions are defined as:
and there is no jump components in the price process.
It is notable that these affine SV models do not possess closed-form solution for their transition density functions. However, for the general affine SV model, closed-form expressions for their conditional Characteristic Functions (CFs) can be derived duffie2000transform, singleton2001estimation, chacko2003spectral. jiang2002estimation extended this work by developing a closed-form CF for a simplified version of the Heston model.
Our work in this paper aims to extend the analytical tractability of these models to include closed-form formulas for moments and covariances of any (positive integer) order, which can then be used to construct moment estimators for the affine SV model.
Since the Heston model is the baseline affine SV model, we will focus on its moment calculations which will lay the foundation for the more complex moment computations of other affine SV models. The methodology developed for the Heston model can be generalized and applied to estimate other affine SV models. The detailed moment-derivation procedures for these other affine SV models are provided in the appendices.
We use the following Heston model heston1993closed as our baseline affine SV model :
where $v(t)$ is a scalar affine diffusion, and $w^s(t)$ and $w^v(t)$ are two Wiener processes with correlation $\rho$. Assume that the initial values $s(0)$ and $v(0)$ are independent of each other and also independent of $w^s(t)$ and $w^v(t)$. The variance process (ref) is a Cox-Ingersoll-Ross process cox1985theory, which is also called square-root diffusion. The parameters $\theta>0$, $k>0$, and $\sigma_v$ determine the long-run mean, the mean reversion velocity, and the volatility of the variance process $v(t)$, respectively, and satisfy the condition $\sigma_v^2 \le 2k\theta$.
Based on It\^{o} formula, the log price process can be written as:
Though modeled as a continuous-time process, the asset price is observed at discrete-time instances. Assume we have observations of $s(t)$ at discrete-time $ih$ ($i=0,1,\cdots,N$), and let $s_i \triangleq s(ih)$. Similarly, let $v_i \triangleq v(ih)$, however, we should note that $v_i$ is not observable. Finally, we define the return $y_i$ as $$ y_i \triangleq \log s_i - \log s_{i-1} . $$
We can decompose $w^s(t)$ as
where $w(t)$ is another Wiener process independent of $w^v(t)$. For notational simplicity, we define:
and
$IV_{s,t}$, $IV_i$, and $IV_{i,t}$ defined above are also referred as integrated variance (volatility). We can then express $y_i$ as
Based on ((ref)), the variance process $v(t)$ can be re-written as:
which is a Markov process that has a steady-state gamma distribution with mean $\theta$ and variance $\theta \sigma_v^2/(2k)$, e.g., see cox1985theory. Without loss of generality, throughout this paper we assume that $v(0)$ is distributed according to the steady-state distribution of $v(t)$, which implies that $v(t)$ is strictly stationary and ergodic overbeck1997. However, all the results we shall derive in the rest of this paper, including the formulas for all the moments/covariances and the MM estimators, remain the same for any non-negative $v(0)$. Since $v(t)$ is stationary, we have
for $m=1,2,\ldots$ and $t>0$, and
for $s<t$. Furthermore, we have
where $\tilde{h} \triangleq (1-e^{-kh})/k$. The above formulas will be useful in calculating the moments and covariances of $y_i$ in the next section. Hereafter, we will use a generic subscript $n$ to denote variance and return (i.e., $v_n, y_n$) in their steady states.
There are five parameters in the Heston model, therefore, to estimate them we need to obtain at least five moments/covariances of $y_n$. In this section, we show specifically how $\operatorname*{\mathbb{E}}[y_n]$, $\operatorname{var}(y_n)$, $\operatorname{cov}(y_n,y_{n+1})$, $\operatorname{cov}(y_n,y_{n+2})$, and $\operatorname{cov}(y_n^2,y_{n+1})$ can be calculated and derive explicit formulas for them in terms of the five parameters involved in the Heston model, which can then be used to estimate these parameters.
Based on ((ref)), together with $\operatorname*{\mathbb{E}}[I_n] = 0$, $\operatorname*{\mathbb{E}}[I^*_n] = 0$, and $\operatorname*{\mathbb{E}}[IV_n] = \theta h$, we have
To calculate $\operatorname{var}(y_n)$, we first note $\operatorname{cov}(IV_n,I_n^*) = 0$ , $\operatorname{cov}(v_{n-1},I_n^*) = 0$, $\operatorname{cov}(\ie_n,I_n^*) = 0$, and $\operatorname{cov}(I_n,I_n^*) = 0$. Hence,
Let us take $\operatorname{var}(IV_n)$ as an example, it can be further expanded as
and we have
Terms $\operatorname{cov}(IV_n,I_n),\operatorname{var}(I_n)$ and $\operatorname{var}(I_n^*)$ can be computed similarly. In summary, we have
For $\operatorname{cov}(y_n,y_{n+1})$, $\operatorname{cov}(y_n,y_{n+2})$, we have
In fact, ((ref)) can be generalized as
for $m = 2, 3,\cdots$.
Finally, we consider $\operatorname{cov}(y_n^2,y_{n+1})$. Since the computation is fairly straightforward but tedious, we omit all the details and just provide the following final result:
All the detailed calculations are provided in Appendix A.
Before closing, we point out that we can in fact compute the moments and covariances of any (positive integer) order and a recursive calculation procedure is provided in Appendix B. And the calculation of the moments and covariances for other affine SV models is presented in Appendix D.
In this section, we derive our MM estimators for the five parameters in the Heston model based on asset price samples and the moments/covariances obtained in the previous section. We prove both the moments/covariances estimators and the MM estimators satisfy the central limit theorem and also derive the explicit formulas for their covariance matrices.
Assume we are given a sequence of asset price samples, $S_0,S_1,\cdots, S_N$, based on which we can calculate return samples:
$i = 1,\cdots, N$. These samples can be used to compute the sample moments and covariances of $y_n$, which can be used to estimate their population counterparts:
where $\overline{Y^2} \triangleq (1/N)\sum_{i=1}^N Y_i^2$. Let
where $T$ denotes the transpose of a vector or matrix. We also define
Let
where $(N_2,N_3,N_4,N_5) = (N,N-1,N-2,N-1)$ and $\boldsymbol{1}_{n} = (1,\cdots,1)^T$ with $n$ elements. We are now ready to present the following central limit theorem for our moment/covariance estimators.
Having obtained the moment and covariance estimates, we can now develop the estimators for the parameters. First, let us consider $k$, the mean reversion velocity parameter in the variance (volatility) process. Note that (ref) leads to
for $m = 2,3,\cdots$. Therefore, we construct the following estimator for $k$:
where $2\leq M < N$ (usually $M$ takes a small value, e.g., $M\leq 10$) and
The estimators for the other four parameters are given by
where $\tilde{h}_{\hat{k}} \triangleq (1-e^{-\hat{k}h})/\hat{k}$ and $\hat{d}_h \triangleq he^{-\hat{k}h} - \tilde{h}_{\hat{k}}$.
In what follows, we let $M=2$ in $\hat{k}$ and denote the corresponding estimators as a continuous mapping $g: \mathbb{R}^{5} \rightarrow \mathbb{R}^5$, i.e.,
(Our analysis remains the same for $M>2$.) We now present the following central limit theorem for the parameter estimators.
In this section, we provide numerical experiments to test the estimators derived in the previous sections. First, we test the accuracy of our estimators under different parameter value settings. Then, we validate the asymptotic behaviors of our estimators as the size of samples increases. Next, we vary sample interval $h$ to analyze its effect on our estimators. Finally, we compare our method with the MCMC method.
In the first set of experiments, we consider six different settings of parameter values, with S0 being the base setting in which $\mu=0.125$, $k=0.1$, $\theta=0.25$, $\sigma_v=0.1$, $\rho=-0.7$. Each of the other five settings differs from S0 in only one parameter value: S1 increases $\mu$'s value to $0.4$, S2 decreases $k$'s value to $0.03$, S3 increases $\theta$'s value to $0.5$, S4 increases $\sigma_v$'s value to $0.2$, and S5 decreases (in absolute value) $\rho$'s value to $-0.3$. The values of the parameters in the volatility process, $k$, $\theta$, and $\sigma_v$, are the same as those in bollerslev2002estimating except for S3.
To generate sample observations for the Heston model, we use the first order Euler approximation method in which we set $h=1$ and partition each interval into 20 smaller segments for approximation. For each parameter setting, we run 400 replications with 400K samples for each replication. The numerical results are presented in Table (ref) with the format “mean $\pm$ standard deviation" based on these 400 replications (the format remains the same for all numerical results in this section). The results show our estimators are fairly accurate. Of course, the estimation accuracy can be improved by increasing $N$ as will be shown by the next set of experiments.
To test the effect of $N$ on the performance of our estimators, we run the second set of experiments in which we increase $N$ from $25K$ to $1600K$ for S0. The results are provided in Table (ref), which show as $N$ increases the accuracy of our estimators improves at the rate of $1/\sqrt{N}$.
We run the third set of experiments to test the effect of $h$ on our estimators. Again, we consider the parameter setting S0 and increase $h$ from 0.5 to 2. The results are given in Table (ref). It is clear that the accuracy of our estimators is not very sensitive to the value of $h$.
Finally, we compare our method with the MCMC method proposed in eraker2003impact and johannes2010mcmc; for the detailed implementation of the MCMC method, see Wu_MCMC_for_Estimating_2023. The comparison is based on the parameter setting S0, again with 400 replications with 400K samples for each replication (the results are very similar for other parameter settings). For the MCMC method, we use $\mu=0.0625,k=0.05,\theta=0.125,\sigma_v=0.05,\rho=-0.35$ as the initial values for the five parameters in the iteration procedure. For each replication, we run MCMC for 10,000 iterations with the first 5000 iterations as warm-up. The results are given in Table (ref), which shows our method performs much better than MCMC, both in terms of accuracy and computational time. MCMC is very time-consuming and it takes more than one hour for each replication while our method takes less than a second.
In this paper, we consider the problem of parameter estimation for affine SV models. We first obtain the moments and covariances of the asset price based on which we develop MM estimators. A key of our method is to develop a recursive procedure to calculate the moments and covariances of the asset price. Our MM estimators are simple and easy to implement. We also establish the central limit theorem for our estimators in which the explicit formulas for the asymptotic covariance matrix are given. Finally, we provide numerical results to test and validate our method. The method proposed in this paper can potentially be applied to other SV models. This is a possible future research direction.
This work was supported by the National Natural Science Foundation of China (NSFC) under Grants 72033003 and 72350710219; the China Postdoctoral Science Foundation under Grant 2023M732054; the Shandong Provincial Natural Science Foundation under Grant ZR2023QG159; and the Shandong Postdoctoral Science Foundation under Grant SDCX-RS-202303004.
Data sharing is not applicable to this article as no new data were created or analyzed in this study.