Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
56,940 characters · 18 sections · 31 citation commands
On a quantile autoregressive conditional duration model applied to high-frequency financial data
\def\spacingset#1{ {#1}} \spacingset{1}
\if00 \fi
\if10 {
} \fi
{\it Keywords:} Skewed Birnbaum-Saunders distribution; Conditional quantile; ECM algorithm; Monte Carlo simulation; Financial transaction data.
\spacingset{1.45}
Transaction-level high-frequency financial data modeling has become increasingly important in the era of big data. In particular, duration data between successive events (trades, price changes, etc.) were initially modeled by the family of autoregressive conditional duration (ACD) models, proposed by er:98. Since the inception of Engle and Russell's ACD model, several works have been published, supplying the literature with generalizations of the original model, associated inferential results, and diverse applications; see the literature review by pa:08 and its extension by bhogalvariyam:19.
One of the most prominent ACD models in terms of fit and forecasting performance is the Birnbaum-Saunders distribution-based ACD (BS-ACD) model. This model was proposed by b:10 and as emphasized by bhogalvariyam:19, it is constructed by taking into consideration the concept of conditional quantile estimation; the BS-ACD model is constructed in terms of a time-varying median duration; see also the recent discussion paper by balakundu:19. Recently, saulolla:19 made a study and have shown that the BS-ACD model outperforms many other existing models in terms of model fitting and forecasting ability.
In this work, we propose a new ACD model by modeling the conditional quantile duration rather than the traditionally employed conditional mean (or median) duration. The quantile approach allows us to capture the influences of conditioning variables on the characteristics of the response distribution; see koenkexiao:06. The proposed model is based on a skewed version of the BS (skew-BS) distribution, which has the main advantage over the classic BS distribution in its capability to fit data that are highly concentrated on the left-tail of the distribution, such is the case for transaction-level high-frequency financial data; the skew-BS distribution was proposed by vl:06 and inferential results for this model have been discussed by vslb:11. We first introduce a reparameterization of the skew-BS model by inserting a quantile parameter, similarly to the work of sanchezetal:19 who introduced a quantile parameter in the BS distribution, and then develop the new ACD model, denoted by skew-QBS-ACD. We then demonstrate that the proposed skew-QBS-ACD model outperforms the BS-ACD model in terms of model fitting and forecasting ability, making it a practical and useful model.
The rest of this paper proceeds as follows. In Section (ref), we describe the usual skew-BS distribution and propose a reparameterization of this distribution in terms of a quantile parameter. In Section (ref), we introduce the skew-QBS-ACD model. In this section, we also describe the expectation conditional maximization (ECM) algorithm for the maximum likelihood (ML) estimation of the model parameters. In Section (ref), we carry out a Monte Carlo simulation study to evaluate the performance of the estimators. In Section (ref), we apply the skew-QBS-ACD model to a real price duration data set, and finally in Section (ref), we provide some concluding remarks.
In this section, we describe briefly the skew-BS distribution. We then introduce a quantile-based reparameterization of this distribution, which will be useful subsequently for developing the new ACD model. We derive some properties of the quantile-based skew-BS distribution; see Appendix A.
A random variable $Y$ follows a skew-BS distribution with shape parameter $\alpha>0$, scale parameter $\beta>0$ and skewness parameter $\lambda$, denoted by $Y\sim\textrm{skew-BS}(\alpha,\beta,\lambda)$, if its probability density function (PDF) is given by
where
and $\phi(\cdot)$ and $\Phi(\cdot)$ are the standard normal PDF and cumulative distribution function (CDF), respectively.
Figure (ref) displays different shapes of the skew-BS PDF for different combinations of parameters. The plot show the clear effect the skewness parameter $\lambda$ has on the density function.
The random variable $Y$ can be stochastically represented as
where $\delta=\lambda/\sqrt{1+\lambda^2}$, $X$ is a random variable following a skew-normal (SN) distribution, $X\sim\textrm{SN}(\lambda)$, $Z$ is a standard normal random variable, $Z\sim\textrm{N}(0,1)$, and $U$ is a standard half-normal (HN) random variable, $U\sim \textrm{HN}(0,1)$; see vslb:11. The $100q$-th quantile of $Y\sim\textrm{skew-BS}(\alpha,\beta,\lambda)$ is given by
Interested readers may refer to the book by azzalinicap:14 for elaborate details on skew-normal distribution and related issues.
Consider a fixed number $q \in (0,1)$ and the following reparameterization $(\alpha, \beta,\lambda) \mapsto (\alpha, \Xi_{q},\lambda)$, where $\Xi_{q}$ is the $100q$-th quantile of $Y\sim\textrm{skew-BS}(\alpha,\beta,\lambda)$, as in (ref). Then, the PDF of $Y$ based on $\Xi_{q}$ is given by
where { \color{black}
with $\eta_{\alpha; \lambda}= \alpha Q_X(q;\lambda) + \sqrt{[\alpha Q_X(q;\lambda) ]^2 +4 } $. } Let us denote $Y\sim\textrm{skew-QBS}(\alpha,\Xi_{q},\lambda)$. Figure (ref) illustrates some shapes provided by different values of $\Xi_{q}$ and $\lambda$ for the skew-QBS PDF.
Note that, as in vslb:11, the extended BS (EBS) distribution is obtained from the PDF of $Y|(U=u)$ with $Y\sim\textrm{skew-QBS}(\alpha,\Xi_{q},\lambda)$ and $U\sim \textrm{HN}(0,1)$, which is given by
where $\alpha_{\delta}=\alpha\sqrt{1-\delta^2}$, $\lambda_{h}=-\delta h /\sqrt{1-\delta^2}$, and $a$ and $a'$ are as presented in (ref) with $\alpha$ and $\lambda$ replaced by $\alpha_{\delta}$ and $\lambda_{h}$, respectively. We use the notation $Y\sim{\rm EBS}(\alpha_{\delta},\Xi_{q},\lambda_{h})$ in this case. Moreover, it is possible to obtain the conditional PDF of $U$, given $Y=t$, that is, the HN standard PDF:
where $a(y;\alpha,\Xi_{q})$ is as given in (ref).
Let $T_1,\ldots, T_n$ be a sequence of successive recorded times at which market events (trade, price change, etc.) occur. Then, define $Y_t=T_t-T_{t-1}$, for $i=1,\ldots,n$, that is, $Y_t$ is the time elapsed between two successive market events. We now propose a skew-QBS ACD model specified in terms of a time varying conditional quantile, $\Xi_{q,t} = F_{Y}^{-1}(q;\alpha,\Xi_{q,t},\lambda|\mathbb{F}_{t-1})$ say, where $F_{Y_t}^{-1}$ is the quantile function of the skew-QBS distribution and $\mathbb{F}_{t-1}$ is the set of all data available until time $T_{t-1}$. The conditional PDF of the skew-QBS ACD is given by
where
and
Here $Q_X(q;\lambda)$ denotes the $100q$-th quantile of $X\sim\textrm{SN}(\lambda)$. Furthermore, the dynamics of the conditional quantile are governed by
implying the notation skew-QBS-ACD($r, s, q$). The ACD model in (ref) included as a special case the BS-ACD($r, s$) model proposed by b:10 when $q=0.5$ and $\lambda=0$. Note that the model in (ref) can be represented as
where $\varrho_{t}$ are independent and identically distributed (iid) random variables following the skew-QBS distribution, that is, $\varrho_{t} \stackrel{\text{iid}}{\sim}\textrm{skew-QBS}(\alpha,\Xi_{q}=1,\lambda)$ (see Proposition (ref)), and then the $Y_t$s are independent (ind) and non-identically distributed, i.e., $Y_t\stackrel{\text{ind}}{\sim} \textrm{skew-QBS}(\alpha,\Xi_{q,t},\lambda)$.
In what follows we establish some results on moments, dispersion index function and autocorrelation function of the skew-QBS ACD model. To this end, let us define the following: \[ \boldsymbol{\Omega}=
, \quad \phi_k=\boldsymbol{\rho}^\top \boldsymbol{\Omega}^{k-p-1} \boldsymbol{\phi}, \ k>p=\max\{r,s\}, \] where $\boldsymbol{\rho}=(\rho_1,\ldots,\rho_p)^\top$ and $\boldsymbol{\phi}=(\phi_1,\ldots,\phi_p)$ such that \[ \phi_0=1, \quad \phi_1=\rho_1, \quad \phi_l=
\] Using the above notation, the next result shows that the moments of the skew-QBS ACD model depend on the moments of $\varrho_{t}$.
Let $\bm Y = (Y_{1}, \ldots, Y_{n})^\top$ be a random sample from $Y_{t}\sim \textrm{skew-QBS}(\alpha,\Xi_{q,t}, \lambda)$, for $t=1,\ldots,n$, with PDF as given in (ref), and ${\bm y} = (y_{1},\ldots, y_{n})^{\top}$ be the corresponding observations. To obtain the ML estimates of the model parameters, $\bm{\theta} = (\alpha, \varpi, \rho_{1},\ldots,\rho_{r}, \sigma_{1}, \ldots, \sigma_{s}, \lambda)^{\top}$ say, we can maximize the log-likelihood function
where $f_{Y|\Xi_{q}}(y_{t};\alpha,\Xi_{q,t},\lambda)$ and $\Xi_{q,t}$ are as in (ref) and (ref), respectively. As usual, it can be done by equating the score vector $\dot{\bm\ell}(\bm{\theta})$ to zero, and by using an iterative procedure for non-linear optimization, such as the Broyden-Fletcher-Goldfarb-Shanno (BFGS) quasi-Newton method.
Let $f_{T|\Xi_{q}}(y_{t};\alpha,\Xi_{q,t},\lambda) = \phi\big[a(y;\alpha,\Xi_{q,t},\lambda)\big]a'(y;\alpha,\Xi_{q,t},\lambda)$ be the conditional PDF of the QBS distribution, where $a(y;\alpha,\Xi_{q,t},\lambda)$ and $a'(y;\alpha,\Xi_{q,t},\lambda)$ are as in (ref) and (ref), respectively. Then, by using (ref), the conditional PDF of the skew-QBS ACD can be written as $$ f_{Y|\Xi_{q}}(y_{t};\alpha,\Xi_{q,t},\lambda) = 2 \Phi\big[\lambda a(y_t;\alpha,\Xi_{q,t},\lambda)\big] f_{T|\Xi_{q}}(y_{t};\alpha,\Xi_{q,t},\lambda). $$ Using the above relation, note that the elements of the score vector $\dot{\bm\ell}(\bm{\theta})$ are given by
for each $\gamma\in\{\alpha, \varpi, \rho_{1},\ldots,\rho_{r}, \sigma_{1}, \ldots, \sigma_{s}, \lambda\}$, where { \scalefont{0.85}
} for each $\gamma\in\{\alpha, \varpi, \rho_{1},\ldots,\rho_{r}, \sigma_{1}, \ldots, \sigma_{s},\lambda\}$. The above first order partial derivatives of functions $a(y_t;\alpha,\Xi_{q,t},\lambda)$ and $a'(y_t;\alpha,\Xi_{q,t},\lambda)$ defined in (ref) and (ref), respectively, are written as
and
where $\eta_{\alpha; \lambda}$ is as in (ref) and
Furthermore, by using (ref), the above partial derivatives of quantile $\Xi_{q,t}$ are given by
Let $I(\boldsymbol{\theta_0})$ denote the expected Fisher information matrix, where $\boldsymbol{\theta_0}$ is the true value of the population parameter vector. Let us consider the following regularity conditions:
Following Davison08, under the above assumptions, it follows that
where $\widehat{\bm \theta}_n$ is the MLE of the parameter of interest ${\bm \theta}$, ${\bm 0}$ is the $(r+s+3)\times 1$ zero vector and $I_{(r+s+3)\times (r+s+3)}$ is the ${(r+s+3)\times (r+s+3)}$ identity matrix. In practice, one may approximate the expected Fisher information matrix by its observed version, which is obtained from the Hessian matrix; see Appendix B. Then, the corresponding diagonal elements of the inverse of the observed Fisher information matrix can be used to approximate the standard errors (SEs) of the estimates.
We can implement the ECM algorithm presented in vslb:11. Specifically, let ${\bm y}=(y_{1},\ldots,y_{n})^{\top}$ and ${\bm u}=(u_{1},\ldots,u_{n})^{\top}$ be the observed and missing data, respectively, with $\bm Y$ and $\bm U$ being their corresponding random vectors. Thus, the complete data vector is written as ${\bm y}_{c}=({\bm y}^{\top}, {\bm u}^{\top})^{\top}$. From (ref), we obtain
Then, the complete data log-likelihood function for the skew-QBS-ACD model, associated with ${\bm y}_{c}=({\bm y}^{\top}, {\bm u}^{\top})^{\top}$, is given by (without the constant)
The implementation of the ECM algorithm requires the computation of the expected value of the complete log-likelihood equation in (ref), conditional on $ \bm U = \bm u$, that is,
where $b(y_t;\Xi_{q,t}) = \left[\sqrt{\frac{y_t \eta_q^2}{4\Xi_{q,t}}}-\sqrt{\frac{4\Xi_{q,t}}{y_t\eta_{q}^2}}\right]$, $\lambda=\delta/\sqrt{1-\delta^2}$, and $$ \widehat{u}_t=E[U_t|Y_t=y_t;\widehat{\bm \theta}]= \widehat{\delta} a(y_t;\widehat{\alpha},{\Xi}_{q,t})+ \sqrt{1-\widehat{\delta}^2}\Upsilon_{\Phi}\left[ \frac{\widehat{\delta} a(y_t;\widehat{\alpha},{\Xi}_{q,t})}{\sqrt{1-\widehat{\delta}^2}} \right] $$
$$ \widehat{u}_t^2=E[U_t^2|Y_t=y_t;\widehat{\bm \theta}]= \widehat{\delta}^2 a^2(y_t;\widehat{\alpha},{\Xi}_{q,t})+ (1-\widehat{\delta}^2)+\widehat{\delta}\sqrt{1-\widehat{\delta}^2} \Upsilon_{\Phi}\left[ \frac{\widehat{\delta} a(y_t;\widehat{\alpha},{\Xi}_{q,t})}{\sqrt{1-\widehat{\delta}^2}} \right] a(y_t;\widehat{\alpha},{\Xi}_{q,t}), $$ $$ \log(\Xi_{q,t}) = \widehat{\varpi} + \sum_{j=1}^{r}\widehat{\rho}_{j} \log(\Xi_{q,t-j}) + \sum_{j=1}^{s}\frac{\widehat{\sigma}_{j}\,y_{t-j}}{\Xi_{q,t-j}}, $$ with $\Upsilon_{\Phi}(x)=\phi(x)/\Phi(x)$ for $x\in\mathbb{R}$; see vslb:11.
The steps to obtain the ML estimates of the skew-QBS-ACD model via the ECM algorithm have been summarized in Algorithm (ref). Starting values are required to initiate the procedure, particularly, $\widehat{\alpha}^{(0)}, \widehat{\varpi}^{(0)}, \widehat{\rho}_{1}^{(0)},\ldots,\widehat{\rho}_{r}^{(0)}, \widehat{\sigma}_{1}^{(0)}, \ldots, \widehat{\sigma}_{s}^{(0)}$. These are obtained from saulolla:19, and then $\widehat{\lambda}^{(0)}$ from vslb:11. The iterations of Algorithm (ref) are repeated until a stopping criterion has been met, such as $\ell(\bm{\theta}^{v+1})-\ell(\bm{\theta}^{v})\leq \epsilon$, where $\epsilon$ is a prescribed small tolerance level.
To assess the goodness of fit and departures from the assumptions of the model, we shall use the generalized Cox-Snell (GCS) residual, given by
where $\widehat{S}(\cdot)$ denotes the survival function fitted to the ACD data. This residual has a standard exponential, EXP(1), distribution when the model is correctly specified. Then, quantile-quantile (QQ) plots with simulated envelope can be used to assess the fit.
We present the results of two Monte Carlo simulation studies for the skew-QBS-ACD($r, s, q$) model. Results are presented only for order of the lags set as $p=1$ and $q=1$, since a higher order for skew-QBS-ACD models did not improve the model fit considerably; see related results in b:10 and saulolla:19. The first study considers the evaluation of the performance of the ML estimation procedure developed in Algorithm (ref), while the second study evaluates the empirical distribution of the residuals.
The first simulation scenario considers the following setting: sample size $n \in \{500, 1000, 2000\}$, vector of true parameters $(\varpi,\rho,\sigma,\alpha,\lambda) = (0.20,0.70,0.10,0.50,-0.50)$ and $(\varpi,\rho,\sigma,\alpha,\lambda) = (0.20,0.70,0.10,1.00,-0.50)$, and $q=\{0.20,0.50,0.80\}$, with $200$ Monte Carlo replications for each sample size. The skew-QBS-ACD samples were generated using the transformation
where $\eta_{q}$ is as in (ref), $\Xi_{q,t}$ is as in (ref), and $X_t$ is a SN random variable generated by the function rsn of the R package sn; see sn:19.
The ML estimation results are presented in Table (ref) wherein the empirical mean, bias, root mean squared error (RMSE), coefficients of skewness, and coefficients of (excess) kurtosis are all reported. A look at the results in Table (ref) allows us to conclude that, in general, as the sample size increases, the bias and $\textrm{RMSE}$ of the estimates decrease, as expected. Moreover, $\widehat{\varpi}$, $\widehat{\rho}$, $\widehat{\sigma}$, $\widehat{\alpha}$ and $\widehat{\lambda}$ seem all to be consistent and marginally asymptotically distributed as normal.
In this subsection, we present a Monte Carlo simulation study to evaluate the performance of the GCS residuals. The sample generation as well as the simulation scenario are the same ($\alpha=0.50$) in Subsection (ref).
Table (ref) presents the empirical mean, standard deviation, coefficient of skewness and coefficient of kurtosis, whose values are expected to be 1, 1, 2 and 9, respectively, for $R^\textrm{GCS}$. Notice that the reference distribution for $R^\textrm{GCS}$ is standard exponential distribution. From Table (ref), we note that, in general, the considered residuals conform well with the reference distribution.
In this section, a real high frequency financial data set corresponding to price durations of BASF-SE stock on 19th April 2016 is used to illustrate the proposed methodology; see saulolla:19. The data were adjusted to remove intraday seasonality, namely, activity is low at lunchtime and high at the beginning and closing of the trading day. Table (ref) provides descriptive statistics for both plain and adjusted price durations. From this table, we note that both series show positive skewness and a high degree of kurtosis.
Table (ref) presents the ML estimates, computed by the ECM algorithm, and SEs for the skew-QBS-ACD model parameters with $q = \{0.13,0.50\}$. On the one hand, the value of $q = 0.13$ was chosen through the profile log-likelihood, that is, for a grid of values of $q$, $0.01, 0.02,\ldots,0.99$, we estimated the model parameters and computed the corresponding log-likelihood functions. Then, the optimal value of $q$ was the one which maximized the log-likelihood function. On the other hand, the value of $q = 0.50$ corresponds to the skew-QBS-ACD in terms of the conditional median duration, such as the BS-ACD model. Table (ref) also reports the log-likelihood value, Akaike (AIC) and Bayesian information (BIC) criteria, and $p$-values of the Ljung-Box (LB) statistic, $Q(l)$ say, for up to $l$-th order serial correlation. The LB statistics evaluate the autocorrelation of the residuals. For comparative purpose, the results of the exponential ACD (EXP-ACD), generalized gamma ACD (GG-ACD) and BS-ACD models, are provided as well. The results of Table (ref) reveal that the skew-QBS-ACD models with $q = \{0.13,0.50\}$ provide better adjustments compared to other models based on the values of log-likelihood, AIC and BIC. However, these values are not substantially different for the two values of $q$. Finally, the LB $p$-values provide no evidence of serial correlation in the residuals for all models considered in this application.
Figure (ref) plots the estimated parameters of the skew-QBS-ACD model across $q\in\{0.01,\ldots, 0.99\}$. Results suggest that this BASF-SE series displays asymmetric dynamics: the estimates of ${\varpi}$ and ${\sigma}$ tend to be increasing across $q$, with a change of sign in the first case. On the other hand, the estimates of $\rho$, $\alpha$ and $\lambda$ appear to vary with the same pattern, but without an explicit trend.
Figure (ref) shows the QQ plots with simulated envelope of the generalized Cox-Snell residuals for the models considered in Table (ref). We see clearly that the skew-QBS-ACD models provide better fit than the EXP-ACD, GG-ACD and BS-ACD models.
Figure (ref) plots the in-sample and out-of-sample quantile forecasts from the adjusted skew-QBS-ACD and BS-ACD models. The construction of the plots is as follows. For each value of $q\in\{0.01,\ldots, 0.99\}$, the skew-QBS-ACD model is estimated and the respective estimates are used to obtain the forecast of the quantile of interest. Since the BS-ACD model is independent of the value of $q$, it is estimated only once and its forecasts are obtained from this single estimation. For in-sample forecasting, all data are used. In the case of out-of-sample forecasting, 2/3 of the data are used for estimation and the other 1/3 for forecasting.
From Figures (ref)(a,b), we observe that the forecasts of the skew-BS-ACD model are much closer to the actual values than those of the BS-ACD model. In fact, the forecasts of this latter model are quite far from the actual values. The superiority of the skew-QBS-ACD model may be due to different estimates obtained for the parameters; see, e.g., Figure (ref). For different quantiles, we obtain different estimates, which makes the forecasts closer to the actual values. The results of the estimated mean square error (MSE) between the actual data and in-sample forecasted values of the adjusted skew-QBS-ACD and BS-ACD models are 0.041 and 1.277, respectively. Moreover, the MSE between the actual data and out-of-sample forecasted values of the adjusted skew-QBS-ACD and BS-ACD models are 0.029 and 1.199, respectively. Thus, the MSE values do confirm the results found in Figures (ref)(a,b).
In this paper, we have proposed a quantile autoregressive conditional duration model for dealing with high frequency financial duration data. The proposed model is based on a reparameterized version of the skewed Birnbaum-Saunders distribution, which has the distribution quantile as one of its parameters. Thus, the proposed conditional duration model is defined in terms of a conditional quantile duration, allowing different estimates for different quantiles of interest. A Monte Carlo simulation is carried out to evaluate the performance of the maximum likelihood estimators, study obtained by the ECM algorithm. We have applied the proposed and some other existing models to a real financial duration data set from the German DAX stock exchange. The results are seen to be quite favorable to the proposed model both in terms of model fitting and forecasting ability. The gain in considering the dynamics in terms of conditional quantile duration seems quite evident, especially in the forecasting superiority of the model. As part of future research, it will be of interest to use the theory of estimating functions for estimating the model parameters ALLEN2013214. Furthermore, quantile autoregressive conditional duration models can be used to estimate the intraday volatility of a stock Tse-2012. Work on these problems is currently in progress and we hope to report these findings in future papers.
Funding: Helton Saulo gratefully acknowledges financial support from FAP-DF.