EconBase
← Back to paper

On a quantile autoregressive conditional duration model applied to high-frequency financial data

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

56,940 characters · 18 sections · 31 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

On a quantile autoregressive conditional duration model applied to high-frequency financial data

\def\spacingset#1{ {#1}} \spacingset{1}

\if00 \fi

\if10 {

center[center omitted — 126 chars of source]

} \fi

abstractAutoregressive conditional duration (ACD) models are primarily used to deal with data arising from times between two successive events. These models are usually specified in terms of a time-varying conditional mean or median duration. In this paper, we relax this assumption and consider a conditional quantile approach to facilitate the modeling of different percentiles. The proposed ACD quantile model is based on a skewed version of Birnbaum-Saunders distribution, which provides better fitting of the tails than the traditional Birnbaum-Saunders distribution, in addition to advancing the implementation of an expectation conditional maximization (ECM) algorithm. A Monte Carlo simulation study is performed to assess the behavior of the model as well as the parameter estimation method and to evaluate a form of residual. A real financial transaction data set is finally analyzed to illustrate the proposed approach.

{\it Keywords:} Skewed Birnbaum-Saunders distribution; Conditional quantile; ECM algorithm; Monte Carlo simulation; Financial transaction data.

\spacingset{1.45}

Introduction

Transaction-level high-frequency financial data modeling has become increasingly important in the era of big data. In particular, duration data between successive events (trades, price changes, etc.) were initially modeled by the family of autoregressive conditional duration (ACD) models, proposed by er:98. Since the inception of Engle and Russell's ACD model, several works have been published, supplying the literature with generalizations of the original model, associated inferential results, and diverse applications; see the literature review by pa:08 and its extension by bhogalvariyam:19.

One of the most prominent ACD models in terms of fit and forecasting performance is the Birnbaum-Saunders distribution-based ACD (BS-ACD) model. This model was proposed by b:10 and as emphasized by bhogalvariyam:19, it is constructed by taking into consideration the concept of conditional quantile estimation; the BS-ACD model is constructed in terms of a time-varying median duration; see also the recent discussion paper by balakundu:19. Recently, saulolla:19 made a study and have shown that the BS-ACD model outperforms many other existing models in terms of model fitting and forecasting ability.

In this work, we propose a new ACD model by modeling the conditional quantile duration rather than the traditionally employed conditional mean (or median) duration. The quantile approach allows us to capture the influences of conditioning variables on the characteristics of the response distribution; see koenkexiao:06. The proposed model is based on a skewed version of the BS (skew-BS) distribution, which has the main advantage over the classic BS distribution in its capability to fit data that are highly concentrated on the left-tail of the distribution, such is the case for transaction-level high-frequency financial data; the skew-BS distribution was proposed by vl:06 and inferential results for this model have been discussed by vslb:11. We first introduce a reparameterization of the skew-BS model by inserting a quantile parameter, similarly to the work of sanchezetal:19 who introduced a quantile parameter in the BS distribution, and then develop the new ACD model, denoted by skew-QBS-ACD. We then demonstrate that the proposed skew-QBS-ACD model outperforms the BS-ACD model in terms of model fitting and forecasting ability, making it a practical and useful model.

The rest of this paper proceeds as follows. In Section (ref), we describe the usual skew-BS distribution and propose a reparameterization of this distribution in terms of a quantile parameter. In Section (ref), we introduce the skew-QBS-ACD model. In this section, we also describe the expectation conditional maximization (ECM) algorithm for the maximum likelihood (ML) estimation of the model parameters. In Section (ref), we carry out a Monte Carlo simulation study to evaluate the performance of the estimators. In Section (ref), we apply the skew-QBS-ACD model to a real price duration data set, and finally in Section (ref), we provide some concluding remarks.

Preliminaries

In this section, we describe briefly the skew-BS distribution. We then introduce a quantile-based reparameterization of this distribution, which will be useful subsequently for developing the new ACD model. We derive some properties of the quantile-based skew-BS distribution; see Appendix A.

The classical skew-BS distribution

A random variable $Y$ follows a skew-BS distribution with shape parameter $\alpha>0$, scale parameter $\beta>0$ and skewness parameter $\lambda$, denoted by $Y\sim\textrm{skew-BS}(\alpha,\beta,\lambda)$, if its probability density function (PDF) is given by

equation[equation omitted — 170 chars of source]

where

align[align omitted — 225 chars of source]

and $\phi(\cdot)$ and $\Phi(\cdot)$ are the standard normal PDF and cumulative distribution function (CDF), respectively.

remNote that the skew-BS distribution (ref) can be approached as a weighted distribution \begin{align*} f_{Y}(y;\alpha,\beta,\lambda) ={w(y) f_{T}(y;\alpha,\beta)\over \mathbb{E}[w(T)]}, \quad y>0, \end{align*} where $w(y)=\Phi\big[\lambda a(y;\alpha,\beta)\big]$ and $T$ follows a classic BS distribution, $T\sim\textrm{BS}(\alpha,\beta)$; see bs:69a for details on the classic BS distribution.

Figure (ref) displays different shapes of the skew-BS PDF for different combinations of parameters. The plot show the clear effect the skewness parameter $\lambda$ has on the density function.

figure[figure omitted — 267 chars of source]

The random variable $Y$ can be stochastically represented as

equation[equation omitted — 148 chars of source]

where $\delta=\lambda/\sqrt{1+\lambda^2}$, $X$ is a random variable following a skew-normal (SN) distribution, $X\sim\textrm{SN}(\lambda)$, $Z$ is a standard normal random variable, $Z\sim\textrm{N}(0,1)$, and $U$ is a standard half-normal (HN) random variable, $U\sim \textrm{HN}(0,1)$; see vslb:11. The $100q$-th quantile of $Y\sim\textrm{skew-BS}(\alpha,\beta,\lambda)$ is given by

equation[equation omitted — 176 chars of source]

Interested readers may refer to the book by azzalinicap:14 for elaborate details on skew-normal distribution and related issues.

The quantile-based skew-QBS distribution

Consider a fixed number $q \in (0,1)$ and the following reparameterization $(\alpha, \beta,\lambda) \mapsto (\alpha, \Xi_{q},\lambda)$, where $\Xi_{q}$ is the $100q$-th quantile of $Y\sim\textrm{skew-BS}(\alpha,\beta,\lambda)$, as in (ref). Then, the PDF of $Y$ based on $\Xi_{q}$ is given by

align[align omitted — 182 chars of source]

where { \color{black}

equation[equation omitted — 346 chars of source]

with $\eta_{\alpha; \lambda}= \alpha Q_X(q;\lambda) + \sqrt{[\alpha Q_X(q;\lambda) ]^2 +4 } $. } Let us denote $Y\sim\textrm{skew-QBS}(\alpha,\Xi_{q},\lambda)$. Figure (ref) illustrates some shapes provided by different values of $\Xi_{q}$ and $\lambda$ for the skew-QBS PDF.

figure[figure omitted — 323 chars of source]

Note that, as in vslb:11, the extended BS (EBS) distribution is obtained from the PDF of $Y|(U=u)$ with $Y\sim\textrm{skew-QBS}(\alpha,\Xi_{q},\lambda)$ and $U\sim \textrm{HN}(0,1)$, which is given by

equation[equation omitted — 191 chars of source]

where $\alpha_{\delta}=\alpha\sqrt{1-\delta^2}$, $\lambda_{h}=-\delta h /\sqrt{1-\delta^2}$, and $a$ and $a'$ are as presented in (ref) with $\alpha$ and $\lambda$ replaced by $\alpha_{\delta}$ and $\lambda_{h}$, respectively. We use the notation $Y\sim{\rm EBS}(\alpha_{\delta},\Xi_{q},\lambda_{h})$ in this case. Moreover, it is possible to obtain the conditional PDF of $U$, given $Y=t$, that is, the HN standard PDF:

equation[equation omitted — 226 chars of source]

where $a(y;\alpha,\Xi_{q})$ is as given in (ref).

The skew-QBS ACD model and ECM estimation

The skew-QBS ACD model

Let $T_1,\ldots, T_n$ be a sequence of successive recorded times at which market events (trade, price change, etc.) occur. Then, define $Y_t=T_t-T_{t-1}$, for $i=1,\ldots,n$, that is, $Y_t$ is the time elapsed between two successive market events. We now propose a skew-QBS ACD model specified in terms of a time varying conditional quantile, $\Xi_{q,t} = F_{Y}^{-1}(q;\alpha,\Xi_{q,t},\lambda|\mathbb{F}_{t-1})$ say, where $F_{Y_t}^{-1}$ is the quantile function of the skew-QBS distribution and $\mathbb{F}_{t-1}$ is the set of all data available until time $T_{t-1}$. The conditional PDF of the skew-QBS ACD is given by

equation[equation omitted — 239 chars of source]

where

align[align omitted — 398 chars of source]

and

align[align omitted — 124 chars of source]

Here $Q_X(q;\lambda)$ denotes the $100q$-th quantile of $X\sim\textrm{SN}(\lambda)$. Furthermore, the dynamics of the conditional quantile are governed by

equation[equation omitted — 191 chars of source]

implying the notation skew-QBS-ACD($r, s, q$). The ACD model in (ref) included as a special case the BS-ACD($r, s$) model proposed by b:10 when $q=0.5$ and $\lambda=0$. Note that the model in (ref) can be represented as

equation[equation omitted — 104 chars of source]

where $\varrho_{t}$ are independent and identically distributed (iid) random variables following the skew-QBS distribution, that is, $\varrho_{t} \stackrel{\text{iid}}{\sim}\textrm{skew-QBS}(\alpha,\Xi_{q}=1,\lambda)$ (see Proposition (ref)), and then the $Y_t$s are independent (ind) and non-identically distributed, i.e., $Y_t\stackrel{\text{ind}}{\sim} \textrm{skew-QBS}(\alpha,\Xi_{q,t},\lambda)$.

In what follows we establish some results on moments, dispersion index function and autocorrelation function of the skew-QBS ACD model. To this end, let us define the following: \[ \boldsymbol{\Omega}=

pmatrix[pmatrix omitted — 139 chars of source]

, \quad \phi_k=\boldsymbol{\rho}^\top \boldsymbol{\Omega}^{k-p-1} \boldsymbol{\phi}, \ k>p=\max\{r,s\}, \] where $\boldsymbol{\rho}=(\rho_1,\ldots,\rho_p)^\top$ and $\boldsymbol{\phi}=(\phi_1,\ldots,\phi_p)$ such that \[ \phi_0=1, \quad \phi_1=\rho_1, \quad \phi_l=

cases\sum_{j=1}^l \rho_j \phi_{l-j}, & l=2,\ldots,p, \\ \sum_{j=1}^p \rho_j \phi_{l-j}, & l>p.

\] Using the above notation, the next result shows that the moments of the skew-QBS ACD model depend on the moments of $\varrho_{t}$.

propAssume that $\mathbb{E}[{\rm e}^{m\theta_j \varrho_{t}} ]$ and $\mu_m=\mathbb{E}[Y_t^m]$ exist for each $m>0$. For the skew-QBS ACD model, $\lambda_{\rm max}<1$ iff $\mu_m$ exists, where $\lambda_{\rm max}$ is the absolute value of the maximum eigenvalue of the matrix $\boldsymbol{\Omega}$. Furthermore, \begin{enumerate} • Moments of the skew-QBS ACD model: \[ \mathbb{E}[Y_t^m] = \mu_m {\rm e}^{m\varpi(1-\sum_{j=1}^p\rho_j)^{-1}} \, \prod_{j=1}^\infty \mathbb{E}[{\rm e}^{m\theta_j \varrho_{t}}], \] where $ \theta_1=\sigma_1 \ \text{and} \ \theta_l= \begin{cases} \sum_{j=1}^l \sigma_j \phi_{l-j}, & l=2,\ldots,p, \\ \sum_{j=1}^p \sigma_j \phi_{l-j}, & l>p. \end{cases} $ • Dispersion index function: \[ 1+\delta_Y = (1+\delta^2) { \prod_{j=1}^\infty \mathbb{E}[ {\rm e}^{2\theta_j \varrho_t } ] \over \big\{ \prod_{j=1}^\infty \mathbb{E}[{\rm e}^{\theta_j \varrho_t }] \big\}^2 } \geqslant 1+\delta^2, \] where $\delta=\sqrt{\rm Var(\varrho_{t})}/\mathbb{E}[\varrho_{t}]$ is the dispersion index of $\varrho_{t}$ and $\delta_Y$ denotes the degree of dispersion of the random variable $Y$ by the coefficient of variation. \end{enumerate}
proofThe proof follows the same steps as those of Theorem 1 and Corollary 1 in bauwensetal:03, by taking \[ \psi_t=\log(\Xi_{q,t}), \ \ \ \varpi=\omega, \ \ \ \alpha_j=\sigma_j, \ \ \ g(\epsilon_t)=\epsilon_t =\varrho_{t}={y_t\over\Xi_{q,t}}, \ \ \ \beta_j=\rho_j \ \ \ \text{and} \ \ \ p=\max\{r,s\} \] in Equation (5) of reference bauwensetal:03.
propAssume that $\mu=\mathbb{E}[\varrho_{t}]<\infty$, $\lambda_{\rm max}<1$, $\mathbb{E}[{{\rm e}^{\delta \varrho_{t}}}]<\infty$ for any $\delta\in\mathbb{R}$, $\mathbb{E}[\varrho_{t-n} {\rm e}^{\theta_n \varrho_{t-n}}]<\infty$ for any $n\in\mathbb{N}_+$, and $\mathbb{E}[{\rm e}^{(\phi_{n-j}\sigma_{j+h}+\theta^*_{hn}) (\varrho_{t-n-h})}]<\infty$ for $j$ and $h$ such that $n\geqslant 1$, where $\lambda_{\rm max}$ is as defined in Proposition (ref). Then, for $n\geqslant 1$, the $n$-th order autocorrelation of $\{Y_t\}$ has the form \begin{align*} \rho_n = { \mu \mathbb{E}[\varrho_{t} {\rm e}^{\theta_n \varrho_{t}}] \prod_{j=1}^{n-1} \mathbb{E}[{{\rm e}^{\theta_j \varrho_{t}}}] \prod_{j=p}^{\infty} \mathbb{E}[{{\rm e}^{\theta^*_{jn} \varrho_{t}}}] M_{n,p} - \mu^2 (\prod_{j=1}^\infty \mathbb{E}[{{\rm e}^{\theta_{j} \varrho_{t}}}])^2 \over \mu_2 \prod_{j=1}^\infty \mathbb{E}[{{\rm e}^{2\theta_{j} \varrho_{t}}}] - \mu^2 (\prod_{j=1}^\infty \mathbb{E}[{{\rm e}^{\theta_{j} \varrho_{t}}}])^2 }, \end{align*} where \begin{align*} M_{n,p} = \begin{cases} \prod_{h=1}^{p-n} \mathbb{E}[{\rm e}^{(\sum_{j=1}^n\phi_{n-j}\sigma_{j+h}+\theta^*_{hn})(\varrho_{t-n-h})}] \prod_{h=1}^{n-1} \mathbb{E}[{\rm e}^{(\sum_{j=1}^h\phi_{n-j}\sigma_{p-h+j}+\theta^*_{p-h,n})(\varrho_{t-n-h})}], & n \leqslant p \\ \prod_{h=1}^{p-1} \mathbb{E}[{\rm e}^{(\sum_{j=1}^{p-h}\phi_{n-j}\sigma_{h+j}+\theta^*_{hn})(\varrho_{t-n-h})}], & n> p \end{cases}, \end{align*} $ \theta^*_{jn} = \begin{cases} \sum_{j=1}^h \sigma_j \phi^*_{h+1-j,n}, & h=1,\ldots,p \\ \sum_{j=1}^p \sigma_j \phi^*_{h+1-j,n}, & h>p \end{cases}, $ $\phi^*$ is as in (56) of bauwensetal:03, and $\mu_2={\rm Var}(\varrho_{t})+\mu^2$.
proofThe proof follows from Theorem 2 of bauwensetal:03 by using the same notation as in the proof of Proposition (ref).

Estimation via the ECM algorithm

ML approach

Let $\bm Y = (Y_{1}, \ldots, Y_{n})^\top$ be a random sample from $Y_{t}\sim \textrm{skew-QBS}(\alpha,\Xi_{q,t}, \lambda)$, for $t=1,\ldots,n$, with PDF as given in (ref), and ${\bm y} = (y_{1},\ldots, y_{n})^{\top}$ be the corresponding observations. To obtain the ML estimates of the model parameters, $\bm{\theta} = (\alpha, \varpi, \rho_{1},\ldots,\rho_{r}, \sigma_{1}, \ldots, \sigma_{s}, \lambda)^{\top}$ say, we can maximize the log-likelihood function

equation[equation omitted — 130 chars of source]

where $f_{Y|\Xi_{q}}(y_{t};\alpha,\Xi_{q,t},\lambda)$ and $\Xi_{q,t}$ are as in (ref) and (ref), respectively. As usual, it can be done by equating the score vector $\dot{\bm\ell}(\bm{\theta})$ to zero, and by using an iterative procedure for non-linear optimization, such as the Broyden-Fletcher-Goldfarb-Shanno (BFGS) quasi-Newton method.

Let $f_{T|\Xi_{q}}(y_{t};\alpha,\Xi_{q,t},\lambda) = \phi\big[a(y;\alpha,\Xi_{q,t},\lambda)\big]a'(y;\alpha,\Xi_{q,t},\lambda)$ be the conditional PDF of the QBS distribution, where $a(y;\alpha,\Xi_{q,t},\lambda)$ and $a'(y;\alpha,\Xi_{q,t},\lambda)$ are as in (ref) and (ref), respectively. Then, by using (ref), the conditional PDF of the skew-QBS ACD can be written as $$ f_{Y|\Xi_{q}}(y_{t};\alpha,\Xi_{q,t},\lambda) = 2 \Phi\big[\lambda a(y_t;\alpha,\Xi_{q,t},\lambda)\big] f_{T|\Xi_{q}}(y_{t};\alpha,\Xi_{q,t},\lambda). $$ Using the above relation, note that the elements of the score vector $\dot{\bm\ell}(\bm{\theta})$ are given by

align*[align* omitted — 220 chars of source]

for each $\gamma\in\{\alpha, \varpi, \rho_{1},\ldots,\rho_{r}, \sigma_{1}, \ldots, \sigma_{s}, \lambda\}$, where { \scalefont{0.85}

align*[align* omitted — 786 chars of source]

} for each $\gamma\in\{\alpha, \varpi, \rho_{1},\ldots,\rho_{r}, \sigma_{1}, \ldots, \sigma_{s},\lambda\}$. The above first order partial derivatives of functions $a(y_t;\alpha,\Xi_{q,t},\lambda)$ and $a'(y_t;\alpha,\Xi_{q,t},\lambda)$ defined in (ref) and (ref), respectively, are written as

align*[align* omitted — 777 chars of source]

and

align*[align* omitted — 785 chars of source]

where $\eta_{\alpha; \lambda}$ is as in (ref) and

align[align omitted — 214 chars of source]

Furthermore, by using (ref), the above partial derivatives of quantile $\Xi_{q,t}$ are given by

align*[align* omitted — 989 chars of source]

Let $I(\boldsymbol{\theta_0})$ denote the expected Fisher information matrix, where $\boldsymbol{\theta_0}$ is the true value of the population parameter vector. Let us consider the following regularity conditions:

itemize• The conditional PDF (ref) of the skew-QBS ACD is an injective function in $\boldsymbol{\theta}$. • The log-likelihood function ${\ell}(\bm{\theta})$ is three times differentiable with respect to the unknown parameter vector $\bm{\theta}$ and the third partial derivatives are continuous in $\bm{\theta}$. • Let $\bm Y = (Y_{1}, \ldots, Y_{n})^\top$ be a random sample from $Y_{t}\sim \textrm{skew-QBS}(\alpha,\Xi_{q,t}, \lambda)$, for $t=1,\ldots,n$. For any $\boldsymbol{\theta_0}\in\Theta$, there exists $\delta> 0$ and functions $H_{uvw}({\bm y})$, for all ${\bm y}$ in $\Sigma$ (the support of the data), such that for $\bm{\theta}= (\alpha, \varpi, \rho_{1},\ldots,\rho_{r}, \sigma_{1}, \ldots, \sigma_{s}, \lambda)^{\top}=(\gamma_1,\gamma_2,\ldots,\gamma_{r+s+3})^{\top}$ and $u,v,w,z = 1, 2,\ldots, r+s+3$, \begin{align*} \left\vert {\partial^3{\ell}(\bm{\theta})\over \partial\gamma_u \partial\gamma_v\partial\gamma_w} \right\vert \leqslant H_{uvw}({\bm y}), \quad for all \, {\bm y}\in\Sigma \, and \, \vert \gamma_z-\gamma_{0,z} \vert<\delta, \end{align*} whenever $\mathbb{E}\big[H_{uvw}({\bm Y})\big]$ is finite. • For all $\boldsymbol{\theta}\in\Theta$ and $u = 1, 2,\ldots, r+s+3$; $\mathbb{E}\big[{\partial{\ell}(\bm{\theta})\over \partial\gamma_u}\big]=0$. • The expected Fisher information matrix $I(\boldsymbol{\theta})$ is finite, symmetric and positive definite. Furthermore, \begin{align*} n\big[I(\boldsymbol{\theta})\big]_{uv} = \mathbb{E}\bigg[{\partial{\ell}(\bm{\theta})\over \partial\gamma_u}\, {\partial{\ell}(\bm{\theta})\over \partial\gamma_v}\bigg] = - \mathbb{E}\bigg[{\partial^2{\ell}(\bm{\theta})\over \partial\gamma_u \partial\gamma_v}\bigg], \end{align*} for each $u, v= 1, 2,\ldots, r+s+3$.

Following Davison08, under the above assumptions, it follows that

align*[align* omitted — 248 chars of source]

where $\widehat{\bm \theta}_n$ is the MLE of the parameter of interest ${\bm \theta}$, ${\bm 0}$ is the $(r+s+3)\times 1$ zero vector and $I_{(r+s+3)\times (r+s+3)}$ is the ${(r+s+3)\times (r+s+3)}$ identity matrix. In practice, one may approximate the expected Fisher information matrix by its observed version, which is obtained from the Hessian matrix; see Appendix B. Then, the corresponding diagonal elements of the inverse of the observed Fisher information matrix can be used to approximate the standard errors (SEs) of the estimates.

ECM algorithm

We can implement the ECM algorithm presented in vslb:11. Specifically, let ${\bm y}=(y_{1},\ldots,y_{n})^{\top}$ and ${\bm u}=(u_{1},\ldots,u_{n})^{\top}$ be the observed and missing data, respectively, with $\bm Y$ and $\bm U$ being their corresponding random vectors. Thus, the complete data vector is written as ${\bm y}_{c}=({\bm y}^{\top}, {\bm u}^{\top})^{\top}$. From (ref), we obtain

equation[equation omitted — 185 chars of source]

Then, the complete data log-likelihood function for the skew-QBS-ACD model, associated with ${\bm y}_{c}=({\bm y}^{\top}, {\bm u}^{\top})^{\top}$, is given by (without the constant)

equation[equation omitted — 264 chars of source]

The implementation of the ECM algorithm requires the computation of the expected value of the complete log-likelihood equation in (ref), conditional on $ \bm U = \bm u$, that is,

eqnarray[eqnarray omitted — 440 chars of source]

where $b(y_t;\Xi_{q,t}) = \left[\sqrt{\frac{y_t \eta_q^2}{4\Xi_{q,t}}}-\sqrt{\frac{4\Xi_{q,t}}{y_t\eta_{q}^2}}\right]$, $\lambda=\delta/\sqrt{1-\delta^2}$, and $$ \widehat{u}_t=E[U_t|Y_t=y_t;\widehat{\bm \theta}]= \widehat{\delta} a(y_t;\widehat{\alpha},{\Xi}_{q,t})+ \sqrt{1-\widehat{\delta}^2}\Upsilon_{\Phi}\left[ \frac{\widehat{\delta} a(y_t;\widehat{\alpha},{\Xi}_{q,t})}{\sqrt{1-\widehat{\delta}^2}} \right] $$

$$ \widehat{u}_t^2=E[U_t^2|Y_t=y_t;\widehat{\bm \theta}]= \widehat{\delta}^2 a^2(y_t;\widehat{\alpha},{\Xi}_{q,t})+ (1-\widehat{\delta}^2)+\widehat{\delta}\sqrt{1-\widehat{\delta}^2} \Upsilon_{\Phi}\left[ \frac{\widehat{\delta} a(y_t;\widehat{\alpha},{\Xi}_{q,t})}{\sqrt{1-\widehat{\delta}^2}} \right] a(y_t;\widehat{\alpha},{\Xi}_{q,t}), $$ $$ \log(\Xi_{q,t}) = \widehat{\varpi} + \sum_{j=1}^{r}\widehat{\rho}_{j} \log(\Xi_{q,t-j}) + \sum_{j=1}^{s}\frac{\widehat{\sigma}_{j}\,y_{t-j}}{\Xi_{q,t-j}}, $$ with $\Upsilon_{\Phi}(x)=\phi(x)/\Phi(x)$ for $x\in\mathbb{R}$; see vslb:11.

The steps to obtain the ML estimates of the skew-QBS-ACD model via the ECM algorithm have been summarized in Algorithm (ref). Starting values are required to initiate the procedure, particularly, $\widehat{\alpha}^{(0)}, \widehat{\varpi}^{(0)}, \widehat{\rho}_{1}^{(0)},\ldots,\widehat{\rho}_{r}^{(0)}, \widehat{\sigma}_{1}^{(0)}, \ldots, \widehat{\sigma}_{s}^{(0)}$. These are obtained from saulolla:19, and then $\widehat{\lambda}^{(0)}$ from vslb:11. The iterations of Algorithm (ref) are repeated until a stopping criterion has been met, such as $\ell(\bm{\theta}^{v+1})-\ell(\bm{\theta}^{v})\leq \epsilon$, where $\epsilon$ is a prescribed small tolerance level.

algorithm[algorithm omitted — 1,396 chars of source]

Residual analysis

To assess the goodness of fit and departures from the assumptions of the model, we shall use the generalized Cox-Snell (GCS) residual, given by

equation[equation omitted — 106 chars of source]

where $\widehat{S}(\cdot)$ denotes the survival function fitted to the ACD data. This residual has a standard exponential, EXP(1), distribution when the model is correctly specified. Then, quantile-quantile (QQ) plots with simulated envelope can be used to assess the fit.

Monte Carlo simulation

We present the results of two Monte Carlo simulation studies for the skew-QBS-ACD($r, s, q$) model. Results are presented only for order of the lags set as $p=1$ and $q=1$, since a higher order for skew-QBS-ACD models did not improve the model fit considerably; see related results in b:10 and saulolla:19. The first study considers the evaluation of the performance of the ML estimation procedure developed in Algorithm (ref), while the second study evaluates the empirical distribution of the residuals.

ML estimation via the ECM algorithm

The first simulation scenario considers the following setting: sample size $n \in \{500, 1000, 2000\}$, vector of true parameters $(\varpi,\rho,\sigma,\alpha,\lambda) = (0.20,0.70,0.10,0.50,-0.50)$ and $(\varpi,\rho,\sigma,\alpha,\lambda) = (0.20,0.70,0.10,1.00,-0.50)$, and $q=\{0.20,0.50,0.80\}$, with $200$ Monte Carlo replications for each sample size. The skew-QBS-ACD samples were generated using the transformation

equation*[equation* omitted — 112 chars of source]

where $\eta_{q}$ is as in (ref), $\Xi_{q,t}$ is as in (ref), and $X_t$ is a SN random variable generated by the function rsn of the R package sn; see sn:19.

The ML estimation results are presented in Table (ref) wherein the empirical mean, bias, root mean squared error (RMSE), coefficients of skewness, and coefficients of (excess) kurtosis are all reported. A look at the results in Table (ref) allows us to conclude that, in general, as the sample size increases, the bias and $\textrm{RMSE}$ of the estimates decrease, as expected. Moreover, $\widehat{\varpi}$, $\widehat{\rho}$, $\widehat{\sigma}$, $\widehat{\alpha}$ and $\widehat{\lambda}$ seem all to be consistent and marginally asymptotically distributed as normal.

table[table omitted — 5,045 chars of source]

Empirical distribution of the residuals

In this subsection, we present a Monte Carlo simulation study to evaluate the performance of the GCS residuals. The sample generation as well as the simulation scenario are the same ($\alpha=0.50$) in Subsection (ref).

Table (ref) presents the empirical mean, standard deviation, coefficient of skewness and coefficient of kurtosis, whose values are expected to be 1, 1, 2 and 9, respectively, for $R^\textrm{GCS}$. Notice that the reference distribution for $R^\textrm{GCS}$ is standard exponential distribution. From Table (ref), we note that, in general, the considered residuals conform well with the reference distribution.

table[table omitted — 927 chars of source]

Application to price duration data

Exploratory data analysis

In this section, a real high frequency financial data set corresponding to price durations of BASF-SE stock on 19th April 2016 is used to illustrate the proposed methodology; see saulolla:19. The data were adjusted to remove intraday seasonality, namely, activity is low at lunchtime and high at the beginning and closing of the trading day. Table (ref) provides descriptive statistics for both plain and adjusted price durations. From this table, we note that both series show positive skewness and a high degree of kurtosis.

table[table omitted — 1,006 chars of source]

Estimation and model validation

Table (ref) presents the ML estimates, computed by the ECM algorithm, and SEs for the skew-QBS-ACD model parameters with $q = \{0.13,0.50\}$. On the one hand, the value of $q = 0.13$ was chosen through the profile log-likelihood, that is, for a grid of values of $q$, $0.01, 0.02,\ldots,0.99$, we estimated the model parameters and computed the corresponding log-likelihood functions. Then, the optimal value of $q$ was the one which maximized the log-likelihood function. On the other hand, the value of $q = 0.50$ corresponds to the skew-QBS-ACD in terms of the conditional median duration, such as the BS-ACD model. Table (ref) also reports the log-likelihood value, Akaike (AIC) and Bayesian information (BIC) criteria, and $p$-values of the Ljung-Box (LB) statistic, $Q(l)$ say, for up to $l$-th order serial correlation. The LB statistics evaluate the autocorrelation of the residuals. For comparative purpose, the results of the exponential ACD (EXP-ACD), generalized gamma ACD (GG-ACD) and BS-ACD models, are provided as well. The results of Table (ref) reveal that the skew-QBS-ACD models with $q = \{0.13,0.50\}$ provide better adjustments compared to other models based on the values of log-likelihood, AIC and BIC. However, these values are not substantially different for the two values of $q$. Finally, the LB $p$-values provide no evidence of serial correlation in the residuals for all models considered in this application.

Figure (ref) plots the estimated parameters of the skew-QBS-ACD model across $q\in\{0.01,\ldots, 0.99\}$. Results suggest that this BASF-SE series displays asymmetric dynamics: the estimates of ${\varpi}$ and ${\sigma}$ tend to be increasing across $q$, with a change of sign in the first case. On the other hand, the estimates of $\rho$, $\alpha$ and $\lambda$ appear to vary with the same pattern, but without an explicit trend.

table[table omitted — 1,927 chars of source]
figure[figure omitted — 631 chars of source]

Figure (ref) shows the QQ plots with simulated envelope of the generalized Cox-Snell residuals for the models considered in Table (ref). We see clearly that the skew-QBS-ACD models provide better fit than the EXP-ACD, GG-ACD and BS-ACD models.

figure[figure omitted — 649 chars of source]

Forecasting performance

Figure (ref) plots the in-sample and out-of-sample quantile forecasts from the adjusted skew-QBS-ACD and BS-ACD models. The construction of the plots is as follows. For each value of $q\in\{0.01,\ldots, 0.99\}$, the skew-QBS-ACD model is estimated and the respective estimates are used to obtain the forecast of the quantile of interest. Since the BS-ACD model is independent of the value of $q$, it is estimated only once and its forecasts are obtained from this single estimation. For in-sample forecasting, all data are used. In the case of out-of-sample forecasting, 2/3 of the data are used for estimation and the other 1/3 for forecasting.

From Figures (ref)(a,b), we observe that the forecasts of the skew-BS-ACD model are much closer to the actual values than those of the BS-ACD model. In fact, the forecasts of this latter model are quite far from the actual values. The superiority of the skew-QBS-ACD model may be due to different estimates obtained for the parameters; see, e.g., Figure (ref). For different quantiles, we obtain different estimates, which makes the forecasts closer to the actual values. The results of the estimated mean square error (MSE) between the actual data and in-sample forecasted values of the adjusted skew-QBS-ACD and BS-ACD models are 0.041 and 1.277, respectively. Moreover, the MSE between the actual data and out-of-sample forecasted values of the adjusted skew-QBS-ACD and BS-ACD models are 0.029 and 1.199, respectively. Thus, the MSE values do confirm the results found in Figures (ref)(a,b).

figure[figure omitted — 391 chars of source]

Concluding remarks

In this paper, we have proposed a quantile autoregressive conditional duration model for dealing with high frequency financial duration data. The proposed model is based on a reparameterized version of the skewed Birnbaum-Saunders distribution, which has the distribution quantile as one of its parameters. Thus, the proposed conditional duration model is defined in terms of a conditional quantile duration, allowing different estimates for different quantiles of interest. A Monte Carlo simulation is carried out to evaluate the performance of the maximum likelihood estimators, study obtained by the ECM algorithm. We have applied the proposed and some other existing models to a real financial duration data set from the German DAX stock exchange. The results are seen to be quite favorable to the proposed model both in terms of model fitting and forecasting ability. The gain in considering the dynamics in terms of conditional quantile duration seems quite evident, especially in the forecasting superiority of the model. As part of future research, it will be of interest to use the theory of estimating functions for estimating the model parameters ALLEN2013214. Furthermore, quantile autoregressive conditional duration models can be used to estimate the intraday volatility of a stock Tse-2012. Work on these problems is currently in progress and we hope to report these findings in future papers.

Funding: Helton Saulo gratefully acknowledges financial support from FAP-DF.

thebibliography\bibitem[Allen et al., 2013]{ALLEN2013214} Allen, D., Ng, K., and Peiris, S. (2013). \newblock Estimating and simulating {W}eibull models of risk or price durations: {A}n application to {ACD} models. \newblock {\em The North American Journal of Economics and Finance}, 25:214--225. \bibitem[Azzalini, 2019]{sn:19} Azzalini, A. (2019). \newblock {\em The {R} package sn: The Skew-Normal and Related Distributions such as the Skew-$t$ (version 1.5-4).} \newblock Universit\`a di Padova, Italia. \bibitem[Azzalini and Capitanio, 2014]{azzalinicap:14} Azzalini, A. and Capitanio, A. (2014). \newblock {\em The Skew-Normal and Related Families}. \newblock Cambridge University Press, Cambridge, England. \bibitem[Balakrishnan and Kundu, 2019]{balakundu:19} Balakrishnan, N. and Kundu, D. (2019). \newblock Birnbaum-saunders distribution: A review of models, analysis, and applications. \newblock {\em Applied Stochastic Models in Business and Industry}, 35:4--49. \bibitem[Bauwens et al., 2003]{bauwensetal:03} Bauwens, L., Galli, F., and Giot, P. (2003). \newblock The moments of log-{ACD} models. \newblock {\em CORE Discussion Paper, Available at SSRN: https://ssrn.com/abstract=375180 or http://dx.doi.org/10.2139/ssrn.375180}. \bibitem[Bhatti, 2010]{b:10} Bhatti, C. (2010). \newblock {The Birnbaum-Saunders autoregressive conditional duration model}. \newblock {\em Mathematics and Computers in Simulation}, 80:2062--2078. \bibitem[Bhogal and Variyam Thekke, 2019]{bhogalvariyam:19} Bhogal, S. K. and Variyam Thekke, R. (2019). \newblock Conditional duration models for high-frequency data: {A} review on recent developments. \newblock {\em Journal of Economic Surveys}, 33(1):252--273. \bibitem[Birnbaum and Saunders, 1969]{bs:69a} Birnbaum, Z. W. and Saunders, S. C. (1969). \newblock A new family of life distributions. \newblock {\em Journal of Applied Probability}, 6:319--327. \bibitem[Davison, 2008]{Davison08} Davison, A. C. (2008). \newblock {\em {Statistical Models}}. \newblock Cambridge Series in Statistical and Probabilistic Mathematics, Cambridge University Press. \bibitem[Engle and Russell, 1998]{er:98} Engle, R. and Russell, J. (1998). \newblock {Autoregressive conditional duration: {A} new method for irregularly spaced transaction data}. \newblock {\em Econometrica}, 66:1127--1162. \bibitem[Glaser, 1980]{gla:80} Glaser, R. E. (1980). \newblock Bathtub and related failure rate characterizations. \newblock {\em Journal of the American Statistical Association}, 75:667--672. \bibitem[Koenker and Xiao, 2006]{koenkexiao:06} Koenker, R. and Xiao, Z. (2006). \newblock Quantile autoregression. \newblock {\em Journal of the American Statistical Association}, 101(475):980--990. \bibitem[Pacurar, 2008]{pa:08} Pacurar, M. (2008). \newblock {Autoregressive conditional durations models in finance: {A} survey of the theoretical and empirical literature}. \newblock {\em Journal of Economic Surveys}, 22:711--751. \bibitem[S\'anchez et al., 2021]{sanchezetal:19} S\'anchez, L., Leiva, V., Galea, M., and Saulo, H. (2021). \newblock Birnbaum-saunders quantile regression and its diagnostics with application to economic data. \newblock {\em Applied Stochastic Models in Business and Industry}, 37:53--73. \bibitem[Saulo et al., 2019]{saulolla:19} Saulo, H., Le\ {a}o, J., Leiva, V., and Aykroyd, R. G. (2019). \newblock Birnbaum-{S}aunders autoregressive conditional duration models applied to high-frequency financial data. \newblock {\em Statistical Papers}, 60:1605--1629. \bibitem[Tse and Yang, 2012]{Tse-2012} Tse, Y.-k. and Yang, T. T. (2012). \newblock Estimation of high-frequency volatility: An autoregressive conditional duration approach. \newblock {\em Journal of Business and Economic Statistics}, 30. \bibitem[Vilca and Leiva, 2006]{vl:06} Vilca, F. and Leiva, V. (2006). \newblock {A new fatigue life model based on the family of skew-elliptical distributions}. \newblock {\em Communications in Statistics: Theory and Methods}, 35:229--244. \bibitem[Vilca et al., 2011]{vslb:11} Vilca, F., Santana, L., Leiva, V., and Balakrishnan, N. (2011). \newblock {Estimation of extreme percentiles in Birnbaum-Saunders distributions}. \newblock {\em Computational Statistics and Data Analysis}, 55:1665--1678.