EconBase
← Back to paper

Forecasting open-high-low-close data contained in candlestick chart

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

66,218 characters · 23 sections · 33 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
center[center omitted — 108 chars of source]

\vskip 0.5cm

center[center omitted — 394 chars of source]

\let \thefootnote \relax \footnotetext{ * Corresponding author. Correspondence to: School of Economics and Management, Beihang University, Beijing 100191, China. E-mail address: [email removed] (S.S. Wang).}

\vskip 0.5cm

center[center omitted — 47 chars of source]

Forecasting the (open-high-low-close)OHLC data contained in candlestick chart is of great practical importance, as exemplified by applications in the field of finance. Typically, the existence of the inherent constraints in OHLC data poses great challenge to its prediction, e.g., forecasting models may yield unrealistic values if these constraints are ignored. To address it, a novel transformation approach is proposed to relax these constraints along with its explicit inverse transformation, which ensures the forecasting models obtain meaningful open-high-low-close values. A flexible and efficient framework for forecasting the OHLC data is also provided. As an example, the detailed procedure of modelling the OHLC data via the vector auto-regression (VAR) model and vector error correction (VEC) model is given. The new approach has high practical utility on account of its flexibility, simple implementation and straightforward interpretation. Extensive simulation studies are performed to assess the effectiveness and stability of the proposed approach. Three financial data sets of the Kweichow Moutai, CSI $100$ index and $50$ ETF of Chinese stock market are employed to document the empirical effect of the proposed methodology.

\vskip 0.5cm { Key Words:} OHLC data; Candlestick chart; Unconstrained transformation; VAR; VEC

Introduction

Nowadays, the candlestick chart analysis has become one of the most intuitive and widely used methods for analyzing the price movements of financial products romeo2015study. And there has been vast literature on statistical methods to analyze the (open-high-low-close)OHLC data contained in candlestick chart, which may broadly be categorized into two main aspects: graphical analysis caginalp1998predictive, goo2007application, Lu2012, Tao2017Further, wan2018hidden and numerical analysis pai2005hybrid, Faria2009Predicting, caporin2013predictability, manurung2018algorithm, kumar2019stock .

From the essence of OHLC data, it represents the prices of certain financial product. Investors continue to make purchases and sells according to accurate predictions of OHLC data and thus earn profits is the fundamental incentive mechanism to maintain its effective operation liu2017integrated. Therefore, forecasting the OHLC data is one of the most important aspect among the various methods of investigating OHLC data, which is also our concern.

The existing literature on forecasting OHLC data have two shortcomings. First, the information contained in OHLC data is always under-utilized. For example, arroyo2011different used exponential smoothing method, multi-layer perception, K-nearest neighbor algorithm, autoregressive integrated moving average model, vector auto-regression (VAR) model and vector error correction (VEC) methods to perform regressions based on the Dow Jones Industrial Average index and Euro-Dollar Exchange Rate. In their study, only the center and range (Center & Range method, CRM), or the high and low prices (Min & Max method) of the OHLC data were used, which is the common practice of modelling OHLC data. However, for OHLC data, in addition to the high and low prices, there are open price and close price in the middle of the two boundaries, which have an explanatory power for the fluctuation of OHLC data and should be carefully considered in the model cheung2007empirical. Regrettably, the open price and close price are beyond the consideration scope of the CRM method and Min & Max method, and few articles fully considered the information contained in OHLC data.

Secondly, existing literature may not guarantee a meaningful forecasting of OHLC data. On the one hand, most machine learning models can only predict part information of OHLC data, such as high-low range 2009Forecasting, von2012forecasting, high and low prices von2014forecasting, or close price liu2012fluctuation. On the other hand, although some literature attempted to forecast the four variables of OHLC data, the inherent constraints of OHLC data were not well handled, see manurung2018algorithm and kumar2019stock. Specifically, the OHLC data requires that all the values are positive, the low price must be smaller than the high price, and the open and close prices must be within the two boundaries. The direct use of the time series modelling approaches for the four variables of OHLC data without accommodating the constraints may yield unrealistic and meaningless results, i.e., (a) the low price of the forecast period becomes negative without practical significance (shown in Fig.(ref)); (b) the predicted open price (or close price) breaks through the high price (or the low price) boundary (shown in Fig.(ref)); (c) the high price of the forecast period is smaller than the low price (shown in Fig.(ref)).

figure[figure omitted — 842 chars of source]

Apart from the aforementioned concerns, it is also necessary to propose a novel method to forecast OHLC data as a whole, from which we may benefit compared with the partial information forecasting methods. First, the prediction of open-high-low-close prices can provide significant help for investors to develop more sophisticated investment plans. Specifically, according to the traditional prediction of the close price, investors can only try to buy in certain financial product on the opening quotation and sell out on the closing quotation if upward trend is forecasted. While if the OHLC data is predicted, one is able to buy in certain financial product with price near the predicted low price, and sell out with price near the predicted high price to gain excess profits. Second, candlestick charts can be drew according to OHLC data, whose pattern can reveal the power of demand and supply in financial markets, as well as reflect market conditions and emotion nison2001japanese, Tsai2014Stock. Graphical analysis can give further investment advise based on the forecasted candlestick chart. Third, full information of OHLC data can be reserved. A full information set of the open-high-low-close prices can enhance the reliability and explanatory ability of researches. As pointed out in cheung2007empirical and fiess2002towards, the open-high-low-close prices as a whole are proved to be of significant power in explaining the price fluctuation. Finally, opening the possibility of applying a wide range of multivariate modelling techniques to explore the dynamic and structural relationship between the four-dimensional variables of OHLC data fiess2002towards.

To this end, we proposed a new transformation method to transform the OHLC data from the original four-dimensional constrained subspace to the four-dimensional real domain full space along with its explicit inverse transformation, which ensures the predicting models obtain meaningful open-high-low-close prices. As an example combining with time series analysis, we illustrated the detailed procedure of the VAR and VEC modelling of OHLC data. Ample simulation experiments under different forecast periods, time period basements as well as signal-noise ratios were conducted to validate the effectiveness and stability of the transformation method. Further, three financial data sets of the Kweichow Moutai, CSI $100$ index and $50$ ETF from their appearance on the Chinese stock market to $14/6/2019$ were used to illustrate the empirical utility of the proposed method. The results showed a satisfying prediction effect.

Compared with existing literature, the advantages of the proposed transformation method are three folded. First, it takes full advantage of the information contained in the OHLC data. The open price, high price, low price and close price are all considered, which enables a more efficient analysis. Second, it can well handle the inherent constraints in the OHLC data. The inherent constraints of OHLC data are satisfied during the whole process of numerical modelling without increasing the complexity of the model, which enables more interpretable results. Third, the proposed method provides a unified framework for a variety of positive interval data which has minimum and maximum boundaries greater than zero, and multi-valued sequences within the two boundaries, such as daily temperature and profit of companies. Meanwhile, the method of dealing with the transformed variables can be generalized to other statistical models and machine learning methods. From this perspective, the proposed method provides a completely new and useful alternative for OHLC data analysis, enriching the existing literature.

The remainder of this paper is organised as follows. In Section (ref), we introduce the mathematical definition of OHLC data and its inherent constraints, and in Section (ref), we propose the transformation and inverse transformation formulas to relax the inherent constraints of OHLC data, and illustrate the VAR and VEC modelling process for OHLC data. Section (ref) presents simulation studies and Section (ref) shows the empirical application of the proposed method in real financial market. Finally, we conclude with a brief discussion in Section (ref).

Preliminaries

To grasp an intuitive picture of the candlestick chart, here we take the daily candlestick chart as an example (in this article, if there is no special explanation, all candlestick charts refer to daily candlestick chart), as shown in Fig.(ref). Obviously, a daily candlestick chart can not only record the open price, high price, low price and close price of a certain stock on the day, but also can visually reflect the difference between any two prices.

figure[figure omitted — 142 chars of source]

Generally, the candlestick chart is divided into two categories as shown in Fig.(ref). Specifically, Fig.(ref)(a) indicates that the close price is higher than the open price, which corresponding to a bull market; while Fig.(ref)(b) corresponds to a bear market with the close price being lower than the open price. In the stock market of the United States, green and red are habitually used to mark the real body of candlestick chart of bull market and bear market respectively. If the daily candlestick charts are arranged in chronological order, a sequence reflecting the historical price changes of a certain financial product is formed, called the candlestick chart series, and the corresponding data is called OHLC series.

Actually, the essence of OHLC series is a four-dimensional time series of stock prices with three inherent constraints. First, all the four prices should be positive due to the limitation that the values of OHLC data in the financial market cannot be less than zero; Second, the high price must be higher than the low price of the same day; Third, the open price and close price should fall into the two boundaries. To represent it in mathematical form, for any time period $t$, we give the following definition of OHLC data.

definitionA four-dimensional vector $\bm{X}_t=(x_t^{(o)}, x_t^{(h)}, x_t^{(l)}, x_t^{(c)})^T$ is a typical OHLC data if it satisfies: \begin{itemize} • $x_t^{(l)} > 0$; • $x_t^{(l)} < x_t^{(h)}$; • $x_t^{(o)}, x_t^{(c)} \in [x_t^{(l)},x_t^{(h)}]$. \end{itemize} Here $x_t^{(o)}$ is the $t$ period daily open price, $x_t^{(h)}$ is the $t$ period daily high price, $x_t^{(l)}$ is the $t$ period daily low price, and $x_t^{(c)}$ is the $t$ period daily close price.

For the time period $\mathcal{T}=[1,T]$, the collection of $\bm{X}_t$ for any $t\in\mathcal{T}$ forms the OHLC time series, denoted by $$\bm{S}=\{\bm{X}_t\}_{t=1}^T.$$ Compared with the ordinary real domain vector, the biggest difference of the vectors in $\bm{S}$ is that there are intrinsic constraint formulas between its four components, which poses great challenge to classical statistical analysis. Specifically, to establish a time series model of OHLC series, the most difficult problem is how to always ensure that the calculation process and the prediction results are also subject to these constraint formulas. That is, after obtaining the prediction results at the prediction period $(T+m) \; (m \in \mathbb{R}^+)$ by time series modelling, it must be ensured that:

eqnarray*[eqnarray* omitted — 147 chars of source]

Obviously, these constraints are not guaranteed to be valid if we directly apply the time series forecasting methods to the original four time series of OHLC data. To address this problem, a common practice is to remove these inherent constraints via some proper data transformation, then we can forecast the transformed time series data freely. Finally, we can obtain the forecaster of the original OHLC data using the corresponding inverse transformation.

Methodology

In this section, we first proposed a flexible transformation method along with its inverse transformation method for OHLC data, as well as a model-independent framework for modelling OHLC data are proposed in Section (ref). And then we use VAR and VEC models as a case of the framework and present the corresponding forecasting procedure in Section (ref).

Data transformation method

Note that from Definition (ref), the first constraint is $x_t^{(l)}>0$, which can be relaxed via the most commonly used logarithm transformation. That is,

equation[equation omitted — 56 chars of source]

It is quite clear that the transformed data $y_t^{(1)}$ in Eq.((ref)) satisfies $-\infty<y_t^{(1)}<+\infty$ with no positive constraint. Moreover, it not only preserves a positive relative relationship between the original data as logarithm transformation is a monotonous increasing function, but also compresses the scale of the data, which reduces the absolute values of the original data and makes the data more stable to some extent.

Secondly, in order to guarantee the second constraint $x_t^{(l)}<x_t^{(h)}$, i.e., $x_t^{(h)}-x_t^{(l)}>0$, the same practice as that in Eq.((ref)) yields

equation[equation omitted — 67 chars of source]

where $y_t^{(2)}$ is also free of any constraint, which can be modelled easily.

Finally, the last constraint is $x_t^{(o)}, x_t^{(c)} \in [x_t^{(l)},x_t^{(h)}]$, implying that both the open price and close price must be greater than the low price and less than the high price. Obviously, without properly processing the raw data, it is very likely that the predicted open/close price is beyond boundaries. To remedy this situation, inquired by the idea of convex combination, here we introduce two proxy data $\lambda_t^{(o)}$ and $\lambda_t^{(c)}$, which are formulated as

equation[equation omitted — 184 chars of source]

Obviously, $0\leqslant\lambda_t^{(o)},\lambda_t^{(c)}\leqslant1$, and the original data $x_t^{(o)}$ and $x_t^{(c)}$ can be obtained as follows, respectively. That is,

equation[equation omitted — 95 chars of source]
equation[equation omitted — 95 chars of source]

Thus, the original constraint $x_t^{(o)}, x_t^{(c)} \in [x_t^{(l)},x_t^{(h)}]$ reduces to $0\leqslant\lambda_t^{(o)},\lambda_t^{(c)}\leqslant1$ if we deal with the proxy data $\lambda_t^{(o)}$ and $\lambda_t^{(c)}$ instead of $x_t^{(o)}$ and $x_t^{(c)}$. Moreover, $\lambda_t^{(o)}$ and $\lambda_t^{(c)}$ make sense. Specifically, if $\lambda_t^{(o)}$ has a growing trend, the open price $x_t^{(o)}$ will be closer to the high price $x_t^{(h)}$; while a downward trend of $\lambda_t^{(o)}$ indicates that the open price $x_t^{(o)}$ will be closer to the low price $x_t^{(l)}$. Similarly, the explanation for $\lambda_t^{(c)}$ can also be obtained.

To further remove the constraint $0\leqslant\lambda_t^{(o)},\lambda_t^{(c)}\leqslant1$ on $\lambda_t^{(o)}$ and $\lambda_t^{(c)}$, following the idea of logistic regression, we propose the logit transformation to obtain the unconstrained data $y_t^{(3)}$ and $y_t^{(4)}$ as follows:

equation[equation omitted — 85 chars of source]
equation[equation omitted — 85 chars of source]

Up to now, via the transformation process, the raw OHLC data $\bm{X}_t=(x_t^{(o)}, x_t^{(h)}, x_t^{(l)}, x_t^{(c)})^T$ is transformed to the unconstraint four-dimensional data $\bm{Y}_t=(y_t^{(1)}, y_t^{(2)},y_t^{(3)},y_t^{(4)})^T.$ In summary, the proposed transformation method can be described as

equation[equation omitted — 539 chars of source]

where $\lambda_t^{(o)}$ and $\lambda_t^{(c)}$ are defined in Eq.((ref)). Not only does the transformation in Eq.((ref)) have ranges from $-\infty$ to $+\infty$ and an explicit inverse for the values in its range, but it also shares the flexibility of the well known log- and logit- transformation.

Therefore, the predictive modelling of the OHLC series $\{\bm{X}_t\}_{t=1}^T$ can be transformed into the predictions for the unconstrained series $\{\bm{Y}_t\}_{t=1}^T$ with the whole real number domain and variance stability, which will provide significant convenience for subsequent statistical modelling. That is, we can apply the classical forecasting models (ARMA, ARIMA, VAR, VEC etc) or machine learning models to $\{\bm{Y}_t\}_{t=1}^T$. After obtaining the forecaster of $\bm{Y}_t$, we can obtain the corresponding forecaster of $\bm{X}_t$ via the inverse transformation as follows:

equation[equation omitted — 630 chars of source]

where

equation[equation omitted — 179 chars of source]

The above transformation Eq.((ref)) and inverse transformation Eq.((ref)) methods provide a new perspective for forecasting the OHLC data, which will make the prediction results obey its three inherent constraints listed in Definition (ref). Furthermore, the proposed method has great feasibility, which can be easily generalized to any type of positive interval data that owns the minimum and maximum boundaries greater than zero, and multi-valued sequences between the two boundaries.

It should be noticed that in the transformation process, we assume that $x^{(o)}, x^{(h)}, x^{(l)}$, and $x^{(c)}$ are not equal to each other $(\text{except for} \; x^{(o)} \equiv x^{(c)})$. In other words, $x^{(h)} \neq x^{(l)} \neq 0$ and $\lambda^{(o)}, \lambda^{(c)} \notin \left\{ 0,1 \right\}$. However, such assumptions are inevitably spoiled sometimes in the real financial market. Here we list the circumstances that makes these assumptions invalid and give the measurement to deal with them accordingly. (1) When the subject is on trade suspension and all prices equal to $0$, namely, $x^{(o)}=x^{(h)}=x^{(l)}=x^{(c)}=0$, we exclude these extreme cases in the raw data. (2) When $\lambda^{(o)}$ or $\lambda^{(c)}$ is equal to $0$, it corresponds to $x^{(o)}=x^{(l)}$ or $x^{(c)}=x^{(l)}$, respectively. We add a random term to $x^{(o)}$ or $x^{(c)}$ and make $\lambda^{(o)}$ or $\lambda^{(c)}$ slightly greater than 0. (3) When $\lambda^{(o)}$ or $\lambda^{(c)}$ is equal to $1$, it indicates that $x^{(o)}=x^{(h)}$ or $x^{(c)}=x^{(h)}$, respectively. We subtract a random term from $x^{(o)}$ or $x^{(c)}$ to make $\lambda^{(o)}$ or $\lambda^{(c)}$ slightly less than $1$. (4) When certain subject reaches limit-up or limit-down as soon as the opening quotation, that is, $x^{(o)}=x^{(h)}=x^{(l)}=x^{(c)}\neq0$. If limit-up(limit-down) happens, we firstly multiply $x^{(c)}(x^{(o)})$ and $x^{(h)}$ by 1.1 to make a relatively large interval. And then conduct measurements given in circumstances (2) and (3).

In summary, the general forecasting framework for OHLC data with $T$ periods is described in Algorithm (ref).

algorithm[algorithm omitted — 767 chars of source]

The VAR and VEC modelling process for OHLC data

Here we employ the VAR and VEC models as an example of the framework proposed in Algorithm (ref) and present the corresponding procedure for forecasting OHLC data.

VAR model for OHLC data

As one of the most widely used multiple time series analysis methods, the VAR model proposed by sims1980macroeconomics has become an important research tool in economic studies, with advantages of capturing the linear interdependencies among multiple time series pesaran1998autoregressive. According to Algorithm (ref), we should build the forecast model for the unconstrained four-dimensional time series data $\{\bm{Y}_t\}_{t=1}^T$. Without loss of generality, we first assume that all time series in $\bm{Y}_t$ are stationary, then a $p$-order ($p\geqslant1$) VAR model, denoted by VAR($p$), can be formulated as:

equation[equation omitted — 1,589 chars of source]

VEC model for OHLC data

Note that the reliability of the VAR model estimation is closely related to the stationarity of the variable sequences. If this assumption does not hold, we may use a restricted VAR model, i.e., the VEC model, in the presence of the co-integration among variables; otherwise the variables have to be differenced by $d$ times firstly until they can be modelled by VAR or VEC model. As evidenced in cheung2007empirical, for the US stock markets, stock prices are typically characterized by $I(1)$ processes and the daily highs and lows follow a co-integration relationship. This implies that perhaps the VEC model is a more practically relevant model than the VAR model in the context of forecasting the OHLC series. Here we use the augmented Dickey-Fuller (ADF) unit root test to examine the stationary of each variable, and the Johansen test johansen1988statistical is employed to determine the presence or absence of the co-integration relationship.

Assume that $\bm{Y}_t$ is integrated of order one, then the corresponding VEC model takes the following form:

equation[equation omitted — 148 chars of source]

where $\Delta$ denotes the first difference, $\sum_{j=1}^{p-1}\bm{\Gamma}_j\Delta\bm{Y}_{t-j}$ and $\bm{\gamma}\bm{\beta}^T\bm{Y}_{t-p}$ are the VAR component of the first difference and error-correction component, respectively. Here $\bm{\Gamma}_j$ is a $4\times 4$ matrix that represents short-term adjustments among variables across four equations at the $j$-th lag; two matrices, $\bm{\gamma}$ and $\bm{\beta}$, are of dimension $4\times r$ with $r$ being the order of co-integration, where $\bm{\gamma}$ denotes the speed of adjustment (loading) and $\bm{\beta}$ represents the co-integrating vector, which can be obtained by the Johansen test johansen1988statistical, cuthbertson1992applied; $\bm{\alpha}$ is a $4\times 1$ constant vector representing a linear trend; $p$ is a lag structure; and $\bm{w}_t$ is the $4\times 1$ vector of white noise error term.

For the VEC model in Eq.((ref)), johansen1991estimation employed the full information maximum likelihood method to implement its estimation. Specifically, the main procedure consists of (i) testing whether all variables are integrated of order one by applying a unit root test Lai1991A, (ii) determining the lag order $p$ such that the residuals from each equation of the VEC model are uncorrelated, (iii) regressing $\Delta\bm{Y}_t$ against the lagged differences of $\Delta\bm{Y}_t$ ($\bm{Y}_{t-p}$) and estimating the eigenvectors (co-integrating vectors) from the Canonical correlations of the set of residuals from these regression equations, and finally (iv) determining the order of co-integration $r$.

Discussion of parameter selection in VAR and VEC models

Finally, we discuss the determination of the lag order $p$ in the VAR model and the order of co-integration $r$ in the VEC model for modelling OHLC data. First, for $p$, on the one hand, it should be made large enough to fully reflect the dynamic characteristics of the constructed model; while on the other hand, an increase of $p$ will cause an increase of the parameters to be estimated, thus the freedom degree of the model decreases. A trade-off must be evaluated to choose $p$, the common used criterions in practice are AIC, BIC and HQ (Hannan-Quinn). In this paper, we prefer AIC because of its conciseness, which is formulated as

equation[equation omitted — 135 chars of source]

where $T$ stands for the total period number of OHLC series, $p$ is VAR lag order, $K$ is the VAR dimension, and $\hat{u}_{ij}=\hat{Y}_{j}^{(i)}-Y_{j}^{(i)}(1\leq i\leq4,1\leq j\leq T)$ represents for the residuals of the VAR model. The optimal $p$ is obtained via minimizing $\mbox{AIC}(p)$.

Second, as the order of co-integration, $r$ indicates the dimension of the co-integrating space, and (i) if the rank of $\bm{\gamma}\bm{\beta}^T$ is $4$, i.e., $r=4$, $\bm{Y}_t$ is already stationary; (ii) if $\bm{\gamma}\bm{\beta}^T$ is a null matrix, i.e., $r=0$, the proper specification of Eq.((ref)) is one without the error correction term and degenerate to a VAR model; and (iii) if the rank of $\bm{\gamma}\bm{\beta}^T$ is between $0$ and $4$, i.e., $0<r<4$, there exist $r$ linearly independent columns in the matrix and $r$ co-integrating relations in the system of equations. Along the line of johansen1991estimation, $r$ is determined by constructing the $``$Trace" or $``$Eigen" test statistics, which are two widely used methods in Johansen test. For more details, please refer to the monograph johansen1995likelihood and lutkepohl2005new.

Unified modelling framework for OHLC data

In summary, as one of the popular econometric forecasting models in Step $4$ of Algorithm (ref), the main implementation of VAR and VEC modelling can be summarized as Algorithm (ref).

algorithm[algorithm omitted — 1,544 chars of source]

Incorporating Algorithm (ref) into Algorithm (ref), we can obtain the unified framework for statistical modelling of the OHLC data, shown in Fig.(ref).

figure[figure omitted — 197 chars of source]

Simulations

We assess the performance of our proposed method via finite sample simulation studies. We firstly describe the data construction in Section (ref); then give five indicators to measure the difference between predicted values and observed ones in Section (ref); and finally report the simulation results in Section (ref).

Data construction

We generate simulation data under the VAR model structure in Eq.((ref)) as follows: (i) Assign the lag period $p$, and the coefficient matrices $\bm{A}_1, \bm{A}_2, \cdots, \bm{A}_p$; (ii) Generate an original four-dimensional vector $\bm{Y}_1=[y^{(1)}_{1}, y^{(2)}_{1}, y^{(3)}_{1}, y^{(4)}_{1}]^T$; (iii) Generates $\{\bm{Y}_t\}_{t=2}^T$ in a sequence via the VAR($p$) model $$\bm{Y}_t=\bm{A}_1\bm{Y}_{t-1}+...+\bm{A}_p\bm{Y}_{t-p}+\bm{w}_t,$$ where $\bm{w}_t$ follows the multivariate normal distribution with zero mean and covariance matrix $\Sigma_{\bm{w}}$. Finally, the simulated OHLC data $\{\bm{X}_t\}_{t=1}^T$ are generated by applying the inverse transformation formula in Eq.((ref)).

In order to evaluate the performance of the proposed method with different variance component levels, we consider the following scenarios:

itemize• Scenario 1: $p=1, T=220, \bm{Y}_1=[4,0.7,-0.85,0]^T$ and \begin{equation*} \bm{A}_1=\begin{pmatrix} 0.55 & 0.12 & 0.12 & 0.12 \\ 0.12 & 0.55 & 0.12 & 0.12 \\ 0.12 & 0.12 & 0.55 & 0.12 \\ 0.12 & 0.12 & 0.12 & 0.55 \end{pmatrix}, \end{equation*} and $\Sigma_{\bm{w}}$ is a $4\times 4$ diagonal matrix with diagonal element being $0.05^2$, i.e., $$\Sigma_{\bm{w}}=\mbox{diag}\{0.05^2,0.05^2,0.05^2,0.05^2\}.$$ • Scenario 2: $p, T, \bm{Y}_1$ and $\bm{A}_1$ follows Scenario $1$ except that $$\Sigma_{\bm{w}}=\mbox{diag}\{0.07^2,0.07^2,0.07^2,0.07^2\}.$$ • Scenario 3: $p, T, \bm{Y}_1$ and $\bm{A}_1$ follows Scenario $1$ except that $$\Sigma_{\bm{w}}=\mbox{diag}\{0.03^2,0.03^2,0.03^2,0.03^2\}.$$

All these scenarios present the transformed unconstrained time series data $\{\bm{Y}_t\}_{t=1}^T$ follows VAR$(1)$, only with different variance component levels according to median (Scenario $1$), low (Scenario $2$) and strong (Scenario $3$) signal-to-noise ratios, respectively. A higher signal-noise ratio means the information contained in the data comes more from the signal rather than the noise, indicating a better quality of data. On the contrary, a lower signal-noise ratio means that the noise carries more interference, indicating a worse quality of data.

Note that the raw simulation data has $220$ periods, we only take $21$ to $220$ periods as the final simulated data set as the data generated initially may be highly volatile. Take Scenario $1$ as an illustration, Fig.(ref) shows the simulated OHLC series $\{\bm{X}_t\}_{t=1}^T$.

figure[figure omitted — 163 chars of source]

Measurements

According to the process illustrated in Fig.(ref), the VAR and VEC models are used to measure the statistical relationships between $\bm{X}_t$ and $\bm{Y}_t$. As corrado1992filter and marshall2008candlestick pointed out, short-term technical analysis can be more helpful to investors than long-term technical analysis. Thus, we focus on relatively short-time analysis here. Specifically, $q$ periods of the simulated data, namely the time period basement, are used to train the model, and make prediction ahead of $m$ periods. And we set $q$ ranges from $30$ to $70$, and $m=1,2,3$. For each setting $(q,m)$, $\bm{Y}_i^{(q)}$ scroll forward one period, and predict $(T-q-m+1)$ times in total, as indicated in Fig.(ref).

figure[figure omitted — 155 chars of source]

The predicted $\widehat{\bm{Y}_i}^{(q)}$ are firstly derived, then the predicted $\widehat{\bm{X}_i}^{(q)}$ are obtained based on the inverse transformation formulas Eq.((ref)). We evaluated the effectiveness of prediction in term of five measurements, which are defined as follows:

itemize• The mean absolute percentage error (MAPE) \begin{equation*} MAPE = \frac{100%}{k}\sum\limits_{i = 1}^{k}\left|\frac{x^{(*)}_i-\widehat{x}^{(*)}_i}{x^{(*)}_i}\right|, \end{equation*} where $x^{(*)}_i$ and $\widehat{x}^{(*)}_i$ are the actual and forecasted values with $x^{(*)}_i$ indicating $x^{(o)}_i$, $x^{(h)}_i$, $x^{(l)}_i$, or $x^{(c)}_i$, respectively; $k$ is the number of forecasted points; • The standard deviation (SD) is defined as the empirical standard derivation of the forecasted values $\{\widehat{x}^{(*)}_i\}_{i=1}^k$, i.e., \begin{equation*} SD =\sqrt{ \frac{1}{k-1}\sum\limits_{i = 1}^{k}\Big(\widehat{x}^{(*)}_i-\bar{\widehat{x}}^{(*)}\Big)^2 }, \end{equation*} where $\bar{\widehat{x}}^{(*)}=\sum_{i=1}^k\widehat{x}^{(*)}_i/k$; • The root mean squared error (RMSE) as defined in neto2008centre \begin{equation*} RMSE=\sqrt{\frac{1}{k}\sum\limits_{i = 1}^{k}\Big(x^{(*)}_i-\widehat{x}^{(*)}_i\Big)^2} \end{equation*} • The RMSE based on the Hausdorff distance (RMSEH) defined in de2006adaptive \begin{equation*} RMSEH=\sqrt{\frac{1}{k}\sum\limits_{i = 1}^{k}\Big(\left|\frac{x_i^{(h)}+x_i^{(l)}}{2}-\frac{\widehat{x}_i^{(h)}+\widehat{x}_i^{(l)}}{2}\right| +\left| \frac{x_i^{(h)}-x_i^{(l)}}{2}-\frac{\widehat{x}_i^{(h)}-\widehat{x}_i^{(l)}}{2}\right|\Big)^2} \end{equation*} • The accuracy ratio (AR) adopted in hu2007application \begin{equation*} AR=\left\{ \begin{array}{lcl} \frac{1}{k}\sum\limits_{i = 1}^{k}\frac{w(SP_i \cap \widehat{SP_i})}{w(SP_i \cup \widehat{SP_i})}, \qquad \qquad if\ \ \ \ (w(SP_i \cap \widehat{SP_i})\not=0)\\ 0, \qquad \qquad \qquad \qquad \qquad \mbox{if}\ \ \ (w(SP_i \cap \widehat{SP_i})=0)\\ \end{array} \right., \end{equation*} where $w(SP_i \cap \widehat{SP_i})$ and $w(SP_i \cup \widehat{SP_i})$ represent the length of the intersection and union between the observation interval $[x_i^{(l)}, x_i^{(h)}]$ and the prediction interval $[\widehat{x}_i^{(l)}, \widehat{x}_i^{(h)}]$, respectively.

Smaller values of MAPE, RMSE, RMSEH and a larger AR indicate a better result; and a smaller SD indicates a more stable result.

Results of simulations

For Scenario $1$, we summarize the results with $q=40, 50, 70$ and $m=1,2,3$ in Table (ref). From which we can see that: (i) the overall performance of the proposed method in terms of these five measurements are satisfying and stable; (ii) for fixed $q$, a smaller forecast period $m$ will make the forecasted results more accurate with smaller values of MAPE, RMSE, RMSEH and larger AR. And there is no obvious pattern in terms of SD.

table[table omitted — 2,249 chars of source]

Moreover, we show more results with $q$ ranging from $30$ to $70$ and $m=1,2,3$ to further demonstrate the performance of the proposed method. Specifically, Fig.(ref) summarize the results in terms of MAPE (the left panel), SD (the middle panel), and RMSE (the right panel), respectively; while Fig.(ref) shows the RMSEH and AR of the forecasted results.

figure[figure omitted — 1,809 chars of source]
figure[figure omitted — 366 chars of source]

Basically from Fig.(ref), under different $q$ and $m$, the MAPE is between $3.08\%$ and $7.13\%$, the standard deviation is between $0.048$ and $0.158$ and the RMSE is between $0.051$ and $0.148$, indicating a good prediction accuracy and stability. As the forecast period $m$ increases, these three indicators will increase synchronously, decreasing the prediction accuracy. While the prediction performance shows a trend of getting better first and then getting worse as $q$ increases. From Fig.(ref), the RMSEH maintains a small value between $0.083$ and $0.152$. Meanwhile, the AR is relatively high, varies from $0.842$ to $0.903$, which illustrates the prediction interval is closely coincided with the observation interval, indicating a satisfying prediction effect.

figure[figure omitted — 1,205 chars of source]

For Scenario $2$ and $3$, we also conduct simulations along the line with Scenario $1$, and the results exhibit the same trend. For the sake of space, we only show the forecasted results in terms of MAPE in Fig.(ref) for Scenario $2$ with low signal-noise ratio (the first row) and Scenario $3$ with high signal-noise ratio (the second row), respectively. The MAPE in the first row of Fig.(ref) is between $4.29\%$ and $9.93\%$, while the corresponding MAPE in the second row of Fig.(ref) is between $1.89\%$ and $4.33\%$. The left panel of Fig.(ref) corresponds to the MAPE with medium signal-noise ratio, whose MAPE (ranges form $3.08\%$ to $7.13\%$) is between those in the second and the first row of Fig.(ref), which indicates the accuracy of the predicted results decreases as the signal-noise ratio decreases.

Empirical analysis

We illustrated the practical utility of the proposed method via three different kinds of real data sets: the OHLC series of the Kweichow Moutai, CSI $100$ index, and $50$ ETF in the financial market of China. For each case, we first briefly describe the data set in Section (ref), then apply the proposed method with different prediction basement period $q$ and prediction period $m$, and report the performance in terms of MAPE, RMSE, RMSEH and AR in Section (ref).

Raw OHLC data set description

itemize• OHLC series of the Kweichow Moutai. The Kweichow Moutai is a well-known company in Chinese liquor industry, which has a long history and its stamp (SH: $600519$) is an important part of the China Securities Index (CSI $100$). Here we studies its OHLC series with the time ranging from $27/8/2001$ to $14/6/2019$, yielding $4243$ data in total. • OHLC series of the CSI $100$ index. The CSI $100$ index is composed by the largest $100$ blue stocks selected from the CSI $300$ index stocks, reflecting the overall situation of the companies with the most market influence power in the Shanghai and Shenzhen stock markets. China Securities Index Co.,Ltd officially issued the CSI $100$ index on $30/12/2005$ and $1000$ being its base data and base point. We collected the OHLC series of the CSI $100$ index from $30/12/2005$ to $14/6/2019$, with a total of $3269$ periods. • OHLC series of the 50 ETF. The SSE $50$ index (code: 510050) is China's first transactional exchange traded fund, compiled by the Shanghai Stock Exchange, whose base date and base point are $31/12/2003$ and $1000$, respectively. The investment objective of the $50$ ETF is to closely track the SSE $50$ index, minimizing tracking deviation and tracking error. This paper collected $3481$ OHLC data samples of the 50 ETF from $23/2/2005$ to $14/6/2019$.

Results of empirical analysis

For each raw data set $\{\bm{X}_t\}_{t=1}^T$, the very beginning operation is conducting the unconstrained transformation. In terms of the unconstrained data set $\{\bm{Y}_t\}_{t=1}^T$, we first apply the ADF test together with the auto-correlation function (ACF) plot to examine the stability of each of the four variables. If any variable is non-stationary, the Johansen test is further employed to examine the co-integration relationship between these four variables. If no co-integration relationship exists, take one-order difference of the non-stationary variables and restart the ADF test. In brief, we operate the proposed method shown in Fig.(ref) with $q$ varying from $50$ to $90$ and $m=1,2,3$.

Specifically, we take the first vector times series $\{\bm{Y}_t\}_{t=1}^{90}$ (i.e., $\bm{Y}_1^{(90)}$ in Algorithm (ref)) of the 50 ETF as an example to illustrate our modelling process. At the significance level of $\alpha=0.1$, the four time series are stationary except $\{\bm{y}_t^{(1)}\}_{t=1}^{90}$ with the p-value of ADF test being $0.628$, indicating that $\{\bm{Y}_t\}_{t=1}^{90}$ cannot be modelled by VAR model. The ACF plots in Fig.(ref) further demonstrate the distinct auto-correlation and non-stationary of $\{\bm{y}_t^{(1)}\}_{t=1}^{90}$. Then, the Johansen test is applied to examine the co-integration relationship between the four variables in $\{\bm{Y}_t\}_{t=1}^{90}$, and the essence of Johansen test based on “Trace" is investigating the number of co-integration vectors, which is recorded as $r$. The results show that the possible of $r\leq2$ is less than $1\%$ and the possible of $r\leq3$ is over than $10\%$, thus the $r$ in VEC model is determined as $3$. Finally, the VEC model with the order of co-integration $r=3$ is established and the prediction values $\widehat{\bm{Y}_1}^{(90)}$ are obtained by the regression function. Through inverse transformation method, we can obtain predicted $\widehat{\bm{X}_1}^{(90)}$. Traverse the following vector time series $\{\widehat{\bm{Y}}_{t+l}\}_{t=1}^{90} (l=0,1,\cdots, T-q-m)$, and the prediction accuracy can be evaluated.

figure[figure omitted — 692 chars of source]

The performance in terms of MAPE for these three data sets of the Kweichow Moutai, CSI 100 index and 50 ETF is reported in the left, middle and right panels of Fig.(ref), respectively. From Fig.(ref), the pattern of the performance exhibits similarly. It is obvious that under different setting $(q,m)$, the prediction results are quite stable with all the values of MAPE below $2.87\%$ for the Kweichow Moutai, $2.30\%$ for the CSI $100$ index, and $2.37\%$ for the the 50 ETF, respectively. Moreover, the fewer the $m$ is, the more accurate the prediction results are.

figure[figure omitted — 1,764 chars of source]
table[table omitted — 1,785 chars of source]

Specifically, take the prediction basement period of $q=90$, and proceed one-step forecast ($m=1$) as an example, we summarized the results in term of MAPE, RMSE, RMSEH and AR in Table (ref), and compared them with the Naive method proposed by arroyo2011different, whose indicators are calculated by taking the price of the previous day as the price of the day. We also give the results of unilateral t-test in Table (ref). The null hypothesis of the t-test of all indicators except AR is that the value of the proposed method is no less than that of the naive method. If the p-value is less than 0.01, it indicates that we can reject the null hypothesis at 99% confidence level (remarked by a superscript $*$ ), and choose the alternative hypothesis that the value of the proposed method is less than that of the naive method. The t-test of AR is just the opposite, if the p-value is less than 0.01, it represents that the value of the proposed method is greater than that of the naive method. From Table (ref), it can be seen that the forecasting effect of the proposed method is significantly accurate and most indicators pass the t-test. More specifically,

itemize• for the Kweichow Moutai, the MAPE of the open price is only $0.821\%$, and the MAPE of the close price, high price, and low price of are also controlled within $1.7\%$. There is no evidence that the MAPE of the predicted close price from the proposed method are smaller than that of the naive method. Whereas, comparing to the naive method, the MAPE of the open price is improved by $46.55\%$, $6.71\%$ for the high price and $14.14\%$ for the low price; the RMSE of the open price is improved by $46.55\%$, $3.51\%$ for the high price, and $11.93\%$ for the low price. As for the RMSEH and the AR, the results of the proposed method are improved by $7.92\%$ and $7.97\%$ respectively. • for the CSI 100 index, the MAPE for the forecasted open price is only $0.663\%$, which is $49.23\%$ optimizer than that of the naive method. Meanwhile, the MAPE for the predicted high price and the low price are $12.18\%$ and $15.35\%$ optimizer than that of the naive method, respectively. The RMSE of the open price is improved by $44.21\%$, $12.83\%$ for the high price, and $13.90\%$ for the low price. As for the RMSEH and the AR, the results of the proposed method are improved by $13.75\%$ and $14.29\%$ respectively. • for the 50 ETF, the MAPE for the forecasted open price is $0.729\%$, which improved by $41.87\%$ than that of the naive method. At the same time, $6.95\%$ and $12.03\%$ improvement on the MAPE are made for the high price and the low price, respectively. The RMSE of the open price is improved by $37.50\%$, $7.50\%$ for the high, and $11.36\%$ for the low price. In terms of the RMSEH and AR, the results of the proposed method are improved by $11.54\%$ and $8.55\%$ respectively.

Table (ref) also gives the counts of establishing VAR model and VEC model when employing the proposed framework. We can conclude that the VEC model is far more frequently adopt comparing to the VAR model, which provides a proof to the view of cheung2007empirical that the daily highs and lows of stocks follow a co-integration relationship.

figure[figure omitted — 990 chars of source]

Finally, to get a clear picture of the forecasted performance, we also compared the realistic and the predicted stock values of the Kweichow Moutai during the period from $5/11/2003$ to $22/6/2004$ (Left panel), the CSI $100$ index from $11/5/2011$ to $16/12/2011$ (Middle panel), and the $50$ ETF during the period from $9/8/2007$ to $24/3/2008$ (Right panel) in Fig.(ref). From it, we can see that the data predicted by the proposed method in this paper fits the reality well. Furthermore, for the Kweichow Moutai, the continuous rising that exists before $9/4/2004$ and the subsequent falling are perfectly forecasted; for the CSI $100$ index, the overall downward trend and the two rebounds around $4/7/2011$ and $9/11/2011$ are also fully reflected; for the $50$ ETF, two spikes around $16/10/2007$ and $15/1/2008$ coincided accurately.

Conclusions

We investigated the forecasting problem of the OHLC data contained in candlestick chart, which plays an important role in the financial market. To address it, we proposed a novel transformation approach to relax the inherent constraints of the OHLC data along with its explicit inverse transformation, which facilitates the subsequent establishment of various prediction models and guarantee a meaningful predicted open-high-low-close prices. The transformation approach not only extends the range of variables to $\left(-\infty, +\infty\right)$, but also shares the flexibility of the well known log- and logit- transformation. Based on the unconstrained transformation, we establish a flexible and efficient framework for modelling the OHLC data with full use of its information. For illustration, we thereby show the detailed procedure via the VAR and VEC modelling.

The new approach has high practical utility because of its flexibility, simple implementation and straightforward interpretation. For example, it is applicable to a variety of positive interval data, and the selected model can be generalized to other types of statistical models and machine learning models. Investigations along this direction may merit further research but beyond the scope of this paper. From this perspective, the proposed method provides a completely new and useful alternative for OHLC data analysis, enriching the existing literatures.

Finally, we documented the finite performance of the proposed method via extensive simulation studies in terms of various measurements. And the analysis of the OHLC data of three different kinds of financial subjects in Chinese financial market: the Kweichow Moutai, CSI $100$ index, and $50$ ETF also illustrated the utility of the new methodology. All of these results fully demonstrate the stability and reasonability of the proposed method.

Compliance with Ethical Standards

Funding

This study was funded by the National Natural Science Foundation of China (grant Numbers. $71420107025$, $11701023$).

Conflict of interest

Author Huiwen Wang declares that she has no conflict of interest. Author Wenyang Huang declares that he has no conflict of interest. Author Shanshan Wang declares that she has no conflict of interest.

Ethical approval

This article does not contain any studies with human participants or animals performed by any of the authors.

Availability of data and material

The data of the Kweichow Moutai, 50 ETF, CSI 100 is downloaded form a financial server called Wind. The data can be uploaded as required.

Code availability

This paper applies custom R code and can be uploaded as required.

\nolinenumbers