EconBase
← Back to paper

Dimension reduction of open-high-low-close data in candlestick chart based on pseudo-PCA

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

54,568 characters · 14 sections · 24 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
center[center omitted — 133 chars of source]

\vskip 0.3cm

center[center omitted — 385 chars of source]

\let \thefootnote \relax \footnotetext{ * Corresponding author. Correspondence to: School of Economics and Management, Beihang University, Beijing 100191, China. E-mail address: [email removed] (S.S. Wang).}

\vskip 0.3cm

center[center omitted — 47 chars of source]

The (open-high-low-close) OHLC data is the most common data form in the field of finance and the investigate object of various technical analysis. With increasing features of OHLC data being collected, the issue of extracting their useful information in a comprehensible way for visualization and easy interpretation must be resolved. The inherent constraints of OHLC data also pose a challenge for this issue. This paper proposes a novel approach to characterize the features of OHLC data in a dataset and then performs dimension reduction, which integrates the feature information extraction method and principal component analysis. We refer to it as the pseudo-PCA method. Specifically, we first propose a new way to represent the OHLC data, which will free the inherent constraints and provide convenience for further analysis. Moreover, there is a one-to-one match between the original OHLC data and its feature-based representations, which means that the analysis of the feature-based data can be reversed to the original OHLC data. Next, we develop the pseudo-PCA procedure for OHLC data, which can effectively identify important information and perform dimension reduction. Finally, the effectiveness and interpretability of the proposed method are investigated through finite simulations and the spot data of China's agricultural product market.

\vskip 0.5cm { Key Words:} OHLC data; pseudo-PCA; dimension reduction; feature information extraction

Introduction

As one of the most classic technical analysis tools, the candlestick chart is used to document the prices of almost all financial products, including spots, stocks, futures, options, etc romeo2015study, chmielewski2015pattern. Investors can judge the long and short trading conditions and roughly predict future price trends by analyzing candlestick chart Tsai2014Stock.

Intuitively, Figure.(ref) shows a daily candlestick chart, which not only records the open price, high price, low price, and close price of certain object on that day, but also reflects the difference between any two prices. According to American convention, a green real body refers to a bull market (the close price $\textgreater$ the open price, see Figure.(ref)(a)), while a red real body refers to a bear market (the open price $\textgreater$ the close price, see Figure.(ref)(b)), respectively.

figure[figure omitted — 160 chars of source]

Mathematically, the essence of the candlestick chart is a $4$-dimensional vector, consisting of the open price, high price, low price, and close price, which are collectively called the OHLC data kazemilari2015correlation. However, unlike ordinary $4$-dimensional vectors, the OHLC data have some important inherent constraints. Specifically,

definitionA $4$-dimensional vector $\bm{x}=(x^{(o)}, x^{(h)}, x^{(l)}, x^{(c)})'$ is considered the OHLC data if $\bm{x}$ satisfies the following three constraints: \begin{center} (1) $x^{(l)} > 0$;\\ (2) $x^{(l)} < x^{(h)}$; \\ (3) $x^{(o)}, x^{(c)} \in [x^{(l)},x^{(h)}]$. \end{center} where $x^{(o)}, x^{(h)},x^{(l)}$ and $x^{(c)}$ represent the open price, high price, low price, and close price of the subject during a certain time period (which can vary from seconds to months), respectively.

The OHLC data $\bm{x}$ in Definition.(ref) can be defined as a generalized interval-valued data from the perspective of symbolic data analysis article. Namely, the OHLC data takes $x^{(h)}$ and $x^{(l)}$ as the upper and lower boundaries of the interval, while $x^{(o)}$ and $x^{(c)}$ lie within the interval. Compared to the interval-valued data, the OHLC data contains more information, resulting in more complicated properties than interval-valued data. The common methods for interval-valued data, such as the Max $\&$ Min method arroyo2011different and Center $\&$ Range brito2012modelling, giordani2015lasso method are mimicked by various works on the OHLC data. For example, yang2012acix, xiong2017interval, and sun2018threshold have investigated the OHLC data via the above methods used for interval-valued data without incorporating information on the open price and close price.

Owing to the ignorance of the open price and close price, these two widely-used methods cannot completely use the OHLC data information, yielding possible inefficient results. As indicated by cheung2007empirical, the open price and close price have an explanatory power for the fluctuation of the OHLC data and should be carefully incorporated into statistical modeling.

Through decades of development, scholars have completed multi-angle studies on the candlestick chart and its data (i.e., OHLC data). Initially, scholars are committed to performing pattern recognition on the candlestick chart kamijo1990stock, juang1992discriminative, lo2000foundations, cervello2015stock, giving priority to finding a candlestick chart pattern that could obtain stable profits. At the same time, traditional statistical methods such as the clustering analysis harris1991stock, brown2002influence, tao2017k and discriminant analysis hopkins1996effect, leung2000forecasting, cohen2019trading have also been gradually applied to the study of the candlestick chart. In addition, other popular researches include predicting the future changes of OHLC data, such as econometric forecasting models fiess2002towards, pai2005hybrid, de2020large and machine learning models cao2003support, liu2012fluctuation, mann2017new, mallqui2019predicting.

With technological advances in many scientific areas, a considerable amount of OHLC data is available. It is crucial to develop a dimension reduction technique to extract the informative features contained in the multivariate OHLC data. Among the dimension reduction methods, the principal component analysis (PCA) has always been a key technique to extract the comprehensive features using the linear combination of the original variables, which are called principal components (PCs). These PCs represent the major mode of data variation and are the characterizing features of typical variables in a dataset wang2015principal,m2016forecasting. Furthermore, the scores on PCs have been used for visualization to help researchers better understand the correlation between the original variables jolliffe2002principal.

Nonetheless, the aforementioned PCA method is not applicable when handling dimension reduction for the OHLC data due to its inherent constraints, which poses new challenges. For example, imposing the classical PCA on the OHLC dataset will destroy its structure, rendering the corresponding PCs meaningless. Furthermore, as a special type of data, the OHLC data is characterized as a $4$-dimensional vector data with constraints defined in Definition.(ref) as a whole, which is completely different from the classical scalar data or the vector data in the Euclidean space. All of these concerns are motivations to propose a new approach.

As a reminder, the OHLC data can be used as generalized interval-valued data. Fortunately, various PCA techniques have been extended from classical scalar data to the interval-valued data, such as the vertices principal component analysis (VPCA; cazes1997extension) , centers principal component analysis (CPCA; chouakria2000symbolic), symbolic object principal component analysis (SPCA; lauro2000principal), midpoints radii principal component analysis (MRPCA; palumbo2003pca), interval principal component analysis (IPCA; gioia2006principal), and complete-information-based principal component analysis (CIPCA; wang2012cipca). After revisiting the aforementioned PCA methods for interval-valued data, we find that the key point is to extract feature-based representation of interval-valued data, such as vertices, midpoint, radii, etc. Nonetheless, these existing approaches are only designed for interval-valued data, and they could not be directly applied to the OHLC data. Apart from the two boundaries contained in the interval-valued data, the OHLC data includes more information and constraints. Consequently, this requires more feature-based representations.

Therefore, this paper aims to explore dimension reduction for the OHLC data based on the PCA method with the idea of feature-based representations, constructing a novel approach to characterize features of OHLC data. First, we propose a new way to represent the OHLC data, which can eliminate the inherent constraints and make further analysis easier. Furthermore, there is a one-to-one match between the original OHLC data and its feature-based representations, which allows the analysis of the feature-based representations to be reversed to the original data. Next, we develop the pseudo-PCA procedure for OHLC data, which can effectively identify important information and perform dimension reduction. Finally, the effectiveness and interpretability of the proposed method are investigated through finite simulations and the spot data of China's agricultural product market.

This paper has two main contributions. First, a novel feature-based representation method is proposed for the OHLC data, which provides a new perspective and facilitates further statistical analysis of the OHLC data. Second, we introduce a pseudo-PCA procedure to achieve dimension reduction for multivariate OHLC data, which enhances existing literature on PCA methods.

The rest of the paper is organized as follows: Section (ref) presents the feature extraction method for the OHLC data and related algebraic space of the feature-based representations. In Section (ref), we describe the proposed pseudo-PCA procedure. Then we examine the effectiveness, robustness and interpretability of the pseudo-PCA by simulations in Section.(ref) and empirical experiment in Section (ref). Finally, we conclude with a brief summary and insights, as well as several possible further research topics in Section (ref).

Feature-based representations and its algebraic space

Feature-based representations of OHLC data

Inspired by the feature extraction method for interval-valued data, we propose a novel feature-based representation approach for the OHLC data in Definition.(ref). This satisfies the following merits: First, each component of feature-based representations has a unique and meaningful interpretation of the OHLC data. Second, it eliminates the inherent constraints and provides convenience for further analysis. Third, there is a one-to-one match between the original OHLC data and its feature-based representations, which allows the analysis of the feature-based representations to be reversed to the original data.

definition(Feature-based representations of OHLC data). For the OHLC data $\bm{x}=(x^{(o)}, x^{(h)}, x^{(l)}, x^{(c)})'$, assume that $x^{(o)}, x^{(h)}, x^{(l)}$, and $x^{(c)}$ are not equal to each other $(\text{except for} \; x^{(o)} \equiv x^{(c)})$, then the feature-based representation of $\bm{x}$ is formulated as $\bm{y}=(y^{(1)},y^{(2)},y^{(3)},y^{(4)})'$, where \begin{equation} \left\{\begin{array}{l} y^{(1)}=\ln x^{(l)} \\ y^{(2)}=\ln (x^{(h)}-x^{(l)}) \\ y^{(3)}=\ln\Big(\frac{\lambda^{(o)}}{1-\lambda^{(o)}}\Big) \\ y^{(4)}=\ln\Big(\frac{\lambda^{(c)}}{1-\lambda^{(c)}}\Big) \\ \end{array}\right., \end{equation} with \begin{equation} \lambda^{(o)} = \frac{x^{(o)}-x^{(l)}}{x^{(h)}-x^{(l)}} \ \ \ and\ \ \ \lambda^{(c)} = \frac{x^{(c)}-x^{(l)}}{x^{(h)}-x^{(l)}}. \end{equation}

From Definition.(ref), each component of the feature-based representation for the OHLC $\bm{x}$ is significant. Specifically, $y^{(1)}$ corresponds to the natural logarithmic value of the low price, reflecting the magnitude of the original OHLC data. $y^{(2)}$ is the natural logarithmic value of the range between the high price and the low price, which can note the fluctuation of raw OHLC data. Obviously, $y^{(1)}$ and $y^{(2)}$ can take their values freely in $[-\infty, +\infty]$. Furthermore, in order to ensure that the constraint $x^{(o)}, x^{(c)} \in [x^{(l)},x^{(h)}]$ always holds, we introduce two convex combination coefficients $\lambda^{(o)}$ and $\lambda^{(c)}$, which link $\lambda^{(o)}$ and $x^{(o)}$ as well as $\lambda^{(c)}$ and $x^{(c)}$ together with $x^{(l)}$ and $x^{(h)}$. That is,

equation[equation omitted — 172 chars of source]

$\lambda^{(o)},\lambda^{(c)} \in (0,1)$ implies $x^{(o)}, x^{(c)} \in [x^{(l)},x^{(h)}]$. Besides, the implication of $\lambda^{(o)}$ and $\lambda^{(c)}$ is very clear. Specifically, when $\lambda^{(o)} \textgreater 0.5$, it implies that $x^{(o)}$ is closer to $x^{(h)}$; and if $\lambda^{(o)} \textless 0.5$, it indicates that $x^{(o)}$ is closer to $x^{(l)}$. The interpretation of $\lambda^{(c)}$ is similar. Note that $\lambda^{(o)}$ and $\lambda^{(c)}$ are limited to $(0,1)$, which mimic the proportion or probability of $x^{(o)}$ and $x^{(c)}$, respectively, which measures the relative distance from $x^{(l)}$. Thus, similar to the idea of logistic regression, we utilize the “log odds ratio" to yield $y^{(3)}$ and $y^{(4)}$, which also removes the finite support $(0,1)$ constraint and free $y^{(3)}$ and $y^{(4)}$ to $[-\infty, +\infty]$.

Furthermore, a one-to-one correspondence between the original OHLC data $\bm{x}$ and its feature-based representations $\bm{y}$ exists, as specified in Proposition.(ref).

propositionFor any two OHLC data $\bm{x}$ and $\tilde{\bm{x}}$ with their corresponding feature-based representations $\bm{y}$ and $\tilde{\bm{y}}$, then it holds $\bm{x}=\tilde{\bm{x}}$ only if $\bm{y}=\tilde{\bm{y}}$ holds.

Note that through feature information extraction, the feature-based representation $\bm{y}$ maps the original OHLC data $\bm{x}$ from the constricted $\mathbb{R}^4$ space with constraints in Definition.(ref) into the full $\mathbb{R}^4$ space, i.e., $$\mathbb{R}^{4}=\left\{\bm{y}=\left(y^{(1)}, y^{(2)}, y^{(3)}, y^{(4)}\right)^{\prime},-\infty<y^{(j)}<+\infty, j=1,2,3,4\right\}.$$

Here, we can directly perform vector linear and inner product operations based on classic linear algebra principles. Furthermore, from Proposition.(ref) and Definition.(ref), we can easily formulate the inverse transformation from the feature-based representations $\bm{y}$ to the original OHLC data $\bm{x}$ according to Equation ((ref)):

equation[equation omitted — 450 chars of source]

In summary, the novel feature-based representations not only eliminate the inherent constraints, providing convenience for further analysis but also are a one-to-one match with explicit inverse transformation in Equation ((ref)), which means that any analysis of the feature-based representations can be reversed to the original data. From this perspective, the proposed method may cast a new insight to investigate the OHLC data.

Finally, we show additional remarks on the assumption in Definition.(ref). Note that assumptions $x^{(o)}, x^{(h)}, x^{(l)}$, and $x^{(c)}$ are not equal to each other $(\text{except for} \; x^{(o)} \equiv x^{(c)})$ implies that $x^{(h)} \neq x^{(l)} \neq 0$ and $\lambda^{(o)}, \lambda^{(c)} \notin \left\{ 0,1 \right\}$. These assumptions always hold under normal circumstances. However, when these assumptions are occasionally invalid, we recommend the following preprocessing procedure. (1) When the subject is on trade suspension and all prices equal to $0$, namely, $x^{(o)}=x^{(h)}=x^{(l)}=x^{(c)}=0$, we exclude these extreme cases in the raw data. (2) When $\lambda^{(o)}$ or $\lambda^{(c)}$ is equal to $0$, it corresponds to $x^{(o)}=x^{(l)}$ or $x^{(c)}=x^{(l)}$, respectively. We add a random term to $x^{(o)}$ or $x^{(c)}$ and make $\lambda^{(o)}$ or $\lambda^{(c)}$ slightly greater than 0. (3) When $\lambda^{(o)}$ or $\lambda^{(c)}$ is equal to $1$, it indicates that $x^{(o)}=x^{(h)}$ or $x^{(c)}=x^{(h)}$, respectively. We subtract a random term from $x^{(o)}$ or $x^{(c)}$ to make $\lambda^{(o)}$ or $\lambda^{(c)}$ slightly less than $1$. (4) When certain subject reaches limit-up or limit-down as soon as the opening quotation, that is, $x^{(o)}=x^{(h)}=x^{(l)}=x^{(c)}\neq0$. If limit-up (or limit-down) happens, we firstly multiply $x^{(c)}$ (or $x^{(o)})$ and $x^{(h)}$ by 1.1 to make a relatively large interval. And then conduct measurements given in circumstances (2) and (3).

The sample version of Feature-based representations

Let $\{\textbf{\textrm{x}}_i\}_{i=1}^n$ be independent and identically distributed (iid) samples with size $n$ from the $p$-dimensional vector of OHLC data $\textbf{\textrm{x}}=(\bm{x}_1,\cdots,\bm{x}_p)^{\prime}$, where $\bm{x}_j$ is the $j$th feature of OHLC data in Definition.(ref). Denote $\textbf{\textrm{X}}$ as the $n\times p$ matrix of OHLC data with elements of the $i$th row being $\textbf{\textrm{x}}^\prime_i$ (i.e., $\textbf{\textrm{X}}=(\textbf{\textrm{x}}_1,\cdots,\textbf{\textrm{x}}_n)^\prime$). Rewrite $\textbf{\textrm{X}}$ in the form of column (i.e., $\textbf{\textrm{X}}=(\bm{X}_1,\cdots,\bm{X}_p)$), where the $j$th column $\bm{X}_j=(\bm{x}_{1j},\cdots,\bm{x}_{nj})^\prime$ is the $j$th sample feature of OHLC data.

Applying the proposed procedure in Definition.(ref) to each OHLC data $\bm{x}_{ij}$ yields the corresponding feature-based representation $\bm{y}_{ij}$. Correspondingly, let $\textbf{\textrm{Y}}$ be the $n\times p$ matrix of feature-based representation for OHLC data with elements of the $i$th row being $\textbf{\textrm{y}}^\prime_i$ (i.e., $\textbf{\textrm{Y}}=(\textbf{\textrm{y}}_1,\cdots,\textbf{\textrm{y}}_n)^\prime$). Rewrite $\textbf{\textrm{Y}}$ in the form of column (i.e., $\textbf{\textrm{Y}}=(\bm{Y}_1,\cdots,\bm{Y}_p)$), where the $j$th column $\bm{Y}_j=(\bm{y}_{1j},\cdots,\bm{y}_{nj})^\prime$ is the $j$th sample feature-based representations.

Next, we describe the vector space structure formed by the sample feature-based representations with size $n$, denoted by $\mathbb{R}^4_n$, together with its basic operations.

definition(Vector space structure of the sample feature-based representations with size $n$ and its basic operations). Suppose that $\textbf{Y}=(\bm{y}_1,\cdots,\bm{y}_n)^\prime$ are derived from the iid sample feature of OHLC data matrix $\textbf{X}=(\bm{x}_1,\cdots,\bm{x}_n)^\prime$, we call the vector space formed by $\textbf{Y}$ as the space of sample feature-based representations with size $n$, that is: \begin{equation*} \mathbb{R}_{n}^{4}=\left\{\bm{Y}=\left(\bm{y}_{1}, \bm{y}_{2}, \ldots, \bm{y}_{n}\right)^{\prime} ; \bm{y}_{i} \in \mathbb{R}^{4} \; (i=1,2, \ldots, n)\right\}. \end{equation*} For $\bm{Y}_j$, $\bm{Y}_k \in \mathbb{R}_n^{4}$ and $\beta\in\mathbb{R}$, we define the corresponding operators: \begin{itemize} • (1) Addition operator $\oplus$: \begin{equation} \boldsymbol{Y}_{j} \oplus \boldsymbol{Y}_{k}=\left(\boldsymbol{y}_{1 j}+\boldsymbol{y}_{1 k}, \boldsymbol{y}_{2 j}+\boldsymbol{y}_{2 k}, \ldots, \boldsymbol{y}_{n j}+\boldsymbol{y}_{n k}\right)^{\prime} \in \mathbb{R}_{n}^{4}, \end{equation} where $\forall i=1,2,\ldots,n$, $\bm{y}_{ij}+\bm{y}_{ik}=\left(y_{ij}^{(1)}+y_{ik}^{(1)}, y_{ij}^{(2)}+y_{ik}^{(2)}, y_{ij}^{(3)}+y_{ik}^{(3)}, y_{ij}^{(4)}+y_{ik}^{(4)}\right)^{\prime} \in \mathbb{R}^{4}$. • (2) Scalar multiplication operator $\otimes$: \begin{equation} \beta \otimes \bm{Y}_{j}=\left(\beta \bm{y}_{1j}, \beta \bm{y}_{2j}, \ldots, \beta \bm{y}_{nj}\right)^{\prime} \in \mathbb{R}_{n}^{4}, \end{equation} where $\forall i=1,2,\ldots,n$, $\beta \bm{y}_{ij}=\left(\beta y_{ij}^{(1)}, \beta y_{ij}^{(2)}, \beta y_{ij}^{(3)}, \beta y_{ij}^{(4)}\right)^{\prime} \in \mathbb{R}^{4}$. • (3) Inner product operator $(\cdot, \cdot)_{\mathbb{R}_n^4}$: \begin{equation} \left(\boldsymbol{Y}_{j}, \boldsymbol{Y}_{k}\right)_{\mathbb{R}_{n}^{4}}=\sum_{i=1}^{n}\left\langle\boldsymbol{y}_{i j}, \boldsymbol{y}_{i k}\right\rangle_{\mathbb{R}^{4}}, \end{equation} where $<\cdot, \cdot>_{\mathbb{R}^{4}}$ represents for the inner product operator in $\mathbb{R}^{4}$, which is defined as $\left\langle \bm{y}_{ij}, \bm{y}_{ik}\right\rangle_{\mathbb{R}^{4}}=\sum_{d=1}^{4} y_{ij}^{(d)} y_{ik}^{(d)}$. \end{itemize}

Furthermore, we can deduce the subtraction operator in $\mathbb{R}_{n}^{4}$ according to the definitions of addition and scalar multiplication operators, that is:

equation[equation omitted — 224 chars of source]

for $\forall i=1,2,\ldots,n$, $\bm{y}_{ij}-\bm{y}_{ik}=\left(y_{i j}^{(1)}-y_{i k}^{(1)}, y_{i j}^{(2)}-y_{i k}^{(2)}, y_{i j}^{(3)}-y_{i k}^{(3)}, y_{i j}^{(4)}-y_{i k}^{(4)}\right)^{\prime}$. The element zero $\bm{0}^{(0)}$ in $\mathbb{R}_{n}^{4}$ is: $\mathbf{0}^{(0)}=(\mathbf{0}, \mathbf{0}, \ldots, \mathbf{0})^{\prime} \in \mathbb{R}_{n}^{4}$, where $\mathbf{0}=(0,0,0,0)^{\prime} \in \mathbb{R}^{4}$.

The following standard properties hold, making them analogous to translation and scalar multiplication in real space.

propositionFor any $\bm{Y}_j, \bm{Y}_k, \bm{Y}_l \in \mathbb{R}_n^{4}$ and $\beta \in \mathbb{R}$, the inner product operation in $\mathbb{R}_{n}^{4}$ satisfies: \begin{itemize} • (1) Non-negative property: $\left(\boldsymbol{Y}_{j}, \boldsymbol{Y}_{j}\right)_{\mathbb{R}_{n}^{4}} \geq 0$, and $``="$ holds if and only if $\bm{Y}_j=\bm{0}^{(0)}$; • (2) Commutative property: $\left(\boldsymbol{Y}_{j}, \boldsymbol{Y}_{k}\right)_{\mathbb{R}_{n}^{4}}=\left(\boldsymbol{Y}_{k}, \boldsymbol{Y}_{j}\right)_{\mathbb{R}_{n}^{4}};$ • (3) Associative property: $\left(\boldsymbol{Y}_{j}, \boldsymbol{Y}_{k} \oplus \boldsymbol{Y}_{l}\right)_{\mathbb{R}_{n}^{4}}=\left(\boldsymbol{Y}_{j}, \boldsymbol{Y}_{k}\right)_{\mathbb{R}_{n}^{4}}+\left(\boldsymbol{Y}_{j}, \boldsymbol{Y}_{l}\right)_{\mathbb{R}_{n}^{4}};$ • (4) Linear property: $\left(\beta \otimes \bm{Y}_{j}, \bm{Y}_{k}\right)_{\mathbb{R}_{n}^{4}}=\beta\left(\bm{Y}_{j}, \boldsymbol{Y}_{k}\right)_{\mathbb{R}_{n}^{4}}.$ \end{itemize}

Conclusions (3) and (4) in Proposition.(ref) will play an important role in the derivation of the subsequent pseudo-PCA approach. Before we present the main procedure, we will give some basic summary statistics based on the sample feature-based representations $\textbf{\textrm{Y}}=\{\bm{y}_{ij}\}_{i=1,\cdots,n; \; j=1,\cdots,p}$.

definition(Sample mean, variance and correlation coefficient for feature-based representations). For any $\bm{Y}_j$, $\bm{Y}_k \in \mathbb{R}_n^{4}$ and $\beta\in\mathbb{R}$: \begin{itemize} • (1) Sample mean for the $j$th feature-based representation is: \begin{equation} \overline{\boldsymbol{Y}}_{j}=\frac{1}{n}\left(\boldsymbol{y}_{1 j}+\boldsymbol{y}_{2 j}+\cdots+\boldsymbol{y}_{n j}\right) \in \mathbb{R}^{4}. \end{equation} • (2) Sample covariance for the $j$th and $k$th feature-based representations is: \begin{equation} S_{j k}=\frac{1}{4n} \sum_{i=1}^{n}\left\langle\boldsymbol{y}_{i j}-\overline{\boldsymbol{Y}}_{j}, \; \boldsymbol{y}_{i k}-\overline{\boldsymbol{Y}}_{k}\right\rangle_{\mathbb{R}^{4}} \in \mathbb{R}. \end{equation} $j=k$ yields the sample variance for the $j$th feature-based representation: \begin{equation} S_{j}^{2}=\frac{1}{4n} \sum_{i=1}^{n}\left\langle\boldsymbol{y}_{i j}-\overline{\boldsymbol{Y}}_{j}, \; \boldsymbol{y}_{i j}-\overline{\boldsymbol{Y}}_{j}\right\rangle_{\mathbb{R}^{4}} \in \mathbb{R}. \end{equation} • (4) Sample correlation coefficient for the $j$th and $k$th feature-based representations is: \begin{equation} r_{j k}=\frac{S_{j k}}{S_{j} S_{k}} \in \mathbb{R}. \end{equation} \end{itemize}

The sample variance-covariance matrix $\bm{\Sigma}$ and correlation coefficient matrix $\bm{W}$ for the feature-based representation matrix $\textbf{Y}$ are:

equation[equation omitted — 358 chars of source]

It is important to note that although each element in the feature information matrix $\textbf{\textrm{Y}}=\{\bm{y}_{ij}\}_{i=1,\cdots,n; \; j=1,\cdots,p}$ belongs to $\mathbb{R}^4$ space, its variance-covariance matrix $\bm{\Sigma}$ and correlation coefficient matrix $\bm{W}$ are $p \times p$ dimensional matrices in the real space. This conclusion will make it convenient to derive the pseudo-PCA for OHLC data in $\mathbb{R}_n^4$ space. In addition, we can also standardize the elements in matrix $\textbf{\textrm{Y}}$ according to

equation[equation omitted — 169 chars of source]

Pseudo-PCA procedure

In multivariate statistics, PCA is a widely used method of dimension reduction. Like the objective of the classic PCA, for the feature information variables $\left \{ \bm{Y}_j \in \mathbb{R}_{n}^{4}, \; j=1,2,\ldots,p \right \}$ of OHLC data in $\mathbb{R}_n^4$ space, the aim of pseudo-PCA is also to reduce the $p$-dimensional space to $m$-dimensional $\left( m \leq p \right )$ under the premise of minimal the loss of the variance information. To clarify, it uses the maximum variation direction of the feature information to find a set of so-called ¡°pseudo-PCs¡±, denoted as $\boldsymbol{F}_{h}=\left(\boldsymbol{f}_{1 h}, \boldsymbol{f}_{2 h}, \ldots, \boldsymbol{f}_{n h}\right)^{\prime} \in \mathbb{R}_{n}^{4} \; (h=1, \ldots, m, m \leq p)$, where each pseudo-PC $\bm{F}_h$ is a linear combination of $\bm{Y}_1, \bm{Y}_2,\ldots, \bm{Y}_p$ with the loading coefficients being $\bm{u}_h=(u_{h1},\cdots,u_{hp})^\prime$, namely:

equation[equation omitted — 198 chars of source]

The pseudo-PCA of OHLC data expects $\sum_{h=1}^{m} \operatorname{Var}\left(\boldsymbol{F}_{h}\right)$ to get the maximum value, and $\operatorname{Var}\left(\boldsymbol{F}_{1}\right) \geq \operatorname{Var}\left(\boldsymbol{F}_{2}\right) \geq \cdots \geq \operatorname{Var}\left(\boldsymbol{F}_{m}\right)$. In addition, the combination coefficients $\boldsymbol{u}_{h}=\left(u_{h 1}, u_{h 2}, \ldots, u_{h p}\right)^{\prime} \in \mathbb{R}^{p} \; (h=1,2, \ldots, m)$ should be orthonormal and normalized.

Assuming that $\bm{Y}_1, \bm{Y}_2,\ldots, \bm{Y}_p$ have been standardized, there is $\overline{\boldsymbol{Y}}_{j}=\mathbf{0} \in \mathbb{R}^{4} \; (j=1,2, \ldots, p)$, and $\overline{\boldsymbol{F}}_{h}=\mathbf{0} \in \mathbb{R}^{4}$. Therefore, using the Proposition.(ref) and the Definition.(ref), the sample variance of $\bm{F}_h$ is:

equation[equation omitted — 441 chars of source]

In Equation ((ref)), $\bm{W}$ is the correlation coefficient matrix of the feature information variable set $\{\bm{Y}_1, \bm{Y}_2,\ldots, \bm{Y}_p\}$, which is a $p \times p$ dimension matrix in the real number domain. Thus, we get Theorem.(ref). \newtheorem{theorem}{\bf Theorem}

theoremIf $\bm{W}$ is the correlation coefficient matrix of the feature information variable set $\{\bm{Y}_1, \bm{Y}_2,\ldots, \bm{Y}_p \; | \; \bm{Y}_j \in \mathbb{R}_n^4, \; j=1,2,\ldots,p \} $, and $\boldsymbol{u}_{h}=\left(u_{h 1}, u_{h 2}, \ldots, u_{h p}\right)^{\prime} \in \mathbb{R}^{p} \; (h=1,2, \ldots, m, \; m \leq p)$, remark the pseudo-principle component as $\boldsymbol{F}_{h}=u_{h 1} \otimes \boldsymbol{Y}_{1} \oplus u_{h 2} \otimes \boldsymbol{Y}_{2} \oplus \cdots \oplus u_{h p} \otimes \boldsymbol{Y}_{p}$. Under the condition of $\boldsymbol{u}_{h}$ is orthonormal and normalized, deriving the first $m$ pseudo-PCs is equivalent to solving the following optimization problem: \begin{centering} $\max \sum_{h=1}^{m} \boldsymbol{u}_{h}^{\prime} \boldsymbol{W u}_{h}$ \\ $$ \emph{\text{s.t.}} \begin{cases} \boldsymbol{u}_{h}^{\prime} \boldsymbol{u}_{k}= \begin{cases} \begin{array}{ll} 1 & h=k \\ 0 & h \neq k \end{array}, \quad h, k=1,2, \ldots, m, \; m \leq p \end{cases}\\ \bm{u}_1' \bm{W} \bm{u}_1 \geq \bm{u}_2' \bm{W} \bm{u}_2 \geq \cdots \geq \bm{u}_m' \bm{W} \bm{u}_m, \; m \leq p \end{cases}$$ \end{centering}

According to linear algebra theory, conducting eigen-decomposition of matrix $\bm{W}$ and deriving the first $m$ eigenvalues $\lambda_1 \geq \lambda_2 \geq \cdots \geq \lambda_m > 0$, the corresponding eigenvectors $\bm{u}_1, \bm{u}_2,..., \bm{u}_m$ are the solutions of the optimization problem in Theorem (ref) abdi2010principal. Then, the first $m$ pseudo-principle components $\bm{F}_{1},\bm{F}_{2},...,\bm{F}_{m}$ can be directly obtained by Equation ((ref)).

In addition, it is not difficult to prove that after the pseudo-PCA, the following important properties are established. The corresponding proof can refer to Appendix B.

propositionIf the feature information variables $\bm{Y}_1, \bm{Y}_2,\ldots, \bm{Y}_p$ are standardized, the first $m \; (m \leq p)$ pseudo-principle components $\bm{F}_h \in \mathbb{R}_n^4 \; ( 1 \leq h \leq m)$ derived from the pseudo-PCA satisfy the following properties:\\ (1)\; The sample mean of $\bm{F}_h$ is equal to $0$, that is, $\overline{\boldsymbol{F}}_{h}=\mathbf{0} \in \mathbb{R}^{4}$; \\ (2)\; The sample variance of $\bm{F}_h$ is equal to $\lambda_h$, namely, $\operatorname{Var}\left(\boldsymbol{F}_{h}\right)=\lambda_{h}$; \\ (3)\; For any $\bm{F}_j, \bm{F}_k \in \mathbb{R}_n^4 ( 1 \leq j,k \leq p)$, if $j \neq k$, their sample covariance $Cov(\bm{F}_j, \bm{F}_k) = 0$; \\ (4)\; $\sum_{h=1}^{p} Var(\bm{F}_h)=p.$

Like the classical PCA, we can define the cumulative contribution rate $Q_m$ based on the conclusions (2) and (4) in Proposition.2 as follows:

equation[equation omitted — 85 chars of source]

There is $0 < Q_m \leq 1$. A larger $Q_m$ indicates that the pseudo-principle components retain more variance information after conducting pseudo-PCA, and the analysis accuracy is higher.

To summarize, a pseudo-PCA modeling framework is illustrated by Figure.(ref) and the specific implementation process is summarized as Algorithm.(ref).

figure[figure omitted — 154 chars of source]
algorithm[algorithm omitted — 1,644 chars of source]

Finally, we realize the pseudo-PCA for OHLC data by extracting its feature information and successfully reduce the dimensionality of an OHLC data set with $n$ observations and $p$ variables to a new OHLC data set with $n$ observations and $m$ new components (pseudo-PCs), and $m \leq p$. In this process, the variance loss of the feature information matrix is minimal.

Simulations

The performance of pseudo-PCA is evaluated via finite sample simulation studies in term of the following two aspects: (1) the ability to remove redundant variables and reduce the dimensionality of raw data, and (2) the consistency of the estimated loadings of pseudo-PCs.

Here we assume that $p=6$ and the unconstrained OHLC data variable vector $\textrm{\textbf{Y}}=( \bm{Y}_1, \bm{Y}_2, \bm{Y}_3, \bm{Y}_4 ,\bm{Y}_5,\bm{Y}_6)$ satisfy that (1) each the unconstrained OHLC data variable $\bm{Y}_j$ is mean zero for $j=1,2,\cdots,6$; and (2) the correlation coefficient matrix $\boldsymbol{W_{\mathrm{Y}}}$ is set as following:

equation[equation omitted — 880 chars of source]

Actually, from Equation ((ref)), there are two redundant variables $\bm{Y}_5$ and $\bm{Y}_6$, i.e. $\bm{Y}_5 = \bm{Y}_1+\bm{Y}_2$ and $\bm{Y}_6 = \bm{Y}_3+\bm{Y}_4$. The sample sizes are $n=50,100,150,200$ with each case $300$ repeats, and we summary the performance of proposed procedure in term of two measurements: (1) the cumulative variance contribution rate of the first four pseudo-PCs; and (2) the mean absolute percentage error (MAPE) of the eigenvectors, i.e., $\operatorname{MAPE}_j=\frac{\left\|\widehat{\boldsymbol{u}}_{j}-\boldsymbol{u}_{j}\right\|}{\left\|\boldsymbol{u}_{j}\right\|} \times 100 \%$ $(1 \leq j \leq 6)$. For each criteria, we also report the empirical standard derivations.

The cumulative variance contribution rate of the first four pseudo-PCs is very stable with the change of sample size and always reaches 100.0%. Take $n=200$ as an example, Figure.(ref) shows the 100.0% explanatory capability of the first four pseudo-PCs. This verified the extraordinary ability of pseudo-PCA to remove redundant variables. Further, Table.(ref) shows the results of MAPE and its empirical standard derivations with different sample size. The estimation accuracy of $\textbf{PC5}$ and $\textbf{PC6}$ is relatively high and that of $\textbf{PC3}$ and $\textbf{PC4}$ is relatively low. The mean values of MAPE range from 13.2% to 27.9%, and its empirical standard derivations varies from 0.136 to 0.269, which are both within the acceptable range. Meanwhile, with the increase of sample size, MAPE and its empirical standard derivations show a decreasing trend, indicating that the estimated eigenvectors gradually converges to the theoretical values. These results identified the consistency between the estimated loadings of pseudo-PCs derived from the sample data and the theoretical values.

figure[figure omitted — 170 chars of source]
table[table omitted — 1,080 chars of source]

Empirical analysis

In order to illustrate the effectiveness of the proposed pseudo-PCA, this paper extracts the OHLC data of he spot prices of $6$ common foods of $20$ important agricultural product markets in China from Wind {\color{red}(\url{https://www.wind.com.cn/})}. The $6$ OHLC data variables include: beef $(CNY/kg)$, lamb $(CNY/kg)$, pork $(CNY/kg)$, cucumber $(CNY/kg)$, potato $(CNY/kg)$, and onion $(CNY/kg)$. The observed dates range from $1/1/2019$ to $31/12/2019$, and we use the first price recorded in the time period as the open price, last price recorded as the close price, and highest/lowest price that appears during the time period as the high/low price.

The original OHLC data matrix $\textrm{\textbf{X}}$ can be referred to Table.(ref) in Appendix.B. First, the standardized feature information matrix $\textrm{\textbf{Y}}$ is derived and shown in Table.(ref) in Appendix.B. Conduct pseudo-PCA modelling on the standardized feature information matrix $\textrm{\textbf{Y}}$, the corresponding eigenvalues (EV), variance contribution rate (VCR), and cumulative contribution to the variance (CVCR) of the $6$ extracted pseudo-PCs are shown in Table.2. The cumulative contribution rate to the variance of the first two pseudo-PCs reaches 57.4$\%$, and that of the first three pseudo-PCs reaches 73.2$\%$. This implies that the pseudo-principle components represent the information of the original data matrix well. The corresponding variance cumulative contribution plot is shown in Figure.(ref).

table[table omitted — 929 chars of source]
figure[figure omitted — 178 chars of source]

Furthermore, the loading matrix of the first two pseudo-PCs is shown in Table.3. The expressions of these two pseudo-PCs are:

eqnarray*[eqnarray* omitted — 421 chars of source]
table[table omitted — 951 chars of source]

The correlation between the first two pseudo-principle components and feature information variables are clearly visualized through Figure.(ref). In Figure.(ref), positively related variables are close to each other, while negatively related variables are divergent. The lengths of the arrows represent the information quality of different variables. It is observed that the loading factors' absolute values of the $\textbf{PC1}$ on cucumbers, potatoes, and onions are relatively large, which can be called a `Vegetable' factor. The loading factors' absolute values of the $\textbf{PC2}$ on beef, lamb, and pork are comparatively large, and it can be called a `Meat' factor.

figure[figure omitted — 192 chars of source]

According to Equation ((ref)), the scores of the first two pseudo-PCs can be calculated, and the corresponding OHLC formed data can be derived by using Equation ((ref)). The results are shown in Figure.(ref) and Figure.(ref), where the abbreviations of the 20 markets are used (the corresponding full name can be referred to Table.(ref) in Abbreviation.C).

Based on the OHLC formed pseudo-PC scores, we can observe the size and fluctuation of the original data during the observation period. For instance, (1) For the OHLC data converted from the first pseudo-PC scores (`Vegetable' factor), it is observed that the open, high, low, and open prices of Zhangjiajie and Xinyu have always been at a high level. This means that the prices of cucumbers, potatoes, and onions in these two cities maintained a high level through 2019. At the same time, the difference between the low price and high price of Zhangjiajie is relatively large, indicating that the prices of cucumbers, potatoes, and onions in Zhangjiajie experienced great fluctuations in 2019. While the difference between the upper and lower bounds of Xinyang is very small, corresponding to the prices of `Vegetable' in Xinyang were very stable in 2019. Markets with green-body OHLC data had a relatively low comprehensive price for cucumbers, potatoes, and onions at the beginning of 2019 and relatively high comprehensive price at the end of 2019. While the comprehensive price for cucumbers, potatoes, and onions of the red-body markets tended to open higher, and finish lower. (2) For the OHLC data transformed by the second pseudo-PC scores (`Meat' factor), we see that the open price of Zhangjiajie is at the medium level. This indicates that at the beginning of 2019, the prices of beef, lamb, and pork in Zhangjiajie were not high. However, its candlestick stretches to the highest and lowest among all the markets, suggesting that the `Meat' price in Zhangjiajie experienced great fluctuations in 2019, resulting in the prices appearing extremely high and extremely low. On the other hand, the four-dimensional data of Jiangsu are very small, indicating that its price of `Meat' remained stable in 2019, which corresponds to the original lamb price of Jiangsu remained at $58 \; CNY/kg$ throughout the year.

figure[figure omitted — 145 chars of source]
figure[figure omitted — 146 chars of source]

It is important to mention that an eigenvector uniquely corresponds to one pseudo-PC and one transformed OHLC data. However, for each eigenvalue of the standardized feature information matrix $\textrm{\textbf{Y}}$, there are two corresponding eigenvectors, which are mutually opposite vectors. The difference in sign does not affect the orthogonal and normalized requirements in Theorem.(ref), they are both feasible solutions. Paradoxically, the pseudo-PCs corresponding to these two eigenvectors are different, and the restored OHLC formed pseudo-principle component scores are also different. In order to enhance the interpretability of the pseudo-PCA for OHLC data, it is recommended that the eigenvector is selected by the following rule. If the eigenvector has the largest absolute loading values on the feature information variables, such as $\bm{Y}_1, \bm{Y}_2$ and $\bm{Y}_3$, then select the eigenvector that has positive loading values on $\bm{Y}_1, \bm{Y}_2$ and $\bm{Y}_3$. This way, a relationship is established. The original data size, feature information size, pseudo-PC scores, and OHLC formed pseudo-PC scores are all positively related to each other. The advantage of this practice lies in the original data size and can be easily analyzed through the final comprehensive OHLC formed pseudo-principle component scores. For instance, in this article, we see that the loading values of the `Vegetable' factor for cucumbers, potatoes, and onions are all positive; the loading values of the `Meat' factor for beef, lamb, and pork are also all positive.

Conclusions

As more OHLC data is collected, the demand for dimension reduction of the OHLC data emerges. Inspired by the various PCA methods applied for interval data, this paper proposed a novel modus for characterizing features of OHLC data and built a unique $\mathbb{R}_n^{4}$ space containing the feature-based representations accordingly. The new method of representing the OHLC data can eliminate the inherent constraints of OHLC data and transition between the original OHLC data and its feature-based representations freely. In the unique $\mathbb{R}_n^{4}$ space, we defined the linear operations and inner product operation and deduced the pseudo-PCA method for OHLC data under this new algebraic framework. The pseudo-PCA has high visibility and high interpretability, which can effectively identify important information and perform dimension reduction. Therefore, reducing the dimensionality of an OHLC data set with $n$ observations and $p$ variables to a new OHLC data set with $n$ observations and $m$ components is achieved, and $m \leq p$. The effectiveness, robustness and interpretability of the pseudo-PCA is examined by finite simulations and empirical experiment.

The OHLC formed pseudo-PCs derived from pseudo-PCA still hold the constraints of OHLC data, which maintains the original format of the raw data and provides significant convenience for the subsequent establishment of econometric or machine learning models. Through the OHLC formed pseudo-PCs, one can observe the size and the fluctuation of certain original data variables during the observation period. Furthermore, the OHLC formed pseudo-PCs can be presented in the format of the candlestick chart, which helps people obtain information easily. Overall, this paper provides a novel feature-based representation method for OHLC data and enriches the existing literature on PCA methods.

For scalar data, the Singular Spectrum Analysis (SSA) is applied instead of PCA when performing dimension reduction of panel data. Similarly, we may further propose the new pseudo-SSA approach to address the time-series OHLC data set, which can be an interesting research direction in the future. Besides, exploring new feature extraction methods for more generalized interval-valued data is significant. For instance, many types of data (i.e., daily temperatures, companies' profits, and stock returns) hold intervals and own values between the upper and lower boundaries (low value may be negative) while not belonging to OHLC data.

Declarations

Funding

This study was funded by the National Natural Science Foundation of China (grant numbers. $71420107025$, $11701023$).

Conflicts of interest

Author Wenyang Huang declares that he has no conflict of interest. Author Huiwen Wang declares that she has no conflict of interest. Author Shanshan Wang declares that she has no conflict of interest.

Ethical approval

This article does not contain any studies with human participants or animals performed by any of the authors.

Availability of data and material

The spot data is downloaded form a financial server called Wind. The data can be uploaded as required.

Code availability

This paper applies custom R code and can be uploaded as required.

\nolinenumbers \appendixpage