EconBase
← All papers

Dimension reduction of open-high-low-close data in candlestick chart based on pseudo-PCA

Wenyang Huang, Huiwen Wang, Shanshan Wang

arXiv 31 Mar 2021 · Econometrics

arXiv:2103.16908 · PDF · DOI · OpenAlex · Extracted main text

Abstract

The (open-high-low-close) OHLC data is the most common data form in the field of finance and the investigate object of various technical analysis. With increasing features of OHLC data being collected, the issue of extracting their useful information in a comprehensible way for visualization and easy interpretation must be resolved. The inherent constraints of OHLC data also pose a challenge for this issue. This paper proposes a novel approach to characterize the features of OHLC data in a dataset and then performs dimension reduction, which integrates the feature information extraction method and principal component analysis. We refer to it as the pseudo-PCA method. Specifically, we first propose a new way to represent the OHLC data, which will free the inherent constraints and provide convenience for further analysis. Moreover, there is a one-to-one match between the original OHLC data and its feature-based representations, which means that the analysis of the feature-based data can be reversed to the original OHLC data. Next, we develop the pseudo-PCA procedure for OHLC data, which can effectively identify important information and perform dimension reduction. Finally, the effectiveness and interpretability of the proposed method are investigated through finite simulations and the spot data of China's agricultural product market.

Citation extraction

39
references
39
in-text mentions
39
distinct cited
4
self-citations
7,933
main-text words

appendix boundary found by appendix_command · 84% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Chih Fong Tsai and Zen-Yu Quan (2014) Stock prediction by searching for similarities in candlestick charts0.40511100%
2Hervé Abdi and Lynne Williams (2010) Principal component analysis0.40511100%
3Javier Arroyo, Rosa Espńola, and Carlos Maté (2011) Different approaches to forecast interval time series: a comparison in finance0.40511100%
4Edwin Diday (1988) The symbolic approach in clustering0.40511100%
5Paula Brito and A Pedro Duarte Silva (2012) Modelling interval data with normal and skew-normal distributions0.40511100%
6Philip Brown, Angeline Chua, and Jason Mitchell (2002) The influence of cultural factors on price clustering: Evidence from asia–pacific stock markets0.40511100%
7Li-Juan Cao and Francis Eng Hock Tay (2003) Support vector machine with adaptive parameters in financial time series forecasting0.40511100%
8Pierre Cazes, Ahlame Chouakria, Edwin Diday, and Yves Schektman (1997) Extension de l'analyse en composantes principales à des données de type intervalle0.40511100%
9Roberto Cervelló-Royo, Francisco Guijarro, and Karolina Michniuk (2015) Stock market trading rule based on pattern recognition and technical analysis: Forecasting the djia index with intraday data0.40511100%
10Yin-Wong Cheung (2007) An empirical model of daily highs and lows0.40511100%

Showing the top 10 of 39 scored citations.