arXiv 31 May 2022 · Econometrics
arXiv:2205.15853 · PDF · DOI · OpenAlex · Extracted main text
The internet has changed the way we live, work and take decisions. As it is the major modern resource for research, detailed data on internet usage exhibits vast amounts of behavioral information. This paper aims to answer the question whether this information can be facilitated to predict future returns of stocks on financial capital markets. In an empirical analysis it implements gradient boosted decision trees to learn relationships between abnormal returns of stocks within the S&P 100 index and lagged predictors derived from historical financial data, as well as search term query volumes on the internet search engine Google. Models predict the occurrence of day-ahead stock returns in excess of the index median. On a time frame from 2005 to 2017, all disparate datasets exhibit valuable information. Evaluated models have average areas under the receiver operating characteristic between 54.2% and 56.7%, clearly indicating a classification better than random guessing. Implementing a simple statistical arbitrage strategy, models are used to create daily trading portfolios of ten stocks and result in annual performances of more than 57% before transaction costs. With ensembles of different data sets topping up the performance ranking, the results further question the weak form and semi-strong form efficiency of modern financial capital markets. Even though transaction costs are not included, the approach adds to the existing literature. It gives guidance on how to use and transform data on internet usage behavior for financial and economic modeling and forecasting.
appendix boundary found by appendix_titled_section at “Appendix” · 94% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Krauss, Christopher, Do, Xuan Anh, Huck, Nicolas (2017) Deep neural networks, gradient-boosted trees, random forests: Statistical arbitrage on the S&P 500 | 1.000 | 18 | 6 | 100% |
| 2 | Choi, Hyunyoung, Varian, Hal R (2012) Predicting the Present with Google Trends | 1.000 | 5 | 4 | 100% |
| 3 | Preis, Tobias, Moat, Helen Susannah, Stanley, H. Eugene, Bishop, Ste… (2012) Quantifying the advantage of looking forward | 1.000 | 5 | 4 | 100% |
| 4 | Huck, Nicolas (2009) Pairs selection and outranking: An application to the S&P 100 index | 1.000 | 5 | 3 | 100% |
| 5 | Ettredge, Michael, Gerdes, John, Karuga, Gilbert (2005) Using web-based search data to predict macroeconomic statistics | 0.928 | 4 | 4 | 100% |
| 6 | Preis, Tobias, Moat, Helen Susannah, Stanley, H. Eugene (2013) Quantifying trading behavior in financial markets using Google Trends | 0.928 | 4 | 4 | 100% |
| 7 | Takeuchi, Lawrence, Lee, Yu-Ying (2013) Applying Deep Learning to Enhance Momentum Trading Strategies in Stocks | 0.928 | 4 | 4 | 100% |
| 8 | Bordino, Ilaria, Battiston, Stefano, Caldarelli, Guido, Cristelli, M… (2012) Web search queries can predict stock market volumes | 0.843 | 3 | 3 | 100% |
| 9 | Hastie, Trevor, Tibshirani, Robert, Friedman, Jerome H (2017) The elements of statistical learning: Data mining, inference, and prediction | 0.811 | 4 | 2 | 100% |
| 10 | Malkiel, Burton G., Fama, Eugene F (1970) Efficient Capital Markets: A Review of Theory and Empirical Work | 0.737 | 3 | 2 | 100% |
Showing the top 10 of 50 scored citations.