EconBase
← All papers

High-Dimensional Learning in Finance

Hasan Fallahgoul

arXiv 4 Jun 2025 · Finance — Statistical Finance · 3 citations (OpenAlex)

arXiv:2506.03780 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Recent advances in machine learning have shown promising results for financial prediction using large, over-parameterized models. This paper provides theoretical foundations and empirical validation for understanding when and how these methods achieve predictive success. I examine two key aspects of high-dimensional learning in finance. First, I prove that within-sample standardization in Random Fourier Features implementations fundamentally alters the underlying Gaussian kernel approximation, replacing shift-invariant kernels with training-set dependent alternatives. Second, I establish information-theoretic lower bounds that identify when reliable learning is impossible no matter how sophisticated the estimator. A detailed quantitative calibration of the polynomial lower bound shows that with typical parameter choices, e.g., 12,000 features, 12 monthly observations, and R-square 2-3%, the required sample size to escape the bound exceeds 25-30 years of data--well beyond any rolling-window actually used. Thus, observed out-of-sample success must originate from lower-complexity artefacts rather than from the intended high-dimensional mechanism.

Citation extraction

28
references
79
in-text mentions
29
distinct cited
0
self-citations
12,325
main-text words

appendix boundary found by appendix_command · 67% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Rahimi \ Recht (2007) Random features for large-scale kernel machines, in `Advances in Neural Information Processing Systems', Vol. 20, pp. 1177–11841.000116100%
2Nagel (2025) `Seemingly virtuous complexity in return prediction', Working paper1.00073100%
3Kelly, Malamud \ Zhou (2024) `The virtue of complexity in return prediction', Journal of Finance 79(1), 459–5030.98118794%
4Welch \ Goyal (2008) `A comprehensive look at the empirical performance of equity premium prediction', Review of Financial Studies 21(4), 1455–15080.81142100%
5Gu, Kelly \ Xiu (2020) `Empirical asset pricing via machine learning', Review of Financial Studies 33(5), 2223–22730.73732100%
6Sutherland \ Schneider (2015) On the error of random fourier features, in `Proceedings of the Thirty-First Conference on Uncertainty in Artificial Intelligenc…0.64422100%
7Vapnik (1998) Statistical Learning Theory, Wiley, New York0.5113233%
8Shalev-Shwartz \ Ben-David (2014) Understanding Machine Learning: From Theory to Algorithms, Cambridge University Press, Cambridge, UK0.5112250%
9Bianchi, Büchner \ Tamoni (2021) `Bond risk premiums with machine learning', Review of Financial Studies 34(2), 1046–10890.51121100%
10Chen, Pelger \ Zhu (2024) `Deep learning in asset pricing', Management Science 70(2), 714–7500.51121100%

Showing the top 10 of 29 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Overparametrized models with posterior drift0.51121