arXiv 13 Apr 2025 · Machine Learning · 1 citations (OpenAlex)
arXiv:2504.09663 · PDF · DOI · OpenAlex · Extracted main text
I show that ordinary least squares (OLS) predictions can be rewritten as the output of a restricted attention module, akin to those forming the backbone of large language models. This connection offers an alternative perspective on attention beyond the conventional information retrieval framework, making it more accessible to researchers and analysts with a background in traditional statistics. It falls into place when OLS is framed as a similarity-based method in a transformed regressor space, distinct from the standard view based on partial correlations. In fact, the OLS solution can be recast as the outcome of an alternative problem: minimizing squared prediction errors by optimizing the embedding space in which training and test vectors are compared via inner products. Rather than estimating coefficients directly, we equivalently learn optimal encoding and decoding operations for predictors. From this vantage point, OLS maps naturally onto the query-key-value structure of attention mechanisms. Building on this foundation, I discuss key elements of Transformer-style attention and draw connections to classic ideas from time series econometrics.
appendix boundary found by appendix_command · 96% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gome… (2017) Attention is all you need | 1.000 | 5 | 3 | 100% |
| 2 | Goulet Coulombe, P., Göbel, M., and Klieber, K (2024) Dual interpretation of machine learning forecasts | 0.843 | 3 | 3 | 100% |
| 3 | Choromanski, K., Likhosherstov, V., Dohan, D., Song, X., Gane, A., S… (2020) Rethinking attention with performers | 0.644 | 2 | 2 | 100% |
| 4 | Fein-Ashley, J., Kannan, R., and Prasanna, V (2025) The fft strikes again: An efficient alternative to self-attention | 0.644 | 2 | 2 | 100% |
| 5 | Garnelo, M. and Czarnecki, W. M (2023) Exploring the space of key-value-query models with intention | 0.644 | 2 | 2 | 100% |
| 6 | Gu, A. and Dao, T (2023) Mamba: Linear-time sequence modeling with selective state spaces | 0.644 | 2 | 2 | 100% |
| 7 | Hastie, T., Tibshirani, R., and Friedman, J (2009) The Elements of Statistical Learning: Data Mining, Inference, and Prediction | 0.644 | 2 | 2 | 100% |
| 8 | Kelly, B. T., Kuznetsov, B., Malamud, S., and Xu, T. A (2025) Artificial intelligence asset pricing models | 0.644 | 2 | 2 | 100% |
| 9 | Peng, H., Pappas, N., Yogatama, D., Schwartz, R., Smith, N. A., and… (2021) Random feature attention | 0.644 | 2 | 2 | 100% |
| 10 | Tsai, Y.-H. H., Bai, S., Yamada, M., Morency, L.-P., and Salakhutdin… (2019) Transformer dissection: a unified understanding of transformer's attention via the lens of kernel | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 39 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Nonlinear Dynamic Factor Analysis With a Transformer Network | 0.405 | 1 | 1 |
| 2 | A Nonlinear Target-Factor Model with Attention Mechanism for Mixed-Frequency Data | 0.405 | 1 | 1 |