Andrea Carriero, Florian Huber, Davide Pettenuzzo
arXiv 14 May 2026 · Econometrics
arXiv:2605.15358 · PDF · DOI · OpenAlex · Extracted main text
We study double descent and benign overfitting in macroeconomic forecasting. We document that double-descent risk curves arise in standard macroeconomic datasets that are driven by a small number of latent factors, and we characterize when the underlying benign-overfitting mechanism holds. The conditions of Bartlett et al. (2020) are satisfied under the exact factor model and can also hold under the more realistic approximate factor model, provided idiosyncratic variances are not too dispersed across series. Because macroeconomic panels have only moderate dimensions, the overparameterization ratio N/T required by the theory is not naturally available. Our solution is to augment the data with synthetic copies from an estimated factor model and we prove that this strategy converges to a kernel ridge regression with a factor-structured kernel. Using monthly (FRED-MD) and quarterly (FRED-QD) US data, the resulting estimator consistently outperforms the Stock-Watson factor model for point forecasting across all series and horizons, with gains that are pervasive, statistically significant, and increasing with the forecast horizon. Our results suggest that benign overfitting, when it works, succeeds because overparameterization implicitly constructs a well-behaved kernel, not because overparameterization is intrinsically desirable.
appendix boundary found by appendix_command · 66% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Bartlett, P. L., Long, P. M., Lugosi, G., and Tsigler, A (2020) Benign overfitting in linear regression | 1.000 | 7 | 3 | 100% |
| 2 | Stock, J. H. and Watson, M. W (2002) Forecasting using principal components from a large number of predictors | 1.000 | 6 | 3 | 100% |
| 3 | McCracken, M. W. and Ng, S (2016) FRED-MD: A monthly database for macroeconomic research | 0.928 | 5 | 3 | 80% |
| 4 | Nakakita, S. and Imaizumi, M (2025) Benign overfitting in time series linear model with over-parameterization | 0.874 | 5 | 2 | 100% |
| 5 | Stock, J. H. and Watson, M. W (2002) Macroeconomic forecasting using diffusion indexes | 0.874 | 5 | 2 | 100% |
| 6 | Bai, J. and Ng, S (2002) Determining the number of factors in approximate factor models | 0.843 | 3 | 3 | 100% |
| 7 | Bunea, F., Strimas-Mackey, S., and Wegkamp, M (2021) Interpolating predictors in high-dimensional factor regression | 0.737 | 3 | 2 | 100% |
| 8 | Coulombe, P. G., Leroux, M., Stevanović, D., and Surprenant, S (2022) How is machine learning useful for macroeconomic forecasting? | 0.644 | 2 | 2 | 100% |
| 9 | Medeiros, M. C., Vasconcelos, G. F., Veiga, Á., and Zilberman, E (2021) Forecasting inflation in a data-rich environment: The benefits of machine learning methods | 0.644 | 2 | 2 | 100% |
| 10 | Gabaix, X. and Ibragimov, R (2011) Rank $-$ 1/2: A simple way to improve the OLS estimation of tail exponents | 0.511 | 2 | 2 | 50% |
Showing the top 10 of 22 scored citations.