arXiv 10 May 2026 · Econometrics
arXiv:2605.09740 · PDF · DOI · OpenAlex · Extracted main text
Needless to say, linear dynamics are pervasive in economic time series, particularly autoregressive ones. While gradient boosting with trees excels at capturing nonlinearities, it is inefficient in small samples when much of the predictive content is linear, expending splits to approximate relationships better captured by simple linear terms. This paper proposes LGB+, a boosting procedure operating on a more inclusive set of basis functions. The idea comes in two flavors. LGB+ evaluates a tree and a linear candidate at each step against out-of-bag data; only the winner advances. The simpler variant, LGB^A+, alternates on a fixed schedule: a block of tree updates, then a greedy linear correction, repeat. Both designs avoid ex ante commitments to any particular functional form or predictor selection. Because the prediction is the sum of a linear and a tree component, forecasts decompose natively into linear and nonlinear contributions, and so does permutation-based variable importance and historical proximity weights. In a quarterly U.S. macroeconomic forecasting exercise, LGB+ delivers strong gains for targets with pronounced autoregressive dynamics or mixed linear-nonlinear signals. Variables dominating the linear channel are those operating through autoregressive persistence or near-accounting relationships to the target (e.g., initial claims for unemployment and building permits for housing starts).
appendix boundary found by appendix_command · 55% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., Ye, Q., an… (2017) LightGBM: A highly efficient gradient boosting decision tree | 0.843 | 4 | 3 | 75% |
| 2 | Goulet Coulombe, P., Göbel, M., and Klieber, K (2024) The dual interpretation of machine learning forecasts | 0.843 | 3 | 3 | 100% |
| 3 | Goulet Coulombe, P (2024) The macroeconomy as a Random Forest | 0.737 | 3 | 2 | 100% |
| 4 | Goulet Coulombe, P (2025) A neural Phillips curve and a deep output gap | 0.644 | 3 | 2 | 67% |
| 5 | Chinn, M. D., Meunier, B., and Stumpner, S (2023) Nowcasting world trade with machine learning: A three-step approach | 0.644 | 2 | 2 | 100% |
| 6 | Geertsema, P. and Lu, H (2023) Instance-based explanations for gradient boosting machine predictions with AXIL weights | 0.644 | 2 | 2 | 100% |
| 7 | Breiman, L (2001) Random forests | 0.511 | 2 | 2 | 50% |
| 8 | Hollmann, N., Müller, S., Eggensperger, K., and Hutter, F (2022) TabPFN: A transformer that solves small tabular classification problems in a second | 0.511 | 2 | 2 | 50% |
| 9 | Stock, J. H. and Watson, M. W (2002) Macroeconomic forecasting using diffusion indexes | 0.511 | 2 | 2 | 50% |
| 10 | Athey, S., Tibshirani, J., and Wager, S (2019) Generalized random forests | 0.511 | 2 | 1 | 100% |
Showing the top 10 of 32 scored citations.