EconBase
← All papers

Nonlinear Boosting with Multiple Testing in High-Dimensional Generalised Linear Models with Binary Responses

Charisios Grivas, George Kapetanios, Zacharias Psaradakis, Vasilis Sarafidis, Marian Vavra, Alexia Ventouri

arXiv 24 Jul 2026 · Econometrics

arXiv:2607.22440 · PDF · Extracted main text

Abstract

This paper proposes a nonlinear boosting with multiple testing (BMT) approach to variable selection in high-dimensional generalised linear models with binary responses. At each stage of the BMT procedure, the model is updated by adding only the most significant covariate, conditional on those already selected in previous stages, while taking into account the multiple testing nature of the problem. It is shown that, under the stated conditions, the BMT procedure selects all covariates whose true coefficients are nonzero, and no other covariates, with probability tending to one. Furthermore, the procedure enjoys an oracle property, in the sense that the post-BMT maximum likelihood estimator of the parameters of the model is asymptotically equivalent to an oracle estimator that knows the correct sparse model in advance. Monte Carlo experiments demonstrate that BMT outperforms competing methods, delivering high covariate-selection accuracy and low parameter estimation error. An empirical example illustrates that BMT delivers a predictive model for the probability that U.S. inflation exceeds a given threshold over a 12-month horizon which has very good out-of-sample performance.

Citation extraction

25
references
32
in-text mentions
26
distinct cited
2
self-citations
15,529
main-text words

appendix boundary found by appendix_command · 50% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Kapetanios, George and Sarafidis, Vasilis and Ventouri, Alexia (2026) Model Selection in High-Dimensional Linear Regression Using Boosting with Multiple Testing self0.84333100%
2McCracken, Michael W and Ng, Serena (2021) FRED-QD: A Quarterly Database for Macroeconomic Research0.73732100%
3White, Halbert (1982) Maximum Likelihood Estimation of Misspecified Models0.51121100%
4White, Halbert Estimation, Inference and Specification Analysis0.51121100%
5Baldi, Pierre and Brunak, Søren and Chauvin, Yves and Andersen, Clau… (2000) Assessing the accuracy of prediction algorithms for classification: an overview0.40511100%
6Bamber, Donald (1975) The area above the ordinal dominance graph and the area below the receiver operating characteristic graph0.40511100%
7Bühlmann, Peter and Hothorn, Torsten (2007) Boosting algorithms: regularization, prediction and model fitting0.40511100%
8Chen, Jiahua and Chen, Zehua (2012) Extended BIC for small-$n$-large-$P$ sparse GLM0.40511100%
9Chen, Yu-chin and Turnovsky, Stephen J. and Zivot, Eric (2014) Forecasting inflation using commodity price aggregates0.40511100%
10Chicco, Davide and Jurman, Giuseppe (2020) The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation0.40511100%

Showing the top 10 of 26 scored citations.