EconBase
← All papers

Forward Variable Selection in Ultra-High Dimensional Linear Regression Using Gram-Schmidt Orthogonalization

Jialuo Chen, Zhaoxing Gao, Ruey S. Tsay

arXiv 7 Jul 2025 · Statistics — Methodology

arXiv:2507.04668 · PDF · DOI · OpenAlex · Extracted main text

Abstract

We investigate forward variable selection for ultra-high dimensional linear regression using a Gram-Schmidt orthogonalization procedure. Unlike the commonly used Forward Regression (FR) method, which computes regression residuals using an increasing number of selected features, or the Orthogonal Greedy Algorithm (OGA), which selects variables based on their marginal correlations with the residuals, our proposed Gram-Schmidt Forward Regression (GSFR) simplifies the selection process by evaluating marginal correlations between the residuals and the orthogonalized new variables. Moreover, we introduce a new model size selection criterion that determines the number of selected variables by detecting the most significant change in their unique contributions, effectively filtering out redundant predictors along the selection path. While GSFR is theoretically equivalent to FR except for the stopping rule, our refinement and the newly proposed stopping rule significantly improve computational efficiency. In ultra-high dimensional settings, where the dimensionality far exceeds the sample size and predictors exhibit strong correlations, we establish that GSFR achieves a convergence rate comparable to OGA and ensures variable selection consistency under mild conditions. We demonstrate the proposed method {using} simulations and real data examples. Extensive numerical studies show that GSFR outperforms commonly used methods in ultra-high dimensional variable selection.

Citation extraction

23
references
38
in-text mentions
23
distinct cited
2
self-citations
10,129
main-text words

appendix boundary found by appendix_titled_section at “Supplementary Material” · 99% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Wang, Hansheng (2009) Forward regression for ultra-high dimensional variable screening1.00064100%
2Ing, Ching-Kang and Lai, Tze Leung (2011) A Stepwise Regression Method and Consistent Model Selection for High-Dimensional Sparse Linear Models1.00063100%
3Borodin, P. A. and Konyagin, S. V (2021) Projection Greedy Algorithm0.64422100%
4Ing, Ching-Kang (2020) Model selection for high-dimensional linear regression with dependent observations0.64422100%
5Gao, Zhaoxing and Tsay, Ruey S (2025) Supervised dynamic pca: Linear dynamic forecasting with many predictors self0.51121100%
6McCracken, Michael W and Ng, Serena (2016) FRED-MD: A monthly database for macroeconomic research0.51121100%
7Stock, James H and Watson, Mark W (2002) Macroeconomic forecasting using diffusion indexes0.51121100%
8Peter J. Bickel and Elizaveta Levina (2008) Regularized estimation of large covariance matrices0.40511100%
9Shuo-Chieh Huang and Ruey S. Tsay (2024) Scalable High-Dimensional Multivariate Linear Regression for Feature-Distributed Data self0.40511100%
10Peter Bühlmann (2006) Boosting for high-dimensional linear models0.40511100%

Showing the top 10 of 23 scored citations.