arXiv 1 May 2024 · Econometrics
arXiv:2405.00424 · PDF · DOI · OpenAlex · Extracted main text
Ridge regression is an indispensable tool in big data analysis. Yet its inherent bias poses a significant and longstanding challenge, compromising both statistical efficiency and scalability across various applications. To tackle this critical issue, we introduce an iterative strategy to correct bias effectively when the dimension $p$ is less than the sample size $n$. For $p>n$, our method optimally mitigates the bias such that any remaining bias in the proposed de-biased estimator is unattainable through linear transformations of the response data. To address the remaining bias when $p>n$, we employ a Ridge-Screening (RS) method, producing a reduced model suitable for bias correction. Crucially, under certain conditions, the true model is nested within our selected one, highlighting RS as a novel variable selection approach. Through rigorous analysis, we establish the asymptotic properties and valid inferences of our de-biased ridge estimators for both $p<n$ and $p>n$, where, both $p$ and $n$ may increase towards infinity, along with the number of iterations. We further validate these results using simulated and real-world data examples. Our method offers a transformative solution to the bias challenge in ridge regression inferences across various disciplines.
appendix boundary found by appendix_command · 84% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Athey and Imbens (2019) Machine learning methods that economists should know about | 0.811 | 4 | 2 | 100% |
| 2 | Giannone et al (2021) Economic predictions with big data: The illusion of sparsity | 0.811 | 4 | 2 | 100% |
| 3 | Abadie and Kasy (2019) Choosing among regularized estimators in empirical economics: The risk of machine learning | 0.737 | 3 | 2 | 100% |
| 4 | Hansen (2022) Econometrics | 0.737 | 3 | 2 | 100% |
| 5 | Shao and Deng (2012) Estimation in high-dimensional linear models with deterministic design matrices | 0.737 | 3 | 2 | 100% |
| 6 | van Wieringen (2023) Lecture notes on ridge regression | 0.737 | 3 | 2 | 100% |
| 7 | Fan and Lv (2008) Sure independence screening for ultrahigh dimensional feature space | 0.644 | 4 | 1 | 100% |
| 8 | Gao and Tsay (2024) Supervised dynamic pca: Linear dynamic forecasting with many predictors | 0.644 | 4 | 1 | 100% |
| 9 | Hansen (2022) A modern Gauss–Markov theorem | 0.644 | 2 | 2 | 100% |
| 10 | Hoerl (1959) Optimum solution of many variables equations | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 35 scored citations.