Xiaorui Zhu, Yichen Qin, Peng Wang
arXiv 14 Jul 2023 · Statistics — Methodology · publishedMetrika (2024) · 1 citations (OpenAlex)
arXiv:2307.07574 · PDF · DOI · OpenAlex · Extracted main text
Statistical inference of the high-dimensional regression coefficients is challenging because the uncertainty introduced by the model selection procedure is hard to account for. A critical question remains unsettled; that is, is it possible and how to embed the inference of the model into the simultaneous inference of the coefficients? To this end, we propose a notion of simultaneous confidence intervals called the sparsified simultaneous confidence intervals. Our intervals are sparse in the sense that some of the intervals' upper and lower bounds are shrunken to zero (i.e., $[0,0]$), indicating the unimportance of the corresponding covariates. These covariates should be excluded from the final model. The rest of the intervals, either containing zero (e.g., $[-1,1]$ or $[0,1]$) or not containing zero (e.g., $[2,3]$), indicate the plausible and significant covariates, respectively. The proposed method can be coupled with various selection procedures, making it ideal for comparing their uncertainty. For the proposed method, we establish desirable asymptotic properties, develop intuitive graphical tools for visualization, and justify its superior performance through simulation and real data analysis.
appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Yang Liu and Peng Wang (1988) Selection by partitioning the solution paths self | 1.000 | 8 | 3 | 100% |
| 2 | Xianyang Zhang and Guang Cheng (2016) Simultaneous Inference for High-Dimensional Linear Models | 1.000 | 8 | 3 | 100% |
| 3 | Yiyun Zhang, Runze Li, and Chih-Ling Tsai (2010) Regularization parameter selections via generalized information criterion | 0.737 | 3 | 2 | 100% |
| 4 | Huai-Kuang Tsai, Henry Horng-Shing Lu, and Wen-Hsiung Li (2005) Statistical methods for identifying yeast cell cycle transcription factors | 0.693 | 13 | 1 | 100% |
| 5 | N. Banerjee (2003) Identifying cooperativity among transcription factors controlling the cell cycle in yeast | 0.693 | 10 | 1 | 100% |
| 6 | Lifeng Wang, Guang Chen, and Hongzhe Li (2007) Group SCAD regression analysis for microarray time course gene expression data | 0.693 | 8 | 1 | 100% |
| 7 | Mu Yue, Jialiang Li, and Ming-Yen Cheng (2019) Two-step sparse boosting for high-dimensional longitudinal data with varying coefficients | 0.644 | 4 | 1 | 100% |
| 8 | A. Chatterjee and S. N. Lahiri (2011) Bootstrapping Lasso Estimators | 0.644 | 2 | 2 | 100% |
| 9 | Ruben Dezeure, Peter Bühlmann, and Cun-Hui Zhang (2017) High-dimensional simultaneous inference with the bootstrap | 0.644 | 2 | 2 | 100% |
| 10 | Jianqing Fan and Runze Li (2001) Variable selection via nonconcave penalized likelihood and its oracle properties | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 57 scored citations.