Lucas Darius Konrad, Nikolas Kuschnig
arXiv 23 Oct 2025 · Statistics — Machine Learning
arXiv:2510.20372 · PDF · Extracted main text
Small influential data subsets can dramatically impact model conclusions, with a few data points overturning key findings. While recent work identifies these most influential sets, there is no formal way to tell when maximum influence is excessive rather than expected under natural random sampling variation. We address this gap by developing a principled framework for most influential sets. Focusing on linear least-squares, we derive a convenient exact influence formula and identify the extreme value distributions of maximal influence - the heavy-tailed Fréchet for constant-size sets and heavy-tailed data, and the well-behaved Gumbel for growing sets or light tails. This allows us to conduct rigorous hypothesis tests for excessive influence. We demonstrate through applications across economics, biology, and machine learning benchmarks, resolving contested findings and replacing ad-hoc heuristics with rigorous inference.
appendix boundary found by appendix_command · 72% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Kuschnig, Nikolas and Zens, Gregor and Crespo Cuaresma, Jesús (2021) Hidden in Plain Sight: Influential Sets in Linear Regression self | 1.000 | 8 | 4 | 100% |
| 2 | Hu, Y. and Hu, P. and Zhao, H. and Ma, J. W (2024) Most Influential Subset Selection: Challenges, Promises, and Beyond | 1.000 | 6 | 4 | 100% |
| 3 | Broderick, Tamara and Giordano, Ryan and Meager, Rachael (2023) An Automatic Finite-Sample Robustness Metric: When Can Dropping a Little Data Make a Big Difference? | 1.000 | 5 | 3 | 100% |
| 4 | Freund, D. and Hopkins, S. B (2023) Towards Practical Robustness Auditing for Linear Regression | 0.843 | 3 | 3 | 100% |
| 5 | Huang, Jenny Y. and Burt, David R. and Shen, Yunyi and Nguyen, Tin D… (2025) Approximations to Worst-Case Data Dropping: Unmasking Failure Modes | 0.843 | 3 | 3 | 100% |
| 6 | Dombry, Clément and Ferreira, Ana (2019) Maximum Likelihood Estimators Based on the Block Maxima Method | 0.737 | 3 | 3 | 67% |
| 7 | Basu, Samyadeep and You, Xuchen and Feizi, Soheil (2020) On Second-Order Group Influence Functions for Black-Box Predictions | 0.644 | 2 | 2 | 100% |
| 8 | Belsley, David A. and Kuh, Edwin and Welsch, Roy E (1980) Regression Diagnostics: Identifying Influential Data and Sources of Collinearity | 0.644 | 2 | 2 | 100% |
| cook1979Influential | unmatched citation key cook1979Influential | 0.644 | 2 | 2 | 100% |
| fisher2023Influence | unmatched citation key fisher2023Influence | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 33 scored citations. 2 of these could not be matched to a bibliography entry, so only the citation key is shown.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Finding Most Influential Sets | 0.843 | 3 | 3 |