EconBase
← All papers

Testing Most Influential Sets

Lucas Darius Konrad, Nikolas Kuschnig

arXiv 23 Oct 2025 · Statistics — Machine Learning

arXiv:2510.20372 · PDF · Extracted main text

Abstract

Small influential data subsets can dramatically impact model conclusions, with a few data points overturning key findings. While recent work identifies these most influential sets, there is no formal way to tell when maximum influence is excessive rather than expected under natural random sampling variation. We address this gap by developing a principled framework for most influential sets. Focusing on linear least-squares, we derive a convenient exact influence formula and identify the extreme value distributions of maximal influence - the heavy-tailed Fréchet for constant-size sets and heavy-tailed data, and the well-behaved Gumbel for growing sets or light tails. This allows us to conduct rigorous hypothesis tests for excessive influence. We demonstrate through applications across economics, biology, and machine learning benchmarks, resolving contested findings and replacing ad-hoc heuristics with rigorous inference.

Citation extraction

31
references
62
in-text mentions
33
distinct cited
1
self-citations
5,961
main-text words

appendix boundary found by appendix_command · 72% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Kuschnig, Nikolas and Zens, Gregor and Crespo Cuaresma, Jesús (2021) Hidden in Plain Sight: Influential Sets in Linear Regression self1.00084100%
2Hu, Y. and Hu, P. and Zhao, H. and Ma, J. W (2024) Most Influential Subset Selection: Challenges, Promises, and Beyond1.00064100%
3Broderick, Tamara and Giordano, Ryan and Meager, Rachael (2023) An Automatic Finite-Sample Robustness Metric: When Can Dropping a Little Data Make a Big Difference?1.00053100%
4Freund, D. and Hopkins, S. B (2023) Towards Practical Robustness Auditing for Linear Regression0.84333100%
5Huang, Jenny Y. and Burt, David R. and Shen, Yunyi and Nguyen, Tin D… (2025) Approximations to Worst-Case Data Dropping: Unmasking Failure Modes0.84333100%
6Dombry, Clément and Ferreira, Ana (2019) Maximum Likelihood Estimators Based on the Block Maxima Method0.7373367%
7Basu, Samyadeep and You, Xuchen and Feizi, Soheil (2020) On Second-Order Group Influence Functions for Black-Box Predictions0.64422100%
8Belsley, David A. and Kuh, Edwin and Welsch, Roy E (1980) Regression Diagnostics: Identifying Influential Data and Sources of Collinearity0.64422100%
cook1979Influentialunmatched citation key cook1979Influential0.64422100%
fisher2023Influenceunmatched citation key fisher2023Influence0.64422100%

Showing the top 10 of 33 scored citations. 2 of these could not be matched to a bibliography entry, so only the citation key is shown.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Finding Most Influential Sets0.84333