EconBase
← All papers

Finding Most Influential Sets

Lucas D. Konrad, Nikolas Kuschnig

arXiv 4 Jun 2026 · Statistics — Machine Learning

arXiv:2606.05919 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Identifying most influential sets (MIS) - size-$k$ subsets whose removal maximally changes a target estimand - is typically infeasible because it requires searching over $\binom{n}{k}$ subsets. For estimands with linear-fractional leave-set-out effects, we show that MIS selection reduces to a one-parameter sequence of top-$k$ problems. Dinkelbach's method yields an algorithm with $\mathcal{O}(n)$ cost per iteration and finite termination. For fixed residualized inputs, the algorithm returns a globally optimal set for the univariate ratio objective, including the oracle-residualized partial linear model. With estimated nuisance functions, uniform denominator and generated-score stability imply approximation to the first-order oracle orthogonal-score objective; exact set recovery follows under a separation condition. Simulations and applications show that the method recovers exact MIS that were previously computationally inaccessible.

Citation extraction

40
references
83
in-text mentions
41
distinct cited
2
self-citations
6,927
main-text words

appendix boundary found by appendix_command · 68% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Kuschnig, Nikolas and Zens, Gregor and Crespo Cuaresma, Jesús (2021) Hidden in Plain Sight: Influential Sets in Linear Regression self0.9416483%
2Broderick, Tamara and Giordano, Ryan and Meager, Rachael (2023) An Automatic Finite-Sample Robustness Metric: When Can Dropping a Little Data Make a Big Difference?0.8746467%
3Lucas Darius Konrad and Nikolas Kuschnig (2026) Testing Most Influential Sets self0.84333100%
4Dinkelbach, Werner (1967) On Nonlinear Fractional Programming0.7373367%
5Hu, Y. and Hu, P. and Zhao, H. and Ma, J. W (2024) Most Influential Subset Selection: Challenges, Promises, and Beyond0.64422100%
6Huang, Jenny Y. and Burt, David R. and Shen, Yunyi and Nguyen, Tin D… (2025) Approximations to Worst-Case Data Dropping: Unmasking Failure Modes0.64422100%
7Meager, Rachael (2019) Understanding the Average Impact of Microcredit Expansions: A Bayesian Hierarchical Analysis of Seven Randomized Experiments0.58531100%
8Attanasio, Orazio and Augsburg, Britta and De Haas, Ralph and Fitzsi… (2015) The Impacts of Microfinance: Evidence from Joint-Liability Lending in Mongolia0.51121100%
9Solon Barocas and Moritz Hardt and Arvind Narayanan (2023) Fairness and Machine Learning: Limitations and Opportunities0.40511100%
10Basu, Samyadeep and You, Xuchen and Feizi, Soheil (2020) On Second-Order Group Influence Functions for Black-Box Predictions0.40511100%

Showing the top 10 of 41 scored citations.