Lucas D. Konrad, Nikolas Kuschnig
arXiv 4 Jun 2026 · Statistics — Machine Learning
arXiv:2606.05919 · PDF · DOI · OpenAlex · Extracted main text
Identifying most influential sets (MIS) - size-$k$ subsets whose removal maximally changes a target estimand - is typically infeasible because it requires searching over $\binom{n}{k}$ subsets. For estimands with linear-fractional leave-set-out effects, we show that MIS selection reduces to a one-parameter sequence of top-$k$ problems. Dinkelbach's method yields an algorithm with $\mathcal{O}(n)$ cost per iteration and finite termination. For fixed residualized inputs, the algorithm returns a globally optimal set for the univariate ratio objective, including the oracle-residualized partial linear model. With estimated nuisance functions, uniform denominator and generated-score stability imply approximation to the first-order oracle orthogonal-score objective; exact set recovery follows under a separation condition. Simulations and applications show that the method recovers exact MIS that were previously computationally inaccessible.
appendix boundary found by appendix_command · 68% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Kuschnig, Nikolas and Zens, Gregor and Crespo Cuaresma, Jesús (2021) Hidden in Plain Sight: Influential Sets in Linear Regression self | 0.941 | 6 | 4 | 83% |
| 2 | Broderick, Tamara and Giordano, Ryan and Meager, Rachael (2023) An Automatic Finite-Sample Robustness Metric: When Can Dropping a Little Data Make a Big Difference? | 0.874 | 6 | 4 | 67% |
| 3 | Lucas Darius Konrad and Nikolas Kuschnig (2026) Testing Most Influential Sets self | 0.843 | 3 | 3 | 100% |
| 4 | Dinkelbach, Werner (1967) On Nonlinear Fractional Programming | 0.737 | 3 | 3 | 67% |
| 5 | Hu, Y. and Hu, P. and Zhao, H. and Ma, J. W (2024) Most Influential Subset Selection: Challenges, Promises, and Beyond | 0.644 | 2 | 2 | 100% |
| 6 | Huang, Jenny Y. and Burt, David R. and Shen, Yunyi and Nguyen, Tin D… (2025) Approximations to Worst-Case Data Dropping: Unmasking Failure Modes | 0.644 | 2 | 2 | 100% |
| 7 | Meager, Rachael (2019) Understanding the Average Impact of Microcredit Expansions: A Bayesian Hierarchical Analysis of Seven Randomized Experiments | 0.585 | 3 | 1 | 100% |
| 8 | Attanasio, Orazio and Augsburg, Britta and De Haas, Ralph and Fitzsi… (2015) The Impacts of Microfinance: Evidence from Joint-Liability Lending in Mongolia | 0.511 | 2 | 1 | 100% |
| 9 | Solon Barocas and Moritz Hardt and Arvind Narayanan (2023) Fairness and Machine Learning: Limitations and Opportunities | 0.405 | 1 | 1 | 100% |
| 10 | Basu, Samyadeep and You, Xuchen and Feizi, Soheil (2020) On Second-Order Group Influence Functions for Black-Box Predictions | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 41 scored citations.