Daniel Freund, Samuel B. Hopkins
arXiv 30 Jul 2023 · Statistics — Methodology · 1 citations (OpenAlex)
arXiv:2307.16315 · PDF · DOI · OpenAlex · Extracted main text
We investigate practical algorithms to find or disprove the existence of small subsets of a dataset which, when removed, reverse the sign of a coefficient in an ordinary least squares regression involving that dataset. We empirically study the performance of well-established algorithmic techniques for this task -- mixed integer quadratically constrained optimization for general linear regression problems and exact greedy methods for special cases. We show that these methods largely outperform the state of the art and provide a useful robustness check for regression problems in a few dimensions. However, significant computational bottlenecks remain, especially for the important task of disproving the existence of such small sets of influential samples for regression problems of dimension $3$ or greater. We make some headway on this challenge via a spectral algorithm using ideas drawn from recent innovations in algorithmic robust statistics. We summarize the limitations of known techniques in several challenge datasets to encourage further algorithmic innovation.
appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Ankur Moitra and Dhruv Rohatgi, Provably auditing ordinary least squ… (2022) | 1.000 | 29 | 4 | 100% |
| 2 | Tamara Broderick, Ryan Giordano, and Rachael Meager, An automatic fi… (2020) | 1.000 | 15 | 3 | 100% |
| 3 | Nikolas Kuschnig, Gregor Zens, and Jesús Crespo Cuaresma, Hidden in… | 1.000 | 8 | 3 | 100% |
| 4 | Nicholas Eubank and Adriane Fresh, Enfranchisement and incarceration… (2022) no. 3, 791–806 | 0.874 | 7 | 2 | 100% |
| 5 | Luis R Martinez, How much should we trust the dictator’s gdp growth… (2022) no. 10, 2731–2769 | 0.874 | 7 | 2 | 100% |
| 6 | Manuela Angelucci, Dean Karlan, and Jonathan Zinman, Microcredit imp… (2015) no. 1, 151–182 | 0.843 | 3 | 3 | 100% |
| 7 | David Card and Alan B Krueger, Minimum wages and employment: A case… (1993) | 0.811 | 4 | 2 | 100% |
| 8 | Ainesh Bakshi and Adarsh Prasad, Robust linear regression: Optimal r… (2021) pp. 102–115 | 0.737 | 3 | 2 | 100% |
| 9 | Tobias Achterberg and Eli Towle, Non-convex quadratic optimization,… (2020) | 0.644 | 2 | 2 | 100% |
| 10 | Orazio Attanasio, Britta Augsburg, Ralph De Haas, Emla Fitzsimons, a… (2015) no. 1, 90–122 | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 28 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Testing Most Influential Sets | 0.843 | 3 | 3 |
| 2 | Finding Most Influential Sets | 0.405 | 1 | 1 |