EconBase
← All papers

Provably Auditing Ordinary Least Squares in Low Dimensions

Ankur Moitra, Dhruv Rohatgi

arXiv 28 May 2022 · Statistics — Machine Learning

arXiv:2205.14284 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Measuring the stability of conclusions derived from Ordinary Least Squares linear regression is critically important, but most metrics either only measure local stability (i.e. against infinitesimal changes in the data), or are only interpretable under statistical assumptions. Recent work proposes a simple, global, finite-sample stability metric: the minimum number of samples that need to be removed so that rerunning the analysis overturns the conclusion, specifically meaning that the sign of a particular coefficient of the estimated regressor changes. However, besides the trivial exponential-time algorithm, the only approach for computing this metric is a greedy heuristic that lacks provable guarantees under reasonable, verifiable assumptions; the heuristic provides a loose upper bound on the stability and also cannot certify lower bounds on it. We show that in the low-dimensional regime where the number of covariates is a constant but the number of samples is large, there are efficient algorithms for provably estimating (a fractional version of) this metric. Applying our algorithms to the Boston Housing dataset, we exhibit regression analyses where we can estimate the stability up to a factor of $3$ better than the greedy heuristic, and analyses where we can certify stability to dropping even a majority of the samples.

Citation extraction

42
references
73
in-text mentions
42
distinct cited
0
self-citations
6,874
main-text words

appendix boundary found by appendix_command · 35% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Tamara Broderick, Ryan Giordano, and Rachael Meager, An automatic fi… (2020) 160.94613585%
2Nikolas Kuschnig, Gregor Zens, and Jes Cuaresma, Hidden in plain sig…0.8746367%
3David A Belsley, Edwin Kuh, and Roy E Welsch, Regression diagnostics… (1980)0.81142100%
4James Renegar, On the computational complexity and geometry of the f… (1992) no. 3, 255–2990.7375260%
5David Harrison Jr and Daniel L Rubinfeld, Hedonic housing prices and… (1978) no. 1, 81–1020.64422100%
6Suyash Gupta and Dominik Rothenhäusler, The $ r $-value: evaluating… (2021)0.58531100%
7John Milnor, On the betti numbers of real varieties, Proceedings of… (1964) no. 2, 275–2800.5112250%
8Samprit Chatterjee and Ali S Hadi, Influential observations, high le… (1986) 379–3930.51121100%
9Adam Klivans, Pravesh K Kothari, and Raghu Meka, Efficient algorithm… (2018) pp. 1420–14300.51121100%
10Edward E Leamer, Global sensitivity results for generalized least sq… (1984) no. 388, 867–8700.51121100%

Showing the top 10 of 42 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Sample Fit Reliability0.73732
2Testing Most Influential Sets0.40511
3Finding Most Influential Sets0.40511