Lingxiao Huang, K. Sudhir, Nisheeth K. Vishnoi
arXiv 2 Nov 2020 · Machine Learning · 9 citations (OpenAlex)
arXiv:2011.00981 · PDF · DOI · OpenAlex · Extracted main text
This paper introduces the problem of coresets for regression problems to panel data settings. We first define coresets for several variants of regression problems with panel data and then present efficient algorithms to construct coresets of size that depend polynomially on 1/$\varepsilon$ (where $\varepsilon$ is the error parameter) and the number of regression parameters - independent of the number of individuals in the panel data or the time units each individual is observed for. Our approach is based on the Feldman-Langberg framework in which a key step is to upper bound the "total sensitivity" that is roughly the sum of maximum influences of all individual-time pairs taken over all possible choices of regression parameters. Empirically, we assess our approach with synthetic and real-world datasets; the coreset sizes constructed using our approach are much smaller than the full dataset and coresets indeed accelerate the running time of computing the regression objective.
appendix boundary found by appendix_command · 88% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Dan Feldman and Michael Langberg (2011) A unified framework for approximating and clustering data | 1.000 | 9 | 3 | 100% |
| 2 | Vladimir Braverman, Dan Feldman, and Harry Lang (2016) New frameworks for offline and streaming coreset constructions | 0.874 | 7 | 2 | 100% |
| 3 | James P LeSage (1999) The theory and practice of spatial econometrics | 0.843 | 3 | 3 | 100% |
| 4 | Christos Boutsidis, Petros Drineas, and Malik Magdon-Ismail (2013) Near-optimal coresets for least-squares regression | 0.693 | 12 | 3 | 33% |
| 5 | Elad Tolochinsky and Dan Feldman (2018) Generic coreset for scalable learning of monotonic kernels: Logistic regression, sigmoid and more | 0.644 | 2 | 2 | 100% |
| 6 | John N Haddad (1998) A simple method for computing the covariance matrix and its inverse of a stationary autoregressive process | 0.644 | 2 | 2 | 100% |
| 7 | Mario Lucic, Matthew Faulkner, Andreas Krause, and Dan Feldman (2017) Training Gaussian mixture models at scale via coresets | 0.644 | 2 | 2 | 100% |
| 8 | Ibrahim Jubran, Alaa Maalouf, and Dan Feldman (2019) Fast and accurate least-mean-squares solvers | 0.567 | 11 | 2 | 27% |
| 9 | Martin Anthony and Peter L Bartlett (2009) Neural network learning: Theoretical foundations | 0.511 | 2 | 1 | 100% |
| 10 | Michael B Cohen, Yin Tat Lee, Cameron Musco, Christopher Musco, Rich… (2015) Uniform sampling for matrix approximation | 0.511 | 2 | 1 | 100% |
Showing the top 10 of 53 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Coresets for Time Series Clustering | 1.000 | 14 | 6 |