Lingxiao Huang, K. Sudhir, Nisheeth K. Vishnoi
arXiv 28 Oct 2021 · Machine Learning · 6 citations (OpenAlex)
arXiv:2110.15263 · PDF · DOI · OpenAlex · Extracted main text
We study the problem of constructing coresets for clustering problems with time series data. This problem has gained importance across many fields including biology, medicine, and economics due to the proliferation of sensors facilitating real-time measurement and rapid drop in storage costs. In particular, we consider the setting where the time series data on $N$ entities is generated from a Gaussian mixture model with autocorrelations over $k$ clusters in $\mathbb{R}^d$. Our main contribution is an algorithm to construct coresets for the maximum likelihood objective for this mixture model. Our algorithm is efficient, and under a mild boundedness assumption on the covariance matrices of the underlying Gaussians, the size of the coreset is independent of the number of entities $N$ and the number of observations for each entity, and depends only polynomially on $k$, $d$ and $1/\varepsilon$, where $\varepsilon$ is the error parameter. We empirically assess the performance of our coreset with synthetic data.
appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Lingxiao Huang, K. Sudhir, and Nisheeth K. Vishnoi (2020) Coresets for regressions with panel data self | 1.000 | 14 | 6 | 100% |
| 2 | Mario Lucic, Matthew Faulkner, Andreas Krause, and Dan Feldman (2017) Training Gaussian mixture models at scale via coresets | 1.000 | 11 | 5 | 100% |
| 3 | Vladimir Braverman, Dan Feldman, and Harry Lang (2016) New frameworks for offline and streaming coreset constructions | 1.000 | 9 | 3 | 100% |
| 4 | Dan Feldman and Michael Langberg (2011) A unified framework for approximating and clustering data | 1.000 | 9 | 3 | 100% |
| 5 | Dan Feldman, Zahi Kfir, and Xuan Wu (2019) Coresets for Gaussian mixture models of any shape | 1.000 | 8 | 3 | 100% |
| 6 | T Warren Liao (2005) Clustering of time series data—a survey | 0.874 | 5 | 2 | 100% |
| 7 | Saeed Aghabozorgi, Ali Seyed Shirkhorshidi, and Teh Ying Wah (2015) Time-series clustering–a decade review | 0.811 | 4 | 2 | 100% |
| 8 | Dan Feldman, Matthew Faulkner, and Andreas Krause (2011) Scalable training of mixture models via coresets | 0.811 | 4 | 2 | 100% |
| 9 | Avrim Blum, John Hopcroft, and Ravindran Kannan (2020) Foundations of data science | 0.737 | 3 | 2 | 100% |
| 10 | Sariel Har-Peled and Soham Mazumdar (2004) On coresets for $k$-means and $k$-median clustering | 0.737 | 3 | 2 | 100% |
Showing the top 10 of 62 scored citations.