EconBase
← All papers

Coresets for Time Series Clustering

Lingxiao Huang, K. Sudhir, Nisheeth K. Vishnoi

arXiv 28 Oct 2021 · Machine Learning · 6 citations (OpenAlex)

arXiv:2110.15263 · PDF · DOI · OpenAlex · Extracted main text

Abstract

We study the problem of constructing coresets for clustering problems with time series data. This problem has gained importance across many fields including biology, medicine, and economics due to the proliferation of sensors facilitating real-time measurement and rapid drop in storage costs. In particular, we consider the setting where the time series data on $N$ entities is generated from a Gaussian mixture model with autocorrelations over $k$ clusters in $\mathbb{R}^d$. Our main contribution is an algorithm to construct coresets for the maximum likelihood objective for this mixture model. Our algorithm is efficient, and under a mild boundedness assumption on the covariance matrices of the underlying Gaussians, the size of the coreset is independent of the number of entities $N$ and the number of observations for each entity, and depends only polynomially on $k$, $d$ and $1/\varepsilon$, where $\varepsilon$ is the error parameter. We empirically assess the performance of our coreset with synthetic data.

Citation extraction

62
references
134
in-text mentions
62
distinct cited
4
self-citations
14,911
main-text words

appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Lingxiao Huang, K. Sudhir, and Nisheeth K. Vishnoi (2020) Coresets for regressions with panel data self1.000146100%
2Mario Lucic, Matthew Faulkner, Andreas Krause, and Dan Feldman (2017) Training Gaussian mixture models at scale via coresets1.000115100%
3Vladimir Braverman, Dan Feldman, and Harry Lang (2016) New frameworks for offline and streaming coreset constructions1.00093100%
4Dan Feldman and Michael Langberg (2011) A unified framework for approximating and clustering data1.00093100%
5Dan Feldman, Zahi Kfir, and Xuan Wu (2019) Coresets for Gaussian mixture models of any shape1.00083100%
6T Warren Liao (2005) Clustering of time series data—a survey0.87452100%
7Saeed Aghabozorgi, Ali Seyed Shirkhorshidi, and Teh Ying Wah (2015) Time-series clustering–a decade review0.81142100%
8Dan Feldman, Matthew Faulkner, and Andreas Krause (2011) Scalable training of mixture models via coresets0.81142100%
9Avrim Blum, John Hopcroft, and Ravindran Kannan (2020) Foundations of data science0.73732100%
10Sariel Har-Peled and Soham Mazumdar (2004) On coresets for $k$-means and $k$-median clustering0.73732100%

Showing the top 10 of 62 scored citations.