paper-with-me

홈 › Papers

Detection and Evaluation of Clusters within Sequential Data

2022-10-04 · Alexander Van Werde, Albert Senen-Cerda, Gianluca Kosmella, Jaron Sanders

Motivated by theoretical advancements in dimensionality reduction techniques we use a recent model, called Block Markov Chains, to conduct a practical study of clustering in real-world sequential data. Clustering algorithms for Block Markov Chains possess theoretical optimality guarantees and can be deployed in sparse data regimes. Despite these favorable theoretical properties, a thorough evaluation of these algorithms in realistic settings has been lacking. We address this issue and investigate the suitability of these clustering algorithms in exploratory data analysis of real-world sequential data. In particular, our sequential data is derived from human DNA, written text, animal movement data and financial markets. In order to evaluate the determined clusters, and the associated Block Markov Chain model, we further develop a set of evaluation tools. These tools include benchmarking, spectral noise analysis and statistical model selection tools. An efficient implementation of the clustering algorithm and the new evaluation tools is made available together with this paper. Practical challenges associated to real-world data are encountered and discussed. It is ultimately found that the Block Markov Chain model assumption, together with the tools developed here, can indeed produce meaningful insights in exploratory data analyses despite the complexity and sparsity of real-world data.

📄 PDF Abstract BibTeX arXiv:2210.01679

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingClusteringDimensionality ReductionModel Selection

Similar Papers 제목 키워드 기반

Estimating the number of clusters of a Block Markov Chain

2024-07-25 · Thomas van Vuren, Thomas Cronk, Jaron Sanders

Clustering algorithms frequently require the number of clusters to be chosen in advance, but it is usually not clear how to do this. To tackle this challenge when clustering within sequential data, we present a method fo…

ClusteringStochastic Block Model

ClusterSeq: Enhancing Sequential Recommender Systems with Clustering based Meta-Learning

2023-07-25 · Mohammmadmahdi Maheri, Reza Abdollahzadeh, Bardia Mohammadi, Mina Rafiei 외

In practical scenarios, the effectiveness of sequential recommendation systems is hindered by the user cold-start problem, which arises due to limited interactions for accurately determining user preferences. Previous st…

ClusteringMeta-LearningRecommendation SystemsSequential Recommendation

Subspace Clustering for Sequential Data

2014-06-01 · CVPR 2014 6 · Stephen Tierney, Junbin Gao, Yi Guo

We propose Ordered Subspace Clustering (OSC) to segment data drawn from a sequentially ordered union of subspaces. Current subspace clustering techniques learn the relationships within a set of data and then use a separa…

Clustering

Streaming, Memory Limited Algorithms for Community Detection

2014-12-01 · NeurIPS 2014 12 · Se-Young Yun, Marc Lelarge, Alexandre Proutiere

In this paper, we consider sparse networks consisting of a finite number of non-overlapping communities, i.e. disjoint clusters, so that there is higher density within clusters than across clusters. Both the intra- and i…

ClusteringCommunity Detection

Weakly Supervised Spatio-Temporal Candidate Discovery of Dairy Farm Sites from Seasonal Satellite Imagery

2026-07-14 · Usman Haider, Fatima Khalid, Karl Mason arxiv

Farm site discovery from satellite imagery is a spatiotemporal candidate ranking problem because farm evidence is distributed across pasture, field boundaries, roads, buildings, and seasonal vegetation patterns. Direct f…

Representation Learning