paper-with-me

홈 › Papers

A Practical Algorithm for Distributed Clustering and Outlier Detection

2018-05-24 · NeurIPS 2018 12 · Jiecao Chen, Erfan Sadeqi Azer, Qin Zhang

We study the classic $k$-means/median clustering, which are fundamental problems in unsupervised learning, in the setting where data are partitioned across multiple sites, and where we are allowed to discard a small portion of the data by labeling them as outliers. We propose a simple approach based on constructing small summary for the original dataset. The proposed method is time and communication efficient, has good approximation guarantees, and can identify the global outliers effectively. To the best of our knowledge, this is the first practical algorithm with theoretical guarantees for distributed clustering with outliers. Our experiments on both real and synthetic data have demonstrated the clear superiority of our algorithm against all the baseline algorithms in almost all metrics.

📄 PDF Abstract BibTeX arXiv:1805.09495

Code (0)

등록된 구현이 없습니다.

Tasks

ClusteringOutlier Detection

Similar Papers 제목 키워드 기반

On Integrated Clustering and Outlier Detection

2014-12-01 · NeurIPS 2014 12 · Lionel Ott, Linsey Pang, Fabio T. Ramos, Sanjay Chawla

We model the joint clustering and outlier detection problem using an extension of the facility location formulation. The advantages of combining clustering and outlier selection include: (i) the resulting clusters tend t…

ClusteringOutlier Detection

Simple KNN-Based Outlier Detection Achieves Robust Clustering

2026-05-08 · Tianle Jiang, Yufa Zhou arxiv

Being robust to the presence of outliers is crucial for applying clustering algorithms in practice. In the $\textit{robust $k$-Means}$ problem (i.e., $k$-Means with outliers), the goal is to remove $z$ outliers and minim…

Outlier Detection

Detecting outliers by clustering algorithms

2024-12-07 · Qi Li, Shuliang Wang

Clustering and outlier detection are two important tasks in data mining. Outliers frequently interfere with clustering algorithms to determine the similarity between objects, resulting in unreliable clustering results. C…

ClusteringOutlier Detection

Linear-time Outlier Detection via Sensitivity

2016-05-02 · Mario Lucic, Olivier Bachem, Andreas Krause

Outliers are ubiquitous in modern data sets. Distance-based techniques are a popular non-parametric approach to outlier detection as they require no prior assumptions on the data generating distribution and are simple to…

ClusteringOutlier DetectionSensitivity

Fast Distributed k-Center Clustering with Outliers on Massive Data

2015-12-01 · NeurIPS 2015 12 · Gustavo Malkomes, Matt J. Kusner, Wenlin Chen, Kilian Q. Weinberger 외

Clustering large data is a fundamental problem with a vast number of applications. Due to the increasing size of data, practitioners interested in clustering have turned to distributed computation methods. In this work…

ClusteringDistributed Computing