paper-with-me

홈 › Papers

Simple KNN-Based Outlier Detection Achieves Robust Clustering

2026-05-08 · Tianle Jiang, Yufa Zhou arxiv

Being robust to the presence of outliers is crucial for applying clustering algorithms in practice. In the $\textit{robust $k$-Means}$ problem (i.e., $k$-Means with outliers), the goal is to remove $z$ outliers and minimize the $k$-Means cost on the remaining points. Despite the close connection between robust $k$-Means and outlier detection, both theoretical and empirical understanding of the effectiveness of $\textit{classic outlier detection heuristics}$ for robust $k$-Means remains limited. In this paper, we prove that under a practical assumption on the optimal cluster sizes, simply removing points with large $K$-Nearest-Neighbor distances achieves performance comparable to prior work in terms of approximation guarantees: it yields a constant-factor reduction from robust $k$-Means to standard $k$-Means, without introducing additional centers or discarding extra outliers, as is commonly required by existing approaches. Empirically, experiments on real-world datasets show that our method outperforms or matches several more sophisticated algorithms in terms of clustering cost and runtime. These results demonstrate that simple KNN-based heuristics can be surprisingly effective for robust clustering, highlighting new opportunities to bridge techniques from outlier detection and clustering.

📄 PDF Abstract BibTeX arXiv:2605.07130

Code (0)

등록된 구현이 없습니다.

Tasks

Outlier Detection

Similar Papers 제목 키워드 기반

Deep Clustering based Fair Outlier Detection

2021-06-09 · Hanyu Song, Peizhao Li, Hongfu Liu

In this paper, we focus on the fairness issues regarding unsupervised outlier detection. Traditional algorithms, without a specific design for algorithmic fairness, could implicitly encode and propagate statistical bias …

AttributeClusteringDeep ClusteringFairness+2

Detecting outliers by clustering algorithms

2024-12-07 · Qi Li, Shuliang Wang

Clustering and outlier detection are two important tasks in data mining. Outliers frequently interfere with clustering algorithms to determine the similarity between objects, resulting in unreliable clustering results. C…

ClusteringOutlier Detection

Linear-time Outlier Detection via Sensitivity

2016-05-02 · Mario Lucic, Olivier Bachem, Andreas Krause

Outliers are ubiquitous in modern data sets. Distance-based techniques are a popular non-parametric approach to outlier detection as they require no prior assumptions on the data generating distribution and are simple to…

ClusteringOutlier DetectionSensitivity

A Practical Algorithm for Distributed Clustering and Outlier Detection

2018-05-24 · NeurIPS 2018 12 · Jiecao Chen, Erfan Sadeqi Azer, Qin Zhang

We study the classic $k$-means/median clustering, which are fundamental problems in unsupervised learning, in the setting where data are partitioned across multiple sites, and where we are allowed to discard a small port…

ClusteringOutlier Detection

Onion-Peeling Outlier Detection in 2-D data Sets

2018-03-12 · Archit Harsh, John E. Ball, Pan Wei

Outlier Detection is a critical and cardinal research task due its array of applications in variety of domains ranging from data mining, clustering, statistical analysis, fraud detection, network intrusion detection and …

ClusteringFraud DetectionIntrusion DetectionNetwork Intrusion Detection+1