paper-with-me

홈 › Papers

A Fast Greedy Algorithm for Outlier Mining

2006-04-01 · Advances in Knowledge Discovery and Data Mining 2006 4 · Zengyou He, Shengchun Deng, Xiaofei Xu, Joshua Zhexue Huang

The task of outlier detection is to find small groups of data objects that are exceptional when compared with rest large amount of data. Recently, the problem of outlier detection in categorical data is defined as an optimization problem and a local-search heuristic based algorithm (LSA) is presented. However, as is the case with most iterative type algorithms, the LSA algorithm is still very time-consuming on very large datasets. In this paper, we present a very fast greedy algorithm for mining outliers under the same optimization model. Experimental results on real datasets and large synthetic datasets show that: (1) Our new algorithm has comparable performance with respect to those state-of-the-art outlier detection algorithms on identifying true outliers and (2) Our algorithm can be an order of magnitude faster than LSA algorithm.

📄 PDF Abstract BibTeX

Code (1)

ducthanhtran/greedy_outlier_entropy

Tasks

Outlier Detection

Similar Papers 제목 키워드 기반

Greedy Sampling for Approximate Clustering in the Presence of Outliers

2019-12-01 · NeurIPS 2019 12 · Aditya Bhaskara, Sharvaree Vadgama, Hong Xu

Greedy algorithms such as adaptive sampling (k-means++) and furthest point traversal are popular choices for clustering problems. One the one hand, they possess good theoretical approximation guarantees, and on the other…

Clustering

MSD-Kmeans: A Novel Algorithm for Efficient Detection of Global and Local Outliers

2019-10-15 · Yuanyuan Wei, Julian Jang-Jaccard, Fariza Sabrina, Timothy McIntosh

Outlier detection is a technique in data mining that aims to detect unusual or unexpected records in the dataset. Existing outlier detection algorithms have different pros and cons and exhibit different sensitivity to no…

ClusteringOutlier Detection

Greedy Strategy Works for $k$-Center Clustering with Outliers and Coreset Construction

2019-01-24 · Hu Ding, Haikuo Yu, Zixiu Wang

We study the problem of $k$-center clustering with outliers in arbitrary metrics and Euclidean space. Though a number of methods have been developed in the past decades, it is still quite challenging to design quality gu…

Clustering

Randomized Greedy Algorithms and Composable Coreset for k-Center Clustering with Outliers

2023-01-07 · Hu Ding, Ruomin Huang, Kai Liu, Haikuo Yu 외

In this paper, we study the problem of {\em $k$-center clustering with outliers}. The problem has many important applications in real world, but the presence of outliers can significantly increase the computational compl…

Clustering

Improved Algorithms for Overlapping and Robust Clustering of Edge-Colored Hypergraphs: An LP-Based Combinatorial Approach

2025-05-23 · Changyeol Lee, Yongho Shin, Hyung-Chan An

Clustering is a fundamental task in both machine learning and data mining. Among various methods, edge-colored clustering (ECC) has emerged as a useful approach for handling categorical data. Given a hypergraph with (hyp…

ClusteringComputational Efficiency