paper-with-me

홈 › Papers

Power to the Points: Validating Data Memberships in Clusterings

2013-05-21 · Parasaran Raman, Suresh Venkatasubramanian

A clustering is an implicit assignment of labels of points, based on proximity to other points. It is these labels that are then used for downstream analysis (either focusing on individual clusters, or identifying representatives of clusters and so on). Thus, in order to trust a clustering as a first step in exploratory data analysis, we must trust the labels assigned to individual data. Without supervision, how can we validate this assignment? In this paper, we present a method to attach affinity scores to the implicit labels of individual points in a clustering. The affinity scores capture the confidence level of the cluster that claims to "own" the point. This method is very general: it can be used with clusterings derived from Euclidean data, kernelized data, or even data derived from information spaces. It smoothly incorporates importance functions on clusters, allowing us to eight different clusters differently. It is also efficient: assigning an affinity score to a point depends only polynomially on the number of clusters and is independent of the number of points in the data. The dimensionality of the underlying space only appears in preprocessing. We demonstrate the value of our approach with an experimental study that illustrates the use of these scores in different data analysis tasks, as well as the efficiency and flexibility of the method. We also demonstrate useful visualizations of these scores; these might prove useful within an interactive analytics framework.

📄 PDF Abstract BibTeX arXiv:1305.4757

Code (0)

등록된 구현이 없습니다.

Tasks

Clustering

Similar Papers 제목 키워드 기반

Graph Sensitive Indices for Comparing Clusterings

2014-11-27 · Zaeem Hussain, Marina Meila

This report discusses two new indices for comparing clusterings of a set of points. The motivation for looking at new ways for comparing clusterings stems from the fact that the existing clustering indices are based on s…

Clustering

How To Overcome Richness Axiom Fallacy

2022-10-27 · Mieczysław A. Kłopotek, Robert A. Kłopotek

The paper points at the grieving problems implied by the richness axiom in the Kleinberg's axiomatic system and suggests resolutions. The richness induces learnability problem in general and leads to conflicts with consi…

Text Classification and Clustering with Annealing Soft Nearest Neighbor Loss

2021-07-23 · Abien Fred Agarap

We define disentanglement as how far class-different data points from each other are, relative to the distances among class-similar data points. When maximizing disentanglement during representation learning, we obtain a…

ClassificationClusteringDisentanglementRepresentation Learning+3

Spectral clustering in the dynamic stochastic block model

2017-05-02 · Marianna Pensky, Teng Zhang

In the present paper, we studied a Dynamic Stochastic Block Model (DSBM) under the assumptions that the connection probabilities, as functions of time, are smooth and that at most $s$ nodes can switch their class members…

ClusteringmodelStochastic Block Model

Contrastive Examples for Addressing the Tyranny of the Majority

2020-04-14 · Viktoriia Sharmanska, Lisa Anne Hendricks, Trevor Darrell, Novi Quadrianto

Computer vision algorithms, e.g. for face recognition, favour groups of individuals that are better represented in the training data. This happens because of the generalization that classifiers have to make. It is simple…

DiversityFace Recognition