paper-with-me

Papers

Clustering Unclustered Data: Unsupervised Binary Labeling of Two Datasets Having Different Class Balances

2013-05-01 · Marthinus Christoffel du Plessis, Masashi Sugiyama

We consider the unsupervised learning problem of assigning labels to unlabeled data. A naive approach is to use clustering methods, but this works well only when data is properly clustered and each cluster corresponds to an underlying class. In this paper, we first show that this unsupervised labeling problem in balanced binary cases can be solved if two unlabeled datasets having different class balances are available. More specifically, estimation of the sign of the difference between probability densities of two unlabeled datasets gives the solution. We then introduce a new method to directly estimate the sign of the density difference without density estimation. Finally, we demonstrate the usefulness of the proposed method against several clustering methods on various toy problems and real-world datasets.

📄 PDF Abstract BibTeX arXiv:1305.0103

Code (0)

등록된 구현이 없습니다.

Tasks

ClusteringDensity Estimation

Similar Papers 제목 키워드 기반

A Dynamic Framework for Semantic Grouping of Common Data Elements (CDE) Using Embeddings and Clustering

2025-06-02 · Madan Krishnamurthy, Daniel Korn, Melissa A Haendel, Christopher J Mungall 외

This research aims to develop a dynamic and scalable framework to facilitate harmonization of Common Data Elements (CDEs) across heterogeneous biomedical datasets by addressing challenges such as semantic heterogeneity, …

Clusteringscientific discovery

Clustering by Attention: Leveraging Prior Fitted Transformers for Data Partitioning

2025-07-27 · Ahmed Shokry, Ayman Khalafallah arxiv

Clustering is a core task in machine learning with wide-ranging applications in data mining and pattern recognition. However, its unsupervised nature makes it inherently challenging. Many existing clustering algorithms s…

SPICE: Semantic Pseudo-labeling for Image Clustering

2021-03-17 · Chuang Niu, Hongming Shan, Ge Wang

The similarity among samples and the discrepancy between clusters are two crucial aspects of image clustering. However, current deep clustering methods suffer from the inaccurate estimation of either feature similarity o…

ClusteringContrastive LearningDeep ClusteringImage Clustering+3

Auto-Dialabel: Labeling Dialogue Data with Unsupervised Learning

2018-10-01 · EMNLP 2018 10 · Chen Shi, Qi Chen, Lei Sha, Sujian Li 외

The lack of labeled data is one of the main challenges when building a task-oriented dialogue system. Existing dialogue datasets usually rely on human labeling, which is expensive, limited in size, and in low coverage. I…

Active LearningClustering

Semantic-Aware Task Clustering for Constructive and Cooperative Multi-Tasking

2026-07-23 · Ahmad Halimi Razlighi, Maximilian H. V. Tillmann, Edgar Beck, Bho Matthiesen 외 arxiv

Cooperative multi-task semantic communication (CMT-SemCom) improves task execution performance by leveraging shared representations. However, as we demonstrated in [1], cooperative multi-tasking can be either constructiv…

Semantic Communication