paper-with-me

Papers

Semi-supervised cross-entropy clustering with information bottleneck constraint

2017-05-03 · Marek Śmieja, Bernhard C. Geiger

In this paper, we propose a semi-supervised clustering method, CEC-IB, that models data with a set of Gaussian distributions and that retrieves clusters based on a partial labeling provided by the user (partition-level side information). By combining the ideas from cross-entropy clustering (CEC) with those from the information bottleneck method (IB), our method trades between three conflicting goals: the accuracy with which the data set is modeled, the simplicity of the model, and the consistency of the clustering with side information. Experiments demonstrate that CEC-IB has a performance comparable to Gaussian mixture models (GMM) in a classical semi-supervised scenario, but is faster, more robust to noisy labels, automatically determines the optimal number of clusters, and performs well when not all classes are present in the side information. Moreover, in contrast to other semi-supervised models, it can be successfully applied in discovering natural subgroups if the partition-level side information is derived from the top levels of a hierarchical clustering.

📄 PDF Abstract BibTeX arXiv:1705.01601

Code (0)

등록된 구현이 없습니다.

Tasks

Clustering

Similar Papers 제목 키워드 기반

Semi-Supervised Clustering via Structural Entropy with Different Constraints

2023-12-18 · Guangjie Zeng, Hao Peng, Angsheng Li, Zhiwei Liu 외

Semi-supervised clustering techniques have emerged as valuable tools for leveraging prior information in the form of constraints to improve the quality of clustering outcomes. Despite the proliferation of such methods, t…

Clustering

Fuzzy Overclustering: Semi-Supervised Classification of Fuzzy Labels with Overclustering and Inverse Cross-Entropy

2021-10-13 · Lars Schmarje, Johannes Brünger, Monty Santarossa, Simon-Martin Schröder 외

Deep learning has been successfully applied to many classification problems including underwater challenges. However, a long-standing issue with deep learning is the need for large and consistently labeled datasets. Alth…

Semi-supervised learning made simple with self-supervised clustering

2023-06-13 · CVPR 2023 1 · Enrico Fini, Pietro Astolfi, Karteek Alahari, Xavier Alameda-Pineda 외

Self-supervised learning models have been shown to learn rich visual representations without requiring human annotations. However, in many real-world scenarios, labels are partially available, motivating a recent line of…

ClusteringSelf-Supervised Learning

Exploring Information-Theoretic Metrics Associated with Neural Collapse in Supervised Training

2024-09-25 · Kun Song, Zhiquan Tan, Bochao Zou, Jiansheng Chen 외

In this paper, we utilize information-theoretic metrics like matrix entropy and mutual information to analyze supervised learning. We explore the information content of data representations and classification head weight…

Classificationcross-modal alignment

Semi-supervised Clustering of Medical Text

2016-12-01 · WS 2016 12 · Pracheta Sahoo, Asif Ekbal, Sriparna Saha, Diego Moll{\'a} 외

Semi-supervised clustering is an attractive alternative for traditional (unsupervised) clustering in targeted applications. By using the information of a small annotated dataset, semi-supervised clustering can produce cl…

Clustering