paper-with-me

홈 › Papers

An Empirical Study into Clustering of Unseen Datasets with Self-Supervised Encoders

2024-06-04 · Scott C. Lowe, Joakim Bruslund Haurum, Sageev Oore, Thomas B. Moeslund, Graham W. Taylor

Can pretrained models generalize to new datasets without any retraining? We deploy pretrained image models on datasets they were not trained for, and investigate whether their embeddings form meaningful clusters. Our suite of benchmarking experiments use encoders pretrained solely on ImageNet-1k with either supervised or self-supervised training techniques, deployed on image datasets that were not seen during training, and clustered with conventional clustering algorithms. This evaluation provides new insights into the embeddings of self-supervised models, which prioritize different features to supervised models. Supervised encoders typically offer more utility than SSL encoders within the training domain, and vice-versa far outside of it, however, fine-tuned encoders demonstrate the opposite trend. Clustering provides a way to evaluate the utility of self-supervised learned representations orthogonal to existing methods such as kNN. Additionally, we find the silhouette score when measured in a UMAP-reduced space is highly correlated with clustering performance, and can therefore be used as a proxy for clustering performance on data with no ground truth labels. Our code implementation is available at \url{https://github.com/scottclowe/zs-ssl-clustering/}.

📄 PDF Abstract BibTeX arXiv:2406.02465

Code (3)

scottclowe/zs-ssl-clustering 공식 구현 pytorch
bioscan-ml/bioscan-5m
zahrag/BIOSCAN-5M

Tasks

BenchmarkingClustering

Similar Papers 제목 키워드 기반

Multiple-Kernel Dictionary Learning for Reconstruction and Clustering of Unseen Multivariate Time-series

2019-03-05 · Babak Hosseini, Barbara Hammer

There exist many approaches for description and recognition of unseen classes in datasets. Nevertheless, it becomes a challenging problem when we deal with multivariate time-series (MTS) (e.g., motion data), where we can…

ClusteringDictionary LearningOnline ClusteringTime Series+1

Meta-Learning to Cluster

2019-10-30 · Yibo Jiang, Nakul Verma

Clustering is one of the most fundamental and wide-spread techniques in exploratory data analysis. Yet, the basic approach to clustering has not really changed: a practitioner hand-picks a task-specific clustering loss t…

ClusteringMeta-Learning

Automatic selection of clustering algorithms using supervised graph embedding

2020-11-16 · Noy Cohen-Shapira, Lior Rokach

The widespread adoption of machine learning (ML) techniques and the extensive expertise required to apply them have led to increased interest in automated ML solutions that reduce the need for human intervention. One of …

AutoMLClusteringGraph EmbeddingMeta-Learning

Clustering Time Series Data through Autoencoder-based Deep Learning Models

2020-04-11 · Neda Tavakoli, Sima Siami-Namini, Mahdi Adl Khanghah, Fahimeh Mirza Soltani 외

Machine learning and in particular deep learning algorithms are the emerging approaches to data analysis. These techniques have transformed traditional data mining-based analysis radically into a learning-based model in …

ClusteringDeep LearningTime SeriesTime Series Analysis

Fair Clustering Through Fairlets

2018-02-15 · NeurIPS 2017 12 · Flavio Chierichetti, Ravi Kumar, Silvio Lattanzi, Sergei Vassilvitskii

We study the question of fair clustering under the {\em disparate impact} doctrine, where each protected class must have approximately equal representation in every cluster. We formulate the fair clustering problem under…

Clustering