paper-with-me

홈 › Papers

ELFS: Enhancing Label-Free Coreset Selection via Clustering-based Pseudo-Labeling

2024-06-06 · Haizhong Zheng, Elisa Tsai, Yifu Lu, Jiachen Sun, Brian R. Bartoldson, Bhavya Kailkhura, Atul Prakash

High-quality human-annotated data is crucial for modern deep learning pipelines, yet the human annotation process is both costly and time-consuming. Given a constrained human labeling budget, selecting an informative and representative data subset for labeling can significantly reduce human annotation effort. Well-performing state-of-the-art (SOTA) coreset selection methods require ground-truth labels over the whole dataset, failing to reduce the human labeling burden. Meanwhile, SOTA label-free coreset selection methods deliver inferior performance due to poor geometry-based scores. In this paper, we introduce ELFS, a novel label-free coreset selection method. ELFS employs deep clustering to estimate data difficulty scores without ground-truth labels. Furthermore, ELFS uses a simple but effective double-end pruning method to mitigate bias on calculated scores, which further improves the performance on selected coresets. We evaluate ELFS on five vision benchmarks and show that ELFS consistently outperforms SOTA label-free baselines. For instance, at a 90% pruning rate, ELFS surpasses the best-performing baseline by 5.3% on CIFAR10 and 7.1% on CIFAR100. Moreover, ELFS even achieves comparable performance to supervised coreset selection at low pruning rates (e.g., 30% and 50%) on CIFAR10 and ImageNet-1K.

📄 PDF Abstract BibTeX arXiv:2406.04273

Code (1)

eltsai/elfs 공식 구현 pytorch

Tasks

ClusteringDeep Clustering

Methods 이 논문이 사용한 방법론

Pruning 설명 없음

Similar Papers 제목 키워드 기반

HyperCore: Coreset Selection under Noise via Hypersphere Models

2025-09-26 · Brian B. Moser, Arundhati S. Shanbhag, Tobias C. Nauen, Stanislav Frolov 외 arxiv

The goal of coreset selection methods is to identify representative subsets of datasets for efficient model training. Yet, existing methods often ignore the possibility of annotation errors and require fixed pruning rati…

Zero-Shot Coreset Selection: Efficient Pruning for Unlabeled Data

2024-11-22 · Brent A. Griffin, Jacob Marks, Jason J. Corso

Deep learning increasingly relies on massive data with substantial costs for storage, annotation, and model training. To reduce these costs, coreset selection aims to find a representative subset of data to train models …

Anti-Backdoor Coreset Selection via Cumulative Entropy

2026-07-28 · Qi Zhao, Christian Wressnegger arxiv

Recent training-time defenses against neural backdoors isolate a benign subset from poisoned training data, to learn a backdoor-free model from it. In this paper, we formulate this defense strategy as a coreset selection…

Extending Contrastive Learning to Unsupervised Coreset Selection

2021-03-05 · Jeongwoo Ju, Heechul Jung, Yoonju Oh, Junmo Kim

Self-supervised contrastive learning offers a means of learning informative features from a pool of unlabeled data. In this paper, we delve into another useful approach -- providing a way of selecting a core-set that is …

Contrastive Learning

GraphSculptor: Sculpting Pre-training Coreset for Graph Self-supervised Learning

2026-05-02 · Chuang Liu, Zelin Yao, Xueqi Ma, Luzhi Wang 외 arxiv

Graph self-supervised learning typically relies on large-scale unlabeled datasets, heavily inflating computational costs. However, empirical evidence suggests that these datasets contain substantial redundancy-our analys…

Self-Supervised Learning