paper-with-me

홈 › Papers

Fast Rates for Semi-Supervised Learning via Data-Augmentation Graph Regularization

2026-07-08 · Adam M. Oberman arxiv

Self-supervised learning matches supervised accuracy from a fraction of the labels, but the labeled-sample efficiency behind this has lacked a theoretical explanation. We provide one. Data augmentation induces a similarity graph on the unlabeled data, so downstream learning on that graph is graph-Laplacian-regularized learning. We prove a fast transductive rate, $O(1/n_L)$ in the number of labels, in place of the supervised $O(1/\sqrt{n_L})$, by carrying the leave-one-out stability apparatus of Johnson and Zhang (JMLR 2007) over to the augmentation graph, and without the unrealistic assumptions of limit-based analyses (exact kernel, generalizing features). The bound makes augmentation quality explicit: the expected error is at most $C/n_L + R_{\mathrm{DA}}(y)$, where the data-augmentation alignment error $R_{\mathrm{DA}}(y)$ is the graph-cut mass of augmentations that cross a label boundary, so good augmentations let few labels suffice. The analysis uses a streamlined loss that drops the projector, negative-sample, and orthogonality overhead of standard objectives yet still recovers the top-$K$ ideal features in the infinite-data limit, the augmentation-kernel eigenspace studied by Zhai et al. The result explains the observed accuracy-versus-label-count curve rather than only bounding a generalization gap.

📄 PDF Abstract BibTeX arXiv:2607.07513

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised LearningData Augmentation

Similar Papers 제목 키워드 기반

ClassMix: Segmentation-Based Data Augmentation for Semi-Supervised Learning

2020-07-15 · Viktor Olsson, Wilhelm Tranheden, Juliano Pinto, Lennart Svensson

The state of the art in semantic segmentation is steadily increasing in performance, resulting in more precise and reliable segmentations in many different applications. However, progress is limited by the cost of genera…

Data AugmentationSegmentationSemantic SegmentationSemi-Supervised Semantic Segmentation

Augmentation Matters: A Simple-yet-Effective Approach to Semi-supervised Semantic Segmentation

2022-12-09 · CVPR 2023 1 · Zhen Zhao, Lihe Yang, Sifan Long, Jimin Pi 외

Recent studies on semi-supervised semantic segmentation (SSS) have seen fast progress. Despite their promising performance, current state-of-the-art methods tend to increasingly complex designs at the cost of introducing…

Semantic SegmentationSemi-Supervised Semantic Segmentation

Camouflaged Chinese Spam Content Detection with Semi-supervised Generative Active Learning

2020-07-01 · ACL 2020 6 · Zhuoren Jiang, Zhe Gao, Yu Duan, Yangyang Kang 외

We propose a Semi-supervIsed GeNerative Active Learning (SIGNAL) model to address the imbalance, efficiency, and text camouflage problems of Chinese text spam detection task. A {``}self-diversity{''} criterion is propose…

Active LearningChinese Spam DetectionData AugmentationDiversity+1

SemiETPicker: Fast and Label-Efficient Particle Picking for CryoET Tomography Using Semi-Supervised Learning

2025-10-25 · Linhan Wang, Jianwen Dou, Wang Li, Shengkun Wang 외 arxiv

Cryogenic Electron Tomography (CryoET) combined with sub-volume averaging (SVA) is the only imaging modality capable of resolving protein structures inside cells at molecular resolution. Particle picking, the task of loc…

Keypoint Detection

Semi-Supervised Few-Shot Intent Classification and Slot Filling

2021-09-17 · Samyadeep Basu, Karine lp Kiun Chong, Amr Sharaf, Alex Fischer 외

Intent classification (IC) and slot filling (SF) are two fundamental tasks in modern Natural Language Understanding (NLU) systems. Collecting and annotating large amounts of data to train deep learning models for such sy…

ClassificationContrastive LearningData Augmentationintent-classification+6