paper-with-me

Papers

How does unlabeled data improve generalization in self-training? A one-hidden-layer theoretical analysis

2022-01-21 · Shuai Zhang, Meng Wang, Sijia Liu, Pin-Yu Chen, JinJun Xiong

Self-training, a semi-supervised learning algorithm, leverages a large amount of unlabeled data to improve learning when the labeled data are limited. Despite empirical successes, its theoretical characterization remains elusive. To the best of our knowledge, this work establishes the first theoretical analysis for the known iterative self-training paradigm and proves the benefits of unlabeled data in both training convergence and generalization ability. To make our theoretical analysis feasible, we focus on the case of one-hidden-layer neural networks. However, theoretical understanding of iterative self-training is non-trivial even for a shallow neural network. One of the key challenges is that existing neural network landscape analysis built upon supervised learning no longer holds in the (semi-supervised) self-training paradigm. We address this challenge and prove that iterative self-training converges linearly with both convergence rate and generalization accuracy improved in the order of $1/\sqrt{M}$, where $M$ is the number of unlabeled samples. Experiments from shallow neural networks to deep neural networks are also provided to justify the correctness of our established theoretical insights on self-training.

📄 PDF Abstract BibTeX arXiv:2201.08514

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

How unlabeled data improve generalization in self-training? A one-hidden-layer theoretical analysis

2021-09-29 · ICLR 2022 4 · Shuai Zhang, Meng Wang, Sijia Liu, Pin-Yu Chen 외

Self-training, a semi-supervised learning algorithm, leverages a large amount of unlabeled data to improve learning when the labeled data are limited. Despite empirical successes, its theoretical characterization remains…

BenchHAR: Benchmarking Self-Supervised Learning for Generalizable Sensor-based Activity Recognition

2026-05-08 · Yize Cai, Rui Feng, Anlan Yu, Baoshen Guo 외 arxiv

Human Activity Recognition (HAR) from wearable sensors supports broad healthcare and behavior science applications. However, data heterogeneity and the scarcity of labeled data limit its real-world generalization. Recent…

Human Activity RecognitionSelf-Supervised Learning

Adversarially Robust Generalization Just Requires More Unlabeled Data

2019-06-03 · Runtian Zhai, Tianle Cai, Di He, Chen Dan 외

Neural network robustness has recently been highlighted by the existence of adversarial examples. Many previous works show that the learned networks do not perform well on perturbed test data, and significantly more labe…

Out-distribution aware Self-training in an Open World Setting

2020-12-21 · Maximilian Augustin, Matthias Hein

Deep Learning heavily depends on large labeled datasets which limits further improvements. While unlabeled data is available in large amounts, in particular in image recognition, it does not fulfill the closed world assu…

Harnessing Unlabeled Data to Improve Generalization of Biometric Gender and Age Classifiers

2021-10-09 · Aakash Varma Nadimpalli, Narsi Reddy, Sreeraj Ramachandran, Ajita Rattani

With significant advances in deep learning, many computer vision applications have reached the inflection point. However, these deep learning models need large amount of labeled data for model training and optimum parame…

Age ClassificationDeep LearningGender Classificationparameter estimation