Learning Consistent Deep Generative Models from Sparsely Labeled Data
We consider training deep generative models toward two simultaneous goals: discriminative classification and generative modeling using an explicit likelihood. While variational autoencoders (VAEs) offer a promising solution, we show that the dominant approach to training semi-supervised VAEs has several key weaknesses: it is fragile as generative modeling capacity increases, it is slow due to a required marginalization over labels, and it incoherently decouples into separate discriminative and generative models when all data is labeled. We remedy these concerns in a new proposed framework for semi-supervised VAE training that considers a more coherent downstream model architecture and a new objective which maximizes generative quality subject to a task-specific prediction constraint that ensures discriminative quality. We further enforce a consistency constraint, derived naturally from the generative model, that requires predictions on reconstructed data to match those on the original data. We show that our contributions -- a downstream architecture with prediction constraints and consistency constraints -- lead to improved generative samples as well as accurate image classification, with consistency particularly crucial for accuracy on sparsely-labeled datasets. Our approach enables advances in generative modeling to directly boost semi-supervised classification, an ability we demonstrate by learning a "very deep" prediction-constrained VAE with many layers of latent variables.
Code (0)
등록된 구현이 없습니다.
Tasks
image-classificationImage ClassificationSimilar Papers 제목 키워드 기반
Sparsely Supervised Diffusion
Diffusion models have shown remarkable success across a wide range of generative tasks. However, they often suffer from spatially inconsistent generation, arguably due to the inherent locality of their denoising mechanis…
Improving Event Detection via Open-domain Trigger Knowledge
Event Detection (ED) is a fundamental task in automatically structuring texts. Due to the small scale of training data, previous methods perform poorly on unseen/sparsely labeled trigger words and are prone to overfittin…
Event DetectionKnowledge DistillationMonoSAOD: Monocular 3D Object Detection with Sparsely Annotated Label
Monocular 3D object detection has achieved impressive performance on densely annotated datasets. However, it struggles when only a fraction of objects are labeled due to the high cost of 3D annotation. This sparsely anno…
Monocular 3D Object DetectionUnsupervised Domain Adaptation: from Simulation Engine to the RealWorld
Large-scale labeled training datasets have enabled deep neural networks to excel on a wide range of benchmark vision tasks. However, in many applications it is prohibitively expensive or time-consuming to obtain large qu…
Domain AdaptationUnsupervised Domain AdaptationSemiMultiPose: A Semi-supervised Multi-animal Pose Estimation Framework
Multi-animal pose estimation is essential for studying animals' social behaviors in neuroscience and neuroethology. Advanced approaches have been proposed to support multi-animal estimation and achieve state-of-the-art p…
Animal Pose EstimationPose Estimation