Identity From Here, Pose From There: Self-Supervised Disentanglement and Generation of Objects Using Unlabeled Videos
We propose a novel approach that disentangles the identity and pose of objects for image generation. Our model takes as input an ID image and a pose image, and generates an output image with the identity of the ID image and the pose of the pose image. Unlike most previous unsupervised work which rely on cyclic constraints, which can often be brittle, we instead propose to learn this in a self-supervised way. Specifically, we leverage unlabeled videos to automatically construct pseudo ground-truth targets to directly supervise our model. To enforce disentanglement, we propose a novel disentanglement loss, and to improve realism, we propose a pixel-verification loss in which the generated image's pixels must trace back to the ID input. We conduct extensive experiments on both synthetic and real images to demonstrate improved realism, diversity, and ID/pose disentanglement compared to existing methods.
Code (0)
등록된 구현이 없습니다.
Tasks
DisentanglementDiversityImage GenerationSimilar Papers 제목 키워드 기반
Identity-Disentangled Adversarial Augmentation for Self-supervised Learning
Data augmentation is critical to contrastive self-supervised learning, whose goal is to distinguish a sample's augmentations (positives) from other samples (negatives). However, strong augmentations may change the sample…
Contrastive LearningData AugmentationSelf-Supervised LearningMandarin-English Code-switching Speech Recognition with Self-supervised Speech Representation Models
Code-switching (CS) is common in daily conversations where more than one language is used within a sentence. The difficulties of CS speech recognition lie in alternating languages and the lack of transcribed data. Theref…
Language IdentificationSelf-Supervised LearningSentencespeech-recognition+1Asymmetric Mask Scheme for Self-Supervised Real Image Denoising
In recent years, self-supervised denoising methods have gained significant success and become critically important in the field of image restoration. Among them, the blind spot network based methods are the most typical …
DenoisingImage DenoisingImage RestorationIntra-Camera Supervised Person Re-Identification: A New Benchmark
Existing person re-identification (re-id) methods rely mostly on a large set of inter-camera identity labelled training data, requiring a tedious data collection and annotation process therefore leading to poor scalabili…
Multi-Label LearningPerson Re-IdentificationSelf-supervised Data Bootstrapping for Deep Optical Character Recognition of Identity Documents
The essential task of verifying person identities at airports and national borders is very time consuming. To accelerate it, optical character recognition for identity documents (IDs) using dictionaries is not appropriat…
Optical Character RecognitionOptical Character Recognition (OCR)