Contrastive Learning for Diverse Disentangled Foreground Generation
We introduce a new method for diverse foreground generation with explicit control over various factors. Existing image inpainting based foreground generation methods often struggle to generate diverse results and rarely allow users to explicitly control specific factors of variation (e.g., varying the facial identity or expression for face inpainting results). We leverage contrastive learning with latent codes to generate diverse foreground results for the same masked input. Specifically, we define two sets of latent codes, where one controls a pre-defined factor (`known''), and the other controls the remaining factors (`unknown''). The sampled latent codes from the two sets jointly bi-modulate the convolution kernels to guide the generator to synthesize diverse results. Experiments demonstrate the superiority of our method over state-of-the-arts in result diversity and generation controllability.
Code (0)
등록된 구현이 없습니다.
Tasks
Contrastive LearningDiversityFacial InpaintingImage InpaintingMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Attribute2Image: Conditional Image Generation from Visual Attributes
This paper investigates a novel problem of generating images from visual attributes. We model the image as a composite of foreground and background and develop a layered generative model with disentangled latent variable…
AttributeConditional Image GenerationImage GenerationImage ReconstructionDisentangled and Controllable Face Image Generation via 3D Imitative-Contrastive Learning
We propose DiscoFaceGAN, an approach for face image generation of virtual people with disentangled, precisely-controllable latent representations for identity of non-existing people, expression, pose, and illumination. W…
Contrastive LearningDisentanglementImage GenerationDisentangled Person Image Generation
Generating novel, yet realistic, images of persons is a challenging task due to the complex interplay between the different image factors, such as the foreground, background and pose information. In this work, we aim at …
Gesture-to-Gesture TranslationImage GenerationPerson Re-IdentificationPose TransferDisentangled Contrastive Image Translation for Nighttime Surveillance
Nighttime surveillance suffers from degradation due to poor illumination and arduous human annotations. It is challengable and remains a security risk at night. Existing methods rely on multi-spectral images to perceive …
Contrastive LearningTranslationDualVAE: Dual Disentangled Variational AutoEncoder for Recommendation
Learning precise representations of users and items to fit observed interaction data is the fundamental task of collaborative filtering. Existing studies usually infer entangled representations to fit such interaction da…
Collaborative FilteringDisentanglementRepresentation LearningVariational Inference