On-the-fly Object Detection using StyleGAN with CLIP Guidance
We present a fully automated framework for building object detectors on satellite imagery without requiring any human annotation or intervention. We achieve this by leveraging the combined power of modern generative models (e.g., StyleGAN) and recent advances in multi-modal learning (e.g., CLIP). While deep generative models effectively encode the key semantics pertinent to a data distribution, this information is not immediately accessible for downstream tasks, such as object detection. In this work, we exploit CLIP's ability to associate image features with text descriptions to identify neurons in the generator network, which are subsequently used to build detectors on-the-fly.
Code (0)
등록된 구현이 없습니다.
Tasks
Objectobject-detectionObject DetectionSimilar Papers 제목 키워드 기반
CLIP2StyleGAN: Unsupervised Extraction of StyleGAN Edit Directions
The success of StyleGAN has enabled unprecedented semantic editing capabilities, on both synthesized and real images. However, such editing operations are either trained with semantic supervision or described using human…
Zero-Shot LearningStyleHumanCLIP: Text-guided Garment Manipulation for StyleGAN-Human
This paper tackles text-guided control of StyleGAN for editing garments in full-body human images. Existing StyleGAN-based methods suffer from handling the rich diversity of garments and body shapes and poses. We propose…
DiversityImage GenerationBridging CLIP and StyleGAN through Latent Alignment for Image Editing
Text-driven image manipulation is developed since the vision-language model (CLIP) has been proposed. Previous work has adopted CLIP to design a text-image consistency-based objective to address this issue. However, thes…
Image GenerationImage ManipulationLanguage ModelingLanguage Modelling+2clip2latent: Text driven sampling of a pre-trained StyleGAN using denoising diffusion and CLIP
We introduce a new method to efficiently create text-to-image models from a pre-trained CLIP and StyleGAN. It enables text driven sampling with an existing generative model without any external data or fine-tuning. This …
DenoisingTräumerAI: Dreaming Music with StyleGAN
The goal of this paper to generate a visually appealing video that responds to music with a neural network so that each frame of the video reflects the musical characteristics of the corresponding audio clip. To achieve …
Music Auto-Tagging