paper-with-me

홈 › Papers

On-the-fly Object Detection using StyleGAN with CLIP Guidance

2022-10-30 · Yuzhe Lu, Shusen Liu, Jayaraman J. Thiagarajan, Wesam Sakla, Rushil Anirudh

We present a fully automated framework for building object detectors on satellite imagery without requiring any human annotation or intervention. We achieve this by leveraging the combined power of modern generative models (e.g., StyleGAN) and recent advances in multi-modal learning (e.g., CLIP). While deep generative models effectively encode the key semantics pertinent to a data distribution, this information is not immediately accessible for downstream tasks, such as object detection. In this work, we exploit CLIP's ability to associate image features with text descriptions to identify neurons in the generator network, which are subsequently used to build detectors on-the-fly.

📄 PDF Abstract BibTeX arXiv:2210.16742

Code (0)

등록된 구현이 없습니다.

Tasks

Objectobject-detectionObject Detection

Similar Papers 제목 키워드 기반

CLIP2StyleGAN: Unsupervised Extraction of StyleGAN Edit Directions

2021-12-09 · Rameen Abdal, Peihao Zhu, John Femiani, Niloy J. Mitra 외

The success of StyleGAN has enabled unprecedented semantic editing capabilities, on both synthesized and real images. However, such editing operations are either trained with semantic supervision or described using human…

Zero-Shot Learning

StyleHumanCLIP: Text-guided Garment Manipulation for StyleGAN-Human

2023-05-26 · Takato Yoshikawa, Yuki Endo, Yoshihiro Kanamori

This paper tackles text-guided control of StyleGAN for editing garments in full-body human images. Existing StyleGAN-based methods suffer from handling the rich diversity of garments and body shapes and poses. We propose…

DiversityImage Generation

Bridging CLIP and StyleGAN through Latent Alignment for Image Editing

2022-10-10 · Wanfeng Zheng, Qiang Li, Xiaoyan Guo, Pengfei Wan 외

Text-driven image manipulation is developed since the vision-language model (CLIP) has been proposed. Previous work has adopted CLIP to design a text-image consistency-based objective to address this issue. However, thes…

Image GenerationImage ManipulationLanguage ModelingLanguage Modelling+2

clip2latent: Text driven sampling of a pre-trained StyleGAN using denoising diffusion and CLIP

2022-10-05 · Justin N. M. Pinkney, Chuan Li

We introduce a new method to efficiently create text-to-image models from a pre-trained CLIP and StyleGAN. It enables text driven sampling with an existing generative model without any external data or fine-tuning. This …

Denoising

TräumerAI: Dreaming Music with StyleGAN

2021-02-09 · Dasaem Jeong, Seungheon Doh, Taegyun Kwon

The goal of this paper to generate a visually appealing video that responds to music with a neural network so that each frame of the video reflects the musical characteristics of the corresponding audio clip. To achieve …

Music Auto-Tagging