Generating images from caption and vice versa via CLIP-Guided Generative Latent Space Search
In this research work we present CLIP-GLaSS, a novel zero-shot framework to generate an image (or a caption) corresponding to a given caption (or image). CLIP-GLaSS is based on the CLIP neural network, which, given an image and a descriptive caption, provides similar embeddings. Differently, CLIP-GLaSS takes a caption (or an image) as an input, and generates the image (or the caption) whose CLIP embedding is the most similar to the input one. This optimal image (or caption) is produced via a generative network, after an exploration by a genetic algorithm. Promising results are shown, based on the experimentation of the image Generators BigGAN and StyleGAN2, and of the text Generator GPT2
Code (3)
Tasks
DescriptiveImage GenerationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Towards Generating Diverse Audio Captions via Adversarial Training
Automated audio captioning is a cross-modal translation task for describing the content of audio clips with natural language sentences. This task has attracted increasing attention and substantial progress has been made …
Audio captioningDiversityGenerative Adversarial NetworkDiverse Audio Captioning via Adversarial Training
Audio captioning aims at generating natural language descriptions for audio clips automatically. Existing audio captioning models have shown promising improvement in recent years. However, these models are mostly trained…
Audio captioningDiversityGenerative Adversarial NetworkSentenceEffectively Leveraging CLIP for Generating Situational Summaries of Images and Videos
Situation recognition refers to the ability of an agent to identify and understand various situations or contexts based on available information and sensory inputs. It involves the cognitive process of interpreting data …
Semantic Role LabelingVideo CaptioningPragmatic Inference with a CLIP Listener for Contrastive Captioning
We propose a simple yet effective and robust method for contrastive captioning: generating discriminative captions that distinguish target images from very similar alternative distractor images. Our approach is built on …
CLIP4IDC: CLIP for Image Difference Captioning
Image Difference Captioning (IDC) aims at generating sentences to describe differences between two similar-looking images. Conventional approaches learn an IDC model with a pre-trained and usually frozen visual feature e…
Domain AdaptationImage Classification