paper-with-me

Papers

Sound-Guided Semantic Image Manipulation

2021-11-30 · CVPR 2022 1 · Seung Hyun Lee, Wonseok Roh, Wonmin Byeon, Sang Ho Yoon, Chan Young Kim, Jinkyu Kim, Sangpil Kim

The recent success of the generative model shows that leveraging the multi-modal embedding space can manipulate an image using text information. However, manipulating an image with other sources rather than text, such as sound, is not easy due to the dynamic characteristics of the sources. Especially, sound can convey vivid emotions and dynamic expressions of the real world. Here, we propose a framework that directly encodes sound into the multi-modal (image-text) embedding space and manipulates an image from the space. Our audio encoder is trained to produce a latent representation from an audio input, which is forced to be aligned with image and text representations in the multi-modal embedding space. We use a direct latent optimization method based on aligned embeddings for sound-guided image manipulation. We also show that our method can mix text and audio modalities, which enrich the variety of the image modification. We verify the effectiveness of our sound-guided image manipulation quantitatively and qualitatively. We also show that our method can mix different modalities, i.e., text and audio, which enrich the variety of the image modification. The experiments on zero-shot audio classification and semantic-level image classification show that our proposed model outperforms other text and sound-guided state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:2112.00007

Code (1)

kuai-lab/sound-guided-semantic-image-manipulation 공식 구현 pytorch

Tasks

Audio Classificationimage-classificationImage ClassificationImage ManipulationZero-shot Audio Classification

Similar Papers 제목 키워드 기반

Robust Sound-Guided Image Manipulation

2022-08-30 · Seung Hyun Lee, Gyeongrok Oh, Wonmin Byeon, Sang Ho Yoon 외

Recent successes suggest that an image can be manipulated by a text prompt, e.g., a landscape scene on a sunny day is manipulated into the same scene on a rainy day driven by a text input "raining". These approaches ofte…

Image Manipulation

Robot Synesthesia: A Sound and Emotion Guided AI Painter

2023-02-09 · Vihaan Misra, Peter Schaldenbrand, Jean Oh

If a picture paints a thousand words, sound may voice a million. While recent robotic painting and image synthesis methods have achieved progress in generating visuals from text inputs, the translation of sound into imag…

Image GenerationImage Manipulation

TextCLIP: Text-Guided Face Image Generation And Manipulation Without Adversarial Training

2023-09-21 · Xiaozhou You, Jian Zhang

Text-guided image generation aimed to generate desired images conditioned on given texts, while text-guided image manipulation refers to semantically edit parts of a given image based on specified texts. For these two si…

Image GenerationImage Manipulationtext-guided-generation

AC-Foley: Reference-Audio-Guided Video-to-Audio Synthesis with Acoustic Transfer

2026-03-16 · Pengjun Fang, Yingqing He, Yazhou Xing, Qifeng Chen 외 arxiv

Existing video-to-audio (V2A) generation methods predominantly rely on text prompts alongside visual information to synthesize audio. However, two critical bottlenecks persist: semantic granularity gaps in training data,…

SoundBrush: Sound as a Brush for Visual Scene Editing

2024-12-31 · Kim Sung-Bin, Kim Jun-Seong, Junseok Ko, Yewon Kim 외

We propose SoundBrush, a model that uses sound as a brush to edit and manipulate visual scenes. We extend the generative capabilities of the Latent Diffusion Model (LDM) to incorporate audio information for editing visua…

Novel View Synthesis