paper-with-me

홈 › Papers

Images that Sound: Composing Images and Sounds on a Single Canvas

2024-05-20 · Ziyang Chen, Daniel Geng, Andrew Owens

Spectrograms are 2D representations of sound that look very different from the images found in our visual world. And natural images, when played as spectrograms, make unnatural sounds. In this paper, we show that it is possible to synthesize spectrograms that simultaneously look like natural images and sound like natural audio. We call these visual spectrograms images that sound. Our approach is simple and zero-shot, and it leverages pre-trained text-to-image and text-to-spectrogram diffusion models that operate in a shared latent space. During the reverse process, we denoise noisy latents with both the audio and image diffusion models in parallel, resulting in a sample that is likely under both models. Through quantitative evaluations and perceptual studies, we find that our method successfully generates spectrograms that align with a desired audio prompt while also taking the visual appearance of a desired image prompt. Please see our project page for video results: https://ificl.github.io/images-that-sound/

📄 PDF Abstract BibTeX arXiv:2405.12221

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Generating Images from Sounds Using Multimodal Features and GANs

2018-09-27 · Jeonghyun Lyu, Takashi Shinozaki, Kaoru Amano

Although generative adversarial networks (GANs) have enabled us to convert images from one domain to another similar one, converting between different sensory modalities, such as images and sounds, has been difficult. Th…

Image Generation

Generating Realistic Images from In-the-wild Sounds

2023-09-05 · ICCV 2023 1 · Taegyeong Lee, Jeonghun Kang, Hyeonyu Kim, Taehwan Kim

Representing wild sounds as images is an important but challenging task due to the lack of paired datasets between sound and images and the significant differences in the characteristics of these two modalities. Previous…

Audio captioningSentence

Seeing Soundscapes: Audio-Visual Generation and Separation from Soundscapes Using Audio-Visual Separator

2025-04-25 · Minjae Kang, Martim Brandão

Recent audio-visual generative models have made substantial progress in generating images from audio. However, existing approaches focus on generating images from single-class audio and fail to generate images from mixed…

Sat2Sound: A Unified Framework for Zero-Shot Soundscape Mapping

2025-05-19 · Subash Khanal, Srikumar Sastry, Aayush Dhakal, Adeel Ahmad 외

We present Sat2Sound, a multimodal representation learning framework for soundscape mapping, designed to predict the distribution of sounds at any location on Earth. Existing methods for this task rely on satellite image…

Contrastive LearningCross-Modal RetrievalDiversityImage Captioning+3

SounDiT: Geo-Contextual Soundscape-to-Landscape Generation

2025-05-19 · JunBo Wang, Haofeng Tan, Bowen Liao, Albert Jiang 외

We present a novel and practically significant problem-Geo-Contextual Soundscape-to-Landscape (GeoS2L) generation-which aims to synthesize geographically realistic landscape images from environmental soundscapes. Prior a…

Image Generation