paper-with-me

홈 › Papers

Generating Images from Sounds Using Multimodal Features and GANs

2018-09-27 · Jeonghyun Lyu, Takashi Shinozaki, Kaoru Amano

Although generative adversarial networks (GANs) have enabled us to convert images from one domain to another similar one, converting between different sensory modalities, such as images and sounds, has been difficult. This study aims to propose a network that reconstructs images from sounds. First, video data with both images and sounds are labeled with pre-trained classifiers. Second, image and sound features are extracted from the data using pre-trained classifiers. Third, multimodal layers are introduced to extract features that are common to both the images and sounds. These layers are trained to extract similar features regardless of the input modality, such as images only, sounds only, and both images and sounds. Once the multimodal layers have been trained, features are extracted from input sounds and converted into image features using a feature-to-feature GAN. Finally, the generated image features are used to reconstruct images. Experimental results show that this method can successfully convert from the sound domain into the image domain. When we applied a pre-trained classifier to both the generated and original images, 31.9% of the examples had at least one of their top 10 labels in common, suggesting reasonably good image generation. Our results suggest that common representations can be learned for different modalities, and that proposed method can be applied not only to sound-to-image conversion but also to other conversions, such as from images to sounds.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Image Generation

Similar Papers 제목 키워드 기반

PMC-GANs: Generating Multi-Scale High-Quality Pedestrian with Multimodal Cascaded GANs

2019-12-30 · Jie Wu, Ying Peng, Chenghao Zheng, Zongbo Hao 외

Recently, generative adversarial networks (GANs) have shown great advantages in synthesizing images, leading to a boost of explorations of using faked images to augment data. This paper proposes a multimodal cascaded gen…

Data AugmentationPedestrian Detection

EnvGAN: Adversarial Synthesis of Environmental Sounds for Data Augmentation

2021-04-15 · Aswathy Madhu, Suresh K

The research in Environmental Sound Classification (ESC) has been progressively growing with the emergence of deep learning algorithms. However, data scarcity poses a major hurdle for any huge advance in this domain. Dat…

Data AugmentationEnvironmental Sound ClassificationSound Classification

Structure from Silence: Learning Scene Structure from Ambient Sound

2021-11-10 · Ziyang Chen, Xixi Hu, Andrew Owens

From whirling ceiling fans to ticking clocks, the sounds that we hear subtly vary as we move through a scene. We ask whether these ambient sounds convey information about 3D scene structure and, if so, whether they provi…

SynthScribe: Deep Multimodal Tools for Synthesizer Sound Retrieval and Exploration

2023-12-07 · Stephen Brade, Bryan Wang, Mauricio Sousa, Gregory Lee Newsome 외

Synthesizers are powerful tools that allow musicians to create dynamic and original sounds. Existing commercial interfaces for synthesizers typically require musicians to interact with complex low-level parameters or to …

Multimodal Deep LearningRetrieval

Generating Realistic Images from In-the-wild Sounds

2023-09-05 · ICCV 2023 1 · Taegyeong Lee, Jeonghun Kang, Hyeonyu Kim, Taehwan Kim

Representing wild sounds as images is an important but challenging task due to the lack of paired datasets between sound and images and the significant differences in the characteristics of these two modalities. Previous…

Audio captioningSentence