paper-with-me

홈 › Papers

Towards Audio to Scene Image Synthesis using Generative Adversarial Network

2018-08-13 · Chia-Hung Wan, Shun-Po Chuang, Hung-Yi Lee

Humans can imagine a scene from a sound. We want machines to do so by using conditional generative adversarial networks (GANs). By applying the techniques including spectral norm, projection discriminator and auxiliary classifier, compared with naive conditional GAN, the model can generate images with better quality in terms of both subjective and objective evaluations. Almost three-fourth of people agree that our model have the ability to generate images related to sounds. By inputting different volumes of the same sound, our model output different scales of changes based on the volumes, showing that our model truly knows the relationship between sounds and images to some extent.

📄 PDF Abstract BibTeX arXiv:1808.04108

Code (0)

등록된 구현이 없습니다.

Tasks

Generative Adversarial NetworkImage Generation

Methods 이 논문이 사용한 방법론

Projection Discriminator A Projection Discriminator is a type of discriminator for generative adversarial networks. It is motivated by a probabilistic model in which the distribution of the…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Dogecoin Customer Service Number +1-833-534-1729 설명 없음

Similar Papers 제목 키워드 기반

Adversarial Audio Synthesis

2018-02-12 · ICLR 2019 5 · Chris Donahue, Julian McAuley, Miller Puckette

Audio signals are sampled at high temporal resolutions, and learning to synthesize audio requires capturing structure across a range of timescales. Generative adversarial networks (GANs) have seen wide success at generat…

Audio GenerationAudio SynthesisImage Generation

Adversarial Generation of Time-Frequency Features with application in audio synthesis

2019-02-11 · 36th International Conference on Machine Learning 2019 6 · Andrés Marafioti, Nicki Holighaus, Nathanaël Perraudin, Piotr Majdak

Time-frequency (TF) representations provide powerful and intuitive features for the analysis of time series such as audio. But still, generative modeling of audio in the TF domain is a subtle matter. Consequently, neural…

Audio GenerationAudio SynthesisGenerative Adversarial NetworkTime Series+1

Sound Scene Synthesis at the DCASE 2024 Challenge

2025-01-15 · Mathieu Lagrange, Junwon Lee, Modan Tailleur, Laurie M. Heller 외

This paper presents Task 7 at the DCASE 2024 Challenge: sound scene synthesis. Recent advances in sound synthesis and generative models have enabled the creation of realistic and diverse audio content. We introduce a sta…

FAD

pi-GAN: Periodic Implicit Generative Adversarial Networks for 3D-Aware Image Synthesis

2020-12-02 · CVPR 2021 1 · Eric R. Chan, Marco Monteiro, Petr Kellnhofer, Jiajun Wu 외

We have witnessed rapid progress on 3D-aware image synthesis, leveraging recent advances in generative visual models and neural rendering. Existing approaches however fall short in two ways: first, they may lack an under…

3D-Aware Image SynthesisImage GenerationNeural RenderingScene Generation

HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis

2020-10-12 · NeurIPS 2020) 2020 10 · Jungil Kong, Jaehyeon Kim, Jaekyoung Bae

Several recent work on speech synthesis have employed generative adversarial networks (GANs) to produce raw waveforms. Although such methods improve the sampling efficiency and memory usage, their sample quality has not …

CPUGPUSpeech Synthesis