EnvGAN: Adversarial Synthesis of Environmental Sounds for Data Augmentation
The research in Environmental Sound Classification (ESC) has been progressively growing with the emergence of deep learning algorithms. However, data scarcity poses a major hurdle for any huge advance in this domain. Data augmentation offers an excellent solution to this problem. While Generative Adversarial Networks (GANs) have been successful in generating synthetic speech and sounds of musical instruments, they have hardly been applied to the generation of environmental sounds. This paper presents EnvGAN, the first ever application of GANs for the adversarial generation of environmental sounds. Our experiments on three standard ESC datasets illustrate that the EnvGAN can synthesize audio similar to the ones in the datasets. The suggested method of augmentation outshines most of the futuristic techniques for audio augmentation.
Code (0)
등록된 구현이 없습니다.
Tasks
Data AugmentationEnvironmental Sound ClassificationSound ClassificationSimilar Papers 제목 키워드 기반
A General Framework for Learning Procedural Audio Models of Environmental Sounds
This paper introduces the Procedural (audio) Variational autoEncoder (ProVE) framework as a general approach to learning Procedural Audio PA models of environmental sounds with an improvement to the realism of the synthe…
FADCAESynth: Real-Time Timbre Interpolation and Pitch Control with Conditional Autoencoders
In this paper, we present a novel audio synthesizer, CAESynth, based on a conditional autoencoder. CAESynth synthesizes timbre in real-time by interpolating the reference sounds in their shared latent feature space, whil…
Audio SynthesisMixed RealityPitch controlTimbre InterpolationDrumGAN VST: A Plugin for Drum Sound Analysis/Synthesis With Autoencoding Generative Adversarial Networks
In contemporary popular music production, drum sound design is commonly performed by cumbersome browsing and processing of pre-recorded samples in sound libraries. One can also use specialized synthesis hardware, typical…
Generative Adversarial NetworkResynthesisDetection of Adversarial Attacks and Characterization of Adversarial Subspace
Adversarial attacks have always been a serious threat for any data-driven model. In this paper, we explore subspaces of adversarial examples in unitary vector domain, and we propose a novel detector for defending our mod…
BenchmarkingEnvironmental Sound ClassificationregressionSound ClassificationMTCRNN: A multi-scale RNN for directed audio texture synthesis
Audio textures are a subset of environmental sounds, often defined as having stable statistical characteristics within an adequately large window of time but may be unstructured locally. They include common everyday soun…
Texture Synthesis