paper-with-me

홈 › Papers

Conditional Sound Generation Using Neural Discrete Time-Frequency Representation Learning

2021-07-21 · Xubo Liu, Turab Iqbal, Jinzheng Zhao, Qiushi Huang, Mark D. Plumbley, Wenwu Wang

Deep generative models have recently achieved impressive performance in speech and music synthesis. However, compared to the generation of those domain-specific sounds, generating general sounds (such as siren, gunshots) has received less attention, despite their wide applications. In previous work, the SampleRNN method was considered for sound generation in the time domain. However, SampleRNN is potentially limited in capturing long-range dependencies within sounds as it only back-propagates through a limited number of samples. In this work, we propose a method for generating sounds via neural discrete time-frequency representation learning, conditioned on sound classes. This offers an advantage in efficiently modelling long-range dependencies and retaining local fine-grained structures within sound clips. We evaluate our approach on the UrbanSound8K dataset, compared to SampleRNN, with the performance metrics measuring the quality and diversity of generated sounds. Experimental results show that our method offers comparable performance in quality and significantly better performance in diversity.

📄 PDF Abstract BibTeX arXiv:2107.09998

Code (1)

liuxubo717/sound_generation 공식 구현 pytorch

Tasks

DiversityMusic GenerationRepresentation LearningSpeech Synthesis

Similar Papers 제목 키워드 기반

Masked Conditional Neural Networks for Automatic Sound Events Recognition

2018-02-15 · Fady Medhat, David Chesmore, John Robinson

Deep neural network architectures designed for application domains other than sound, especially image recognition, may not optimally harness the time-frequency representation when adapted to the sound recognition problem…

Masked Conditional Neural Networks for Environmental Sound Classification

2018-05-25 · Fady Medhat, David Chesmore, John Robinson

The ConditionaL Neural Network (CLNN) exploits the nature of the temporal sequencing of the sound signal represented in a spectrogram, and its variant the Masked ConditionaL Neural Network (MCLNN) induces the network to …

ClassificationEnvironmental Sound ClassificationGeneral ClassificationSound Classification

Parallel Complex Diffusion for Scalable Time Series Generation

2026-02-10 · Rongyao Cai, Yuxi Wan, Kexin Zhang, Ming Jin 외 arxiv

Diffusion models learn data distributions indirectly through denoising, making the difficulty of generative modeling closely tied to the dependency structure of data. For time series, strong temporal dependence forces th…

Computational Efficiency

Fast frequency discrimination and phoneme recognition using a biomimetic membrane coupled to a neural network

2020-04-09 · Woo Seok Lee, Hyunjae Kim, Andrew N. Cleland, Kang-Hun Ahn

In the human ear, the basilar membrane plays a central role in sound recognition. When excited by sound, this membrane responds with a frequency-dependent displacement pattern that is detected and identified by the audit…

Phoneme Recognition

Recognition of Acoustic Events Using Masked Conditional Neural Networks

2018-02-07 · Fady Medhat, David Chesmore, John Robinson

Automatic feature extraction using neural networks has accomplished remarkable success for images, but for sound recognition, these models are usually modified to fit the nature of the multi-dimensional temporal represen…