paper-with-me

Papers

Full-band General Audio Synthesis with Score-based Diffusion

2022-10-26 · Santiago Pascual, Gautam Bhattacharya, Chunghsin Yeh, Jordi Pons, Joan Serrà

Recent works have shown the capability of deep generative models to tackle general audio synthesis from a single label, producing a variety of impulsive, tonal, and environmental sounds. Such models operate on band-limited signals and, as a result of an autoregressive approach, they are typically conformed by pre-trained latent encoders and/or several cascaded modules. In this work, we propose a diffusion-based generative model for general audio synthesis, named DAG, which deals with full-band signals end-to-end in the waveform domain. Results show the superiority of DAG over existing label-conditioned generators in terms of both quality and diversity. More specifically, when compared to the state of the art, the band-limited and full-band versions of DAG achieve relative improvements that go up to 40 and 65%, respectively. We believe DAG is flexible enough to accommodate different conditioning schemas while providing good quality synthesis.

📄 PDF Abstract BibTeX arXiv:2210.14661

Code (0)

등록된 구현이 없습니다.

Tasks

Audio SynthesisDiversity

Similar Papers 제목 키워드 기반

FlowDec: A flow-based full-band general audio codec with high perceptual quality

2025-03-03 · Simon Welker, Matthew Le, Ricky T. Q. Chen, Wei-Ning Hsu 외

We propose FlowDec, a neural full-band audio codec for general audio sampled at 48 kHz that combines non-adversarial codec training with a stochastic postfilter based on a novel conditional flow matching method. Compared…

FAD

PTQ4ADM: Post-Training Quantization for Efficient Text Conditional Audio Diffusion Models

2024-09-20 · Jayneel Vora, Aditya Krishnan, Nader Bouacida, Prabhu RV Shankar 외

Denoising diffusion models have emerged as state-of-the-art in generative tasks across image, audio, and video domains, producing high-quality, diverse, and contextually relevant data. However, their broader adoption is …

Audio GenerationAudio SynthesisDenoisingQuantization

ScoreDec: A Phase-preserving High-Fidelity Audio Codec with A Generalized Score-based Diffusion Post-filter

2024-01-22 · Yi-Chiao Wu, Dejan Marković, Steven Krenn, Israel D. Gebru 외

Although recent mainstream waveform-domain end-to-end (E2E) neural audio codecs achieve impressive coded audio quality with a very low bitrate, the quality gap between the coded and natural audio is still significant. A …

Generative Adversarial Network

MIDI-VALLE: Improving Expressive Piano Performance Synthesis Through Neural Codec Language Modelling

2025-07-11 · Jingjing Tang, Xin Wang, Zhe Zhang, Junichi Yamagishi 외

Generating expressive audio performances from music scores requires models to capture both instrument acoustics and human interpretation. Traditional music performance synthesis pipelines follow a two-stage approach, fir…

Audio SynthesisLanguage Modellingtext-to-speechText to Speech

Discriminating real and synthetic super-resolved audio samples using embedding-based classifiers

2026-01-06 · Mikhail Silaev, Konstantinos Drossos, Tuomas Virtanen arxiv

Generative adversarial networks (GANs) and diffusion models have recently achieved state-of-the-art performance in audio super-resolution (ADSR), producing perceptually convincing wideband audio from narrowband inputs. H…

Audio Super-Resolution