paper-with-me

홈 › Papers

Differentiable WORLD Synthesizer-based Neural Vocoder With Application To End-To-End Audio Style Transfer

2022-08-15 · Shahan Nercessian

In this paper, we propose a differentiable WORLD synthesizer and demonstrate its use in end-to-end audio style transfer tasks such as (singing) voice conversion and the DDSP timbre transfer task. Accordingly, our baseline differentiable synthesizer has no model parameters, yet it yields adequate synthesis quality. We can extend the baseline synthesizer by appending lightweight black-box postnets which apply further processing to the baseline output in order to improve fidelity. An alternative differentiable approach considers extraction of the source excitation spectrum directly, which can improve naturalness albeit for a narrower class of style transfer applications. The acoustic feature parameterization used by our approaches has the added benefit that it naturally disentangles pitch and timbral information so that they can be modeled separately. Moreover, as there exists a robust means of estimating these acoustic features from monophonic audio sources, it allows for parameter loss terms to be added to an end-to-end objective function, which can help convergence and/or further stabilize (adversarial) training.

📄 PDF Abstract BibTeX arXiv:2208.07282

Code (0)

등록된 구현이 없습니다.

Tasks

Style TransferVoice Conversion

Methods 이 논문이 사용한 방법론

DDSP 설명 없음

Similar Papers 제목 키워드 기반

Unsupervised Harmonic Parameter Estimation Using Differentiable DSP and Spectral Optimal Transport

2023-12-22 · Bernardo Torres, Geoffroy Peeters, Gaël Richard

In neural audio signal processing, pitch conditioning has been used to enhance the performance of synthesizers. However, jointly training pitch estimators and synthesizers is a challenge when using standard audio-to-audi…

Audio Signal Processingparameter estimation

SING: Symbol-to-Instrument Neural Generator

2018-10-23 · NeurIPS 2018 12 · Alexandre Défossez, Neil Zeghidour, Nicolas Usunier, Léon Bottou 외

Recent progress in deep learning for audio synthesis opens the way to models that directly produce the waveform, shifting away from the traditional paradigm of relying on vocoders or MIDI synthesizers for speech or music…

Audio SynthesisDecoderMusic Generation

Differentiable Time-Varying Linear Prediction in the Context of End-to-End Analysis-by-Synthesis

2024-06-07 · Chin-Yun Yu, György Fazekas

Training the linear prediction (LP) operator end-to-end for audio synthesis in modern deep learning frameworks is slow due to its recursive formulation. In addition, frame-wise approximation as an acceleration method can…

Audio Synthesis

DDSP: Differentiable Digital Signal Processing

2020-01-14 · ICLR 2020 1 · Jesse Engel, Lamtharn Hantrakul, Chenjie Gu, Adam Roberts

Most generative models of audio directly generate samples in one of two domains: time or frequency. While sufficient to express any signal, these representations are inefficient, as they do not utilize existing knowledge…

Audio GenerationAudio Synthesis

DDX7: Differentiable FM Synthesis of Musical Instrument Sounds

2022-08-12 · Franco Caspe, Andrew McPherson, Mark Sandler

FM Synthesis is a well-known algorithm used to generate complex timbre from a compact set of design primitives. Typically featuring a MIDI interface, it is usually impractical to control it from an audio source. On the o…

continuous-controlContinuous ControlResynthesisSpectral Reconstruction