Differentiable Digital Signal Processing Mixture Model for Synthesis Parameter Extraction from Mixture of Harmonic Sounds
A differentiable digital signal processing (DDSP) autoencoder is a musical sound synthesizer that combines a deep neural network (DNN) and spectral modeling synthesis. It allows us to flexibly edit sounds by changing the fundamental frequency, timbre feature, and loudness (synthesis parameters) extracted from an input sound. However, it is designed for a monophonic harmonic sound and cannot handle mixtures of harmonic sounds. In this paper, we propose a model (DDSP mixture model) that represents a mixture as the sum of the outputs of multiple pretrained DDSP autoencoders. By fitting the output of the proposed model to the observed mixture, we can directly estimate the synthesis parameters of each source. Through synthesis parameter extraction experiments, we show that the proposed method has high and stable performance compared with a straightforward method that applies the DDSP autoencoder to the signals separated by an audio source separation method.
Code (0)
등록된 구현이 없습니다.
Tasks
Audio Source SeparationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
DDSP-SFX: Acoustically-guided sound effects generation with differentiable digital signal processing
Controlling the variations of sound effects using neural audio synthesis models has been a difficult task. Differentiable digital signal processing (DDSP) provides a lightweight solution that achieves high-quality sound …
AttributeAudio SynthesisSinging Voice Synthesis Using Differentiable LPC and Glottal-Flow-Inspired Wavetables
This paper introduces GlOttal-flow LPC Filter (GOLF), a novel method for singing voice synthesis (SVS) that exploits the physical characteristics of the human voice using differentiable digital signal processing. GOLF em…
Singing Voice SynthesisTopology in Sound Synthesis and Digital Signal Processing -- DAFx2022 Lecture Notes
Lecture notes of a tutorial on topology in sound synthesis and digital signal processing held at international conference for digital audio effects (DAFx-22) in Vienna, Austria.
DDX7: Differentiable FM Synthesis of Musical Instrument Sounds
FM Synthesis is a well-known algorithm used to generate complex timbre from a compact set of design primitives. Typically featuring a MIDI interface, it is usually impractical to control it from an audio source. On the o…
continuous-controlContinuous ControlResynthesisSpectral ReconstructionNaturalL2S: End-to-End High-quality Multispeaker Lip-to-Speech Synthesis with Differential Digital Signal Processing
Recent advancements in visual speech recognition (VSR) have promoted progress in lip-to-speech synthesis, where pre-trained VSR models enhance the intelligibility of synthesized speech by providing valuable semantic info…
Lip to Speech Synthesisspeech-recognitionSpeech RecognitionSpeech Synthesis+3