paper-with-me

Papers

Learning and controlling the source-filter representation of speech with a variational autoencoder

2022-04-14 · Samir Sadok, Simon Leglaive, Laurent Girin, Xavier Alameda-Pineda, Renaud Séguier

Understanding and controlling latent representations in deep generative models is a challenging yet important problem for analyzing, transforming and generating various types of data. In speech processing, inspiring from the anatomical mechanisms of phonation, the source-filter model considers that speech signals are produced from a few independent and physically meaningful continuous latent factors, among which the fundamental frequency $f_0$ and the formants are of primary importance. In this work, we start from a variational autoencoder (VAE) trained in an unsupervised manner on a large dataset of unlabeled natural speech signals, and we show that the source-filter model of speech production naturally arises as orthogonal subspaces of the VAE latent space. Using only a few seconds of labeled speech signals generated with an artificial speech synthesizer, we propose a method to identify the latent subspaces encoding $f_0$ and the first three formant frequencies, we show that these subspaces are orthogonal, and based on this orthogonality, we develop a method to accurately and independently control the source-filter speech factors within the latent subspaces. Without requiring additional information such as text or human-labeled data, this results in a deep generative model of speech spectrograms that is conditioned on $f_0$ and the formant frequencies, and which is applied to the transformation speech signals. Finally, we also propose a robust $f_0$ estimation method that exploits the projection of a speech signal onto the learned latent subspace associated with $f_0$.

📄 PDF Abstract BibTeX arXiv:2204.07075

Code (1)

samsad35/source-filter-vae 공식 구현 pytorch

Similar Papers 제목 키워드 기반

FastPitchFormant: Source-filter based Decomposed Modeling for Speech Synthesis

2021-06-29 · Taejun Bak, Jae-Sung Bae, Hanbin Bae, Young-Ik Kim 외

Methods for modeling and controlling prosody with acoustic features have been proposed for neural text-to-speech (TTS) models. Prosodic speech can be generated by conditioning acoustic features. However, synthesized spee…

Speech Synthesistext-to-speechText to Speech

Towards speech enhancement using a variational U-Net architecture

2020-12-07 · Eike J. Nustede, Jörn Anemüller

We investigate the viability of a variational U-Net architecture for denoising of single-channel audio data. Deep network speech enhancement systems commonly aim to estimate filter masks, or opt to work on the waveform s…

DenoisingSpectral ReconstructionSpeech Enhancement

Insights Into Deep Non-linear Filters for Improved Multi-channel Speech Enhancement

2022-06-27 · Kristina Tesch, Timo Gerkmann

The key advantage of using multiple microphones for speech enhancement is that spatial filtering can be used to complement the tempo-spectral processing. In a traditional setting, linear spatial filtering (beamforming) a…

Speech Enhancement

Unsupervised Domain Adaptation for Robust Speech Recognition via Variational Autoencoder-Based Data Augmentation

2017-07-19 · Wei-Ning Hsu, Yu Zhang, James Glass

Domain mismatch between training and testing can lead to significant degradation in performance in many machine learning scenarios. Unfortunately, this is not a rare situation for automatic speech recognition deployments…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationDomain Adaptation+4

ACVAE-VC: Non-parallel many-to-many voice conversion with auxiliary classifier variational autoencoder

2018-08-13 · Hirokazu Kameoka, Takuhiro Kaneko, Kou Tanaka, Nobukatsu Hojo

This paper proposes a non-parallel many-to-many voice conversion (VC) method using a variant of the conditional variational autoencoder (VAE) called an auxiliary classifier VAE (ACVAE). The proposed method has three key …

AttributeDecoderVoice Conversion