paper-with-me

Papers

Multi-speaker Emotion Conversion via Latent Variable Regularization and a Chained Encoder-Decoder-Predictor Network

2020-07-25 · Ravi Shankar, Hsi-Wei Hsieh, Nicolas Charon, Archana Venkataraman

We propose a novel method for emotion conversion in speech based on a chained encoder-decoder-predictor neural network architecture. The encoder constructs a latent embedding of the fundamental frequency (F0) contour and the spectrum, which we regularize using the Large Diffeomorphic Metric Mapping (LDDMM) registration framework. The decoder uses this embedding to predict the modified F0 contour in a target emotional class. Finally, the predictor uses the original spectrum and the modified F0 contour to generate a corresponding target spectrum. Our joint objective function simultaneously optimizes the parameters of three model blocks. We show that our method outperforms the existing state-of-the-art approaches on both, the saliency of emotion conversion and the quality of resynthesized speech. In addition, the LDDMM regularization allows our model to convert phrases that were not present in training, thus providing evidence for out-of-sample generalization.

📄 PDF Abstract BibTeX arXiv:2007.12937

Code (0)

등록된 구현이 없습니다.

Tasks

Decoder

Similar Papers 제목 키워드 기반

TargetSEC: Plug-and-Play In-the-Wild Speech Emotion Conversion via Arousal-Conditioned Latent Style Diffusion

2026-06-05 · Constantin Alexander Auga arxiv

Speech Emotion Conversion (SEC) aims to transform the emotion of a source utterance into a target emotion while preserving content and speaker identity. SEC on in-the-wild data is challenging due to the non-parallel natu…

Provable Speech Attributes Conversion via Latent Independence

2025-10-06 · Jonathan Svirsky, Ofir Lindenbaum, Uri Shaham arxiv

While signal conversion and disentangled representation learning have shown promise for manipulating data attributes across domains such as audio, image, and multimodal generation, existing approaches, especially for spe…

Representation Learningmultimodal generation

StarGAN-VC++: Towards Emotion Preserving Voice Conversion Using Deep Embeddings

2023-09-14 · Arnab Das, Suhita Ghosh, Tim Polzehl, Sebastian Stober

Voice conversion (VC) transforms an utterance to sound like another person without changing the linguistic content. A recently proposed generative adversarial network-based VC method, StarGANv2-VC is very successful in g…

Generative Adversarial NetworkVoice Conversion

Converting Anyone's Emotion: Towards Speaker-Independent Emotional Voice Conversion

2020-05-13 · Kun Zhou, Berrak Sisman, Mingyang Zhang, Haizhou Li

Emotional voice conversion aims to convert the emotion of speech from one state to another while preserving the linguistic content and speaker identity. The prior studies on emotional voice conversion are mostly carried …

DecoderVoice Conversion

Disentanglement of Emotional Style and Speaker Identity for Expressive Voice Conversion

2021-10-20 · Zongyang Du, Berrak Sisman, Kun Zhou, Haizhou Li

Expressive voice conversion performs identity conversion for emotional speakers by jointly converting speaker identity and emotional style. Due to the hierarchical structure of speech emotion, it is challenging to disent…

DisentanglementVoice Conversion