paper-with-me

홈 › Papers

Converting Anyone's Voice: End-to-End Expressive Voice Conversion with a Conditional Diffusion Model

2024-05-02 · Zongyang Du, Junchen Lu, Kun Zhou, Lakshmish Kaushik, Berrak Sisman

Expressive voice conversion (VC) conducts speaker identity conversion for emotional speakers by jointly converting speaker identity and emotional style. Emotional style modeling for arbitrary speakers in expressive VC has not been extensively explored. Previous approaches have relied on vocoders for speech reconstruction, which makes speech quality heavily dependent on the performance of vocoders. A major challenge of expressive VC lies in emotion prosody modeling. To address these challenges, this paper proposes a fully end-to-end expressive VC framework based on a conditional denoising diffusion probabilistic model (DDPM). We utilize speech units derived from self-supervised speech models as content conditioning, along with deep features extracted from speech emotion recognition and speaker verification systems to model emotional style and speaker identity. Objective and subjective evaluations show the effectiveness of our framework. Codes and samples are publicly available.

📄 PDF Abstract BibTeX arXiv:2405.01730

Code (0)

등록된 구현이 없습니다.

Tasks

DenoisingEmotion RecognitionSpeaker VerificationSpeech Emotion RecognitionVoice Conversion

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Converting Anyone's Emotion: Towards Speaker-Independent Emotional Voice Conversion

2020-05-13 · Kun Zhou, Berrak Sisman, Mingyang Zhang, Haizhou Li

Emotional voice conversion aims to convert the emotion of speech from one state to another while preserving the linguistic content and speaker identity. The prior studies on emotional voice conversion are mostly carried …

DecoderVoice Conversion

Disentanglement of Emotional Style and Speaker Identity for Expressive Voice Conversion

2021-10-20 · Zongyang Du, Berrak Sisman, Kun Zhou, Haizhou Li

Expressive voice conversion performs identity conversion for emotional speakers by jointly converting speaker identity and emotional style. Due to the hierarchical structure of speech emotion, it is challenging to disent…

DisentanglementVoice Conversion

Voice Conversion for Stuttered Speech, Instruments, Unseen Languages and Textually Described Voices

2023-10-12 · Matthew Baas, Herman Kamper

Voice conversion aims to convert source speech into a target voice using recordings of the target speaker as a reference. Newer models are producing increasingly realistic output. But what happens when models are fed wit…

Voice Conversion

Using joint training speaker encoder with consistency loss to achieve cross-lingual voice conversion and expressive voice conversion

2023-07-01 · Houjian Guo, Chaoran Liu, Carlos Toshinori Ishi, Hiroshi Ishiguro

Voice conversion systems have made significant advancements in terms of naturalness and similarity in common voice conversion tasks. However, their performance in more complex tasks such as cross-lingual voice conversion…

speech-recognitionSpeech RecognitionVoice Conversion

TTS Skins: Speaker Conversion via ASR

2019-04-18 · Adam Polyak, Lior Wolf, Yaniv Taigman

We present a fully convolutional wav-to-wav network for converting between speakers' voices, without relying on text. Our network is based on an encoder-decoder architecture, where the encoder is pre-trained for the task…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Decoderspeech-recognition+1