paper-with-me

Papers

Stylebook: Content-Dependent Speaking Style Modeling for Any-to-Any Voice Conversion using Only Speech Data

2023-09-06 · Hyungseob Lim, Kyungguen Byun, Sunkuk Moon, Erik Visser

While many recent any-to-any voice conversion models succeed in transferring some target speech's style information to the converted speech, they still lack the ability to faithfully reproduce the speaking style of the target speaker. In this work, we propose a novel method to extract rich style information from target utterances and to efficiently transfer it to source speech content without requiring text transcriptions or speaker labeling. Our proposed approach introduces an attention mechanism utilizing a self-supervised learning (SSL) model to collect the speaking styles of a target speaker each corresponding to the different phonetic content. The styles are represented with a set of embeddings called stylebook. In the next step, the stylebook is attended with the source speech's phonetic content to determine the final target style for each source content. Finally, content information extracted from the source speech and content-dependent target style embeddings are fed into a diffusion-based decoder to generate the converted speech mel-spectrogram. Experiment results show that our proposed method combined with a diffusion-based generative model can achieve better speaker similarity in any-to-any voice conversion tasks when compared to baseline models, while the increase in computational complexity with longer utterances is suppressed.

📄 PDF Abstract BibTeX arXiv:2309.02730

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderSelf-Supervised LearningVoice Conversion

Similar Papers 제목 키워드 기반

Style Tokens: Unsupervised Style Modeling, Control and Transfer in End-to-End Speech Synthesis

2018-03-23 · ICML 2018 7 · Yuxuan Wang, Daisy Stanton, Yu Zhang, RJ Skerry-Ryan 외

In this work, we propose "global style tokens" (GSTs), a bank of embeddings that are jointly trained within Tacotron, a state-of-the-art end-to-end speech synthesis system. The embeddings are trained with no explicit lab…

Speech SynthesisStyle TransferText-To-Speech Synthesis

Improving the quality of neural TTS using long-form content and multi-speaker multi-style modeling

2022-12-20 · Tuomo Raitio, Javier Latorre, Andrea Davis, Tuuli Morrill 외

Neural text-to-speech (TTS) can provide quality close to natural speech if an adequate amount of high-quality speech material is available for training. However, acquiring speech data for TTS training is costly and time-…

Formtext-to-speechText to Speech

Mimic: Speaking Style Disentanglement for Speech-Driven 3D Facial Animation

2023-12-18 · Hui Fu, Zeqing Wang, Ke Gong, Keze Wang 외

Speech-driven 3D facial animation aims to synthesize vivid facial animations that accurately synchronize with speech and match the unique speaking style. However, existing works primarily focus on achieving precise lip s…

DisentanglementRepresentation Learning

MSM-VC: High-fidelity Source Style Transfer for Non-Parallel Voice Conversion by Multi-scale Style Modeling

2023-09-03 · Zhichao Wang, Xinsheng Wang, Qicong Xie, Tao Li 외

In addition to conveying the linguistic content from source speech to converted speech, maintaining the speaking style of source speech also plays an important role in the voice conversion (VC) task, which is essential i…

Data AugmentationDisentanglementEmotion RecognitionSpeech Emotion Recognition+2

Advancing Large Language Models to Capture Varied Speaking Styles and Respond Properly in Spoken Conversations

2024-02-20 · Guan-Ting Lin, Cheng-Han Chiang, Hung-Yi Lee

In spoken dialogue, even if two current turns are the same sentence, their responses might still differ when they are spoken in different styles. The spoken styles, containing paralinguistic and prosodic information, mar…

Sentence