paper-with-me

홈 › Papers

Investigation into Target Speaking Rate Adaptation for Voice Conversion

2022-09-05 · Michael Kuhlmann, Fritz Seebauer, Janek Ebbers, Petra Wagner, Reinhold Haeb-Umbach

Disentangling speaker and content attributes of a speech signal into separate latent representations followed by decoding the content with an exchanged speaker representation is a popular approach for voice conversion, which can be trained with non-parallel and unlabeled speech data. However, previous approaches perform disentanglement only implicitly via some sort of information bottleneck or normalization, where it is usually hard to find a good trade-off between voice conversion and content reconstruction. Further, previous works usually do not consider an adaptation of the speaking rate to the target speaker or they put some major restrictions to the data or use case. Therefore, the contribution of this work is two-fold. First, we employ an explicit and fully unsupervised disentanglement approach, which has previously only been used for representation learning, and show that it allows to obtain both superior voice conversion and content reconstruction. Second, we investigate simple and generic approaches to linearly adapt the length of a speech signal, and hence the speaking rate, to a target speaker and show that the proposed adaptation allows to increase the speaking rate similarity with respect to the target speaker.

📄 PDF Abstract BibTeX arXiv:2209.01978

Code (0)

등록된 구현이 없습니다.

Tasks

DisentanglementRepresentation LearningVoice Conversion

Similar Papers 제목 키워드 기반

Cultural Adaptation of Recipes

2023-10-26 · Yong Cao, Yova Kementchedjhieva, Ruixiang Cui, Antonia Karamolegkou 외

Building upon the considerable advances in Large Language Models (LLMs), we are now equipped to address more sophisticated tasks demanding a nuanced understanding of cross-cultural contexts. A key example is recipe adapt…

Information RetrievalMachine TranslationTranslation

Meta-Learning Empowered Meta-Face: Personalized Speaking Style Adaptation for Audio-Driven 3D Talking Face Animation

2024-08-18 · Xukun Zhou, Fengxin Li, Ziqiao Peng, Kejian Wu 외

Audio-driven 3D face animation is increasingly vital in live streaming and augmented reality applications. While remarkable progress has been observed, most existing approaches are designed for specific individuals with …

3D Face AnimationMeta-LearningModel Optimization

Lexical Constraints on the Acquisition of English Split Intransitivity-An Experimental Investigation of Chinese-speaking L2 Learners

2018-12-01 · PACLIC 2018 12 · Lili Wu

StyleTalk: One-shot Talking Head Generation with Controllable Speaking Styles

2023-01-03 · Yifeng Ma, Suzhen Wang, Zhipeng Hu, Changjie Fan 외

Different people speak with diverse personalized speaking styles. Although existing one-shot talking head methods have made significant progress in lip sync, natural facial expressions, and stable head motions, they stil…

DecoderFace GenerationTalking Face GenerationTalking Head Generation

Learning to mirror speaking styles incrementally

2020-03-05 · Siyi Liu, Ziang Leng, Derry Wijaya

Mirroring is the behavior in which one person subconsciously imitates the gesture, speech pattern, or attitude of another. In conversations, mirroring often signals the speakers enjoyment and engagement in their communic…