paper-with-me

홈 › Papers

Speaker disentanglement in video-to-speech conversion

2021-05-20 · Dan Oneata, Adriana Stan, Horia Cucu

The task of video-to-speech aims to translate silent video of lip movement to its corresponding audio signal. Previous approaches to this task are generally limited to the case of a single speaker, but a method that accounts for multiple speakers is desirable as it allows to i) leverage datasets with multiple speakers or few samples per speaker; and ii) control speaker identity at inference time. In this paper, we introduce a new video-to-speech architecture and explore ways of extending it to the multi-speaker scenario: we augment the network with an additional speaker-related input, through which we feed either a discrete identity or a speaker embedding. Interestingly, we observe that the visual encoder of the network is capable of learning the speaker identity from the lip region of the face alone. To better disentangle the two inputs -- linguistic content and speaker identity -- we add adversarial losses that dispel the identity from the video embeddings. To the best of our knowledge, the proposed method is the first to provide important functionalities such as i) control of the target voice and ii) speech synthesis for unseen identities over the state-of-the-art, while still maintaining the intelligibility of the spoken output.

📄 PDF Abstract BibTeX arXiv:2105.09652

Code (1)

danoneata/xts pytorch

Tasks

DisentanglementSpeech Synthesis

Similar Papers 제목 키워드 기반

Disentangled Speech Representation Learning for One-Shot Cross-lingual Voice Conversion Using $β$-VAE

2022-10-25 · Hui Lu, Disong Wang, Xixin Wu, Zhiyong Wu 외

We propose an unsupervised learning method to disentangle speech into content representation and speaker identity representation. We apply this method to the challenging one-shot cross-lingual voice conversion task to de…

DisentanglementRepresentation LearningSpeech Representation LearningVoice Conversion

VQMIVC: Vector Quantization and Mutual Information-Based Unsupervised Speech Representation Disentanglement for One-shot Voice Conversion

2021-06-18 · Disong Wang, Liqun Deng, Yu Ting Yeung, Xiao Chen 외

One-shot voice conversion (VC), which performs conversion across arbitrary speakers with only a single target-speaker utterance for reference, can be effectively achieved by speech representation disentanglement. Existin…

DisentanglementQuantizationVoice Conversion

Zero-Shot Accent Conversion using Pseudo Siamese Disentanglement Network

2022-12-12 · Dongya Jia, Qiao Tian, Kainan Peng, Jiaxin Li 외

The goal of accent conversion (AC) is to convert the accent of speech into the target accent while preserving the content and speaker identity. AC enables a variety of applications, such as language learning, speech cont…

Data AugmentationDisentanglement

Speaker and Style Disentanglement of Speech Based on Contrastive Predictive Coding Supported Factorized Variational Autoencoder

2024-09-05 · Yuying Xie, Michael Kuhlmann, Frederik Rautenberg, Zheng-Hua Tan 외

Speech signals encompass various information across multiple levels including content, speaker, and style. Disentanglement of these information, although challenging, is important for applications such as voice conversio…

DisentanglementVoice Conversion

Disentanglement of Emotional Style and Speaker Identity for Expressive Voice Conversion

2021-10-20 · Zongyang Du, Berrak Sisman, Kun Zhou, Haizhou Li

Expressive voice conversion performs identity conversion for emotional speakers by jointly converting speaker identity and emotional style. Due to the hierarchical structure of speech emotion, it is challenging to disent…

DisentanglementVoice Conversion