Latent linguistic embedding for cross-lingual text-to-speech and voice conversion
As the recently proposed voice cloning system, NAUTILUS, is capable of cloning unseen voices using untranscribed speech, we investigate the feasibility of using it to develop a unified cross-lingual TTS/VC system. Cross-lingual speech generation is the scenario in which speech utterances are generated with the voices of target speakers in a language not spoken by them originally. This type of system is not simply cloning the voice of the target speaker, but essentially creating a new voice that can be considered better than the original under a specific framing. By using a well-trained English latent linguistic embedding to create a cross-lingual TTS and VC system for several German, Finnish, and Mandarin speakers included in the Voice Conversion Challenge 2020, we show that our method not only creates cross-lingual VC with high speaker similarity but also can be seamlessly used for cross-lingual TTS without having to perform any extra steps. However, the subjective evaluations of perceived naturalness seemed to vary between target speakers, which is one aspect for future improvement.
Code (0)
등록된 구현이 없습니다.
Tasks
text-to-speechText to SpeechVoice CloningVoice ConversionSimilar Papers 제목 키워드 기반
Exploring Alignment in Shared Cross-lingual Spaces
Despite their remarkable ability to capture linguistic nuances across diverse languages, questions persist regarding the degree of alignment between languages in multilingual embeddings. Drawing inspiration from research…
Machine Translationnamed-entity-recognitionNamed Entity RecognitionSentiment Analysis+1Cross-lingual Word Embeddings in Hyperbolic Space
Cross-lingual word embeddings can be applied to several natural language processing applications across multiple languages. Unlike prior works that use word embeddings based on the Euclidean space, this short paper prese…
Cross-Lingual Word EmbeddingsWord EmbeddingsCross-lingual Word Embeddings in Hyperbolic Space
Cross-lingual word embeddings can be applied to several natural language processing applications across multiple languages. Unlike prior works that use word embeddings based on the Euclidean space, this short paper prese…
Cross-Lingual Word EmbeddingsWord EmbeddingsA Simple Geometric Method for Cross-Lingual Linguistic Transformations with Pre-trained Autoencoders
Powerful sentence encoders trained for multiple languages are on the rise. These systems are capable of embedding a wide range of linguistic properties into vector representations. While explicit probing tasks can be use…
DecoderSentenceExamining Cross-lingual Contextual Embeddings with Orthogonal Structural Probes
State-of-the-art contextual embeddings are obtained from large language models available only for a few languages. For others, we need to learn representations using a multilingual model. There is an ongoing debate on wh…