paper-with-me

홈 › Papers

Unsupervised Polyglot Text To Speech

2019-02-06 · Eliya Nachmani, Lior Wolf

We present a TTS neural network that is able to produce speech in multiple languages. The proposed network is able to transfer a voice, which was presented as a sample in a source language, into one of several target languages. Training is done without using matching or parallel data, i.e., without samples of the same speaker in multiple languages, making the method much more applicable. The conversion is based on learning a polyglot network that has multiple per-language sub-networks and adding loss terms that preserve the speaker's identity in multiple languages. We evaluate the proposed polyglot neural network for three languages with a total of more than 400 speakers and demonstrate convincing conversion capabilities.

📄 PDF Abstract BibTeX arXiv:1902.02263

Code (0)

등록된 구현이 없습니다.

Tasks

text-to-speechText to Speech

Similar Papers 제목 키워드 기반

Mix and Match: An Empirical Study on Training Corpus Composition for Polyglot Text-To-Speech (TTS)

2022-07-04 · Ziyao Zhang, Alessio Falai, Ariadna Sanchez, Orazio Angelini 외

Training multilingual Neural Text-To-Speech (NTTS) models using only monolingual corpora has emerged as a popular way for building voice cloning based Polyglot NTTS systems. In order to train these models, it is essentia…

Speech Synthesistext-to-speechText to SpeechVoice Cloning

PolyGlotFake: A Novel Multilingual and Multimodal DeepFake Dataset

2024-05-14 · Yang Hou, Haitao Fu, Chuankai Chen, Zida Li 외

With the rapid advancement of generative AI, multimodal deepfakes, which manipulate both audio and visual modalities, have drawn increasing public concern. Currently, deepfake detection has emerged as a crucial strategy …

DeepFake DetectionFace Swappingtext-to-speechText to Speech+1

Polyglot-Lion: Efficient Multilingual ASR for Singapore via Balanced Fine-Tuning of Qwen3-ASR

2026-03-17 · Quy-Anh Dang, Chris Ngo arxiv

We present Polyglot-Lion, a family of compact multilingual automatic speech recognition (ASR) models tailored for the linguistic landscape of Singapore, covering English, Mandarin, Tamil, and Malay. Our models are obtain…

Speech Recognition

Discovering Bilingual Lexicons in Polyglot Word Embeddings

2020-08-31 · Ashiqur R. KhudaBukhsh, Shriphani Palakodety, Tom M. Mitchell

Bilingual lexicons and phrase tables are critical resources for modern Machine Translation systems. Although recent results show that without any seed lexicon or parallel data, highly accurate bilingual lexicons can be l…

Machine TranslationTranslationWord Embeddings

Polyglot: Multilingual Style Preserving Speech-Driven Facial Animation

2026-04-17 · Federico Nocentini, Kwanggyoon Seo, Qingju Liu, Claudio Ferrari 외 arxiv

Speech-Driven Facial Animation (SDFA) has gained significant attention due to its applications in movies, video games, and virtual reality. However, most existing models are trained on single-language data, limiting thei…

Self-Supervised Learning