paper-with-me

홈 › Papers

TTS Skins: Speaker Conversion via ASR

2019-04-18 · Adam Polyak, Lior Wolf, Yaniv Taigman

We present a fully convolutional wav-to-wav network for converting between speakers' voices, without relying on text. Our network is based on an encoder-decoder architecture, where the encoder is pre-trained for the task of Automatic Speech Recognition, and a multi-speaker waveform decoder is trained to reconstruct the original signal in an autoregressive manner. We train the network on narrated audiobooks, and demonstrate multi-voice TTS in those voices, by converting the voice of a TTS robot.

📄 PDF Abstract BibTeX arXiv:1904.08983

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Decoderspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Many-to-Many Voice Conversion with Out-of-Dataset Speaker Support

2019-04-30 · Gokce Keskin, Tyler Lee, Cory Stephenson, Oguz H. Elibol

We present a Cycle-GAN based many-to-many voice conversion method that can convert between speakers that are not in the training set. This property is enabled through speaker embeddings generated by a neural network that…

Speaker IdentificationVoice Conversion

Identifying Source Speakers for Voice Conversion based Spoofing Attacks on Speaker Verification Systems

2022-06-18 · Danwei Cai, Zexin Cai, Ming Li

An automatic speaker verification system aims to verify the speaker identity of a speech signal. However, a voice conversion system could manipulate a person's speech signal to make it sound like another speaker's voice …

Speaker IdentificationSpeaker VerificationVoice Conversion

Converting Anyone's Emotion: Towards Speaker-Independent Emotional Voice Conversion

2020-05-13 · Kun Zhou, Berrak Sisman, Mingyang Zhang, Haizhou Li

Emotional voice conversion aims to convert the emotion of speech from one state to another while preserving the linguistic content and speaker identity. The prior studies on emotional voice conversion are mostly carried …

DecoderVoice Conversion

Building Bilingual and Code-Switched Voice Conversion with Limited Training Data Using Embedding Consistency Loss

2021-04-22 · Yaogen Yang, Haozhe Zhang, Xiaoyi Qin, Shanshan Liang 외

Building cross-lingual voice conversion (VC) systems for multiple speakers and multiple languages has been a challenging task for a long time. This paper describes a parallel non-autoregressive network to achieve bilingu…

Voice CloningVoice Conversion

Style Modeling for Multi-Speaker Articulation-to-Speech

2023-12-21 · Miseul Kim, Zhenyu Piao, Jihyun Lee, Hong-Goo Kang

In this paper, we propose a neural articulation-to-speech (ATS) framework that synthesizes high-quality speech from articulatory signal in a multi-speaker situation. Most conventional ATS approaches only focus on modelin…