paper-with-me

홈 › Papers

Machine Speech Chain with One-shot Speaker Adaptation

2018-03-28 · Andros Tjandra, Sakriani Sakti, Satoshi Nakamura

In previous work, we developed a closed-loop speech chain model based on deep learning, in which the architecture enabled the automatic speech recognition (ASR) and text-to-speech synthesis (TTS) components to mutually improve their performance. This was accomplished by the two parts teaching each other using both labeled and unlabeled data. This approach could significantly improve model performance within a single-speaker speech dataset, but only a slight increase could be gained in multi-speaker tasks. Furthermore, the model is still unable to handle unseen speakers. In this paper, we present a new speech chain mechanism by integrating a speaker recognition model inside the loop. We also propose extending the capability of TTS to handle unseen speakers by implementing one-shot speaker adaptation. This enables TTS to mimic voice characteristics from one speaker to another with only a one-shot speaker sample, even from a text without any speaker information. In the speech chain loop mechanism, ASR also benefits from the ability to further learn an arbitrary speaker's characteristics from the generated speech waveform, resulting in a significant improvement in the recognition rate.

📄 PDF Abstract BibTeX arXiv:1803.10525

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speaker Recognitionspeech-recognitionSpeech RecognitionSpeech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Similar Papers 제목 키워드 기반

Exploring Machine Speech Chain for Domain Adaptation and Few-Shot Speaker Adaptation

2021-04-08 · Fengpeng Yue, Yan Deng, Lei He, Tom Ko

Machine Speech Chain, which integrates both end-to-end (E2E) automatic speech recognition (ASR) and text-to-speech (TTS) into one circle for joint training, has been proven to be effective in data augmentation by leverag…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationDomain Adaptation+4

USAT: A Universal Speaker-Adaptive Text-to-Speech Approach

2024-04-28 · Wenbin Wang, Yang song, Sanjay Jha

Conventional text-to-speech (TTS) research has predominantly focused on enhancing the quality of synthesized speech for speakers in the training dataset. The challenge of synthesizing lifelike speech for unseen, out-of-d…

Decodertext-to-speechText to Speech

GC-TTS: Few-shot Speaker Adaptation with Geometric Constraints

2021-08-16 · Ji-Hoon Kim, Sang-Hoon Lee, Ji-Hyun Lee, Hong-Gyu Jung 외

Few-shot speaker adaptation is a specific Text-to-Speech (TTS) system that aims to reproduce a novel speaker's voice with a few training data. While numerous attempts have been made to the few-shot speaker adaptation sys…

text-to-speechText to Speech

Meta-TTS: Meta-Learning for Few-Shot Speaker Adaptive Text-to-Speech

2021-11-07 · Sung-Feng Huang, Chyi-Jiunn Lin, Da-Rong Liu, Yi-Chen Chen 외

Personalizing a speech synthesis system is a highly desired application, where the system can generate speech with the user's voice with rare enrolled recordings. There are two main approaches to build such a system in r…

Meta-LearningSpeech Synthesistext-to-speechText to Speech

One-shot Voice Conversion For Style Transfer Based On Speaker Adaptation

2021-11-24 · Zhichao Wang, Qicong Xie, Tao Li, Hongqiang Du 외

One-shot style transfer is a challenging task, since training on one utterance makes model extremely easy to over-fit to training data and causes low speaker similarity and lack of expressiveness. In this paper, we build…

Style TransferVoice Conversion