paper-with-me

Papers

A Multi-Discriminator CycleGAN for Unsupervised Non-Parallel Speech Domain Adaptation

2018-03-27 · Ehsan Hosseini-Asl, Yingbo Zhou, Caiming Xiong, Richard Socher

Domain adaptation plays an important role for speech recognition models, in particular, for domains that have low resources. We propose a novel generative model based on cyclic-consistent generative adversarial network (CycleGAN) for unsupervised non-parallel speech domain adaptation. The proposed model employs multiple independent discriminators on the power spectrogram, each in charge of different frequency bands. As a result we have 1) better discriminators that focus on fine-grained details of the frequency features, and 2) a generator that is capable of generating more realistic domain-adapted spectrogram. We demonstrate the effectiveness of our method on speech recognition with gender adaptation, where the model only has access to supervised data from one gender during training, but is evaluated on the other at test time. Our model is able to achieve an average of $7.41\%$ on phoneme error rate, and $11.10\%$ word error rate relative performance improvement as compared to the baseline, on TIMIT and WSJ dataset, respectively. Qualitatively, our model also generates more natural sounding speech, when conditioned on data from the other domain.

📄 PDF Abstract BibTeX arXiv:1804.00522

Code (0)

등록된 구현이 없습니다.

Tasks

Domain AdaptationGenerative Adversarial Networkspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Improving Speech Recognition on Noisy Speech via Speech Enhancement with Multi-Discriminators CycleGAN

2021-12-12 · Chia-Yu Li, Ngoc Thang Vu

This paper presents our latest investigations on improving automatic speech recognition for noisy speech via speech enhancement. We propose a novel method named Multi-discriminators CycleGAN to reduce noise of input spee…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech Enhancementspeech-recognition+1

CycleGAN-VC2: Improved CycleGAN-based Non-parallel Voice Conversion

2019-04-09 · Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, Nobukatsu Hojo

Non-parallel voice conversion (VC) is a technique for learning the mapping from source to target speech without relying on parallel data. This is an important task, but it has been challenging due to the disadvantages of…

Voice Conversion

Generative Adversarial Networks for Unpaired Voice Transformation on Impaired Speech

2018-10-30 · Li-Wei Chen, Hung-Yi Lee, Yu Tsao

This paper focuses on using voice conversion (VC) to improve the speech intelligibility of surgical patients who have had parts of their articulators removed. Due to the difficulty of data collection, VC without parallel…

Speech RecognitionVoice Conversion

Deep Feature CycleGANs: Speaker Identity Preserving Non-parallel Microphone-Telephone Domain Adaptation for Speaker Verification

2021-04-03 · Saurabh Kataria, Jesús Villalba, Piotr Żelasko, Laureano Moro-Velázquez 외

With the increase in the availability of speech from varied domains, it is imperative to use such out-of-domain data to improve existing speech systems. Domain adaptation is a prominent pre-processing approach for this. …

Domain AdaptationSpeaker VerificationTranslation

Cycle-free CycleGAN using Invertible Generator for Unsupervised Low-Dose CT Denoising

2021-04-17 · Taesung Kwon, Jong Chul Ye

Recently, CycleGAN was shown to provide high-performance, ultra-fast denoising for low-dose X-ray computed tomography (CT) without the need for a paired training dataset. Although this was possible thanks to cycle consis…

Computed Tomography (CT)DenoisingGPU