A Cycle-GAN Approach to Model Natural Perturbations in Speech for ASR Applications
Naturally introduced perturbations in audio signal, caused by emotional and physical states of the speaker, can significantly degrade the performance of Automatic Speech Recognition (ASR) systems. In this paper, we propose a front-end based on Cycle-Consistent Generative Adversarial Network (CycleGAN) which transforms naturally perturbed speech into normal speech, and hence improves the robustness of an ASR system. The CycleGAN model is trained on non-parallel examples of perturbed and normal speech. Experiments on spontaneous laughter-speech and creaky-speech datasets show that the performance of four different ASR systems improve by using speech obtained from CycleGAN based front-end, as compared to directly using the original perturbed speech. Visualization of the features of the laughter perturbed speech and those generated by the proposed front-end further demonstrates the effectiveness of our approach.
Code (0)
등록된 구현이 없습니다.
Tasks
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Generative Adversarial Networkspeech-recognitionSpeech RecognitionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Cycle-Consistent GAN Front-End to Improve ASR Robustness to Perturbed Speech
Automatic Speech Recognition (ASR) systems, which perform well on regular speech, are found to be vulnerable to adversarial examples generated by small perturbations in the audio signal. Even naturally introduced pertur…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Generative Adversarial Networkspeech-recognition+1WaveCycleGAN2: Time-domain Neural Post-filter for Speech Waveform Generation
WaveCycleGAN has recently been proposed to bridge the gap between natural and synthesized speech waveforms in statistical parametric speech synthesis and provides fast inference with a moving average model rather than an…
Speech SynthesisWaveCycleGAN: Synthetic-to-natural speech waveform conversion using cycle-consistent adversarial networks
We propose a learning-based filter that allows us to directly modify a synthetic speech waveform into a natural speech waveform. Speech-processing systems using a vocoder framework such as statistical parametric speech s…
Speech SynthesisVoice ConversionA Little Fog for a Large Turn
Small, carefully crafted perturbations called adversarial perturbations can easily fool neural networks. However, these perturbations are largely additive and not naturally found. We turn our attention to the field of Au…
Adversarial AttackAutonomous NavigationSafety Perception RecognitionRoVo: Robust Voice Protection Against Unauthorized Speech Synthesis with Embedding-Level Perturbations
With the advancement of AI-based speech synthesis technologies such as Deep Voice, there is an increasing risk of voice spoofing attacks, including voice phishing and fake news, through unauthorized use of others' voices…
Speaker VerificationSpeech EnhancementSpeech Synthesis