paper-with-me

Papers

Improving Distortion Robustness of Self-supervised Speech Processing Tasks with Domain Adaptation

2022-03-30 · Kuan Po Huang, Yu-Kuan Fu, Yu Zhang, Hung-Yi Lee

Speech distortions are a long-standing problem that degrades the performance of supervisely trained speech processing models. It is high time that we enhance the robustness of speech processing models to obtain good performance when encountering speech distortions while not hurting the original performance on clean speech. In this work, we propose to improve the robustness of speech processing models by domain adversarial training (DAT). We conducted experiments based on the SUPERB framework on five different speech processing tasks. In case we do not always have knowledge of the distortion types for speech data, we analyzed the binary-domain and multi-domain settings, where the former treats all distorted speech as one domain, and the latter views different distortions as different domains. In contrast to supervised training methods, we obtained promising results in target domains where speech data is distorted with different distortions including new unseen distortions introduced during testing.

📄 PDF Abstract BibTeX arXiv:2203.16104

Code (0)

등록된 구현이 없습니다.

Tasks

Domain Adaptation

Similar Papers 제목 키워드 기반

An empirical study on speech restoration guided by self supervised speech representation

2023-05-30 · Jaeuk Byun, Youna Ji, Soo Whan Chung, Soyeon Choe 외

Enhancing speech quality is an indispensable yet difficult task as it is often complicated by a range of degradation factors. In addition to additive noise, reverberation, clipping, and speech attenuation can all adverse…

Representation LearningSpeech Representation Learning

Joint Training of Speech Enhancement and Self-supervised Model for Noise-robust ASR

2022-05-26 · Qiu-Shi Zhu, Jie Zhang, Zi-Qiang Zhang, Li-Rong Dai

Speech enhancement (SE) is usually required as a front end to improve the speech quality in noisy environments, while the enhanced speech might not be optimal for automatic speech recognition (ASR) systems due to speech …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech Enhancementspeech-recognition+1

Wav2code: Restore Clean Speech Representations via Codebook Lookup for Noise-Robust ASR

2023-04-11 · Yuchen Hu, Chen Chen, Qiushi Zhu, Eng Siong Chng

Automatic speech recognition (ASR) has gained remarkable successes thanks to recent advances of deep learning, but it usually degrades significantly under real-world noisy conditions. Recent works introduce speech enhanc…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Self-Supervised LearningSpeech Enhancement+2

Improving generalizability of distilled self-supervised speech processing models under distorted settings

2022-10-14 · Kuan-Po Huang, Yu-Kuan Fu, Tsu-Yuan Hsu, Fabian Ritter Gutierrez 외

Self-supervised learned (SSL) speech pre-trained models perform well across various speech processing tasks. Distilled versions of SSL models have been developed to match the needs of on-device speech applications. Thoug…

Knowledge Distillation

Rep2wav: Noise Robust text-to-speech Using self-supervised representations

2023-08-28 · Qiushi Zhu, Yu Gu, Rilin Chen, Chao Weng 외

Benefiting from the development of deep learning, text-to-speech (TTS) techniques using clean speech have achieved significant performance improvements. The data collected from real scenes often contains noise and genera…

Speech Enhancementtext-to-speechText to Speech