paper-with-me

Papers

Multi-Accent Adaptation based on Gate Mechanism

2020-11-05 · Han Zhu, Li Wang, Pengyuan Zhang, Yonghong Yan

When only a limited amount of accented speech data is available, to promote multi-accent speech recognition performance, the conventional approach is accent-specific adaptation, which adapts the baseline model to multiple target accents independently. To simplify the adaptation procedure, we explore adapting the baseline model to multiple target accents simultaneously with multi-accent mixed data. Thus, we propose using accent-specific top layer with gate mechanism (AST-G) to realize multi-accent adaptation. Compared with the baseline model and accent-specific adaptation, AST-G achieves 9.8% and 1.9% average relative WER reduction respectively. However, in real-world applications, we can't obtain the accent category label for inference in advance. Therefore, we apply using an accent classifier to predict the accent label. To jointly train the acoustic model and the accent classifier, we propose the multi-task learning with gate mechanism (MTL-G). As the accent label prediction could be inaccurate, it performs worse than the accent-specific adaptation. Yet, in comparison with the baseline model, MTL-G achieves 5.1% average relative WER reduction.

📄 PDF Abstract BibTeX arXiv:2011.02774

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-Task Learningspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Remap, warp and attend: Non-parallel many-to-many accent conversion with Normalizing Flows

2022-11-10 · Abdelhamid Ezzerg, Thomas Merritt, Kayoko Yanagisawa, Piotr Bilinski 외

Regional accents of the same language affect not only how words are pronounced (i.e., phonetic content), but also impact prosodic aspects of speech such as speaking rate and intonation. This paper investigates a novel fl…

Don't Stop Self-Supervision: Accent Adaptation of Speech Representations via Residual Adapters

2023-07-02 · Anshu Bhatia, Sanchit Sinha, Saket Dingliwal, Karthik Gopalakrishnan 외

Speech representations learned in a self-supervised fashion from massive unlabeled speech corpora have been adapted successfully toward several downstream tasks. However, such representations may be skewed toward canonic…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

AccentFold: A Journey through African Accents for Zero-Shot ASR Adaptation to Target Accents

2024-02-02 · Abraham Toluwase Owodunni, Aditya Yadavalli, Chris Chinenye Emezue, Tobi Olatunji 외

Despite advancements in speech recognition, accented speech remains challenging. While previous approaches have focused on modeling techniques or creating accented speech datasets, gathering sufficient data for the multi…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Diversityspeech-recognition+1

Improving Self-supervised Pre-training using Accent-Specific Codebooks

2024-07-04 · Darshan Prabhu, Abhishek Gupta, Omkar Nitsure, Preethi Jyothi 외

Speech accents present a serious challenge to the performance of state-of-the-art end-to-end Automatic Speech Recognition (ASR) systems. Even with self-supervised learning and pre-training of ASR models, accent invarianc…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Self-Supervised Learningspeech-recognition+1

CLARITY: Contextual Linguistic Adaptation and Accent Retrieval for Dual-Bias Mitigation in Text-to-Speech Generation

2025-11-14 · Crystal Min Hui Poon, Pai Chet Ng, Xiaoxiao Miao, Immanuel Jun Kai Loh 외 arxiv

Instruction-guided text-to-speech (TTS) research has reached a maturity level where excellent speech generation quality is possible on demand, yet two coupled biases persist in reducing perceived quality: accent bias, wh…