paper-with-me

Papers

Error-driven Fixed-Budget ASR Personalization for Accented Speakers

2021-03-04 · Abhijeet Awasthi, Aman Kansal, Sunita Sarawagi, Preethi Jyothi

We consider the task of personalizing ASR models while being constrained by a fixed budget on recording speaker-specific utterances. Given a speaker and an ASR model, we propose a method of identifying sentences for which the speaker's utterances are likely to be harder for the given ASR model to recognize. We assume a tiny amount of speaker-specific data to learn phoneme-level error models which help us select such sentences. We show that speaker's utterances on the sentences selected using our error model indeed have larger error rates when compared to speaker's utterances on randomly selected sentences. We find that fine-tuning the ASR model on the sentence utterances selected with the help of error models yield higher WER improvements in comparison to fine-tuning on an equal number of randomly selected sentence utterances. Thus, our method provides an efficient way of collecting speaker utterances under budget constraints for personalizing ASR models.

📄 PDF Abstract BibTeX arXiv:2103.03142

Code (1)

awasthiabhijeet/Error-Driven-ASR-Personalization 공식 구현 pytorch

Tasks

Sentence

Similar Papers 제목 키워드 기반

Residual Adapters for Parameter-Efficient ASR Adaptation to Atypical and Accented Speech

2021-09-14 · EMNLP 2021 11 · Katrin Tomanek, Vicky Zayats, Dirk Padfield, Kara Vaillancourt 외

Automatic Speech Recognition (ASR) systems are often optimized to work best for speakers with canonical speech patterns. Unfortunately, these systems perform poorly when tested on atypical speech and heavily accented spe…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Improving Accented Speech Recognition using Data Augmentation based on Unsupervised Text-to-Speech Synthesis

2024-07-04 · Cong-Thanh Do, Shuhei Imai, Rama Doddipatla, Thomas Hain

This paper investigates the use of unsupervised text-to-speech synthesis (TTS) as a data augmentation method to improve accented speech recognition. TTS systems are trained with a small amount of accented speech training…

Accented Speech RecognitionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Data Augmentation+7

Multi-pass Training and Cross-information Fusion for Low-resource End-to-end Accented Speech Recognition

2023-06-20 · Xuefei Wang, Yanhua Long, Yijie Li, Haoran Wei

Low-resource accented speech recognition is one of the important challenges faced by current ASR technology in practical applications. In this study, we propose a Conformer-based architecture, called Aformer, to leverage…

Accented Speech Recognitionspeech-recognitionSpeech Recognition

Domain Adversarial Training for Accented Speech Recognition

2018-06-07 · Sining Sun, Ching-Feng Yeh, Mei-Yuh Hwang, Mari Ostendorf 외

In this paper, we propose a domain adversarial training (DAT) algorithm to alleviate the accented speech recognition problem. In order to reduce the mismatch between labeled source domain data ("standard" accent) and unl…

Accented Speech RecognitionMulti-Task Learningspeech-recognitionSpeech Recognition

AccentFold: A Journey through African Accents for Zero-Shot ASR Adaptation to Target Accents

2024-02-02 · Abraham Toluwase Owodunni, Aditya Yadavalli, Chris Chinenye Emezue, Tobi Olatunji 외

Despite advancements in speech recognition, accented speech remains challenging. While previous approaches have focused on modeling techniques or creating accented speech datasets, gathering sufficient data for the multi…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Diversityspeech-recognition+1