paper-with-me

홈 › Papers

Personalization of End-to-end Speech Recognition On Mobile Devices For Named Entities

2019-12-14 · Khe Chai Sim, Françoise Beaufays, Arnaud Benard, Dhruv Guliani, Andreas Kabel, Nikhil Khare, Tamar Lucassen, Petr Zadrazil, Harry Zhang, Leif Johnson, Giovanni Motta, Lillian Zhou

We study the effectiveness of several techniques to personalize end-to-end speech models and improve the recognition of proper names relevant to the user. These techniques differ in the amounts of user effort required to provide supervision, and are evaluated on how they impact speech recognition performance. We propose using keyword-dependent precision and recall metrics to measure vocabulary acquisition performance. We evaluate the algorithms on a dataset that we designed to contain names of persons that are difficult to recognize. Therefore, the baseline recall rate for proper names in this dataset is very low: 2.4%. A data synthesis approach we developed brings it to 48.6%, with no need for speech input from the user. With speech input, if the user corrects only the names, the name recall rate improves to 64.4%. If the user corrects all the recognition errors, we achieve the best recall of 73.5%. To eliminate the need to upload user data and store personalized models on a server, we focus on performing the entire personalization workflow on a mobile device.

📄 PDF Abstract BibTeX arXiv:1912.09251

Code (0)

등록된 구현이 없습니다.

Tasks

speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

UserLibri: A Dataset for ASR Personalization Using Only Text

2022-07-02 · Theresa Breiner, Swaroop Ramaswamy, Ehsan Variani, Shefali Garg 외

Personalization of speech models on mobile devices (on-device personalization) is an active area of research, but more often than not, mobile devices have more text-only data than paired audio-text data. We explore train…

Language ModelingLanguage Modellingspeech-recognitionSpeech Recognition

An Investigation Into On-device Personalization of End-to-end Automatic Speech Recognition Models

2019-09-14 · Khe Chai Sim, Petr Zadrazil, Françoise Beaufays

Speaker-independent speech recognition systems trained with data from many users are generally robust against speaker variability and work well for a large population of speakers. However, these systems do not always gen…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Mobile Keyboard Input Decoding with Finite-State Transducers

2017-04-13 · Tom Ouyang, David Rybach, Françoise Beaufays, Michael Riley

We propose a finite-state transducer (FST) representation for the models used to decode keyboard inputs on mobile devices. Drawing from learnings from the field of speech recognition, we describe a decoding framework tha…

Decoderspeech-recognitionSpeech Recognition

ValSub: Subsampling Validation Data to Mitigate Forgetting during ASR Personalization

2025-03-12 · Haaris Mehmood, Karthikeyan Saravanan, Pablo Peso Parada, David Tuckey 외

Automatic Speech Recognition (ASR) is widely used within consumer devices such as mobile phones. Recently, personalization or on-device model fine-tuning has shown that adaptation of ASR models towards target user speech…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Low-rank Gradient Approximation For Memory-Efficient On-device Training of Deep Neural Network

2020-01-24 · Mary Gooneratne, Khe Chai Sim, Petr Zadrazil, Andreas Kabel 외

Training machine learning models on mobile devices has the potential of improving both privacy and accuracy of the models. However, one of the major obstacles to achieving this goal is the memory limitation of mobile dev…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition