paper-with-me

홈 › Papers

Contextual Speech Recognition with Difficult Negative Training Examples

2018-10-29 · Uri Alon, Golan Pundak, Tara N. Sainath

Improving the representation of contextual information is key to unlocking the potential of end-to-end (E2E) automatic speech recognition (ASR). In this work, we present a novel and simple approach for training an ASR context mechanism with difficult negative examples. The main idea is to focus on proper nouns (e.g., unique entities such as names of people and places) in the reference transcript, and use phonetically similar phrases as negative examples, encouraging the neural model to learn more discriminative representations. We apply our approach to an end-to-end contextual ASR model that jointly learns to transcribe and select the correct context items, and show that our proposed method gives up to $53.1\%$ relative improvement in word error rate (WER) across several benchmarks.

📄 PDF Abstract BibTeX arXiv:1810.12170

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

DyPCL: Dynamic Phoneme-level Contrastive Learning for Dysarthric Speech Recognition

2025-01-31 · Wonjun Lee, Solee Im, Heejin Do, Yunsu Kim 외

Dysarthric speech recognition often suffers from performance degradation due to the intrinsic diversity of dysarthric severity and extrinsic disparity from normal speech. To bridge these gaps, we propose a Dynamic Phonem…

Contrastive LearningDiversityspeech-recognitionSpeech Recognition

Context-Aware Transformer Transducer for Speech Recognition

2021-11-05 · Feng-Ju Chang, Jing Liu, Martin Radfar, Athanasios Mouchtaris 외

End-to-end (E2E) automatic speech recognition (ASR) systems often have difficulty recognizing uncommon words, that appear infrequently in the training data. One promising method, to improve the recognition accuracy on su…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Transferable Positive/Negative Speech Emotion Recognition via Class-wise Adversarial Domain Adaptation

2018-10-30 · Hao Zhou, Ke Chen

Speech emotion recognition plays an important role in building more intelligent and human-like agents. Due to the difficulty of collecting speech emotional data, an increasingly popular solution is leveraging a related a…

Domain AdaptationEmotion RecognitionSpeech Emotion Recognition

Contextualized End-to-End Speech Recognition with Contextual Phrase Prediction Network

2023-05-21 · Kaixun Huang, Ao Zhang, Zhanheng Yang, Pengcheng Guo 외

Contextual information plays a crucial role in speech recognition technologies and incorporating it into the end-to-end speech recognition models has drawn immense interest recently. However, previous deep bias methods l…

speech-recognitionSpeech Recognition

Approximate Nearest Neighbour Phrase Mining for Contextual Speech Recognition

2023-04-18 · Maurits Bleeker, Pawel Swietojanski, Stefan Braun, Xiaodan Zhuang

This paper presents an extension to train end-to-end Context-Aware Transformer Transducer ( CATT ) models by using a simple, yet efficient method of mining hard negative phrases from the latent space of the context encod…

speech-recognitionSpeech Recognition