paper-with-me

Papers

A practical framework for multi-domain speech recognition and an instance sampling method to neural language modeling

2022-03-09 · Yike Zhang, Xiaobing Feng, Yi Liu, Songjun Cao, Long Ma

Automatic speech recognition (ASR) systems used on smart phones or vehicles are usually required to process speech queries from very different domains. In such situations, a vanilla ASR system usually fails to perform well on every domain. This paper proposes a multi-domain ASR framework for Tencent Map, a navigation app used on smart phones and in-vehicle infotainment systems. The proposed framework consists of three core parts: a basic ASR module to generate n-best lists of a speech query, a text classification module to determine which domain the speech query belongs to, and a reranking module to rescore n-best lists using domain-specific language models. In addition, an instance sampling based method to training neural network language models (NNLMs) is proposed to address the data imbalance problem in multi-domain ASR. In experiments, the proposed framework was evaluated on navigation domain and music domain, since navigating and playing music are two main features of Tencent Map. Compared to a general ASR system, the proposed framework achieves a relative 13% $\sim$ 22% character error rate reduction on several test sets collected from Tencent Map and our in-car voice assistant.

📄 PDF Abstract BibTeX arXiv:2203.04767

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage ModellingRerankingspeech-recognitionSpeech Recognitiontext-classificationText Classification

Similar Papers 제목 키워드 기반

Noise-robust Speech Recognition with 10 Minutes Unparalleled In-domain Data

2022-03-29 · Chen Chen, Nana Hou, Yuchen Hu, Shashank Shirol 외

Noise-robust speech recognition systems require large amounts of training data including noisy speech data and corresponding transcripts to achieve state-of-the-art performances in face of various practical environments.…

Generative Adversarial NetworkRobust Speech Recognitionspeech-recognitionSpeech Recognition

Exploiting Cross Domain Acoustic-to-articulatory Inverted Features For Disordered Speech Recognition

2022-03-19 · Shujie Hu, Shansong Liu, Xurong Xie, Mengzhe Geng 외

Articulatory features are inherently invariant to acoustic signal distortion and have been successfully incorporated into automatic speech recognition (ASR) systems for normal speech. Their practical application to disor…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data Augmentationspeech-recognition+1

Automatic Speech Recognition for Biomedical Data in Bengali Language

2024-06-16 · Shariar Kabir, Nazmun Nahar, Shyamasree Saha, Mamunur Rashid

This paper presents the development of a prototype Automatic Speech Recognition (ASR) system specifically designed for Bengali biomedical data. Recent advancements in Bengali ASR are encouraging, but a lack of domain-spe…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

MIMO-SPEECH: End-to-End Multi-Channel Multi-Speaker Speech Recognition

2019-10-15 · Xuankai Chang, Wangyou Zhang, Yanmin Qian, Jonathan Le Roux 외

Recently, the end-to-end approach has proven its efficacy in monaural multi-speaker speech recognition. However, high word error rates (WERs) still prevent these systems from being used in practical applications. On the …

speech-recognitionSpeech RecognitionSpeech Separation

Ensembling Multilingual Pre-Trained Models for Predicting Multi-Label Regression Emotion Share from Speech

2023-09-20 · Bagus Tris Atmaja, Akira Sasou

Speech emotion recognition has evolved from research to practical applications. Previous studies of emotion recognition from speech have focused on developing models on certain datasets like IEMOCAP. The lack of data in …

Emotion RecognitionEnsemble LearningSpeech Emotion Recognition