paper-with-me

홈 › Papers

Context-based out-of-vocabulary word recovery for ASR systems in Indian languages

2022-06-09 · Arun Baby, Saranya Vinnaitherthan, Akhil Kerhalkar, Pranav Jawale, Sharath Adavanne, Nagaraj Adiga

Detecting and recovering out-of-vocabulary (OOV) words is always challenging for Automatic Speech Recognition (ASR) systems. Many existing methods focus on modeling OOV words by modifying acoustic and language models and integrating context words cleverly into models. To train such complex models, we need a large amount of data with context words, additional training time, and increased model size. However, after getting the ASR transcription to recover context-based OOV words, the post-processing method has not been explored much. In this work, we propose a post-processing technique to improve the performance of context-based OOV recovery. We created an acoustically boosted language model with a sub-graph made at phone level with an OOV words list. We proposed two methods to determine a suitable cost function to retrieve the OOV words based on the context. The cost function is defined based on phonetic and acoustic knowledge for matching and recovering the correct context words in the decode. The effectiveness of the proposed cost function is evaluated at both word-level and sentence-level. The evaluation results show that this approach can recover an average of 50% context-based OOV words across multiple categories.

📄 PDF Abstract BibTeX arXiv:2206.04305

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage ModellingSentencespeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Enhancing Out-of-Vocabulary Performance of Indian TTS Systems for Practical Applications through Low-Effort Data Strategies

2024-07-18 · Srija Anand, Praveen Srinivasa Varadhan, Ashwin Sankar, Giri Raju 외

Publicly available TTS datasets for low-resource languages like Hindi and Tamil typically contain 10-20 hours of data, leading to poor vocabulary coverage. This limitation becomes evident in downstream applications where…

IE-CPS Lexicon: An Automatic Speech Recognition Oriented Indian-English Pronunciation Dictionary

2021-12-01 · ICON 2021 12 · Shelly Jain, Aditya Yadavalli, Ganesh Mirishkar, Chiranjeevi Yarra 외

Indian English (IE), on the surface, seems quite similar to standard English. However, closer observation shows that it has actually been influenced by the surrounding vernacular languages at several levels from phonolog…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Learning to retrieve out-of-vocabulary words in speech recognition

2015-11-17 · Imran Sheikh, Irina Illina, Dominique Fohr, Georges Linarès

Many Proper Names (PNs) are Out-Of-Vocabulary (OOV) words for speech recognition systems used to process diachronic audio data. To help recovery of the PNs missed by the system, relevant OOV PNs can be retrieved out of t…

Retrievalspeech-recognitionSpeech Recognition

Unsupervised Stemming based Language Model for Telugu Broadcast News Transcription

2019-08-10 · Mythili Sharan Pala, Parayitam Laxminarayana, A. V. Ramana

In Indian Languages , native speakers are able to understand new words formed by either combining or modifying root words with tense and / or gender. Due to data insufficiency, Automatic Speech Recognition system (ASR) m…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+2

Samāsa-Kartā: An Online Tool for Producing Compound Words using IndoWordNet

2016-01-01 · GWC 2016 1 · Hanumant Redkar, Nilesh Joshi, Sandhya Singh, Irawati Kulkarni 외

Samāsa or compounds are a regular feature of Indian Languages. They are also found in other languages like German, Italian, French, Russian, Spanish, etc. Compound word is constructed from two or more words to form a sin…

Morphological Analysis