paper-with-me

홈 › Papers

Homophone-based Label Smoothing in End-to-End Automatic Speech Recognition

2020-04-07 · Yi Zheng, Xianjie Yang, Xuyong Dang

A new label smoothing method that makes use of prior knowledge of a language at human level, homophone, is proposed in this paper for automatic speech recognition (ASR). Compared with its forerunners, the proposed method uses pronunciation knowledge of homophones in a more complex way. End-to-end ASR models that learn acoustic model and language model jointly and modelling units of characters are necessary conditions for this method. Experiments with hybrid CTC sequence-to-sequence model show that the new method can reduce character error rate (CER) by 0.4% absolutely.

📄 PDF Abstract BibTeX arXiv:2004.03437

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modellingspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…

Similar Papers 제목 키워드 기반

Improving Rare Words Recognition through Homophone Extension and Unified Writing for Low-resource Cantonese Speech Recognition

2023-02-02 · Holam Chung, Junan Li, Pengfei Liu1, Wai-Kim Leung 외

Homophone characters are common in tonal syllable-based languages, such as Mandarin and Cantonese. The data-intensive end-to-end Automatic Speech Recognition (ASR) systems are more likely to mis-recognize homophone chara…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+2

PAC: Pronunciation-Aware Contextualized Large Language Model-based Automatic Speech Recognition

2025-09-16 · Li Fu, Yu Xin, Sunlu Zeng, Lu Fan 외 arxiv

This paper presents a Pronunciation-Aware Contextualized (PAC) framework to address two key challenges in Large Language Model (LLM)-based Automatic Speech Recognition (ASR) systems: effective pronunciation modeling and …

Reinforcement LearningSpeech Recognition

Cross-lingual studies of ASR errors: paradigms for perceptual evaluations

2012-05-01 · LREC 2012 5 · Ioana Vasilescu, Martine Adda-Decker, Lori Lamel

It is well-known that human listeners significantly outperform machines when it comes to transcribing speech. This paper presents a progress report of the joint research in the automatic vs human speech transcription and…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Information Retrievalspeech-recognition+1

Pronunciation-aware unique character encoding for RNN Transducer-based Mandarin speech recognition

2022-07-29 · Peng Shen, Xugang Lu, Hisashi Kawai

For Mandarin end-to-end (E2E) automatic speech recognition (ASR) tasks, compared to character-based modeling units, pronunciation-based modeling units could improve the sharing of modeling units in model training but mee…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Evaluation of Automated Speech Recognition Systems for Conversational Speech: A Linguistic Perspective

2022-11-05 · Hannaneh B. Pasandi, Haniyeh B. Pasandi

Automatic speech recognition (ASR) meets more informal and free-form input data as voice user interfaces and conversational agents such as the voice assistants such as Alexa, Google Home, etc., gain popularity. Conversat…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition