paper-with-me

Papers

Multilingual Approach to Joint Speech and Accent Recognition with DNN-HMM Framework

2020-10-22 · Yizhou Peng, Jicheng Zhang, Haobo Zhang, HaiHua Xu, Hao Huang, Eng Siong Chng

Human can recognize speech, as well as the peculiar accent of the speech simultaneously. However, present state-of-the-art ASR system can rarely do that. In this paper, we propose a multilingual approach to recognizing English speech, and related accent that speaker conveys using DNN-HMM framework. Specifically, we assume different accents of English as different languages. We then merge them together and train a multilingual ASR system. During decoding, we conduct two experiments. One is a monolingual ASR-based decoding, with the accent information embedded at phone level, realizing word-based accent recognition (AR), and the other is a multilingual ASR-based decoding, realizing an approximated utterance-based AR. Experimental results on an 8-accent English speech recognition show both methods can yield WERs close to the conventional ASR systems that completely ignore the accent, as well as desired AR accuracy. Besides, we conduct extensive analysis for the proposed method, such as transfer learning without-domain data exploitation, cross-accent recognition confusion, as well as characteristics of accented-word.

📄 PDF Abstract BibTeX arXiv:2010.11483

Code (0)

등록된 구현이 없습니다.

Tasks

speech-recognitionSpeech RecognitionTransfer Learning

Similar Papers 제목 키워드 기반

Dyn-ASR: Compact, Multilingual Speech Recognition via Spoken Language and Accent Identification

2021-08-04 · Sangeeta Ghangam, Daniel Whitenack, Joshua Nemecek

Running automatic speech recognition (ASR) on edge devices is non-trivial due to resource constraints, especially in scenarios that require supporting multiple languages. We propose a new approach to enable multilingual …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language Identificationspeech-recognition+1

Accent-Invariant Automatic Speech Recognition via Saliency-Driven Spectrogram Masking

2025-10-10 · Mohammad Hossein Sameti, Sepehr Harfi Moridani, Ali Zarean, Hossein Sameti arxiv

Pre-trained transformer-based models have significantly advanced automatic speech recognition (ASR), yet they remain sensitive to accent and dialectal variations, resulting in elevated word error rates (WER) in linguisti…

Speech RecognitionData Augmentation

CommonAccent: Exploring Large Acoustic Pretrained Models for Accent Classification Based on Common Voice

2023-05-29 · Juan Zuluaga-Gomez, Sara Ahmed, Danielius Visockas, Cem Subakan

Despite the recent advancements in Automatic Speech Recognition (ASR), the recognition of accented speech still remains a dominant problem. In order to create more inclusive ASR systems, research has shown that the integ…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Classificationspeech-recognition+1

Pitch Accent Detection improves Pretrained Automatic Speech Recognition

2025-08-06 · David Sasu, Natalie Schluter arxiv

We show the performance of Automatic Speech Recognition (ASR) systems that use semi-supervised speech representations can be boosted by a complimentary pitch accent detection module, by introducing a joint ASR and pitch …

Speech Recognition

Supervised Contrastive Learning for Accented Speech Recognition

2021-07-02 · Tao Han, Hantao Huang, Ziang Yang, Wei Han

Neural network based speech recognition systems suffer from performance degradation due to accented speech, especially unfamiliar accents. In this paper, we study the supervised contrastive learning framework for accente…

Accented Speech RecognitionContrastive LearningData AugmentationSentence+2