paper-with-me

홈 › Papers

Language-specific Effects on Automatic Speech Recognition Errors for World Englishes

2022-10-01 · COLING 2022 10 · June Choe, Yiran Chen, May Pik Yu Chan, Aini Li, Xin Gao, Nicole Holliday

Despite recent advancements in automated speech recognition (ASR) technologies, reports of unequal performance across speakers of different demographic groups abound. At the same time, the focus on performance metrics such as the Word Error Rate (WER) in prior studies limit the specificity and scope of recommendations that can be offered for system engineering to overcome these challenges. The current study bridges this gap by investigating the performance of Otter’s automatic captioning system on native and non-native English speakers of different language background through a linguistic analysis of segment-level errors. By examining language-specific error profiles for vowels and consonants motivated by linguistic theory, we find that certain categories of errors can be predicted from the phonological structure of a speaker’s native language.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Specificityspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

VILAS: Exploring the Effects of Vision and Language Context in Automatic Speech Recognition

2023-05-31 · Ziyi Ni, Minglun Han, Feilong Chen, Linghui Meng 외

Enhancing automatic speech recognition (ASR) performance by leveraging additional multimodal information has shown promising results in previous studies. However, most of these works have primarily focused on utilizing v…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Lombard Effect for Bilingual Speakers in Cantonese and English: importance of spectro-temporal features

2022-04-14 · Maximilian Karl Scharf, Sabine Hochmuth, Lena L. N. Wong, Birger Kollmeier 외

For a better understanding of the mechanisms underlying speech perception and the contribution of different signal features, computational models of speech recognition have a long tradition in hearing research. Due to th…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Effects of Layer Freezing on Transferring a Speech Recognition System to Under-resourced Languages

2021-02-08 · KONVENS (WS) 2021 9 · Onno Eberhard, Torsten Zesch

In this paper, we investigate the effect of layer freezing on the effectiveness of model transfer in the area of automatic speech recognition. We experiment with Mozilla's DeepSpeech architecture on German and Swiss Germ…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Are ASR foundation models generalized enough to capture features of regional dialects for low-resource languages?

2025-10-27 · Tawsif Tashwar Dipto, Azmol Hossain, Rubayet Sabbir Faruque, Md. Rezuwan Hassan 외 arxiv

Conventional research on speech recognition modeling relies on the canonical form for most low-resource languages while automatic speech recognition (ASR) for regional dialects is treated as a fine-tuning task. To invest…

Speech Recognition

Large-Scale End-to-End Multilingual Speech Recognition and Language Identification with Multi-Task Learning

2020-10-25 · Wenxin Hou, Yue Dong, Bairong Zhuang, Longfei Yang 외

In this paper, we report a large-scale end-to-end language-independent multilingual model for joint automatic speech recognition (ASR) and language identification (LID). This model adopts hybrid CTC/attention architectur…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language IdentificationMulti-Task Learning+2