paper-with-me

Papers

CommonAccent: Exploring Large Acoustic Pretrained Models for Accent Classification Based on Common Voice

2023-05-29 · Juan Zuluaga-Gomez, Sara Ahmed, Danielius Visockas, Cem Subakan

Despite the recent advancements in Automatic Speech Recognition (ASR), the recognition of accented speech still remains a dominant problem. In order to create more inclusive ASR systems, research has shown that the integration of accent information, as part of a larger ASR framework, can lead to the mitigation of accented speech errors. We address multilingual accent classification through the ECAPA-TDNN and Wav2Vec 2.0/XLSR architectures which have been proven to perform well on a variety of speech-related downstream tasks. We introduce a simple-to-follow recipe aligned to the SpeechBrain toolkit for accent classification based on Common Voice 7.0 (English) and Common Voice 11.0 (Italian, German, and Spanish). Furthermore, we establish new state-of-the-art for English accent classification with as high as 95% accuracy. We also study the internal categorization of the Wav2Vev 2.0 embeddings through t-SNE, noting that there is a level of clustering based on phonological similarity. (Our recipe is open-source in the SpeechBrain toolkit, see: https://github.com/speechbrain/speechbrain/tree/develop/recipes)

📄 PDF Abstract BibTeX arXiv:2305.18283

Code (1)

speechbrain/speechbrain 공식 구현 pytorch

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Classificationspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Contrastive Regularization for Accent-Robust ASR

2026-05-05 · Van-Phat Thai, Aradhya Dhruv, Duc-Thinh Pham, Sameer Alam arxiv

ASR systems based on self-supervised acoustic pretraining and CTC fine-tuning achieve strong performance on native speech but remain sensitive to accent variability. We investigate supervised contrastive learning (SupCon…

Contrastive Learning

Accented Text-to-Speech Synthesis with Limited Data

2023-05-08 · Xuehao Zhou, Mingyang Zhang, Yi Zhou, Zhizheng Wu 외

This paper presents an accented text-to-speech (TTS) synthesis framework with limited training data. We study two aspects concerning accent rendering: phonetic (phoneme difference) and prosodic (pitch pattern and phoneme…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Joint Modeling of Accents and Acoustics for Multi-Accent Speech Recognition

2018-02-07 · Xuesong Yang, Kartik Audhkhasi, Andrew Rosenberg, Samuel Thomas 외

The performance of automatic speech recognition systems degrades with increasing mismatch between the training and testing scenarios. Differences in speaker accents are a significant source of such mismatch. The traditio…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Multi-pass Training and Cross-information Fusion for Low-resource End-to-end Accented Speech Recognition

2023-06-20 · Xuefei Wang, Yanhua Long, Yijie Li, Haoran Wei

Low-resource accented speech recognition is one of the important challenges faced by current ASR technology in practical applications. In this study, we propose a Conformer-based architecture, called Aformer, to leverage…

Accented Speech Recognitionspeech-recognitionSpeech Recognition

Accent Conversion with Articulatory Representations

2024-06-10 · Yashish M. Siriwardena, Nathan Swedlow, Audrey Howard, Evan Gitterman 외

Conversion of non-native accented speech to native (American) English has a wide range of applications such as improving intelligibility of non-native speech. Previous work on this domain has used phonetic posteriograms …

Multi-Task Learning