paper-with-me

홈 › Papers

Effects of Speaker Count, Duration, and Accent Diversity on Zero-Shot Accent Robustness in Low-Resource ASR

2025-06-04 · Zheng-Xin Yong, Vineel Pratap, Michael Auli, Jean Maillard

To build an automatic speech recognition (ASR) system that can serve everyone in the world, the ASR needs to be robust to a wide range of accents including unseen accents. We systematically study how three different variables in training data -- the number of speakers, the audio duration per each individual speaker, and the diversity of accents -- affect ASR robustness towards unseen accents in a low-resource training regime. We observe that for a fixed number of ASR training hours, it is more beneficial to increase the number of speakers (which means each speaker contributes less) than the number of hours contributed per speaker. We also observe that more speakers enables ASR performance gains from scaling number of hours. Surprisingly, we observe minimal benefits to prioritizing speakers with different accents when the number of speakers is controlled. Our work suggests that practitioners should prioritize increasing the speaker count in ASR training data composition for new languages.

📄 PDF Abstract BibTeX arXiv:2506.04364

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Diversityspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

TokAN: Accent Normalization Using Self-Supervised Speech Tokens

2026-07-04 · Qibing Bai, Shuai Wang, Yuhan Du, Bohan Li 외 arxiv

Accent normalization (AN) seeks to convert non-native (L2) accented speech into standard (L1) speech while preserving speaker identity. The current techniques either require naturally recorded parallel L1-L2 speech for t…

Reinforcement Learning

Improving Pronunciation and Accent Conversion through Knowledge Distillation And Synthetic Ground-Truth from Native TTS

2024-10-19 · Tuan Nam Nguyen, Seymanur Aktı, Ngoc Quan Pham, Alexander Waibel

Previous approaches on accent conversion (AC) mainly aimed at making non-native speech sound more native while maintaining the original content and speaker identity. However, non-native speakers sometimes have pronunciat…

Knowledge Distillation

Total-Duration-Aware Duration Modeling for Text-to-Speech Systems

2024-06-06 · Sefik Emre Eskimez, Xiaofei Wang, Manthan Thakker, Chung-Hsien Tsai 외

Accurate control of the total duration of generated speech by adjusting the speech rate is crucial for various text-to-speech (TTS) applications. However, the impact of adjusting the speech rate on speech quality, such a…

Diversitytext-to-speechText to Speech

Learning-free L2-Accented Speech Generation using Phonological Rules

2026-03-08 · Thanathai Lertpetchpun, Yoonjeong Lee, Jihwan Lee, Tiantian Feng 외 arxiv

Accent plays a crucial role in speaker identity and inclusivity in speech technologies. Existing accented text-to-speech (TTS) systems either require large-scale accented datasets or lack fine-grained phoneme-level contr…

Accented Speech Recognition: Benchmarking, Pre-training, and Diverse Data

2022-05-16 · Alëna Aksënova, Zhehuai Chen, Chung-Cheng Chiu, Daan van Esch 외

Building inclusive speech recognition systems is a crucial step towards developing technologies that speakers of all language varieties can use. Therefore, ASR systems must work for everybody independently of the way the…

Accented Speech RecognitionBenchmarkingDiversityspeech-recognition+1