paper-with-me

홈 › Papers

Applying Wav2vec2.0 to Speech Recognition in Various Low-resource Languages

2020-12-22 · Cheng Yi, Jianzhong Wang, Ning Cheng, Shiyu Zhou, Bo Xu

There are several domains that own corresponding widely used feature extractors, such as ResNet, BERT, and GPT-x. These models are usually pre-trained on large amounts of unlabeled data by self-supervision and can be effectively applied to downstream tasks. In the speech domain, wav2vec2.0 starts to show its powerful representation ability and feasibility of ultra-low resource speech recognition on the Librispeech corpus, which belongs to the audiobook domain. However, wav2vec2.0 has not been examined on real spoken scenarios and languages other than English. To verify its universality over languages, we apply pre-trained models to solve low-resource speech recognition tasks in various spoken languages. We achieve more than 20% relative improvements in six languages compared with previous work. Among these languages, English achieves a gain of 52.4%. Moreover, using coarse-grained modeling units, such as subword or character, achieves better results than fine-grained modeling units, such as phone or letter.

📄 PDF Abstract BibTeX arXiv:2012.12121

Code (0)

등록된 구현이 없습니다.

Tasks

speech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Average Pooling 설명 없음
Attention 설명 없음
1x1 Convolution A 1 x 1 Convolution is a convolution with some special properties in that it can be used for dimensionality reduction,…
Linear Warmup With Linear Decay Linear Warmup With Linear Decay is a learning rate schedule in which we increase the learning rate linearly for $n$ updates and then linearly decay afterwards.
Batch Normalization 설명 없음
Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…

Similar Papers 제목 키워드 기반

Weighted Cross-entropy for Low-Resource Languages in Multilingual Speech Recognition

2024-09-25 · Andrés Piñeiro-Martín, Carmen García-Mateo, Laura Docío-Fernández, María del Carmen López-Pérez 외

This paper addresses the challenge of integrating low-resource languages into multilingual automatic speech recognition (ASR) systems. We introduce a novel application of weighted cross-entropy, typically used for unbala…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data Augmentationspeech-recognition+1

Meta Learning for End-to-End Low-Resource Speech Recognition

2019-10-26 · Jui-Yang Hsu, Yuan-Jui Chen, Hung-Yi Lee

In this paper, we proposed to apply meta learning approach for low-resource automatic speech recognition (ASR). We formulated ASR for different languages as different tasks, and meta-learned the initialization parameters…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Meta-Learningspeech-recognition+1

An Investigation of Hybrid architectures for Low Resource Multilingual Speech Recognition system in Indian context

2021-12-01 · ICON 2021 12 · Ganesh Mirishkar, Aditya Yadavalli, Anil Kumar Vuppala

India is a land of language diversity. There are approximately 2000 languages spoken around, and among which officially registered are 23. In those, there are very few with Automatic Speech Recognition (ASR) capability. …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DiversityLanguage Modeling+3

Cloud-based Automatic Speech Recognition Systems for Southeast Asian Languages

2022-10-07 · Lei Wang, Rong Tong, Cheung Chi Leung, Sunil Sivadas 외

This paper provides an overall introduction of our Automatic Speech Recognition (ASR) systems for Southeast Asian languages. As not much existing work has been carried out on such regional languages, a few difficulties s…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Findings of the 2023 ML-SUPERB Challenge: Pre-Training and Evaluation over More Languages and Beyond

2023-10-09 · Jiatong Shi, William Chen, Dan Berrebbi, Hsiu-Hsuan Wang 외

The 2023 Multilingual Speech Universal Performance Benchmark (ML-SUPERB) Challenge expands upon the acclaimed SUPERB framework, emphasizing self-supervised models in multilingual speech recognition and language identific…

Language Identificationspeech-recognitionSpeech Recognition