paper-with-me

홈 › Papers

UniSpeech at scale: An Empirical Study of Pre-training Method on Large-Scale Speech Recognition Dataset

2021-07-12 · Chengyi Wang, Yu Wu, Shujie Liu, Jinyu Li, Yao Qian, Kenichi Kumatani, Furu Wei

Recently, there has been a vast interest in self-supervised learning (SSL) where the model is pre-trained on large scale unlabeled data and then fine-tuned on a small labeled dataset. The common wisdom is that SSL helps resource-limited tasks in which only a limited amount of labeled data is available. The benefit of SSL keeps diminishing when the labeled training data amount increases. To our best knowledge, at most a few thousand hours of labeled data was used in the study of SSL. In contrast, the industry usually uses tens of thousands of hours of labeled data to build high-accuracy speech recognition (ASR) systems for resource-rich languages. In this study, we take the challenge to investigate whether and how SSL can improve the ASR accuracy of a state-of-the-art production-scale Transformer-Transducer model, which was built with 65 thousand hours of anonymized labeled EN-US data.

📄 PDF Abstract BibTeX arXiv:2107.05233

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised Learningspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

UniSpeech: Unified Speech Representation Learning with Labeled and Unlabeled Data

2021-01-19 · Chengyi Wang, Yu Wu, Yao Qian, Kenichi Kumatani 외

In this paper, we propose a unified pre-training approach called UniSpeech to learn speech representations with both unlabeled and labeled data, in which supervised phonetic CTC learning and phonetically-aware contrastiv…

Multi-Task LearningRepresentation LearningSelf-Supervised Learningspeech-recognition+3

A Comparative Study of Pre-trained Speech and Audio Embeddings for Speech Emotion Recognition

2023-04-22 · Orchid Chetia Phukan, Arun Balaji Buduru, Rajesh Sharma

Pre-trained models (PTMs) have shown great promise in the speech and audio domain. Embeddings leveraged from these models serve as inputs for learning algorithms with applications in various downstream tasks. One such cr…

Emotion RecognitionSpeaker RecognitionSpeech Emotion Recognition

TranUSR: Phoneme-to-word Transcoder Based Unified Speech Representation Learning for Cross-lingual Speech Recognition

2023-05-23 · Hongfei Xue, Qijie Shao, Peikun Chen, Pengcheng Guo 외

UniSpeech has achieved superior performance in cross-lingual automatic speech recognition (ASR) by explicitly aligning latent representations to phoneme units using multi-task self-supervised learning. While the learned …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Representation LearningSelf-Supervised Learning+3

Yet Another Model for Arabic Dialect Identification

2023-10-20 · Ajinkya Kulkarni, Hanan Aldarmaki

In this paper, we describe a spoken Arabic dialect identification (ADI) model for Arabic that consistently outperforms previously published results on two benchmark datasets: ADI-5 and ADI-17. We explore two architectura…

Dialect Identificationmodel

UniSpeech-SAT: Universal Speech Representation Learning with Speaker Aware Pre-Training

2021-10-12 · Sanyuan Chen, Yu Wu, Chengyi Wang, Zhengyang Chen 외

Self-supervised learning (SSL) is a long-standing goal for speech processing, since it utilizes large-scale unlabeled data and avoids extensive human labeling. Recent years witness great successes in applying self-superv…

Data AugmentationMulti-Task LearningRepresentation LearningSelf-Supervised Learning+4