paper-with-me

홈 › Papers

UniSpeech: Unified Speech Representation Learning with Labeled and Unlabeled Data

2021-01-19 · Chengyi Wang, Yu Wu, Yao Qian, Kenichi Kumatani, Shujie Liu, Furu Wei, Michael Zeng, Xuedong Huang

In this paper, we propose a unified pre-training approach called UniSpeech to learn speech representations with both unlabeled and labeled data, in which supervised phonetic CTC learning and phonetically-aware contrastive self-supervised learning are conducted in a multi-task learning manner. The resultant representations can capture information more correlated with phonetic structures and improve the generalization across languages and domains. We evaluate the effectiveness of UniSpeech for cross-lingual representation learning on public CommonVoice corpus. The results show that UniSpeech outperforms self-supervised pretraining and supervised transfer learning for speech recognition by a maximum of 13.4% and 17.8% relative phone error rate reductions respectively (averaged over all testing languages). The transferability of UniSpeech is also demonstrated on a domain-shift speech recognition task, i.e., a relative word error rate reduction of 6% against the previous approach.

📄 PDF Abstract BibTeX arXiv:2101.07597

Code (5)

cywang97/unispeech 공식 구현 pytorch
MS-P3/code7/tree/main/unispeech_sat mindspore
MindCode-4/code-5/tree/main/unispeech mindspore
facebookresearch/data2vec_vision pytorch
microsoft/unispeech pytorch

Tasks

Multi-Task LearningRepresentation LearningSelf-Supervised Learningspeech-recognitionSpeech RecognitionSpeech Representation LearningTransfer Learning

Similar Papers 제목 키워드 기반

TranUSR: Phoneme-to-word Transcoder Based Unified Speech Representation Learning for Cross-lingual Speech Recognition

2023-05-23 · Hongfei Xue, Qijie Shao, Peikun Chen, Pengcheng Guo 외

UniSpeech has achieved superior performance in cross-lingual automatic speech recognition (ASR) by explicitly aligning latent representations to phoneme units using multi-task self-supervised learning. While the learned …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Representation LearningSelf-Supervised Learning+3

UniSpeech at scale: An Empirical Study of Pre-training Method on Large-Scale Speech Recognition Dataset

2021-07-12 · Chengyi Wang, Yu Wu, Shujie Liu, Jinyu Li 외

Recently, there has been a vast interest in self-supervised learning (SSL) where the model is pre-trained on large scale unlabeled data and then fine-tuned on a small labeled dataset. The common wisdom is that SSL helps …

Self-Supervised Learningspeech-recognitionSpeech Recognition

UniSpeech-SAT: Universal Speech Representation Learning with Speaker Aware Pre-Training

2021-10-12 · Sanyuan Chen, Yu Wu, Chengyi Wang, Zhengyang Chen 외

Self-supervised learning (SSL) is a long-standing goal for speech processing, since it utilizes large-scale unlabeled data and avoids extensive human labeling. Recent years witness great successes in applying self-superv…

Data AugmentationMulti-Task LearningRepresentation LearningSelf-Supervised Learning+4

VATLM: Visual-Audio-Text Pre-Training with Unified Masked Prediction for Speech Representation Learning

2022-11-21 · Qiushi Zhu, Long Zhou, Ziqiang Zhang, Shujie Liu 외

Although speech is a simple and effective way for humans to communicate with the outside world, a more realistic speech interaction contains multimodal information, e.g., vision, text. How to design a unified framework t…

Audio-Visual Speech RecognitionLanguage ModellingRepresentation Learningspeech-recognition+3

Enhancing Quranic Learning: A Multimodal Deep Learning Approach for Arabic Phoneme Recognition

2025-11-21 · Ayhan Kucukmanisa, Derya Gelmez, Sukru Selim Calik, Zeynep Hilal Kilimci arxiv

Recent advances in multimodal deep learning have greatly enhanced the capability of systems for speech analysis and pronunciation assessment. Accurate pronunciation detection remains a key challenge in Arabic, particular…

Multimodal Deep Learning