paper-with-me

Papers

Injecting Text in Self-Supervised Speech Pretraining

2021-08-27 · Zhehuai Chen, Yu Zhang, Andrew Rosenberg, Bhuvana Ramabhadran, Gary Wang, Pedro Moreno

Self-supervised pretraining for Automated Speech Recognition (ASR) has shown varied degrees of success. In this paper, we propose to jointly learn representations during pretraining from two different modalities: speech and text. The proposed method, tts4pretrain complements the power of contrastive learning in self-supervision with linguistic/lexical representations derived from synthesized speech, effectively learning from untranscribed speech and unspoken text. Lexical learning in the speech encoder is enforced through an additional sequence loss term that is coupled with contrastive loss during pretraining. We demonstrate that this novel pretraining method yields Word Error Rate (WER) reductions of 10% relative on the well-benchmarked, Librispeech task over a state-of-the-art baseline pretrained with wav2vec2.0 only. The proposed method also serves as an effective strategy to compensate for the lack of transcribed speech, effectively matching the performance of 5000 hours of transcribed speech with just 100 hours of transcribed speech on the AMI meeting transcription task. Finally, we demonstrate WER reductions of up to 15% on an in-house Voice Search task over traditional pretraining. Incorporating text into encoder pretraining is complimentary to rescoring with a larger or in-domain language model, resulting in additional 6% relative reduction in WER.

📄 PDF Abstract BibTeX arXiv:2108.12226

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningLanguage Modellingspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

Semi-Supervised Spoken Language Understanding via Self-Supervised Speech and Language Model Pretraining

2020-10-26 · Cheng-I Lai, Yung-Sung Chuang, Hung-Yi Lee, Shang-Wen Li 외

Much recent work on Spoken Language Understanding (SLU) is limited in at least one of three ways: models were trained on oracle text input and neglected ASR errors, models were trained to predict only intents without the…

Language ModelingLanguage ModellingSpoken Language Understanding

Speech-XLNet: Unsupervised Acoustic Model Pretraining For Self-Attention Networks

2019-10-23 · Xingchen Song, Guangsen Wang, Zhiyong Wu, Yiheng Huang 외

Self-attention network (SAN) can benefit significantly from the bi-directional representation learning through unsupervised pretraining paradigms such as BERT and XLNet. In this paper, we present an XLNet-like pretrainin…

Representation LearningSpeech Representation Learning

Speech Representation Learning Through Self-supervised Pretraining And Multi-task Finetuning

2021-10-18 · Yi-Chen Chen, Shu-wen Yang, Cheng-Kuang Lee, Simon See 외

Speech representation learning plays a vital role in speech processing. Among them, self-supervised learning (SSL) has become an important research direction. It has been shown that an SSL pretraining model can achieve e…

Multi-Task LearningRepresentation LearningSelf-Supervised LearningSpeech Representation Learning

Large-Scale Self- and Semi-Supervised Learning for Speech Translation

2021-04-14 · Changhan Wang, Anne Wu, Juan Pino, Alexei Baevski 외

In this paper, we improve speech translation (ST) through effectively leveraging large quantities of unlabeled speech and text data in different and complementary ways. We explore both pretraining and self-training by us…

Language ModelingLanguage ModellingTranslation

SPADE: Self-supervised Pretraining for Acoustic DisEntanglement

2023-02-03 · John Harvill, Jarred Barber, Arun Nair, Ramin Pishehvar

Self-supervised representation learning approaches have grown in popularity due to the ability to train models on large amounts of unlabeled data and have demonstrated success in diverse fields such as natural language p…

DisentanglementRepresentation LearningRhythm