paper-with-me

Papers

Large-Scale Self- and Semi-Supervised Learning for Speech Translation

2021-04-14 · Changhan Wang, Anne Wu, Juan Pino, Alexei Baevski, Michael Auli, Alexis Conneau

In this paper, we improve speech translation (ST) through effectively leveraging large quantities of unlabeled speech and text data in different and complementary ways. We explore both pretraining and self-training by using the large Libri-Light speech audio corpus and language modeling with CommonCrawl. Our experiments improve over the previous state of the art by 2.6 BLEU on average on all four considered CoVoST 2 language pairs via a simple recipe of combining wav2vec 2.0 pretraining, a single iteration of self-training and decoding with a language model. Different to existing work, our approach does not leverage any other supervision than ST data. Code and models will be publicly released.

📄 PDF Abstract BibTeX arXiv:2104.06678

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingTranslation

Similar Papers 제목 키워드 기반

Large scale weakly and semi-supervised learning for low-resource video ASR

2020-05-16 · Kritika Singh, Vimal Manohar, Alex Xiao, Sergey Edunov 외

Many semi- and weakly-supervised approaches have been investigated for overcoming the labeling cost of building high quality speech recognition systems. On the challenging task of transcribing social media videos in low-…

Decoderspeech-recognitionSpeech Recognition

VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

2021-01-02 · ACL 2021 5 · Changhan Wang, Morgane Rivière, Ann Lee, Anne Wu 외

We introduce VoxPopuli, a large-scale multilingual corpus providing 100K hours of unlabelled speech data in 23 languages. It is the largest open data to date for unsupervised representation learning as well as semi-super…

Representation Learningspeech-recognitionSpeech Recognition

Semi-Supervised Spoken Language Understanding via Self-Supervised Speech and Language Model Pretraining

2020-10-26 · Cheng-I Lai, Yung-Sung Chuang, Hung-Yi Lee, Shang-Wen Li 외

Much recent work on Spoken Language Understanding (SLU) is limited in at least one of three ways: models were trained on oracle text input and neglected ASR errors, models were trained to predict only intents without the…

Language ModelingLanguage ModellingSpoken Language Understanding

KinSPEAK: Improving speech recognition for Kinyarwanda via semi-supervised learning methods

2023-08-23 · Antoine Nzeyimana

Despite recent availability of large transcribed Kinyarwanda speech data, achieving robust speech recognition for Kinyarwanda is still challenging. In this work, we show that using self-supervised pre-training, following…

Robust Speech Recognitionspeech-recognitionSpeech Recognition

Speech SIMCLR: Combining Contrastive and Reconstruction Objective for Self-supervised Speech Representation Learning

2020-10-27 · Dongwei Jiang, Wubo Li, Miao Cao, Wei Zou 외

Self-supervised visual pretraining has shown significant progress recently. Among those methods, SimCLR greatly advanced the state of the art in self-supervised and semi-supervised learning on ImageNet. The input feature…

Emotion RecognitionRepresentation LearningSpeech Emotion Recognitionspeech-recognition+2