paper-with-me

홈 › Papers

CSTNet: Contrastive Speech Translation Network for Self-Supervised Speech Representation Learning

2020-06-04 · Sameer Khurana, Antoine Laurent, James Glass

More than half of the 7,000 languages in the world are in imminent danger of going extinct. Traditional methods of documenting language proceed by collecting audio data followed by manual annotation by trained linguists at different levels of granularity. This time consuming and painstaking process could benefit from machine learning. Many endangered languages do not have any orthographic form but usually have speakers that are bi-lingual and trained in a high resource language. It is relatively easy to obtain textual translations corresponding to speech. In this work, we provide a multimodal machine learning framework for speech representation learning by exploiting the correlations between the two modalities namely speech and its corresponding text translation. Here, we construct a convolutional neural network audio encoder capable of extracting linguistic representations from speech. The audio encoder is trained to perform a speech-translation retrieval task in a contrastive learning framework. By evaluating the learned representations on a phone recognition task, we demonstrate that linguistic representations emerge in the audio encoder's internal representations as a by-product of learning to perform the retrieval task.

📄 PDF Abstract BibTeX arXiv:2006.02814

Code (0)

등록된 구현이 없습니다.

Tasks

BIG-bench Machine LearningContrastive LearningRepresentation LearningRetrievalSpeech Representation LearningTranslation

Similar Papers 제목 키워드 기반

CoLLD: Contrastive Layer-to-layer Distillation for Compressing Multilingual Pre-trained Speech Encoders

2023-09-14 · Heng-Jui Chang, Ning Dong, Ruslan Mavlyutov, Sravya Popuri 외

Large-scale self-supervised pre-trained speech encoders outperform conventional approaches in speech recognition and translation tasks. Due to the high cost of developing these large models, building new encoders for new…

Contrastive LearningKnowledge DistillationModel Compressionspeech-recognition+4

Enhanced Direct Speech-to-Speech Translation Using Self-supervised Pre-training and Data Augmentation

2022-04-06 · Sravya Popuri, Peng-Jen Chen, Changhan Wang, Juan Pino 외

Direct speech-to-speech translation (S2ST) models suffer from data scarcity issues as there exists little parallel S2ST data, compared to the amount of data available for conventional cascaded systems that consist of aut…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationDecoder+9

Unified Speech-Text Pre-training for Speech Translation and Recognition

2022-04-11 · ACL 2022 5 · Yun Tang, Hongyu Gong, Ning Dong, Changhan Wang 외

We describe a method to jointly pre-train speech and text in an encoder-decoder modeling framework for speech translation and recognition. The proposed method incorporates four self-supervised and supervised subtasks for…

Decoderspeech-recognitionSpeech RecognitionTranslation

Speech SIMCLR: Combining Contrastive and Reconstruction Objective for Self-supervised Speech Representation Learning

2020-10-27 · Dongwei Jiang, Wubo Li, Miao Cao, Wei Zou 외

Self-supervised visual pretraining has shown significant progress recently. Among those methods, SimCLR greatly advanced the state of the art in self-supervised and semi-supervised learning on ImageNet. The input feature…

Emotion RecognitionRepresentation LearningSpeech Emotion Recognitionspeech-recognition+2

Unified Speech-Text Pre-training for Speech Translation and Recognition

2021-11-16 · ACL ARR November 2021 11 · Anonymous

In this work, we describe a method to jointly pre-train speech and text in an encoder-decoder modeling framework for speech translation and recognition. The proposed method utilizes multi-task learning to integrate fou…

DecoderMulti-Task Learningspeech-recognitionSpeech Recognition+1