paper-with-me

홈 › Papers

Self-supervised Representation Learning for Speech Processing

2022-07-01 · NAACL (ACL) 2022 7 · Hung-Yi Lee, Abdelrahman Mohamed, Shinji Watanabe, Tara Sainath, Karen Livescu, Shang-Wen Li, Shu-wen Yang, Katrin Kirchhoff

There is a trend in the machine learning community to adopt self-supervised approaches to pre-train deep networks. Self-supervised representation learning (SSL) utilizes proxy supervised learning tasks, for example, distinguishing parts of the input signal from distractors, or generating masked input segments conditioned on the unmasked ones, to obtain training data from unlabeled corpora. BERT and GPT in NLP and SimCLR and BYOL in CV are famous examples in this direction. These approaches make it possible to use a tremendous amount of unlabeled data available on the web to train large networks and solve complicated tasks. Thus, SSL has the potential to scale up current machine learning technologies, especially for low-resourced, under-represented use cases, and democratize the technologies. Recently self-supervised approaches for speech processing are also gaining popularity. There are several workshops in relevant topics hosted at ICML 2020 (https://icml-sas.gitlab.io/), NeurIPS 2020 (https://neurips-sas-2020.github.io/), and AAAI 2022 (https://aaai-sas-2022.github.io/). However, there is no previous tutorial about a similar topic based on the authors’ best knowledge. Due to the growing popularity of SSL, and the shared mission of the areas in bringing speech and language technologies to more use cases with better quality and scaling the technologies for under-represented languages, we propose this tutorial to systematically survey the latest SSL techniques, tools, datasets, and performance achievement in speech processing. The proposed tutorial is highly relevant to the special theme of ACL about language diversity. One of the main focuses of the tutorial is leveraging SSL to reduce the dependence of speech technologies on labeled data, and to scale up the technologies especially for under-represented languages and use cases.

📄 PDF Abstract BibTeX

Code (1)

s3prl/s3prl 공식 구현 pytorch

Tasks

Representation Learning

Similar Papers 제목 키워드 기반

Speech Representation Learning Through Self-supervised Pretraining And Multi-task Finetuning

2021-10-18 · Yi-Chen Chen, Shu-wen Yang, Cheng-Kuang Lee, Simon See 외

Speech representation learning plays a vital role in speech processing. Among them, self-supervised learning (SSL) has become an important research direction. It has been shown that an SSL pretraining model can achieve e…

Multi-Task LearningRepresentation LearningSelf-Supervised LearningSpeech Representation Learning

Towards a Common Speech Analysis Engine

2022-03-01 · Hagai Aronowitz, Itai Gat, Edmilson Morais, Weizhong Zhu 외

Recent innovations in self-supervised representation learning have led to remarkable advances in natural language processing. That said, in the speech processing domain, self-supervised representation learning-based syst…

Emotion RecognitionLanguage IdentificationRepresentation Learning

Efficiency-oriented approaches for self-supervised speech representation learning

2023-12-18 · Luis Lugo, Valentin Vielzeuf

Self-supervised learning enables the training of large neural models without the need for large, labeled datasets. It has been generating breakthroughs in several fields, including computer vision, natural language proce…

Automatic Speech RecognitionRepresentation LearningSelf-Supervised LearningSpeaker Identification+3

Self-Supervised Speech Representation Learning: A Review

2022-05-21 · Abdelrahman Mohamed, Hung-Yi Lee, Lasse Borgholt, Jakob D. Havtorn 외

Although supervised deep learning has revolutionized speech and audio processing, it has necessitated the building of specialist models for individual tasks and application scenarios. It is likewise difficult to apply th…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)BenchmarkingRepresentation Learning+3

Self-supervised models of audio effectively explain human cortical responses to speech

2022-05-27 · Aditya R. Vaidya, Shailee Jain, Alexander G. Huth

Self-supervised language models are very effective at predicting high-level cortical responses during language comprehension. However, the best current models of lower-level auditory processing in the human brain rely on…

Representation LearningSpeech Representation Learning