paper-with-me

홈 › Papers

Similarity Analysis of Self-Supervised Speech Representations

2020-10-22 · Yu-An Chung, Yonatan Belinkov, James Glass

Self-supervised speech representation learning has recently been a prosperous research topic. Many algorithms have been proposed for learning useful representations from large-scale unlabeled data, and their applications to a wide range of speech tasks have also been investigated. However, there has been little research focusing on understanding the properties of existing approaches. In this work, we aim to provide a comparative study of some of the most representative self-supervised algorithms. Specifically, we quantify the similarities between different self-supervised representations using existing similarity measures. We also design probing tasks to study the correlation between the models' pre-training loss and the amount of specific speech information contained in their learned representations. In addition to showing how various self-supervised models behave differently given the same input, our study also finds that the training objective has a higher impact on representation similarity than architectural choices such as building blocks (RNN/Transformer/CNN) and directionality (uni/bidirectional). Our results also suggest that there exists a strong correlation between pre-training loss and downstream performance for some self-supervised algorithms.

📄 PDF Abstract BibTeX arXiv:2010.11481

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningSpeech Representation Learning

Similar Papers 제목 키워드 기반

Refining Self-Supervised Learnt Speech Representation using Brain Activations

2024-06-12 · Hengyu Li, Kangdi Mei, Zhaoci Liu, Yang Ai 외

It was shown in literature that speech representations extracted by self-supervised pre-trained models exhibit similarities with brain activations of human for speech perception and fine-tuning speech representation mode…

Automatic Speech RecognitionSpeaker Verificationspeech-recognitionSpeech Recognition

Layer-Wise Analysis of Self-Supervised Acoustic Word Embeddings: A Study on Speech Emotion Recognition

2024-02-04 · Alexandra Saliba, Yuanchao Li, Ramon Sanabria, Catherine Lai

The efficacy of self-supervised speech models has been validated, yet the optimal utilization of their representations remains challenging across diverse tasks. In this study, we delve into Acoustic Word Embeddings (AWEs…

Emotion RecognitionSpeech Emotion RecognitionWord Embeddings

Phone and speaker spatial organization in self-supervised speech representations

2023-02-24 · Pablo Riera, Manuela Cerdeiro, Leonardo Pepino, Luciana Ferrer

Self-supervised representations of speech are currently being widely used for a large number of applications. Recently, some efforts have been made in trying to analyze the type of information present in each of these re…

S3PRL-VC: Open-source Voice Conversion Framework with Self-supervised Speech Representations

2021-10-12 · Wen-Chin Huang, Shu-wen Yang, Tomoki Hayashi, Hung-Yi Lee 외

This paper introduces S3PRL-VC, an open-source voice conversion (VC) framework based on the S3PRL toolkit. In the context of recognition-synthesis VC, self-supervised speech representation (S3R) is valuable in its potent…

BenchmarkingVoice Conversion

Self-Supervised Speech Representations are More Phonetic than Semantic

2024-06-12 · Kwanghee Choi, Ankita Pasad, Tomohiko Nakamura, Satoru Fukayama 외

Self-supervised speech models (S3Ms) have become an effective backbone for speech applications. Various analyses suggest that S3Ms encode linguistic properties. In this work, we seek a more fine-grained analysis of the w…

intent-classificationIntent ClassificationSemantic SimilaritySemantic Textual Similarity