paper-with-me

Papers

Refining Self-Supervised Learnt Speech Representation using Brain Activations

2024-06-12 · Hengyu Li, Kangdi Mei, Zhaoci Liu, Yang Ai, Liping Chen, Jie Zhang, ZhenHua Ling

It was shown in literature that speech representations extracted by self-supervised pre-trained models exhibit similarities with brain activations of human for speech perception and fine-tuning speech representation models on downstream tasks can further improve the similarity. However, it still remains unclear if this similarity can be used to optimize the pre-trained speech models. In this work, we therefore propose to use the brain activations recorded by fMRI to refine the often-used wav2vec2.0 model by aligning model representations toward human neural responses. Experimental results on SUPERB reveal that this operation is beneficial for several downstream tasks, e.g., speaker verification, automatic speech recognition, intent classification.One can then consider the proposed method as a new alternative to improve self-supervised speech models.

📄 PDF Abstract BibTeX arXiv:2406.08266

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionSpeaker Verificationspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

MAESTRO: Matched Speech Text Representations through Modality Matching

2022-04-07 · Zhehuai Chen, Yu Zhang, Andrew Rosenberg, Bhuvana Ramabhadran 외

We present Maestro, a self-supervised training method to unify representations learnt from speech and text modalities. Self-supervised learning from speech signals aims to learn the latent structure inherent in the signa…

Language ModellingSelf-Supervised Learningspeech-recognitionSpeech Recognition+1

Perfect match: Improved cross-modal embeddings for audio-visual synchronisation

2018-09-21 · Soo-Whan Chung, Joon Son Chung, Hong-Goo Kang

This paper proposes a new strategy for learning powerful cross-modal embeddings for audio-to-video synchronization. Here, we set up the problem as one of cross-modal retrieval, where the objective is to find the most rel…

Binary ClassificationCross-Modal RetrievalRetrievalspeech-recognition+3

What Do Self-Supervised Speech and Speaker Models Learn? New Findings From a Cross Model Layer-Wise Analysis

2024-01-31 · Takanori Ashihara, Marc Delcroix, Takafumi Moriya, Kohei Matsuura 외

Self-supervised learning (SSL) has attracted increased attention for learning meaningful speech representations. Speech SSL models, such as WavLM, employ masked prediction training to encode general-purpose representatio…

Self-Supervised Learning

Leveraging Hidden Structure in Self-Supervised Learning

2021-06-30 · Emanuele Sansone

This work considers the problem of learning structured representations from raw images using self-supervised learning. We propose a principled framework based on a mutual information objective, which integrates self-supe…

Self-Supervised Learning

Towards objective and interpretable speech disorder assessment: a comparative analysis of CNN and transformer-based models

2024-06-07 · Malo Maisonneuve, Corinne Fredouille, Muriel Lalain, Alain Ghio 외

Head and Neck Cancers (HNC) significantly impact patients' ability to speak, affecting their quality of life. Commonly used metrics for assessing pathological speech are subjective, prompting the need for automated and u…