paper-with-me

Papers

Self-supervised models of audio effectively explain human cortical responses to speech

2022-05-27 · Aditya R. Vaidya, Shailee Jain, Alexander G. Huth

Self-supervised language models are very effective at predicting high-level cortical responses during language comprehension. However, the best current models of lower-level auditory processing in the human brain rely on either hand-constructed acoustic filters or representations from supervised audio neural networks. In this work, we capitalize on the progress of self-supervised speech representation learning (SSL) to create new state-of-the-art models of the human auditory system. Compared against acoustic baselines, phonemic features, and supervised models, representations from the middle layers of self-supervised models (APC, wav2vec, wav2vec 2.0, and HuBERT) consistently yield the best prediction performance for fMRI recordings within the auditory cortex (AC). Brain areas involved in low-level auditory processing exhibit a preference for earlier SSL model layers, whereas higher-level semantic areas prefer later layers. We show that these trends are due to the models' ability to encode information at multiple linguistic levels (acoustic, phonetic, and lexical) along their representation depth. Overall, these results show that self-supervised models effectively capture the hierarchy of information relevant to different stages of speech processing in human cortex.

📄 PDF Abstract BibTeX arXiv:2205.14252

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningSpeech Representation Learning

Similar Papers 제목 키워드 기반

Understanding Self-Attention of Self-Supervised Audio Transformers

2020-06-05 · Shu-wen Yang, Andy T. Liu, Hung-Yi Lee

Self-supervised Audio Transformers (SAT) enable great success in many downstream speech applications like ASR, but how they work has not been widely explored yet. In this work, we present multiple strategies for the anal…

Aligning Text-to-Music Evaluation with Human Preferences

2025-03-20 · Yichen Huang, Zachary Novack, Koichi Saito, Jiatong Shi 외

Despite significant recent advances in generative acoustic text-to-music (TTM) modeling, robust evaluation of these models lags behind, relying in particular on the popular Fr\'echet Audio Distance (FAD). In this work, w…

FAD

SSAMBA: Self-Supervised Audio Representation Learning with Mamba State Space Model

2024-05-20 · Siavash Shams, Sukru Samet Dindar, Xilin Jiang, Nima Mesgarani

Transformers have revolutionized deep learning across various tasks, including audio representation learning, due to their powerful modeling capabilities. However, they often suffer from quadratic complexity in both GPU …

Audio ClassificationGPUKeyword SpottingMamba+3

Exploring bat song syllable representations in self-supervised audio encoders

2024-09-19 · Marianne de Heer Kloots, Mirjam Knörnschild

How well can deep learning models trained on human-generated sounds distinguish between another species' vocalization types? We analyze the encoding of bat song syllables in several self-supervised audio encoders, and fi…

Transfer Learning

Visually Exploring Multi-Purpose Audio Data

2021-10-09 · David Heise, Helen L. Bear

We analyse multi-purpose audio using tools to visualise similarities within the data that may be observed via unsupervised methods. The success of machine learning classifiers is affected by the information contained wit…

Acoustic Scene ClassificationClassificationScene Classification