paper-with-me

Papers

Do self-supervised speech models develop human-like perception biases?

2022-05-31 · ACL 2022 5 · Juliette Millet, Ewan Dunbar

Self-supervised models for speech processing form representational spaces without using any external labels. Increasingly, they appear to be a feasible way of at least partially eliminating costly manual annotations, a problem of particular concern for low-resource languages. But what kind of representational spaces do these models construct? Human perception specializes to the sounds of listeners' native languages. Does the same thing happen in self-supervised models? We examine the representational spaces of three kinds of state-of-the-art self-supervised models: wav2vec 2.0, HuBERT and contrastive predictive coding (CPC), and compare them with the perceptual spaces of French-speaking and English-speaking human listeners, both globally and taking account of the behavioural differences between the two language groups. We show that the CPC model shows a small native language effect, but that wav2vec 2.0 and HuBERT seem to develop a universal speech perception space which is not language specific. A comparison against the predictions of supervised phone recognisers suggests that all three self-supervised models capture relatively fine-grained perceptual phenomena, while supervised models are better at capturing coarser, phone-level, effects of listeners' native language, on perception.

📄 PDF Abstract BibTeX arXiv:2205.15819

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

InfoNCE 설명 없음
Contrastive Predictive Coding Contrastive Predictive Coding (CPC) learns self-supervised representations by predicting the future in latent space by using powerful autoregressive models. The model uses a…

Similar Papers 제목 키워드 기반

Text-guided HuBERT: Self-Supervised Speech Pre-training via Generative Adversarial Networks

2024-02-24 · Duo Ma, Xianghu Yue, Junyi Ao, Xiaoxue Gao 외

Human language can be expressed in either written or spoken form, i.e. text or speech. Humans can acquire knowledge from text to improve speaking and listening. However, the quest for speech pre-trained models to leverag…

Pseudo LabelSelf-Supervised Learning

The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models

2025-06-25 · Yi Wang, Oli Danyi Liu, Peter Bell

Human speech perception is multimodal. In natural speech, lip movements can precede corresponding voicing by a non-negligible gap of 100-300 ms, especially for specific consonants, affecting the time course of neural pho…

Self-Supervised Learning

Self-supervised reinforcement learning for speaker localisation with the iCub humanoid robot

2020-11-12 · Jonas Gonzalez-Billandon, Lukas Grasse, Matthew Tata, Alessandra Sciutti 외

In the future robots will interact more and more with humans and will have to communicate naturally and efficiently. Automatic speech recognition systems (ASR) will play an important role in creating natural interactions…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)reinforcement-learningReinforcement Learning (RL)+2

data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

2022-02-07 · Preprint 2022 1 · Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu 외

While the general idea of self-supervised learning is identical across modalities, the actual algorithms and objectives differ widely because they were developed with a single modality in mind. To get us closer to genera…

image-classificationImage ClassificationLinguistic AcceptabilityNatural Language Inference+4

Improved Self-Supervised Multilingual Speech Representation Learning Combined with Auxiliary Language Information

2022-12-07 · Fenglin Ding, Genshun Wan, Pengcheng Li, Jia Pan 외

Multilingual end-to-end models have shown great improvement over monolingual systems. With the development of pre-training methods on speech, self-supervised multilingual speech representation learning like XLSR has show…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Representation Learningspeech-recognition+2