paper-with-me

Papers

Predicting non-native speech perception using the Perceptual Assimilation Model and state-of-the-art acoustic models

2022-05-31 · CoNLL (EMNLP) 2021 11 · Juliette Millet, Ioana Chitoran, Ewan Dunbar

Our native language influences the way we perceive speech sounds, affecting our ability to discriminate non-native sounds. We compare two ideas about the influence of the native language on speech perception: the Perceptual Assimilation Model, which appeals to a mental classification of sounds into native phoneme categories, versus the idea that rich, fine-grained phonetic representations tuned to the statistics of the native language, are sufficient. We operationalize this idea using representations from two state-of-the-art speech models, a Dirichlet process Gaussian mixture model and the more recent wav2vec 2.0 model. We present a new, open dataset of French- and English-speaking participants' speech perception behaviour for 61 vowel sounds from six languages. We show that phoneme assimilation is a better predictor than fine-grained phonetic modelling, both for the discrimination behaviour as a whole, and for predicting differences in discriminability associated with differences in native language background. We also show that wav2vec 2.0, while not good at capturing the effects of native language on speech perception, is complementary to information about native phoneme assimilation, and provides a good model of low-level phonetic representations, supporting the idea that both categorical and fine-grained perception are used during speech perception.

📄 PDF Abstract BibTeX arXiv:2205.15823

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Perception of Phonological Assimilation by Neural Speech Recognition Models

2024-06-21 · Charlotte Pouw, Marianne de Heer Kloots, Afra Alishahi, Willem Zuidema

Human listeners effortlessly compensate for phonological changes during speech perception, often unconsciously inferring the intended sounds. For example, listeners infer the underlying /n/ when hearing an utterance such…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Do self-supervised speech models develop human-like perception biases?

2022-05-31 · ACL 2022 5 · Juliette Millet, Ewan Dunbar

Self-supervised models for speech processing form representational spaces without using any external labels. Increasingly, they appear to be a feasible way of at least partially eliminating costly manual annotations, a p…

Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR

2025-01-24 · Hao Ma, Rujin Chen, Xiao-Lei Zhang, Ju Liu 외

Target speech extraction (TSE) isolates the speech of a specific speaker from a multi-talker overlapped speech mixture. Most existing TSE models rely on discriminative methods, typically predicting a time-frequency spect…

Speech Extraction

Evaluating computational models of infant phonetic learning across languages

2020-08-06 · Yevgen Matusevych, Thomas Schatz, Herman Kamper, Naomi H. Feldman 외

In the first year of life, infants' speech perception becomes attuned to the sounds of their native language. Many accounts of this early phonetic learning exist, but computational models predicting the attunement patter…

Modeling L1 Influence on L2 Pronunciation: An MFCC-Based Framework for Explainable Machine Learning and Pedagogical Feedback

2025-04-18 · Peyman Jahanbin

This study investigates the extent to which Mel-Frequency Cepstral Coefficients (MFCCs) capture first language (L1) transfer in extended second language (L2) English speech. Speech samples from Mandarin and American Engl…