paper-with-me

홈 › Papers

The Perceptimatic English Benchmark for Speech Perception Models

2020-05-07 · Juliette Millet, Ewan Dunbar

We present the Perceptimatic English Benchmark, an open experimental benchmark for evaluating quantitative models of speech perception in English. The benchmark consists of ABX stimuli along with the responses of 91 American English-speaking listeners. The stimuli test discrimination of a large number of English and French phonemic contrasts. They are extracted directly from corpora of read speech, making them appropriate for evaluating statistical acoustic models (such as those used in automatic speech recognition) trained on typical speech data sets. We show that phone discrimination is correlated with several types of models, and give recommendations for researchers seeking easily calculated norms of acoustic distance on experimental stimuli. We show that DeepSpeech, a standard English speech recognizer, is more specialized on English phoneme discrimination than English listeners, and is poorly correlated with their behaviour, even though it yields a low error on the decision task given to humans.

📄 PDF Abstract BibTeX arXiv:2005.03418

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

American 설명 없음

Similar Papers 제목 키워드 기반

Perceptimatic: A human speech perception benchmark for unsupervised subword modelling

2020-10-12 · Juliette Millet, Ewan Dunbar

In this paper, we present a data set and methods to compare speech processing models and human behaviour on a phone discrimination task. We provide Perceptimatic, an open data set which consists of French and English spe…

Mmm whatcha say? Uncovering distal and proximal context effects in first and second-language word perception using psychophysical reverse correlation

2024-06-08 · Paige Tuttösí, H. Henny Yeung, Yue Wang, Fenqi Wang 외

Acoustic context effects, where surrounding changes in pitch, rate or timbre influence the perception of a sound, are well documented in speech perception, but how they interact with language background remains unclear. …

"Listen, Understand and Translate": Triple Supervision Decouples End-to-end Speech-to-text Translation

2020-09-21 · Qianqian Dong, Rong Ye, Mingxuan Wang, Hao Zhou 외

An end-to-end speech-to-text translation (ST) takes audio in a source language and outputs the text in a target language. Existing methods are limited by the amount of parallel corpus. Can we build a system to fully util…

Speech-to-TextSpeech-to-Text TranslationTranslation

Probing for Phonology in Self-Supervised Speech Representations: A Case Study on Accent Perception

2025-06-21 · Nitin Venkateswaran, Kevin Tang, Ratree Wayland

Traditional models of accent perception underestimate the role of gradient variations in phonological features which listeners rely upon for their accent judgments. We investigate how pretrained representations from curr…

Self-Supervised Learning

An Inferential Phonological Connectionist Approach to the perception of Assimilated-English Connected Speech

2019-09-01 · NSURL 2019 9 · Hiba Zaidi