paper-with-me

홈 › Papers

Evaluating Low-Level Speech Features Against Human Perceptual Data

2017-01-01 · TACL 2017 1 · Caitlin Richter, Naomi H. Feldman, Harini Salgado, Aren Jansen

We introduce a method for measuring the correspondence between low-level speech features and human perception, using a cognitive model of speech perception implemented directly on speech recordings. We evaluate two speaker normalization techniques using this method and find that in both cases, speech features that are normalized across speakers predict human data better than unnormalized speech features, consistent with previous research. Results further reveal differences across normalization methods in how well each predicts human data. This work provides a new framework for evaluating low-level representations of speech on their match to human perception, and lays the groundwork for creating more ecologically valid models of speech perception.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech Recognition (ASR)Representation LearningSpeech Recognitionvalid

Similar Papers 제목 키워드 기반

Evaluating Automatic Speech Recognition Systems in Comparison With Human Perception Results Using Distinctive Feature Measures

2016-12-13 · Xiang Kong, Jeung-Yoon Choi, Stefanie Shattuck-Hufnagel

This paper describes methods for evaluating automatic speech recognition (ASR) systems in comparison with human perception results, using measures derived from linguistic distinctive features. Error patterns in terms of …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Mi-Go: Test Framework which uses YouTube as Data Source for Evaluating Speech Recognition Models like OpenAI's Whisper

2023-09-01 · Tomasz Wojnar, Jaroslaw Hryszko, Adam Roman

This article introduces Mi-Go, a novel testing framework aimed at evaluating the performance and adaptability of general-purpose speech recognition machine learning models across diverse real-world scenarios. The framewo…

speech-recognitionSpeech Recognition

TTSDS2: Resources and Benchmark for Evaluating Human-Quality Text to Speech Systems

2025-06-24 · Christoph Minixhofer, Ondrej Klejch, Peter Bell

Evaluation of Text to Speech (TTS) systems is challenging and resource-intensive. Subjective metrics such as Mean Opinion Score (MOS) are not easily comparable between works. Objective metrics are frequently used, but ra…

text-to-speechText to Speech

Defense Against Synthetic Speech: Real-Time Detection of RVC Voice Conversion Attacks

2025-12-31 · Prajwal Chinchmalatpure, Suyash Chinchmalatpure, Siddharth Chavan arxiv

Generative audio technologies now enable highly realistic voice cloning and real-time voice conversion, increasing the risk of impersonation, fraud, and misinformation in communication channels such as phone and video ca…

Voice Conversion

Self-supervised models of audio effectively explain human cortical responses to speech

2022-05-27 · Aditya R. Vaidya, Shailee Jain, Alexander G. Huth

Self-supervised language models are very effective at predicting high-level cortical responses during language comprehension. However, the best current models of lower-level auditory processing in the human brain rely on…

Representation LearningSpeech Representation Learning