paper-with-me

홈 › Papers

Investigating Human-Model Discrepancies in Speech Quality Assessment via Acoustic and Prosodic Perturbations

2026-06-18 · Masato Takagi, Masaya Kawamura, Reo Shimizu, Yuma Shirahata arxiv

Mean opinion score (MOS) prediction models are widely used as proxy metrics in text-to-speech (TTS) research, yet their ability to capture quality differences beyond acoustic fidelity remains unclear. We investigate this via controlled perturbations on speech: acoustic degradation, prosodic errors, and manipulation of speaker-specific characteristics such as pitch and speaking rate. We obtained MOS predictions for these speech samples from both human listeners and the model, and analyzed the differences in their perceptual characteristics. Results show that most models track acoustic degradation well, while all are insensitive to prosodic errors despite large subjective score drops. For speaker characteristics, models exhibit a double dissociation: strong mean fundamental frequency (F0) biases absent in human ratings, yet insensitivity to speaking rate and F0 variability that humans notice. These findings highlight limitations of scalar MOS prediction beyond acoustic fidelity.

📄 PDF Abstract BibTeX arXiv:2606.19951

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

WER is Unaware: Assessing How ASR Errors Distort Clinical Understanding in Patient Facing Dialogue

2025-11-20 · Zachary Ellis, Jared Joselowitz, Yash Deo, Yajie He 외 arxiv

As Automatic Speech Recognition (ASR) is increasingly deployed in clinical dialogue, standard evaluations still rely heavily on Word Error Rate (WER). This paper challenges that standard, investigating whether WER or oth…

Speech Recognition

EvalTalker: Learning to Evaluate Real-Portrait-Driven Multi-Subject Talking Humans

2025-12-01 · Yingjie Zhou, Xilei Zhu, Siyu Ren, Ziyi Zhao 외 arxiv

Speech-driven Talking Human (TH) generation, commonly known as "Talker," currently faces limitations in multi-subject driving capabilities. Extending this paradigm to "Multi-Talker," capable of animating multiple subject…

AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation

2025-07-17 · Potsawee Manakul, Woody Haosheng Gan, Michael J. Ryan, Ali Sartaz Khan 외 arxiv

Current speech evaluation suffers from two critical limitations: the need and difficulty of designing specialized systems targeting individual audio characteristics, and poor correlation between automatic evaluation meth…

Speaker IdentificationPrompt Engineering

Deep MOS Predictor for Synthetic Speech Using Cluster-Based Modeling

2020-08-09 · Yeunju Choi, Youngmoon Jung, Hoirin Kim

While deep learning has made impressive progress in speech synthesis and voice conversion, the assessment of the synthesized speech is still carried out by human participants. Several recent papers have proposed deep-lea…

Deep LearningSpeech SynthesisVoice Conversion

Quality-Net: An End-to-End Non-intrusive Speech Quality Assessment Model based on BLSTM

2018-08-16 · Szu-Wei Fu, Yu Tsao, Hsin-Te Hwang, Hsin-Min Wang

Nowadays, most of the objective speech quality assessment tools (e.g., perceptual evaluation of speech quality (PESQ)) are based on the comparison of the degraded/processed speech with its clean counterpart. The need of …

Speech Enhancement