paper-with-me

홈 › Papers

From Scores to Preferences: Redefining MOS Benchmarking for Speech Quality Reward Modeling

2025-10-01 · Yifei Cao, Changhao Jiang, Jiabao Zhuang, Jiajun Sun, Ming Zhang, Zhiheng Xi, Hui Li, Shihan Dou, Yuran Wang, Yunke Zhang, Tao Ji, Tao Gui, Qi Zhang, Xuanjing Huang arxiv

Assessing the perceptual quality of synthetic speech is crucial for guiding the development and refinement of speech generation models. However, it has traditionally relied on human subjective ratings such as the Mean Opinion Score (MOS), which depend on manual annotations and often suffer from inconsistent rating standards and poor reproducibility. To address these limitations, we introduce MOS-RMBench, a unified benchmark that reformulates diverse MOS datasets into a preference-comparison setting, enabling rigorous evaluation across different datasets. Building on MOS-RMBench, we systematically construct and evaluate three paradigms for reward modeling: scalar reward models, semi-scalar reward models, and generative reward models (GRMs). Our experiments reveal three key findings: (1) scalar models achieve the strongest overall performance, consistently exceeding 74% accuracy; (2) most models perform considerably worse on synthetic speech than on human speech; and (3) all models struggle on pairs with very small MOS differences. To improve performance on these challenging pairs, we propose a MOS-aware GRM that incorporates an MOS-difference-based reward function, enabling the model to adaptively scale rewards according to the difficulty of each sample pair. Experimental results show that the MOS-aware GRM significantly improves fine-grained quality discrimination and narrows the gap with scalar models on the most challenging cases. We hope this work will establish both a benchmark and a methodological framework to foster more rigorous and scalable research in automatic speech quality assessment.

📄 PDF Abstract BibTeX arXiv:2510.00743

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Preference-based training framework for automatic speech quality assessment using deep neural network

2023-08-29 · Cheng-Hung Hu, Yusuke Yasuda, Tomoki Toda

One objective of Speech Quality Assessment (SQA) is to estimate the ranks of synthetic speech systems. However, recent SQA models are typically trained using low-precision direct scores such as mean opinion scores (MOS) …

AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation

2025-07-17 · Potsawee Manakul, Woody Haosheng Gan, Michael J. Ryan, Ali Sartaz Khan 외 arxiv

Current speech evaluation suffers from two critical limitations: the need and difficulty of designing specialized systems targeting individual audio characteristics, and poor correlation between automatic evaluation meth…

Speaker IdentificationPrompt Engineering

Preference-ASR: A Preference-Aware Test Set for Benchmarking ASR in the Era of Speech LLMs

2026-06-28 · Nithin Rao Koluguri, Sasha Meister, Nikolay Karpov, Piotr Zelasko 외 arxiv

Popular ASR test sets adopt inconsistent conventions for numbers, disfluencies, entities, and casing, while standard normalizers erase the format distinctions users care about. Current benchmarks therefore cannot measure…

On Crowdsourcing-design with Comparison Category Rating for Evaluating Speech Enhancement Algorithms

2023-06-02 · Angélica S. Z. Suárez, Clément Laroche, Line H. Clemmensen, Sneha Das

Speech enhancement techniques improve the quality or the intelligibility of an audio signal by removing unwanted noise. It is used as preprocessing in numerous applications such as speech recognition, hearing aids, broad…

Speech Enhancementspeech-recognitionSpeech Recognition

Emotion-Aligned Generation in Diffusion Text to Speech Models via Preference-Guided Optimization

2025-09-29 · Jiacheng Shi, Hongfei Du, Yangfan He, Y. Alicia Hong 외 arxiv

Emotional text-to-speech seeks to convey affect while preserving intelligibility and prosody, yet existing methods rely on coarse labels or proxy classifiers and receive only utterance-level feedback. We introduce Emotio…

Text to Speech