paper-with-me

Papers

QualiSpeech: A Speech Quality Assessment Dataset with Natural Language Reasoning and Descriptions

2025-03-26 · Siyin Wang, Wenyi Yu, Xianzhao Chen, Xiaohai Tian, Jun Zhang, Lu Lu, Yu Tsao, Junichi Yamagishi, Yuxuan Wang, Chao Zhang

This paper explores a novel perspective to speech quality assessment by leveraging natural language descriptions, offering richer, more nuanced insights than traditional numerical scoring methods. Natural language feedback provides instructive recommendations and detailed evaluations, yet existing datasets lack the comprehensive annotations needed for this approach. To bridge this gap, we introduce QualiSpeech, a comprehensive low-level speech quality assessment dataset encompassing 11 key aspects and detailed natural language comments that include reasoning and contextual insights. Additionally, we propose the QualiSpeech Benchmark to evaluate the low-level speech understanding capabilities of auditory large language models (LLMs). Experimental results demonstrate that finetuned auditory LLMs can reliably generate detailed descriptions of noise and distortion, effectively identifying their types and temporal characteristics. The results further highlight the potential for incorporating reasoning to enhance the accuracy and reliability of quality assessments. The dataset will be released at https://huggingface.co/datasets/tsinghua-ee/QualiSpeech.

📄 PDF Abstract BibTeX arXiv:2503.20290

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Augmenting Dysarthric Speech Severity Assessment with MOS Supervision

2026-06-17 · Kaimeng Jia, Minzhu Tu, Zengrui Jin, Siyin Wang 외 arxiv

Dysarthria is a speech disorder marked by reduced intelligibility and communicative effectiveness. Automatic utterance-level assessment of dysarthric speech can support scalable speech monitoring and therapy-related anal…

Speech Synthesis

Calibration-Reasoning Framework for Descriptive Speech Quality Assessment

2026-03-10 · Elizaveta Kostenok, Mathieu Salzmann, Milos Cernak arxiv

Explainable speech quality assessment requires moving beyond Mean Opinion Scores (MOS) to analyze underlying perceptual dimensions. To address this, we introduce a novel post-training method that tailors the foundational…

Reinforcement Learning

NVMOS: Non-Verbal Vocalization Quality Assessment in Speech

2026-06-14 · Jialong Mai, Jinxin Ji, Xiaofen Xing, Wencui Liu 외 arxiv

Non-verbal vocalizations (NVs), such as laughter, sighs, and coughs, are important acoustic cues for emotion and intent. Existing speech quality assessment methods typically focus on overall naturalness, while non-verbal…

Deep Learning Based Assessment of Synthetic Speech Naturalness

2021-04-23 · Gabriel Mittag, Sebastian Möller

In this paper, we present a new objective prediction model for synthetic speech naturalness. It can be used to evaluate Text-To-Speech or Voice Conversion systems and works language independently. The model is trained en…

Deep LearningPredictionSpeech Synthesistext-to-speech+3

Enabling Auditory Large Language Models for Automatic Speech Quality Evaluation

2024-09-25 · Siyin Wang, Wenyi Yu, Yudong Yang, Changli Tang 외

Speech quality assessment typically requires evaluating audio from multiple aspects, such as mean opinion score (MOS) and speaker similarity (SIM) \etc., which can be challenging to cover using one small model designed f…

text-to-speechText to Speech