paper-with-me

Papers

NVMOS: Non-Verbal Vocalization Quality Assessment in Speech

2026-06-14 · Jialong Mai, Jinxin Ji, Xiaofen Xing, Wencui Liu, Xiangmin Xu arxiv

Non-verbal vocalizations (NVs), such as laughter, sighs, and coughs, are important acoustic cues for emotion and intent. Existing speech quality assessment methods typically focus on overall naturalness, while non-verbal TTS evaluations mainly examine whether a target NV appears with the correct type and position. However, the perceptual quality of NV events themselves remains underexplored. To address this gap, we construct an NV-MOS dataset containing outputs from multiple NV-TTS systems and naturally occurring NV samples, with ratings collected from three acoustic experts on a perceptual quality scale. We further analyze audio-capable multimodal large language models such as Gemini and find clear inconsistencies between their scores and expert ratings. These results suggest that general-purpose multimodal models cannot reliably replace human judgments for NV quality assessment. We then propose NVMOS, to our knowledge the first model that can reliably predict the perceptual quality of NV events in speech. Experimental results show that, with a local NV-event focusing module, NVMOS reaches expert-level or stronger agreement with human MOS.

📄 PDF Abstract BibTeX arXiv:2606.15888

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

voc2vec: A Foundation Model for Non-Verbal Vocalization

2025-02-22 · Alkis Koudounas, Moreno La Quatra, Marco Sabato Siniscalchi, Elena Baralis

Speech foundation models have demonstrated exceptional capabilities in speech-related tasks. Nevertheless, these models often struggle with non-verbal audio data, such as vocalizations, baby crying, etc., which are criti…

model

MNV-17: A High-Quality Performative Mandarin Dataset for Nonverbal Vocalization Recognition in Speech

2025-09-19 · Jialong Mai, Jinxin Ji, Xiaofen Xing, Chen Yang 외 arxiv

Mainstream Automatic Speech Recognition (ASR) systems excel at transcribing lexical content, but largely fail to recognize nonverbal vocalizations (NVs) embedded in speech, such as sighs, laughs, and coughs. This capabil…

Speech Recognition

Speaker Identity in Non-Verbal Vocalizations: Conditional Distillation and Mixture of Experts Approach

2026-06-19 · Tzu-Chieh Wei, Yi-Cheng Lin, Huang-Cheng Chou, Kuan-Yu Chen 외 arxiv

As expressive text-to-speech (TTS) and voice conversion (VC) systems increasingly generate non-verbal vocalizations (NVVs) to enhance naturalness, reliable speaker verification (SV) becomes essential to objectively asses…

Speaker VerificationVoice Conversion

NonverbalTTS: A Public English Corpus of Text-Aligned Nonverbal Vocalizations with Emotion Annotations for Text-to-Speech

2025-07-17 · Maksim Borisov, Egor Spirin, Daria Diatlova

Current expressive speech synthesis models are constrained by the limited availability of open-source datasets containing diverse nonverbal vocalizations (NVs). In this work, we introduce NonverbalTTS (NVTTS), a 17-hour …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion ClassificationExpressive Speech Synthesis+5

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations

2025-08-06 · Huan Liao, Qinke Ni, Yuancheng Wang, Yiheng Lu 외 arxiv

Paralinguistic vocalizations-including non-verbal sounds like laughter and breathing, as well as lexicalized interjections such as "uhm" and "oh"-are integral to natural spoken communication. Despite their importance in …

Speech RecognitionSpeech Synthesis