paper-with-me

Papers

Deep Learning Based Assessment of Synthetic Speech Naturalness

2021-04-23 · Gabriel Mittag, Sebastian Möller

In this paper, we present a new objective prediction model for synthetic speech naturalness. It can be used to evaluate Text-To-Speech or Voice Conversion systems and works language independently. The model is trained end-to-end and based on a CNN-LSTM network that previously showed to give good results for speech quality estimation. We trained and tested the model on 16 different datasets, such as from the Blizzard Challenge and the Voice Conversion Challenge. Further, we show that the reliability of deep learning-based naturalness prediction can be improved by transfer learning from speech quality prediction models that are trained on objective POLQA scores. The proposed model is made publicly available and can, for example, be used to evaluate different TTS system configurations.

📄 PDF Abstract BibTeX arXiv:2104.11673

Code (1)

gabrielmittag/NISQA 공식 구현 pytorch

Tasks

Deep LearningPredictionSpeech Synthesistext-to-speechText to SpeechTransfer LearningVoice Conversion

Similar Papers 제목 키워드 기반

Augmenting Dysarthric Speech Severity Assessment with MOS Supervision

2026-06-17 · Kaimeng Jia, Minzhu Tu, Zengrui Jin, Siyin Wang 외 arxiv

Dysarthria is a speech disorder marked by reduced intelligibility and communicative effectiveness. Automatic utterance-level assessment of dysarthric speech can support scalable speech monitoring and therapy-related anal…

Speech Synthesis

Evaluating Speech Synthesis by Training Recognizers on Synthetic Speech

2023-10-01 · Dareen Alharthi, Roshan Sharma, Hira Dhamyal, Soumi Maiti 외

Modern speech synthesis systems have improved significantly, with synthetic speech being indistinguishable from real speech. However, efficient and holistic evaluation of synthetic speech still remains a significant chal…

speech-recognitionSpeech RecognitionSpeech Synthesistext-to-speech+1

PASQA: Pitch-Accent-Focused Speech Quality Assessment Model Trained on Synthetic Speech with Accent Errors

2026-06-18 · Masaya Kawamura, Yuma Shirahata, Kentaro Mitsui, Reo Shimizu arxiv

Existing mean opinion score (MOS) prediction models typically predict utterance-level naturalness MOS and can be insensitive to localized pitch-accent errors. We propose Pitch-Accent-focused Speech Quality Assessment (PA…

The T05 System for The VoiceMOS Challenge 2024: Transfer Learning from Deep Image Classifier to Naturalness MOS Prediction of High-Quality Synthetic Speech

2024-09-14 · Kaito Baba, Wataru Nakata, Yuki Saito, Hiroshi Saruwatari

We present our system (denoted as T05) for the VoiceMOS Challenge (VMC) 2024. Our system was designed for the VMC 2024 Track 1, which focused on the accurate prediction of naturalness mean opinion score (MOS) for high-qu…

Self-Supervised LearningTransfer Learning

SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models

2026-07-06 · Thomas Thebaud, Yuzhe Wang, Hao Zhang, Sathvik Manikantan Napa Ugandhar 외 arxiv

Streaming speech-to-speech language models aim to answer spoken queries directly with synthetic speech. However, standard speech and text benchmarks do not capture whether these systems behave naturally in conversations,…