paper-with-me

홈 › Papers

AutoMOS: Learning a non-intrusive assessor of naturalness-of-speech

2016-11-28 · Brian Patton, Yannis Agiomyrgiannakis, Michael Terry, Kevin Wilson, Rif A. Saurous, D. Sculley

Developers of text-to-speech synthesizers (TTS) often make use of human raters to assess the quality of synthesized speech. We demonstrate that we can model human raters' mean opinion scores (MOS) of synthesized speech using a deep recurrent neural network whose inputs consist solely of a raw waveform. Our best models provide utterance-level estimates of MOS only moderately inferior to sampled human ratings, as shown by Pearson and Spearman correlations. When multiple utterances are scored and averaged, a scenario common in synthesizer quality assessment, AutoMOS achieves correlations approaching those of human raters. The AutoMOS model has a number of applications, such as the ability to explore the parameter space of a speech synthesizer without requiring a human-in-the-loop.

📄 PDF Abstract BibTeX arXiv:1611.09207

Code (0)

등록된 구현이 없습니다.

Tasks

text-to-speechText to Speech

Similar Papers 제목 키워드 기반

A Study on Zero-shot Non-intrusive Speech Assessment using Large Language Models

2024-09-16 · Ryandhimas E. Zezario, Sabato M. Siniscalchi, Hsin-Min Wang, Yu Tsao

This work investigates two strategies for zero-shot non-intrusive speech assessment leveraging large language models. First, we explore the audio analysis capabilities of GPT-4o. Second, we propose GPT-Whisper, which use…

Automatic Speech RecognitionPrompt Engineeringspeech-recognitionSpeech Recognition

MFCCGAN: A Novel MFCC-Based Speech Synthesizer Using Adversarial Learning

2023-06-22 · Mohammad Reza Hasanabadi Majid Behdad Davood Gharavian

In this paper, we introduce MFCCGAN as a novel speech synthesizer based on adversarial learning that adopts MFCCs as input and generates raw speech waveforms. Benefiting the GAN model capabilities, it produces speech wit…

MetricNet: Towards Improved Modeling For Non-Intrusive Speech Quality Assessment

2021-04-02 · Meng Yu, Chunlei Zhang, Yong Xu, ShiXiong Zhang 외

The objective speech quality assessment is usually conducted by comparing received speech signal with its clean reference, while human beings are capable of evaluating the speech quality without any reference, such as in…

LEIQ-Assessor: Multi-dimensional Quality Assessment of Low-light Enhanced Images via Multi-task Learning

2026-06-29 · Wei Sun, Yanwei Jiang, Dandan Zhu, Jinqiu Sang 외 arxiv

Low-light image enhancement algorithms (LIEAs) aim to improve the visibility of images captured under poor illumination. However, the enhancement process often introduces artifacts such as noise amplification, color shif…

Low-Light Image EnhancementImage Quality AssessmentMulti-Task Learning

HASA-net: A non-intrusive hearing-aid speech assessment network

2021-11-10 · Hsin-Tien Chiang, Yi-Chiao Wu, Cheng Yu, Tomoki Toda 외

Without the need of a clean reference, non-intrusive speech assessment methods have caught great attention for objective evaluations. Recently, deep neural network (DNN) models have been applied to build non-intrusive sp…