paper-with-me

홈 › Papers

A Pyramid Recurrent Network for Predicting Crowdsourced Speech-Quality Ratings of Real-World Signals

2020-07-31 · Xuan Dong, Donald S. Williamson

The real-world capabilities of objective speech quality measures are limited since current measures (1) are developed from simulated data that does not adequately model real environments; or they (2) predict objective scores that are not always strongly correlated with subjective ratings. Additionally, a large dataset of real-world signals with listener quality ratings does not currently exist, which would help facilitate real-world assessment. In this paper, we collect and predict the perceptual quality of real-world speech signals that are evaluated by human listeners. We first collect a large quality rating dataset by conducting crowdsourced listening studies on two real-world corpora. We further develop a novel approach that predicts human quality ratings using a pyramid bidirectional long short term memory (pBLSTM) network with an attention mechanism. The results show that the proposed model achieves statistically lower estimation errors than prior assessment approaches, where the predicted scores strongly correlate with human judgments.

📄 PDF Abstract BibTeX arXiv:2007.15797

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A comparison of recent waveform generation and acoustic modeling methods for neural-network-based speech synthesis

2018-04-07 · Xin Wang, Jaime Lorenzo-Trueba, Shinji Takaki, Lauri Juvela 외

Recent advances in speech synthesis suggest that limitations such as the lossy nature of the amplitude spectrum with minimum phase approximation and the over-smoothing effect in acoustic modeling can be overcome by using…

Speech Synthesis

Enhancing Crowdsourced Audio for Text-to-Speech Models

2024-10-17 · José Giraldo, Martí Llopart-Font, Alex Peiró-Lilja, Carme Armentano-Oller 외

High-quality audio data is a critical prerequisite for training robust text-to-speech models, which often limits the use of opportunistic or crowdsourced datasets. This paper presents an approach to overcome this limitat…

Denoisingtext-to-speechText to Speech

Human Transcription Quality Improvement

2023-09-24 · Jian Gao, Hanbo Sun, Cheng Cao, Zheng Du

High quality transcription data is crucial for training automatic speech recognition (ASR) systems. However, the existing industry-level data collection pipelines are expensive to researchers, while the quality of crowds…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Prediction of Listener Perception of Argumentative Speech in a Crowdsourced Dataset Using (Psycho-)Linguistic and Fluency Features

2021-11-13 · Yu Qiao, Sourabh Zanwar, Rishab Bhattacharyya, Daniel Wiechmann 외

One of the key communicative competencies is the ability to maintain fluency in monologic speech and the ability to produce sophisticated language to argue a position convincingly. In this paper we aim to predict TED tal…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

A Flexible Recurrent Residual Pyramid Network for Video Frame Interpolation

2020-08-01 · ECCV 2020 8 · Haoxian Zhang, Yang Zhao, Ronggang Wang

Video frame interpolation (VFI) aims at synthesizing new video frames in-between existing frames to generate smoother high frame rate videos. Current methods usually use the fixed pre-trained networks to generate interpo…

Optical Flow EstimationVideo Frame Interpolation