paper-with-me

Papers

TTSDS -- Text-to-Speech Distribution Score

2024-07-17 · Christoph Minixhofer, Ondřej Klejch, Peter Bell

Many recently published Text-to-Speech (TTS) systems produce audio close to real speech. However, TTS evaluation needs to be revisited to make sense of the results obtained with the new architectures, approaches and datasets. We propose evaluating the quality of synthetic speech as a combination of multiple factors such as prosody, speaker identity, and intelligibility. Our approach assesses how well synthetic speech mirrors real speech by obtaining correlates of each factor and measuring their distance from both real speech datasets and noise datasets. We benchmark 35 TTS systems developed between 2008 and 2024 and show that our score computed as an unweighted average of factors strongly correlates with the human evaluations from each time period.

📄 PDF Abstract BibTeX arXiv:2407.12707

Code (1)

ttsds/ttsds 공식 구현 jax

Tasks

text-to-speechText to Speech

Similar Papers 제목 키워드 기반

TTSDS2: Resources and Benchmark for Evaluating Human-Quality Text to Speech Systems

2025-06-24 · Christoph Minixhofer, Ondrej Klejch, Peter Bell

Evaluation of Text to Speech (TTS) systems is challenging and resource-intensive. Subjective metrics such as Mean Opinion Score (MOS) are not easily comparable between works. Objective metrics are frequently used, but ra…

text-to-speechText to Speech

SimulSpeech: End-to-End Simultaneous Speech to Text Translation

2020-07-01 · ACL 2020 6 · Yi Ren, Jinglin Liu, Xu Tan, Chen Zhang 외

In this work, we develop SimulSpeech, an end-to-end simultaneous speech to text translation system which translates speech in source language to text in target language concurrently. SimulSpeech consists of a speech enco…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderKnowledge Distillation+9

Grad-TTS: A Diffusion Probabilistic Model for Text-to-Speech

2021-05-13 · Vadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova 외

Recently, denoising diffusion probabilistic models and generative score matching have shown high potential in modelling complex data distributions while stochastic calculus has provided a unified point of view on these t…

DecoderSpeech Synthesistext-to-speechText to Speech+1

Statistical Analysis of Perspective Scores on Hate Speech Detection

2021-06-22 · Hadi Mansourifar, Dana Alsagheer, Weidong Shi, Lan Ni 외

Hate speech detection has become a hot topic in recent years due to the exponential growth of offensive language in social media. It has proven that, state-of-the-art hate speech classifiers are efficient only when teste…

Hate Speech Detection

EmoSpeech: Guiding FastSpeech2 Towards Emotional Text to Speech

2023-06-28 · Daria Diatlova, Vitaly Shutov

State-of-the-art speech synthesis models try to get as close as possible to the human voice. Hence, modelling emotions is an essential part of Text-To-Speech (TTS) research. In our work, we selected FastSpeech2 as the st…

Emotion RecognitionSpeech Synthesistext-to-speechText to Speech