paper-with-me

Text-To-Speech Synthesis 벤치마크

Text-To-Speech Synthesis on LJSpeech

16개 결과 · ⬇ CSV · JSON

Audio Quality MOS

2.4 2.94 3.48 4.02 4.56 2018-09 2026-09 Transformer TTS (Mel + WaveGlow) — 3.88 (2018-09-19) FastSpeech (Mel + WaveGlow) — 3.84 (2019-05-22) Merlin — 2.4 (2019-05-22) Glow-TTS + HiFiGAN — 4.34 (2020-05-22) FastSpeech 2 + HiFiGAN — 4.32 (2020-06-08) Grad-TTS + HiFiGAN (1000 steps) — 4.37 (2021-05-13) FastDiff (4 steps) — 4.28 (2022-04-21) FastDiff-TTS — 4.03 (2022-04-21) NaturalSpeech — 4.56 (2022-05-09) VITS — 4.43 (2022-05-09) FastSpeech 2 + HiFiGAN — 4.34 (2022-05-09) OverFlow — 3.37 (2022-11-13) Transformer TTS (Mel + WaveGlow) — 3.88 (2018-09-19) Glow-TTS + HiFiGAN — 4.34 (2020-05-22) Grad-TTS + HiFiGAN (1000 steps) — 4.37 (2021-05-13) NaturalSpeech — 4.56 (2022-05-09)
RankModel Audio Quality MOSPleasantness MOSWord Error Rate (WER)MOSWER (%) Extra Training Data PaperCodeYear
1 NaturalSpeech 4.56 NaturalSpeech: End-to-End Text to Speech Synthesis with Human-Level Quality microsoft/NeuralSpeech · daniilrobnikov/vits2 · heatz123/naturalspeech 2022
2 VITS 4.43 NaturalSpeech: End-to-End Text to Speech Synthesis with Human-Level Quality microsoft/NeuralSpeech · daniilrobnikov/vits2 · heatz123/naturalspeech 2022
3 Grad-TTS + HiFiGAN (1000 steps) 4.37 Grad-TTS: A Diffusion Probabilistic Model for Text-to-Speech huawei-noah/Speech-Backbones · keonlee9420/DiffGAN-TTS · keonlee9420/DiffSinger · +3 2021
4 Glow-TTS + HiFiGAN 4.34 Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment Search coqui-ai/TTS · jaywalnut310/glow-tts · supertone-inc/super-monotonic-align · +3 2020
4 FastSpeech 2 + HiFiGAN 4.34 NaturalSpeech: End-to-End Text to Speech Synthesis with Human-Level Quality microsoft/NeuralSpeech · daniilrobnikov/vits2 · heatz123/naturalspeech 2022
6 FastSpeech 2 + HiFiGAN 4.32 FastSpeech 2: Fast and High-Quality End-to-End Text to Speech coqui-ai/TTS · PaddlePaddle/PaddleSpeech · TensorSpeech/TensorflowTTS · +34 2020
7 FastDiff (4 steps) 4.28 FastDiff: A Fast Conditional Diffusion Model for High-Quality Speech Synthesis Rongjiehuang/ProDiff · Rongjiehuang/FastDiff 2022
8 FastDiff-TTS 4.03 FastDiff: A Fast Conditional Diffusion Model for High-Quality Speech Synthesis Rongjiehuang/ProDiff · Rongjiehuang/FastDiff 2022
9 Transformer TTS (Mel + WaveGlow) 3.88 Neural Speech Synthesis with Transformer Network PaddlePaddle/PaddleSpeech · as-ideas/TransformerTTS · soobinseo/transformer-tts · +3 2018
10 FastSpeech (Mel + WaveGlow) 3.84 FastSpeech: Fast, Robust and Controllable Text to Speech coqui-ai/TTS · PaddlePaddle/PaddleSpeech · ming024/FastSpeech2 · +19 2019
11 OverFlow 3.372.30 OverFlow: Putting flows on top of neural transducers for better TTS coqui-ai/TTS · shivammehta25/OverFlow 2022
12 Merlin 2.4 FastSpeech: Fast, Robust and Controllable Text to Speech coqui-ai/TTS · PaddlePaddle/PaddleSpeech · ming024/FastSpeech2 · +19 2019
13 temp 1.25
14 Flowtron 3.665 Flowtron: an Autoregressive Flow-based Generative Network for Text-to-Speech Synthesis NVIDIA/flowtron · NVIDIA/radtts · KathyReid/opensource-voice-tools 2020
15 Tacotron 2 3.521 Flowtron: an Autoregressive Flow-based Generative Network for Text-to-Speech Synthesis NVIDIA/flowtron · NVIDIA/radtts · KathyReid/opensource-voice-tools 2020
16 Matcha-TTS 3.842.09 Matcha-TTS: A fast TTS architecture with conditional flow matching shivammehta25/Matcha-TTS 2023
1–16 / 16 페이지당 10 20 50 100