paper-with-me

홈 › Papers

Review of end-to-end speech synthesis technology based on deep learning

2021-04-20 · Zhaoxi Mu, Xinyu Yang, Yizhuo Dong

As an indispensable part of modern human-computer interaction system, speech synthesis technology helps users get the output of intelligent machine more easily and intuitively, thus has attracted more and more attention. Due to the limitations of high complexity and low efficiency of traditional speech synthesis technology, the current research focus is the deep learning-based end-to-end speech synthesis technology, which has more powerful modeling ability and a simpler pipeline. It mainly consists of three modules: text front-end, acoustic model, and vocoder. This paper reviews the research status of these three parts, and classifies and compares various methods according to their emphasis. Moreover, this paper also summarizes the open-source speech corpus of English, Chinese and other languages that can be used for speech synthesis tasks, and introduces some commonly used subjective and objective speech quality evaluation method. Finally, some attractive future research directions are pointed out.

📄 PDF Abstract BibTeX arXiv:2104.09995

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Synthesis

Similar Papers 제목 키워드 기반

A Review of Human Emotion Synthesis Based on Generative Technology

2024-12-10 · Fei Ma, Yukan Li, Yifan Xie, Ying He 외

Human emotion synthesis is a crucial aspect of affective computing. It involves using computational methods to mimic and convey human emotions through various modalities, with the goal of enabling more natural and effect…

A review-based study on different Text-to-Speech technologies

2023-12-17 · Md. Jalal Uddin Chowdhury, Ashab Hussan

This research paper presents a comprehensive review-based study on various Text-to-Speech (TTS) technologies. TTS technology is an important aspect of human-computer interaction, enabling machines to convert written text…

text-to-speechText to Speech

Neural Speech Embeddings for Speech Synthesis Based on Deep Generative Networks

2023-12-10 · Seo-Hyun Lee, Young-Eun Lee, Soowon Kim, Byung-Kwan Ko 외

Brain-to-speech technology represents a fusion of interdisciplinary applications encompassing fields of artificial intelligence, brain-computer interfaces, and speech synthesis. Neural representation learning based inten…

Representation LearningSpeech Synthesis

Macedonian Speech Synthesis for Assistive Technology Applications

2022-05-18 · Bojan Sofronievski, Elena Velovska, Martin Velichkovski, Violeta Argirova 외

Speech technology is becoming ever more ubiquitous with the advance of speech enabled devices and services. The use of speech synthesis in Augmentative and Alternative Communication tools, has facilitated inclusion of in…

Deep LearningPitch controlSpeech Synthesis

Good practices for evaluation of synthesized speech

2025-03-05 · Erica Cooper, Sébastien Le Maguer, Esther Klabbers, Junichi Yamagishi

This document is provided as a guideline for reviewers of papers about speech synthesis. We outline some best practices and common pitfalls for papers about speech synthesis, with a particular focus on evaluation. We als…

Speech Synthesis