paper-with-me

Papers

A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

2023-03-23 · Chenshuang Zhang, Chaoning Zhang, Sheng Zheng, Mengchun Zhang, Maryam Qamar, Sung-Ho Bae, In So Kweon

Generative AI has demonstrated impressive performance in various fields, among which speech synthesis is an interesting direction. With the diffusion model as the most popular generative model, numerous works have attempted two active tasks: text to speech and speech enhancement. This work conducts a survey on audio diffusion model, which is complementary to existing surveys that either lack the recent progress of diffusion-based speech synthesis or highlight an overall picture of applying diffusion model in multiple fields. Specifically, this work first briefly introduces the background of audio and diffusion model. As for the text-to-speech task, we divide the methods into three categories based on the stage where diffusion model is adopted: acoustic model, vocoder and end-to-end framework. Moreover, we categorize various speech enhancement tasks by either certain signals are removed or added into the input speech. Comparisons of experimental results and discussions are also covered in this survey.

📄 PDF Abstract BibTeX arXiv:2303.13336

Code (0)

등록된 구현이 없습니다.

Tasks

Speech EnhancementSpeech SynthesisSurveytext-to-speechText to SpeechText-To-Speech Synthesis

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

CoMoSpeech: One-Step Speech and Singing Voice Synthesis via Consistency Model

2023-05-11 · Zhen Ye, Wei Xue, Xu Tan, Jie Chen 외

Denoising diffusion probabilistic models (DDPMs) have shown promising performance for speech synthesis. However, a large number of iterative steps are required to achieve high sample quality, which restricts the inferenc…

DenoisingGPUSinging Voice SynthesisSpeech Synthesis+2

Cosh-DiT: Co-Speech Gesture Video Synthesis via Hybrid Audio-Visual Diffusion Transformers

2025-03-13 · Yasheng Sun, Zhiliang Xu, Hang Zhou, Jiazhi Guan 외

Co-speech gesture video synthesis is a challenging task that requires both probabilistic modeling of human gestures and the synthesis of realistic images that align with the rhythmic nuances of speech. To address these c…

A Survey on Audio Synthesis and Audio-Visual Multimodal Processing

2021-08-01 · Zhaofeng Shi

With the development of deep learning and artificial intelligence, audio synthesis has a pivotal role in the area of machine learning and shows strong applicability in the industry. Meanwhile, significant efforts have be…

Audio SynthesisMusic GenerationSurveytext-to-speech+1

Sample-Efficient Diffusion for Text-To-Speech Synthesis

2024-09-01 · Justin Lovelace, Soham Ray, Kwangyoun Kim, Kilian Q. Weinberger 외

This work introduces Sample-Efficient Speech Diffusion (SESD), an algorithm for effective speech synthesis in modest data regimes through latent diffusion. It is based on a novel diffusion architecture, that we call U-Au…

Language ModelingLanguage ModellingSpeech Synthesistext-to-speech+2

A Survey of Voice Translation Methodologies - Acoustic Dialect Decoder

2016-10-13 · Hans Krupakar, Keerthika Rajvel, Bharathi B, Angel Deborah S 외

Speech Translation has always been about giving source text or audio input and waiting for system to give translated output in desired form. In this paper, we present the Acoustic Dialect Decoder (ADD) - a voice to voice…

DecoderSentenceSpeech SynthesisSurvey+1