paper-with-me

홈 › Papers

Speaking in Wavelet Domain: A Simple and Efficient Approach to Speed up Speech Diffusion Model

2024-02-16 · Xiangyu Zhang, Daijiao Liu, Hexin Liu, Qiquan Zhang, Hanyu Meng, Leibny Paola Garcia, Eng Siong Chng, Lina Yao

Recently, Denoising Diffusion Probabilistic Models (DDPMs) have attained leading performances across a diverse range of generative tasks. However, in the field of speech synthesis, although DDPMs exhibit impressive performance, their long training duration and substantial inference costs hinder practical deployment. Existing approaches primarily focus on enhancing inference speed, while approaches to accelerate training a key factor in the costs associated with adding or customizing voices often necessitate complex modifications to the model, compromising their universal applicability. To address the aforementioned challenges, we propose an inquiry: is it possible to enhance the training/inference speed and performance of DDPMs by modifying the speech signal itself? In this paper, we double the training and inference speed of Speech DDPMs by simply redirecting the generative target to the wavelet domain. This method not only achieves comparable or superior performance to the original model in speech synthesis tasks but also demonstrates its versatility. By investigating and utilizing different wavelet bases, our approach proves effective not just in speech synthesis, but also in speech enhancement.

📄 PDF Abstract BibTeX arXiv:2402.10642

Code (0)

등록된 구현이 없습니다.

Tasks

DenoisingSpeech EnhancementSpeech Synthesis

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
Focus 설명 없음
SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

Expressive Text-to-Speech using Style Tag

2021-04-01 · Minchan Kim, Sung Jun Cheon, Byoung Jin Choi, Jong Jin Kim 외

As recent text-to-speech (TTS) systems have been rapidly improved in speech quality and generation speed, many researchers now focus on a more challenging issue: expressive TTS. To control speaking styles, existing expre…

Language ModelingLanguage ModellingTAGtext-to-speech+1

Speaking rate attention-based duration prediction for speed control TTS

2023-10-13 · Jesuraj Bandekar, Sathvik Udupa, Abhayjeet Singh, Anjali Jayakumar 외

With the advent of high-quality speech synthesis, there is a lot of interest in controlling various prosodic attributes of speech. Speaking rate is an essential attribute towards modelling the expressivity of speech. In …

AttributeSpeech Synthesis

Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing

2025-05-27 · Jeongsoo Choi, Jaehun Kim, Joon Son Chung

This paper introduces a cross-lingual dubbing system that translates speech from one language to another while preserving key characteristics such as duration, speaker identity, and speaking speed. Despite the strong tra…

Speech-to-Speech TranslationTranslation

E-ffective: A Visual Analytic System for Exploring the Emotion and Effectiveness of Inspirational Speeches

2021-10-28 · Kevin Maher, Zeyuan Huang, Jiancheng Song, Xiaoming Deng 외

What makes speeches effective has long been a subject for debate, and until today there is broad controversy among public speaking experts about what factors make a speech effective as well as the roles of these factors …

A Fully Time-domain Neural Model for Subband-based Speech Synthesizer

2018-10-12 · Azam Rabiee, Soo-Young Lee

This paper introduces a deep neural network model for subband-based speech synthesizer. The model benefits from the short bandwidth of the subband signals to reduce the complexity of the time-domain speech generator. We …

text-to-speechText to Speech