paper-with-me

Papers

Diff-TTSG: Denoising probabilistic integrated speech and gesture synthesis

2023-06-15 · Shivam Mehta, Siyang Wang, Simon Alexanderson, Jonas Beskow, Éva Székely, Gustav Eje Henter

With read-aloud speech synthesis achieving high naturalness scores, there is a growing research interest in synthesising spontaneous speech. However, human spontaneous face-to-face conversation has both spoken and non-verbal aspects (here, co-speech gestures). Only recently has research begun to explore the benefits of jointly synthesising these two modalities in a single system. The previous state of the art used non-probabilistic methods, which fail to capture the variability of human speech and motion, and risk producing oversmoothing artefacts and sub-optimal synthesis quality. We present the first diffusion-based probabilistic model, called Diff-TTSG, that jointly learns to synthesise speech and gestures together. Our method can be trained on small datasets from scratch. Furthermore, we describe a set of careful uni- and multi-modal subjective tests for evaluating integrated speech and gesture synthesis systems, and use them to validate our proposed approach. Please see https://shivammehta25.github.io/Diff-TTSG/ for video examples, data, and code.

📄 PDF Abstract BibTeX arXiv:2306.09417

Code (0)

등록된 구현이 없습니다.

Tasks

DenoisingSpeech Synthesis

Methods 이 논문이 사용한 방법론

fail 설명 없음

Similar Papers 제목 키워드 기반

Unified speech and gesture synthesis using flow matching

2023-10-08 · Shivam Mehta, Ruibo Tu, Simon Alexanderson, Jonas Beskow 외

As text-to-speech technologies achieve remarkable naturalness in read-aloud tasks, there is growing interest in multimodal synthesis of verbal and non-verbal communicative behaviour, such as spontaneous speech and associ…

Audio SynthesisMotion Synthesistext-to-speechText to Speech+1

DiffGAN-TTS: High-Fidelity and Efficient Text-to-Speech with Denoising Diffusion GANs

2022-01-28 · Songxiang Liu, Dan Su, Dong Yu

Denoising diffusion probabilistic models (DDPMs) are expressive generative models that have been used to solve a variety of speech synthesis problems. However, because of their high sampling costs, DDPMs are difficult to…

DenoisingSpeech Synthesistext-to-speechText to Speech

Mandarin Singing Voice Synthesis with Denoising Diffusion Probabilistic Wasserstein GAN

2022-09-21 · Yin-Ping Cho, Yu Tsao, Hsin-Min Wang, Yi-Wen Liu

Singing voice synthesis (SVS) is the computer production of a human-like singing voice from given musical scores. To accomplish end-to-end SVS effectively and efficiently, this work adopts the acoustic model-neural vocod…

DenoisingGenerative Adversarial NetworkSinging Voice Synthesis

Grad-TTS: A Diffusion Probabilistic Model for Text-to-Speech

2021-05-13 · Vadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova 외

Recently, denoising diffusion probabilistic models and generative score matching have shown high potential in modelling complex data distributions while stochastic calculus has provided a unified point of view on these t…

DecoderSpeech Synthesistext-to-speechText to Speech+1

ProSE: Diffusion Priors for Speech Enhancement

2025-03-09 · Sonal Kumar, Sreyan Ghosh, Utkarsh Tyagi, Anton Jeran Ratnarajah 외

Speech enhancement (SE) is the foundational task of enhancing the clarity and quality of speech in the presence of non-stationary additive noise. While deterministic deep learning models have been commonly employed for S…

DenoisingregressionSpeech Enhancement