paper-with-me

Papers

Discrete-Time Diffusion-Like Models for Speech Synthesis

2025-09-22 · Xiaozhou Tan, Minghui Zhao, Anton Ragni arxiv

Diffusion models have attracted a lot of attention in recent years. These models view speech generation as a continuous-time process. For efficient training, this process is typically restricted to additive Gaussian noising, which is limiting. For inference, the time is typically discretized, leading to the mismatch between continuous training and discrete sampling conditions. Recently proposed discrete-time processes, on the other hand, usually do not have these limitations, may require substantially fewer inference steps, and are fully consistent between training/inference conditions. This paper explores some diffusion-like discrete-time processes and proposes some new variants. These include processes applying additive Gaussian noise, multiplicative Gaussian noise, blurring noise and a mixture of blurring and Gaussian noises. The experimental results suggest that discrete-time processes offer comparable subjective and objective speech quality to their widely popular continuous counterpart, with more efficient and consistent training and inference schemas.

📄 PDF Abstract BibTeX arXiv:2509.18470

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Synthesis

Similar Papers 제목 키워드 기반

Cosh-DiT: Co-Speech Gesture Video Synthesis via Hybrid Audio-Visual Diffusion Transformers

2025-03-13 · Yasheng Sun, Zhiliang Xu, Hang Zhou, Jiazhi Guan 외

Co-speech gesture video synthesis is a challenging task that requires both probabilistic modeling of human gestures and the synthesis of realistic images that align with the rhythmic nuances of speech. To address these c…

Continuous Speech Synthesis using per-token Latent Diffusion

2024-10-21 · Arnon Turetzky, Nimrod Shabtay, Slava Shechtman, Hagai Aronowitz 외

The success of autoregressive transformer models with discrete tokens has inspired quantization-based approaches for continuous modalities, though these often limit reconstruction quality. We therefore introduce SALAD, a…

Image GenerationQuantizationSpeech Synthesistext-to-speech+1

Boosting Diffusion Model for Spectrogram Up-sampling in Text-to-speech: An Empirical Study

2024-06-07 · Chong Zhang, Yanqing Liu, Yang Zheng, Sheng Zhao

Scaling text-to-speech (TTS) with autoregressive language model (LM) to large-scale datasets by quantizing waveform into discrete speech tokens is making great progress to capture the diversity and expressiveness in huma…

DiversityLanguage ModelingLanguage Modellingtext-to-speech+1

Towards Universal Speech Discrete Tokens: A Case Study for ASR and TTS

2023-09-14 · Yifan Yang, Feiyu Shen, Chenpeng Du, Ziyang Ma 외

Self-supervised learning (SSL) proficiency in speech-related tasks has driven research into utilizing discrete tokens for speech tasks like recognition and translation, which offer lower storage requirements and great po…

Self-Supervised Learningspeech-recognitionSpeech RecognitionSpeech Synthesis

TransFusion: Transcribing Speech with Multinomial Diffusion

2022-10-14 · Matthew Baas, Kevin Eloff, Herman Kamper

Diffusion models have shown exceptional scaling properties in the image synthesis domain, and initial attempts have shown similar benefits for applying diffusion to unconditional text synthesis. Denoising diffusion model…

DenoisingImage GenerationSentencespeech-recognition+1