paper-with-me

홈 › Papers

Fine-Tuning Text-to-Speech Diffusion Models Using Reinforcement Learning with Human Feedback

2025-08-05 · Jingyi Chen, Ju Seung Byun, Micha Elsner, Pichao Wang, Andrew Perrault arxiv

Diffusion models produce high-fidelity speech but are inefficient for real-time use due to long denoising steps and challenges in modeling intonation and rhythm. To improve this, we propose Diffusion Loss-Guided Policy Optimization (DLPO), an RLHF framework for TTS diffusion models. DLPO integrates the original training loss into the reward function, preserving generative capabilities while reducing inefficiencies. Using naturalness scores as feedback, DLPO aligns reward optimization with the diffusion model's structure, improving speech quality. We evaluate DLPO on WaveGrad 2, a non-autoregressive diffusion-based TTS model. Results show significant improvements in objective metrics (UTMOS 3.65, NISQA 4.02) and subjective evaluations, with DLPO audio preferred 67\% of the time. These findings demonstrate DLPO's potential for efficient, high-quality diffusion TTS in real-time, resource-limited settings.

📄 PDF Abstract BibTeX arXiv:2508.03123

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

DLPO: Diffusion Model Loss-Guided Reinforcement Learning for Fine-Tuning Text-to-Speech Diffusion Models

2024-05-23 · Jingyi Chen, Ju-Seung Byun, Micha Elsner, Andrew Perrault

Recent advancements in generative models have sparked a significant interest within the machine learning community. Particularly, diffusion models have demonstrated remarkable capabilities in synthesizing images and spee…

Image Generationreinforcement-learningReinforcement LearningSpeech Synthesis+3

Speech Recognition With LLMs Adapted to Disordered Speech Using Reinforcement Learning

2024-12-25 · Chirag Nagpal, Subhashini Venugopalan, Jimmy Tobin, Marilyn Ladewig 외

We introduce a large language model (LLM) capable of processing speech inputs and show that tuning it further with reinforcement learning on human preference (RLHF) enables it to adapt better to disordered speech than tr…

Language ModelingLanguage ModellingLarge Language Modelreinforcement-learning+3

DPOK: Reinforcement Learning for Fine-tuning Text-to-Image Diffusion Models

2023-05-25 · Ying Fan, Olivia Watkins, Yuqing Du, Hao liu 외

Learning from human feedback has been shown to improve text-to-image models. These techniques first learn a reward function that captures what humans care about in the task and then improve the models based on the learne…

reinforcement-learningReinforcement Learning (RL)

Reinforcement Learning for Fine-tuning Text-to-Image Diffusion Models

2023-09-21 · NeurIPS 2023 11

Learning from human feedback has been shown to improve text-to-image models. These techniques first learn a reward function that captures what humans care about in the task and then improve the models based on the learne…

Enhancing Diffusion Models with Text-Encoder Reinforcement Learning

2023-11-27 · Chaofeng Chen, Annan Wang, HaoNing Wu, Liang Liao 외

Text-to-image diffusion models are typically trained to optimize the log-likelihood objective, which presents challenges in meeting specific requirements for downstream tasks, such as image aesthetics and image-text alig…

reinforcement-learningReinforcement Learning